EDBT 2026 Demo / reviewers in the wild / expert
Diego Marcos
dblp:171/0518 · also Diego Marcos Gonzalez
· DBLP profile ↗
39ranked-venue papers
7as first author
21since 2021 · last 2026
0000-0001-5607-4445ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 19 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 15 · 5 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 5 first-author · 9 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Two-Stage Vision Transformers and Hard Masking Offer Robust Object Representations
Ananthu Aniraj, Cássio Fraga Dantas, Dino Ienco, Diego Marcos |
ICPR (6) | 4 |
| 2026 | CVGlobal and ZeSCO: Geographically Balanced Cross-View Zero-Shot Orientation Estimation
Leonardo Russo, Diego Marcos, Cássio Fraga Dantas, Dino Ienco |
ICPR (12) | 2 |
| 2025 | Hybrid Phenology Modeling for Predicting Temperature Effects on Tree DormancyabstractBiophysical models offer valuable insights into climate-phenology relationships in both natural and agricultural settings. However, there are substantial structural discrepancies across models which require site-specific recalibration, often yielding inconsistent predictions under similar climate scenarios. Machine learning methods offer data-driven solutions, but often lack interpretability and alignment with existing knowledge. We present a phenology model describing dormancy in fruit trees, integrating conventional biophysical models with a neural network to address their structural disparities. We evaluate our hybrid model in an extensive case study predicting cherry tree phenology in Japan, South Korea and Switzerland. Our approach consistently outperforms both traditional biophysical and machine learning models in predicting blooming dates across years. Additionally, the neural network's adaptability facilitates parameter learning for specific tree varieties, enabling robust generalization to new sites without site-specific recalibration. This hybrid model leverages both biophysical constraints and data-driven flexibility, offering a promising avenue for accurate and interpretable phenology modeling. Ron van Bree, Diego Marcos, Ioannis N. Athanasiadis |
AAAI | 2 |
| 2025 | LifeCLEF 2025 Teaser: Challenges on Species Presence Prediction and Identification, and Individual Animal Identification
Alexis Joly, Lukás Picek, Stefan Kahl, Hervé Goëau, Lukás Adam, Christophe Botella, Maximilien Servajean, Diego Marcos, César Leblanc, Théo Larcher, Jiri Matas, Klára Janousková, Vojtech Cermák, Kostas Papafitsoros, Robert Planqué, Willem-Pier Vellinga, Holger Klinck, Tom Denton, Pierre Bonnet, Henning Müller |
ECIR (5) | 8 |
| 2025 | SenCLIP: Enhancing Zero-Shot Land-Use Mapping for Sentinel-2 with Ground-Level PromptingabstractPre-trained vision-language models (VLMs), such as CLIP, demonstrate impressive zero-shot classification capabilities with free-form prompts and even show some generalization in specialized domains. However, their performance on satellite imagery is limited due to the underrepresentation of such data in their training sets, which predominantly consist of ground-level images. Existing prompting techniques for satellite imagery are often restricted to generic phrases like “a satellite image of …”, limiting their effectiveness for zero-shot land-use/land-cover (LULC) mapping. To address these challenges, we introduce SenCLIP, which transfers CLIP's representation to Sentinel-2 imagery by leveraging a large dataset of Sentinel-2 images paired with geotagged ground-level photos from across Europe. We evaluate SenCLIP alongside other state-of-the-art remote sensing VLMs on zero-shot LULC mapping tasks using the EuroSAT and BigEarthNet datasets with both aerial and ground-level prompting styles. Our approach, which aligns ground-level representations with satellite imagery, demonstrates significant improvements in classification accuracy across both prompt styles, opening new possibilities for applying free-form textual descriptions in zero-shot LULC mapping. Code, dataset and pretrained models are available at https://github.com/pallavijain-pj/SenCLIP Pallavi Jain 0004, Dino Ienco, Roberto Interdonato, Tristan Berchoux, Diego Marcos |
WACV | 5 |
| 2025 | Multi-Scale Grouped Prototypes for Interpretable Semantic SegmentationabstractPrototypical part learning is emerging as a promising approach for making semantic segmentation interpretable. The model selects real patches seen during training as prototypes and constructs the dense prediction map based on the similarity between parts of the test image and the prototypes. This improves interpretability since the user can inspect the link between the predicted output and the patterns learned by the model in terms of prototypical information. In this paper, we propose a method for inter-pretable semantic segmentation that leverages multi-scale image representation for prototypical part learning. First, we introduce a prototype layer that explicitly learns diverse prototypical parts at several scales, leading to multi-scale representations in the prototype activation output. Then, we propose a sparse grouping mechanism that produces multi-scale sparse groups of these scale-specific prototypical parts. This provides a deeper understanding of the interactions between multi-scale object representations while enhancing the interpretability of the segmentation model. The experiments conducted on Pascal VOC, Cityscapes, and ADE20K demonstrate that the proposed method increases model sparsity, improves interpretability over existing prototype-based methods, and narrows the performance gap with the non-interpretable counterpart models. Code is available at github.com/eceo-epfl/ScaleProtoSeg. Hugo Porta, Emanuele Dalsasso, Diego Marcos, Devis Tuia |
WACV | 3 |
| 2024 | PDiscoFormer: Relaxing Part Discovery Constraints with Vision Transformers
Ananthu Aniraj, Cássio Fraga Dantas, Dino Ienco, Diego Marcos |
ECCV (85) | 4 |
| 2024 | LifeCLEF 2024 Teaser: Challenges on Species Distribution Prediction and Identification
Alexis Joly, Lukás Picek, Stefan Kahl, Hervé Goëau, Vincent Espitalier, Christophe Botella, Benjamin Deneu, Diego Marcos, Joaquim Estopinan, César Leblanc, Théo Larcher, Milan Sulc, Marek Hrúz, Maximilien Servajean, Jiri Matas, Hervé Glotin, Robert Planqué, Willem-Pier Vellinga, Holger Klinck, Tom Denton, Andrew Durso, Ivan Eggel, Pierre Bonnet, Henning Müller |
ECIR (6) | 8 |
| 2024 | Aligning Geo-Tagged Clip Representations and Satellite Imagery for Few-Shot Land Use ClassificationabstractA major difference between ground-level and satellite imagery of landscapes lies in their semantic granularity: ground-level images tend to offer details on objects and human activities, while satellite images provide broader geographic context but, typically, with coarser semantics. This study aims to leverage this complementary information by integrating fine-grained insights from a ground-level view into the analysis of satellite image data. To achieve this integration, we propose to align a satellite image representation with co-located geo-tagged ground-level image CLIP representations. This method focuses on enriching satellite image visual features by leveraging the inherent visual characteristics found in ground-level images as a reference in a contrastive manner, without relying on additional textual information to guide the learning process. We evaluate the quality of the learned representations on the EuroSAT benchmark in various few-shot settings. Pallavi Jain 0004, Diego Marcos, Dino Ienco, Roberto Interdonato, Aayush Dhakal, Nathan Jacobs, Tristan Berchoux |
IGARSS | 2 |
| 2024 | Coarse-to-Fine Concept Bottleneck ModelsabstractDeep learning algorithms have recently gained significant attention due to their impressive performance. However, their high complexity and un-interpretable mode of operation hinders their confident deployment in real-world safety-critical tasks. This work targets ante hoc interpretability, and specifically Concept Bottleneck Models (CBMs). Our goal is to design a framework that admits a highly interpretable decision making process with respect to human understandable concepts, on two levels of granularity. To this end, we propose a novel two-level concept discovery formulation leveraging: (i) recent advances in vision-language models, and (ii) an innovative formulation for coarse-to-fine concept selection via data-driven and sparsity inducing Bayesian arguments. Within this framework, concept information does not solely rely on the similarity between the whole image and general unstructured concepts; instead, we introduce the notion of concept hierarchy to uncover and exploit more granular concept information residing in patch-specific regions of the image scene. As we experimentally show, the proposed construction not only outperforms recent CBM approaches, but also yields a principled framework towards interpetability. Konstantinos P. Panousis, Dino Ienco, Diego Marcos |
NeurIPS | 3 |
| 2024 | GeoPlant: Spatial Plant Species Prediction DatasetabstractThe difficulty of monitoring biodiversity at fine scales and over large areas limits ecological knowledge and conservation efforts. To fill this gap, Species Distribution Models (SDMs) predict species across space from spatially explicit features. Yet, they face the challenge of integrating the rich but heterogeneous data made available over the past decade, notably millions of opportunistic species observations and standardized surveys, as well as multi-modal remote sensing data.In light of that, we have designed and developed a new European-scale dataset for SDMs at high spatial resolution (10--50m), including more than 10k species (i.e., most of the European flora). The dataset comprises 5M heterogeneous Presence-Only records and 90k exhaustive Presence-Absence survey records, all accompanied by diverse environmental rasters (e.g., elevation, human footprint, and soil) traditionally used in SDMs. In addition, it provides Sentinel-2 RGB and NIR satellite images with 10 m resolution, a 20-year time series of climatic variables, and satellite time series from the Landsat program.In addition to the data, we provide an openly accessible SDM benchmark (hosted on Kaggle), which has already attracted an active community and a set of strong baselines for single predictor/modality and multimodal approaches.All resources, e.g., the dataset, pre-trained models, and baseline methods (in the form of notebooks), are available on Kaggle, allowing one to start with our dataset literally with two mouse clicks. Lukás Picek, Christophe Botella, Maximilien Servajean, César Leblanc, Rémi Palard, Théo Larcher, Benjamin Deneu, Diego Marcos, Pierre Bonnet, Alexis Joly |
NeurIPS | 8 |
| 2023 | LifeCLEF 2023 Teaser: Species Identification and Prediction Challenges
Alexis Joly, Hervé Goëau, Stefan Kahl, Lukás Picek, Christophe Botella, Diego Marcos, Milan Sulc, Marek Hrúz, Titouan Lorieul, Sara Si-Moussi, Maximilien Servajean, Benjamin Kellenberger, Elijah Cole, Andrew Durso, Hervé Glotin, Robert Planqué, Willem-Pier Vellinga, Holger Klinck, Tom Denton, Ivan Eggel, Pierre Bonnet, Henning Müller |
ECIR (3) | 6 |
| 2023 | PDiscoNet: Semantically consistent part discovery for fine-grained recognitionabstractFine-grained classification often requires recognizing specific object parts, such as beak shape and wing patterns for birds. Encouraging a fine-grained classification model to first detect such parts and then using them to infer the class could help us gauge whether the model is indeed looking at the right details better than with interpretability methods that provide a single attribution map. We propose PDiscoNet to discover object parts by using only image-level class labels along with priors encouraging the parts to be: discriminative, compact, distinct from each other, equivariant to rigid transforms, and active in at least some of the images. In addition to using the appropriate losses to encode these priors, we propose to use part-dropout, where full part feature vectors are dropped at once to prevent a single part from dominating in the classification, and part feature vector modulation, which makes the information coming from each part distinct from the perspective of the classifier. Our results on CUB, CelebA, and PartImageNet show that the proposed method provides substantially better part discovery performance than previous methods while not requiring any additional hyper-parameter tuning and without penalizing the classification performance. The code is available at https://github.com/robertdvdk/part_detection Robert van der Klis, Stephan Alaniz, Massimiliano Mancini, Cássio Fraga Dantas, Dino Ienco, Zeynep Akata, Diego Marcos |
ICCV | 7 |
| 2022 | Abstracting Sketches Through Simple Primitives
Stephan Alaniz, Massimiliano Mancini, Anjan Dutta 0001, Diego Marcos, Zeynep Akata |
ECCV (29) | 4 |
| 2022 | Training Techniques for Presence-Only Habitat Suitability Mapping with Deep LearningabstractThe goal of habitat suitability mapping is to predict the lo-cations in which a given species could be present. This is typically accomplished by statistical models which use envi-ronmental variables to predict species observation data. The relationship between the environmental characteristics of a location and the species that live there is likely to be quite complex, so deep learning models would seem natural to use. In practice, there are biases in the training data which present obstacles to standard deep learning approaches. First, large-scale species observation collections typically consist of presence-only data, which means we only have locations where a species has been observed (not where it has been confirmed to be absent). Second, the class distribution tends to be long-tailed. In this work we examine training tech-niques to mitigate these challenges: (i) a method for sharing species information between nearby observations and (ii) a curriculum learning strategy to reduce class imbalance early in training. These methods enable us to outperform state-of-the-art results on the GeoLifeCLEF 2020 dataset and suggest fruitful directions for future work. Benjamin Kellenberger, Elijah Cole, Diego Marcos, Devis Tuia |
IGARSS | 3 |
| 2022 | Semantic Segmentation of Remote Sensing Images With Sparse AnnotationsabstractTraining convolutional neural networks (CNNs) for very high-resolution images requires a large quantity of high-quality pixel-level annotations, which is extremely labor-intensive and time-consuming to produce. Moreover, professional photograph interpreters might have to be involved in guaranteeing the correctness of annotations. To alleviate such a burden, we propose a framework for semantic segmentation of aerial images based on incomplete annotations, where annotators are asked to label a few pixels with easy-to-draw scribbles. To exploit these sparse scribbled annotations, we propose the FEature and Spatial relaTional regulArization (FESTA) method to complement the supervised task with an unsupervised learning signal that accounts for neighborhood structures both in spatial and feature terms. For the evaluation of our framework, we perform experiments on two remote sensing image segmentation data sets involving aerial and satellite imagery, respectively. Experimental results demonstrate that the exploitation of sparse annotations can significantly reduce labeling costs, while the proposed method can help improve the performance of semantic segmentation when training on such annotations. The sparse labels and codes are publicly available for reproducibility purposes.https://github.com/Hua-YS/Semantic-Segmentation-with-Sparse-Labels Yuansheng Hua, Diego Marcos, Lichao Mou, Xiao Xiang Zhu 0001, Devis Tuia |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | A Semisupervised CRF Model for CNN-Based Semantic Segmentation With Sparse Ground TruthabstractConvolutional neural networks (CNNs) represent the new reference approach for semantic segmentation of very-high-resolution (VHR) images, due to their ability to automatically capture semantic information while learning relevant features. However, as for most supervised methods, the map accuracy depends on the quantity and quality of ground truth (GT) used to train them. The use of densely annotated data (i.e., a detailed, exhaustive, pixel-level GT) allows to obtain effective CNN models but normally implies high efforts in annotation. Such ground truth is often available in benchmark datasets on which new methods are tested, but not on real data for land-cover applications, where only sparse annotations might be sufficiently cost effective. A CNN model trained with such incomplete GT maps has the tendency to smooth object boundaries because they are never precisely delineated in the GT. To cope with those shortcomings, we propose to exploit the intermediate activation maps of the CNN and to deploy a semisupervised fully connected conditional random field (CRF). In comparison with competitors using the same sparse annotations, the proposed method is able to better fill part of the performance gap compared to a CNN trained on the densely annotated, but generally unavailable, GTs. Luca Maggiolo, Diego Marcos, Gabriele Moser, Sebastiano B. Serpico, Devis Tuia |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | deSpeckNet: Generalizing Deep Learning-Based SAR Image DespecklingabstractDeep learning (DL) has proven to be a suitable approach for despeckling synthetic aperture radar (SAR) images. So far, most DL models are trained to reduce speckle that follows a particular distribution, either using simulated noise or a specific set of real SAR images, limiting the applicability of these methods for real SAR images with unknown noise statistics. In this article, we present a DL method, deSpeckNet,1that estimates the speckle noise distribution and the despeckled image simultaneously. Since it does not depend on a specific noise model, deSpeckNet generalizes well across SAR acquisitions in a variety of landcover conditions. We evaluated the performance of deSpeckNet on single polarized Sentinel-1 images acquired in Indonesia, The Democratic Republic of Congo, and The Netherlands, a single polarized ALOS-2/PALSAR-2 image acquired in Japan and an Iceye X2 image acquired in Germany. In all cases, deSpeckNet was able to effectively reduce speckle and restore the images in high quality with respect to the state of the art. Adugna G. Mullissa, Diego Marcos, Devis Tuia, Martin Herold 0001, Johannes Reiche |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Learning Decision Trees Recurrently Through CommunicationabstractIntegrated interpretability without sacrificing the prediction accuracy of decision making algorithms has the potential of greatly improving their value to the user. Instead of assigning a label to an image directly, we propose to learn iterative binary sub-decisions, inducing sparsity and transparency in the decision making process. The key aspect of our model is its ability to build a decision tree whose structure is encoded into the memory representation of a Recurrent Neural Network jointly learned by two models communicating through message passing. In addition, our model assigns a semantic meaning to each decision in the form of binary attributes, providing concise, semantic and relevant rationalizations to the user. On three benchmark image classification datasets, including the large-scale ImageNet, our model generates human interpretable binary decision sequences explaining the predictions of the network while maintaining state-of-the-art accuracy. Stephan Alaniz, Diego Marcos, Bernt Schiele, Zeynep Akata |
CVPR | 2 |
| 2021 | Geo-Data for Mapping Scenic Beauty: Exploring the Potential of Remote Sensing and Social MediaabstractScenic beauty is an important contributing factor to peoples' well-being. Modelling scenic beauty has been made possible at large scales with the availability of open-source remote sensing products. At the same time, the metadata available through social media, including tags and descriptions, offer a novel modelling alternative with a personalised view from the ground. This is especially relevant to policy applications. Using a crowdsourced landscape aesthetics dataset called ScenicOrNot as ground truth, we develop and test models to predict scenic beauty based on remotely sensed indicators and image metadata from social media (Flickr). Initial results show that both model types generate strong predictions of scenic beauty and model accuracy is maximised when the two are combined. Our research shows that both a top-view measurement using remote sensing and a social media-based measurement from the ground can be used to model landscape aesthetics in support of sustainable policy goals. Ilan Havinga, Diego Marcos, Patrick W. Bogaart, Lars Hein, Devis Tuia |
IGARSS | 2 |
| 2021 | Liveability from Above: Understanding Quality of Life with Overhead Imagery and Deep Neural NetworksabstractUrban planners are increasingly interested in understanding what makes a neighbourhood pleasant and liveable. In this paper, we use the overhead perspective as a new way to describe and understand liveability of city neighborhoods. We predict building quality scores from aerial images using deep neural networks and demonstrate that liveability can be predicted from overhead aerial images of a neighbourhood. We make our model interpretable by adding the intermediate task of predicting a list of housing factors, but found this to substantially degrade the results. This suggests that the unconstrained model used visual cues that are unrelated to the housing variables, and shows the difficulty of housing variable prediction from above due to the absence of visual cues such as facades. Alex Levering, Diego Marcos, Devis Tuia |
IGARSS | 2 |
| 2020 | Contextual Semantic Interpretability
Diego Marcos, Ruth Fong, Sylvain Lobry, Rémi Flamary, Nicolas Courty, Devis Tuia |
ACCV (4) | 1 |
| 2020 | Interpretable Scenicness from Sentinel-2 ImageryabstractLandscape aesthetics, or scenicness, has been identified as an important ecosystem service that contribute to human health and well-being. Currently there are no methods to inventorize landscape scenicness on a large scale. In this paper we study how to upscale local assessments of scenicness provided by human observers, and we do so by using satellite images. Moreover, we develop an explicitly interpretable CNN model that allows assessing the connections between landscape scenicness and the presence of specific landcover types. To generate the landscape scenicness ground truth, we use the ScenicOrNot crowdsourcing database, which provides geo-referenced, human-based scenicness estimates for ground based photos in Great Britain. Our results show that it is feasible to predict landscape scenicness based on satellite imagery. The interpretable model performs comparably to an unconstrained model, suggesting that it is possible to learn a semantic bottleneck that represents well the present landcover classes and still contains enough information to accurately predict the location's scenicness. Alex Levering, Diego Marcos, Sylvain Lobry, Devis Tuia |
IGARSS | 2 |
| 2020 | Dual Polarimetric SAR Covariance Matrix Estimation Using Deep LearningabstractA polarimetric Synthetic Aperture Radar (PoISAR) image is able to capture target backscattering properties in different polarimetric states, making it a rich source of information for target characterization. However, as with any SAR image, PolSAR images are affected by speckle. Therefore, to extract useful information about targets, the polarimetric covariance matrix has to be first estimated by reducing speckle. In this paper, we use a deep neural network to estimate the dual PolSAR covariance matrix. This application was compared against the state of the art PolSAR despeckling methods. Even if the method is agnostic on the structure of the covariance matrix, the deep learning based PolSAR covariance matrix estimation performed better than the state of the art PolSAR despeckling methods. These results showcase the potential of supervised deep learning for the improvement of PolSAR despeckling pipelines. Adugna G. Mullissa, Diego Marcos, Martin Herold 0001, Johannes Reiche |
IGARSS | 2 |
| 2020 | RSVQA: Visual Question Answering for Remote Sensing DataabstractThis article introduces the task of visual question answering for remote sensing data (RSVQA). Remote sensing images contain a wealth of information, which can be useful for a wide range of tasks, including land cover classification, object counting, or detection. However, most of the available methodologies are task-specific, thus inhibiting generic and easy access to the information contained in remote sensing data. As a consequence, accurate remote sensing product generation still requires expert knowledge. With RSVQA, we propose a system to extract information from remote sensing data that is accessible to every user: we use questions formulated in natural language and use them to interact with the images. With the system, images can be queried to obtain high-level information specific to the image content or relational dependencies between objects visible in the images. Using an automatic method introduced in this article, we built two data sets (using low- and high-resolution data) of image/question/answer triplets. The information required to build the questions and answers is queried from OpenStreetMap (OSM). The data sets can be used to train (when using supervised methods) and evaluate models to solve the RSVQA task. We report the results obtained by applying a model based on convolutional neural networks (CNNs) for the visual part and a recurrent neural network (RNN) for the natural language part of this task. The model is trained on the two data sets, yielding promising results in both cases. Sylvain Lobry, Diego Marcos, Jesse Murray, Devis Tuia |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | Visual Question Answering From Remote Sensing ImagesabstractRemote sensing images carry wide amounts of information beyond land cover or land use. Images contain visual and structural information that can be queried to obtain high level information about specific image content or relational dependencies between the objects sensed. This paper explores the possibility to use questions formulated in natural language as a generic and accessible way to extract this type of information from remote sensing images, i.e. visual question answering. We introduce an automatic way to create a dataset using OpenStreetMap1data and present some preliminary results. Our proposed approach is based on deep learning, and is trained using our new dataset. Sylvain Lobry, Jesse Murray, Diego Marcos, Devis Tuia |
IGARSS | 3 |
| 2019 | Zoom In, Zoom Out: Injecting Scale Invariance into Landuse Classification CNNsabstractWe propose a Convolutional Neural Network (CNN), which encodes local scale invariance and equivariance in a multiresolution, multi-sensor image classification task. We show that the locally scale invariant model achieves results that are in line with state-of-the-art. The scale invariant and equivariant models also prove to be more robust to reductions in training data and number of filters used in each convolutional layer. These results demonstrate the benefit of disentangling scale within the learned features of CNNs, in particular when processing multi-resolution imagery. This is beneficial in the two studied cases: when training data is limited, or when the number of model parameters must be kept to a minimum. Jesse Murray, Diego Marcos, Devis Tuia |
IGARSS | 2 |
| 2019 | Half a Percent of Labels is Enough: Efficient Animal Detection in UAV Imagery Using Deep CNNs and Active LearningabstractWe present an Active Learning (AL) strategy for reusing a deep Convolutional Neural Network (CNN)-based object detector on a new data set. This is of particular interest for wildlife conservation: given a set of images acquired with an Unmanned Aerial Vehicle (UAV) and manually labeled ground truth, our goal is to train an animal detector that can be reused for repeated acquisitions, e.g., in follow-up years. Domain shifts between data sets typically prevent such a direct model application. We thus propose to bridge this gap using AL and introduce a new criterion called Transfer Sampling (TS). TS uses Optimal Transport (OT) to find corresponding regions between the source and the target data sets in the space of CNN activations. The CNN scores in the source data set are used to rank the samples according to their likelihood of being animals, and this ranking is transferred to the target data set. Unlike conventional AL criteria that exploit model uncertainty, TS focuses on very confident samples, thus allowing quick retrieval of true positives in the target data set, where positives are typically extremely rare and difficult to find by visual inspection. We extend TS with a new window cropping strategy that further accelerates sample retrieval. Our experiments show that with both strategies combined, less than half a percent of oracle-provided labels are enough to find almost 80% of the animals in challenging sets of UAV images, beating all baselines by a margin. Benjamin Kellenberger, Diego Marcos, Sylvain Lobry, Devis Tuia |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2018 | Learning Deep Structured Active Contours End-to-EndabstractThe world is covered with millions of buildings, and precisely knowing each instance's position and extents is vital to a multitude of applications. Recently, automated building footprint segmentation models have shown superior detection accuracy thanks to the usage of Convolutional Neural Networks (CNN). However, even the latest evolutions struggle to precisely delineating borders, which often leads to geometric distortions and inadvertent fusion of adjacent building instances. We propose to overcome this issue by exploiting the distinct geometric properties of buildings. To this end, we present Deep Structured Active Contours (DSAC), a novel framework that integrates priors and constraints into the segmentation process, such as continuous boundaries, smooth edges, and sharp corners. To do so, DSAC employs Active Contour Models (ACM), a family of constraint- and prior-based polygonal models. We learn ACM parameterizations per instance using a CNN, and show how to incorporate all components in a structured output model, making DSAC trainable end-to-end. We evaluate DSAC on three challenging building instance segmentation datasets, where it compares favorably against state-of-the-art. Code will be made available on https://github.com/dmarcosg/DSAC. Diego Marcos, Devis Tuia, Benjamin Kellenberger, Lisa Zhang 0003, Min Bai, Renjie Liao 0001, Raquel Urtasun |
CVPR | 1 |
| 2018 | Detecting Animals in Repeated UAV Image Acquisitions by Matching CNN Activations with Optimal TransportabstractRepeated animal censuses are crucial for wildlife parks to ensure ecological equilibriums. They are increasingly conducted using images generated by Unmanned Aerial Vehicles (UAVs), often coupled to semi-automatic object detection methods. Such methods have shown great progress also thanks to the employment of Convolutional Neural Networks (CNNs), but even the best models trained on the data acquired in one year struggle predicting animal abundances in subsequent campaigns due to the inherent shift between the datasets. In this paper we adapt a CNN-based animal detector to a follow-up UAV dataset by employing an unsupervised domain adaptation method based on Optimal Transport. We show how to infer updated labels from the source dataset by means of an ensemble of bootstraps. Our method increases the precision compared to the unmodified CNN, while not requiring additional labels from the target set. Benjamin Kellenberger, Diego Marcos, Nicolas Courty, Devis Tuia |
IGARSS | 2 |
| 2018 | Improving Maps from CNNs Trained with Sparse, Scribbled Ground Truths Using Fully Connected CRFsabstractConvolutional Neural Networks (CNNs) have become the new standard for semantic segmentation of very high resolution images. But as for other methods, the map accuracy depends on the quantity and quality of ground truth used to train them. Having densely annotated data, i.e. a detailed, pixel-level ground truth (GT), allows obtaining effective models, but requires high efforts in annotation. For this reason, it is more common and efficient to work with point or scribbled annotations rather than with dense ones. A CNN model trained with such incomplete ground truths tends to mischaracterize the shapes of the objects and to be inaccurate near their boundaries. We propose to use an approximation of a fully connected Conditional Random Field (CRF) to solve these issues, in which long range connections are accounted for through auxiliary nodes based on clustering of CNN activation features. Experiments on the ISPRS Vaihingen benchmark, where a CNN is trained only with a non-dense, scribbled ground truth, show that the proposed method can fill part of the performance gap with respect to models trained on the densely annotated, but unrealistic, ground truth. Luca Maggiolo, Diego Marcos, Gabriele Moser, Devis Tuia |
IGARSS | 2 |
| 2018 | Correcting Misaligned Rural Building Annotations in Open Street Map Using Convolutional Neural Networks EvidenceabstractMapping rural buildings in developing countries is crucial to monitor and plan in those vulnerable areas. Despite the existence of some rural building annotations in OpenStreetMap (OSM), those are of insufficient quantity and quality to train models able to map large areas accurately. In particular, these annotations are very often misaligned with respect to the buildings that are present in updated aerial imagery. We propose a Markov Random Field (MRF) method to correct misaligned rural building annotations. To do so, our method uses i) the correlation between candidate aligned OSM annotations and buildings roughly detected on aerial images and ii) the local consistency of the alignment vectors. John E. Vargas-Munoz, Diego Marcos, Sylvain Lobry, Jefersson A. dos Santos, Alexandre X. Falcão, Devis Tuia |
IGARSS | 2 |
| 2018 | Best Practices to Train Deep Models on Imbalanced Datasets - A Case Study on Animal Detection in Aerial Imagery
Benjamin Kellenberger, Diego Marcos, Devis Tuia |
ECML/PKDD (3) | 2 |
| 2017 | Rotation Equivariant Vector Field NetworksabstractIn many computer vision tasks, we expect a particular behavior of the output with respect to rotations of the input image. If this relationship is explicitly encoded, instead of treated as any other variation, the complexity of the problem is decreased, leading to a reduction in the size of the required model. In this paper, we propose the Rotation Equivariant Vector Field Networks (RotEqNet), a Convolutional Neural Network (CNN) architecture encoding rotation equivariance, invariance and covariance. Each convolutional filter is applied at multiple orientations and returns a vector field representing magnitude and angle of the highest scoring orientation at every spatial location. We develop a modified convolution operator relying on this representation to obtain deep architectures. We test RotEqNet on several problems requiring different responses with respect to the inputs' rotation: image classification, biomedical image segmentation, orientation estimation and patch matching. In all cases, we show that RotEqNet offers extremely compact models in terms of number of parameters and provides results in line to those of networks orders of magnitude larger. Diego Marcos, Michele Volpi, Nikos Komodakis, Devis Tuia |
ICCV | 1 |
| 2016 | Geospatial Correspondences for Multimodal RegistrationabstractThe growing availability of very high resolution (<;1 m/pixel) satellite and aerial images has opened up unprecedented opportunities to monitor and analyze the evolution of land-cover and land-use across the world. To do so, images of the same geographical areas acquired at different times and, potentially, with different sensors must be efficiently parsed to update maps and detect land-cover changes. However, a naϊve transfer of ground truth labels from one location in the source image to the corresponding location in the target image is generally not feasible, as these images are often only loosely registered (with up to ± 50m of non-uniform errors). Furthermore, land-cover changes in an area over time must be taken into account for an accurate ground truth transfer. To tackle these challenges, we propose a mid-level sensor-invariant representation that encodes image regions in terms of the spatial distribution of their spectral neighbors. We incorporate this representation in a Markov Random Field to simultaneously account for nonlinear mis-registrations and enforce locality priors to find matches between multi-sensor images. We show how our approach can be used to assist in several multimodal land-cover update and change detection problems. Diego Marcos, Raffay Hamid, Devis Tuia |
CVPR | 1 |
| 2016 | Learning rotation invariant convolutional filters for texture classificationabstractWe present a method for learning discriminative filters using a shallow Convolutional Neural Network (CNN). We encode rotation invariance directly in the model by tying the weights of groups of filters to several rotated versions of the canonical filter in the group. These filters can be used to extract rotation invariant features well-suited for image classification. We test this learning procedure on a texture classification benchmark, where the orientations of the training images differ from those of the test images. We obtain results comparable to the state-of-the-art. Compared to standard shallow CNNs, the proposed method obtains higher classification performance while reducing by an order of magnitude the number of parameters to be learned. Diego Marcos, Michele Volpi, Devis Tuia |
ICPR | 1 |
| 2016 | Solving structured segmentation of aerial images as puzzlesabstractTraditional approaches to structured semantic segmentation employ appearance-based classifiers to provide a class-likelihood at each spatial location and then post-process it with Markov Random Fields (MRF) to enforce label smoothness and structure in the output space. The spatial support for such techniques is usually a patch of pixels, which makes the prediction over-smoothed because the borders of objects are not explicitly taken into account. This is further exacerbated by MRF post-processing employing the standard Potts model, which tends to further over-smooth predictions at boundaries. In this paper, we propose a different but related approach: we optimize an energy function finding the optimal combination of small ground truth (GT) tiles from training data over predictions at test time, effectively solving a puzzle. We optimize over a first configuration given by a Convolutional Neural Network (CNN) output. Diego Marcos, Michele Volpi, Devis Tuia |
IGARSS | 1 |
| 2015 | Weakly supervised alignment of multisensor imagesabstractManifold alignment has become very popular in recent literature. Aligning data distributions prior to product generation is an appealing strategy, since it allows to provide data spaces that are more similar to each other, regardless of the subsequent use of the transformed data. We propose a methodology that finds a common representation among data spaces from different sensors using geographic image correspondences, or semantic ties. To cope with the strong deformations between the data spaces considered, we propose to add nonlinearities by expanding the input space with Gaussian Radial Basis Function (RBF) features with respect to the centroids of a partitioning of the data. Such features allow us to cope with nonlinear transformations, while keeping a simple and efficient linear formulation. The proposed method is multi-domain and does not require co-registration, rather only a partial degree of spatial overlap. We test it on a challenging problem of multisensor classification transferring a model trained on a WorldView 2 image to predict land cover of a 3-bands orthophoto and show that we can transfer the model with an accuracy comparable to the one that would have been obtained by a model trained on the target image with an image-specific ground truth. Diego Marcos, Gustau Camps-Valls, Devis Tuia |
IGARSS | 1 |
| 2015 | Large-scale random features for kernel regressionabstractKernel methods constitute a family of powerful machine learning algorithms, which have found wide use in remote sensing and geosciences. However, kernel methods are still not widely adopted because of the high computational cost when dealing with large scale problems, such as the inversion of radiative transfer models. This paper introduces the method of random kitchen sinks (RKS) for fast statistical retrieval of bio-geo-physical parameters. The RKS method allows to approximate a kernel matrix with a set of random bases sampled from the Fourier domain. We extend their use to other bases, such as wavelets, stumps, and Walsh expansions. We show that kernel regression is now possible for datasets with millions of examples and high dimensionality. Examples on atmospheric parameter retrieval from infrared sounders and biophysical parameter retrieval by inverting PROSAIL radiative transfer models with simulated Sentinel-2 data show the effectiveness of the technique. Valero Laparra, Diego Marcos, Devis Tuia, Gustau Camps-Valls |
IGARSS | 2 |