VLDB 2026 Research / reviewers in the wild / expert
Sylvain Lobry
dblp:189/3781
· DBLP profile ↗
21ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0003-4738-2416ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 6 first-author · 4 since 2021Artificial intelligence and machine learning · 6 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | IC-EO: Interpretable Code-Based Assistant for Earth Observation
Lamia Lahouel, Laurynas Lopata, Simon Gruening, Gabriele Meoni, Gaetan Petit, Sylvain Lobry |
ICPR (7) | 6 |
| 2026 | Checkmate: Interpretable and Explainable RSVQA is the Endgame
Lucrezia Tosato, Christel Tartini-Chappuis, Syrielle Montariol, Flora Weissgerber, Sylvain Lobry, Devis Tuia |
ICPR (7) | 5 |
| 2025 | Automating Geospatial Vision Tasks with a Large Language Model Agent
Camille Kurtz, Sylvain Lobry |
ECML/PKDD (8) | 4 |
| 2024 | Segmentation-Guided Attention for Visual Question Answering from Remote Sensing ImagesabstractVisual Question Answering for Remote Sensing (RSVQA) is a task that aims at answering natural language questions about the content of a remote sensing image. The visual features extraction is therefore an essential step in a VQA pipeline. By incorporating attention mechanisms into this process, models gain the ability to focus selectively on salient regions of the image, prioritizing the most relevant visual information for a given question. In this work, we propose to embed an attention mechanism guided by segmentation into a RSVQA pipeline. We argue that segmentation plays a crucial role in guiding attention by providing a contextual understanding of the visual information, underlying specific objects or areas of interest. To evaluate this methodology, we provide a new VQA dataset that exploits very high-resolution RGB orthophotos annotated with 16 segmentation classes and question/answer pairs. Our study shows promising results of our new methodology, gaining almost 10% of overall accuracy compared to a classical method on the proposed dataset. Lucrezia Tosato, Hichem Boussaid, Flora Weissgerber, Camille Kurtz, Laurent Wendling, Sylvain Lobry |
IGARSS | 6 |
| 2023 | Transforming Multidimensional Data into Images to Overcome the Curse of DimensionalityabstractWhen dealing with high-dimensional multivariate time series classification problems, a well-known difficulty is the curse of dimensionality. In this article, we propose an original approach of transposition of multidimensional data into images to tackle the task of classification. We propose a lightweight hybrid model that take this transposed data as an input. This model contains convolutional layers as a feature extractor followed by a recurrent neural network. We apply our method to a large dataset consisting of individual patient medical records. We show that our approach allows us to significantly reduce the size of a network and increase its performance by opting for a transformation of the input data. Rebecca Leygonie, Sylvain Lobry, Guillaume Vimont, Laurent Wendling |
ICIP | 2 |
| 2023 | Automatic Simulation of SAR Images: Comparing a Deep-Learning Based Method to a Hybrid MethodabstractThis study compares two approaches for simulating synthetic aperture radar (SAR) images. The first approach uses a conditional Generative Adversarial Network (cGAN) to learn statistical image distributions from optical images. In a second approach, we generate SAR images using a electromagnetic simulator taking into input material maps obtained by segmenting optical images. We propose two metrics to evaluate the quality of the simulation. We evaluate the methods on existing Sentinel-1 SAR images of France using the DREAM database. The results suggest that the physical simulator with automatically created material maps is better suited for generating realistic SAR images compared to the cGAN approach, even if a lot of work remains to be done on the complexity of the description of the scene. Nathan Letheule, Flora Weissgerber, Sylvain Lobry, Elise Colin |
IGARSS | 3 |
| 2022 | Embedding Spatial Relations in Visual Question Answering for Remote SensingabstractRemote sensing images carry a wealth of information that is not easily accessible to end-users as it requires strong technical skills and knowledge. Visual Question Answering (VQA), a task that aims at answering an open-ended question in natural language from an image, can provide an easier access to this information. Considering the geographical information contained in remote sensing images, questions often embed an important spatial aspect, for instance regarding the relative position of two objects. Our objective is to better model the spatial relations in the construction of a ground-truth database of image/question/answer triplets and to assess the capacity a VQA model has to answer these questions. In this article, we propose to use histograms of forces to model the directional spatial relations between geo-localized objects. This allows a finer modeling of ambiguous relationships between objects and to provide different levels of assessment of a relation (e.g. object A is slightly/strictly to the west of object B). Using this new dataset, we evaluate the performances of a classical VQA model and propose a curriculum learning strategy to better take into account the varying difficulty of questions embedding spatial relations. With this approach, we show an improvement in the performances of our model, highlighting the interest of embedding spatial relations in VQA for remote sensing applications. Maxime Faure, Sylvain Lobry, Camille Kurtz, Laurent Wendling |
ICPR | 2 |
| 2022 | Language Transformers for Remote Sensing Visual Question AnsweringabstractRemote sensing visual question answering (RSVQA) opens new avenues to promote the use of satellites data, by interfacing satellite image analysis with natural language processing. Capitalizing on the remarkable advances in natural language processing and computer vision, RSVQA aims at finding an answer to a question formulated by a human user about a remote sensing image. This is achieved by extracting representations from images and questions, and then fusing them in a joint representation. Focusing on the language part of the architecture, this study compares and evaluates the adequacy to the RSVQA task of two language models, a traditional recurrent neural network (Skip-thoughts) and a recent attentionbased Transformer (BERT). We study whether large transformer models are beneficial to the task and whether fine-tuning is needed for these models to perform at their best. Our findings show that the models benefit from fine-tuning language models and that RSVQA with BERT is slightly but consistently better when properly fine-tuned. Christel Chappuis, Vincent Mendez, Eliot Walt, Sylvain Lobry, Bertrand Le Saux, Devis Tuia |
IGARSS | 4 |
| 2022 | Wasserstein Adversarial Regularization for Learning With Label NoiseabstractNoisy labels often occur in vision datasets, especially when they are obtained from crowdsourcing or Web scraping. We propose a new regularization method, which enables learning robust classifiers in presence of noisy data. To achieve this goal, we propose a new adversarial regularization scheme based on the Wasserstein distance. Using this distance allows taking into account specific relations between classes by leveraging the geometric properties of the labels space. Our Wasserstein Adversarial Regularization (WAR) encodes a selective regularization, which promotes smoothness of the classifier between some classes, while preserving sufficient complexity of the decision boundary between others. We first discuss how and why adversarial regularization can be used in the context of noise and then show the effectiveness of our method on five datasets corrupted with noisy labels: in both benchmarks and real datasets, WAR outperforms the state-of-the-art competitors. Kilian Fatras, Bharath Bhushan Damodaran, Sylvain Lobry, Rémi Flamary, Devis Tuia, Nicolas Courty |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | RSVQA Meets Bigearthnet: A New, Large-Scale, Visual Question Answering Dataset for Remote SensingabstractVisual Question Answering is a new task that can facilitate the extraction of information from images through textual queries: it aims at answering an open-ended question formulated in natural language about a given image. In this work, we introduce a new dataset to tackle the task of visual question answering on remote sensing images: this large-scale, open access dataset extracts image/question/answer triplets from the BigEarthNet dataset. This new dataset contains close to 15 millions samples and is openly available. We present the dataset construction procedure, its characteristics and first results using a deep-learning based methodology. These first results show that the task of visual question answering is challenging and opens new interesting research avenues at the interface of remote sensing and natural language processing. The dataset and the code to create and process it are open and freely available on https://rsvqa.sylvainlobry.com/ Sylvain Lobry, Begüm Demir, Devis Tuia |
IGARSS | 1 |
| 2020 | Contextual Semantic Interpretability
Diego Marcos, Ruth Fong, Sylvain Lobry, Rémi Flamary, Nicolas Courty, Devis Tuia |
ACCV (4) | 3 |
| 2020 | Learning Multi-Label Aerial Image Classification Under Label Noise: A Regularization Approach Using Word EmbeddingsabstractTraining deep neural networks requires well-annotated datasets. However, real world datasets are often noisy, especially in a multi-label scenario, i.e. where each data point can be attributed to more than one class. To this end, we propose a regularization method to learn multi-label classification networks from noisy data. This regularization is based on the assumption that semantically close classes are more likely to appear together in a given image. Hereby, we encode label correlations with prior knowledge and regularize noisy network predictions using label correlations. To evaluate its effectiveness, we perform experiments on a mutli-label aerial image dataset contaminated with controlled levels of label noise. Results indicate that networks trained using the proposed method outperform those directly learned from noisy labels and that the benefits increase proportionally to the amount of noise present. Yuansheng Hua, Sylvain Lobry, Lichao Mou, Devis Tuia, Xiao Xiang Zhu 0001 |
IGARSS | 2 |
| 2020 | Interpretable Scenicness from Sentinel-2 ImageryabstractLandscape aesthetics, or scenicness, has been identified as an important ecosystem service that contribute to human health and well-being. Currently there are no methods to inventorize landscape scenicness on a large scale. In this paper we study how to upscale local assessments of scenicness provided by human observers, and we do so by using satellite images. Moreover, we develop an explicitly interpretable CNN model that allows assessing the connections between landscape scenicness and the presence of specific landcover types. To generate the landscape scenicness ground truth, we use the ScenicOrNot crowdsourcing database, which provides geo-referenced, human-based scenicness estimates for ground based photos in Great Britain. Our results show that it is feasible to predict landscape scenicness based on satellite imagery. The interpretable model performs comparably to an unconstrained model, suggesting that it is possible to learn a semantic bottleneck that represents well the present landcover classes and still contains enough information to accurately predict the location's scenicness. Alex Levering, Diego Marcos, Sylvain Lobry, Devis Tuia |
IGARSS | 3 |
| 2020 | Fine-grained landuse characterization using ground-based pictures: a deep learning solution based on globally available dataabstractWe study the problem of landuse characterization at the urban-object level using deep learning algorithms. Traditionally, this task is performed by surveys or manual photo interpretation, which are expensive and difficult to update regularly. We seek to characterize usages at the single object level and to differentiate classes such as educational institutes, hospitals and religious places by visual cues contained in side-view pictures from Google Street View (GSV). These pictures provide geo-referenced information not only about the material composition of the objects but also about their actual usage, which otherwise is difficult to capture using other classical sources of data such as aerial imagery. Since the GSV database is regularly updated, this allows to consequently update the landuse maps, at lower costs than those of authoritative surveys. Because every urban-object is imaged from a number of viewpoints with street-level pictures, we propose a deep-learning based architecture that accepts arbitrary number of GSV pictures to predict the fine-grained landuse classes at the object level. These classes are taken from OpenStreetMap. A quantitative evaluation of the area of Île-de-France, France shows that our model outperforms other deep learning-based methods, making it a suitable alternative to manual landuse characterization. Shivangi Srivastava, John E. Vargas-Munoz, Sylvain Lobry, Devis Tuia |
Int. J. Geogr. Inf. Sci. | 3 |
| 2020 | RSVQA: Visual Question Answering for Remote Sensing DataabstractThis article introduces the task of visual question answering for remote sensing data (RSVQA). Remote sensing images contain a wealth of information, which can be useful for a wide range of tasks, including land cover classification, object counting, or detection. However, most of the available methodologies are task-specific, thus inhibiting generic and easy access to the information contained in remote sensing data. As a consequence, accurate remote sensing product generation still requires expert knowledge. With RSVQA, we propose a system to extract information from remote sensing data that is accessible to every user: we use questions formulated in natural language and use them to interact with the images. With the system, images can be queried to obtain high-level information specific to the image content or relational dependencies between objects visible in the images. Using an automatic method introduced in this article, we built two data sets (using low- and high-resolution data) of image/question/answer triplets. The information required to build the questions and answers is queried from OpenStreetMap (OSM). The data sets can be used to train (when using supervised methods) and evaluate models to solve the RSVQA task. We report the results obtained by applying a model based on convolutional neural networks (CNNs) for the visual part and a recurrent neural network (RNN) for the natural language part of this task. The model is trained on the two data sets, yielding promising results in both cases. Sylvain Lobry, Diego Marcos, Jesse Murray, Devis Tuia |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2019 | Visual Question Answering From Remote Sensing ImagesabstractRemote sensing images carry wide amounts of information beyond land cover or land use. Images contain visual and structural information that can be queried to obtain high level information about specific image content or relational dependencies between the objects sensed. This paper explores the possibility to use questions formulated in natural language as a generic and accessible way to extract this type of information from remote sensing images, i.e. visual question answering. We introduce an automatic way to create a dataset using OpenStreetMap1data and present some preliminary results. Our proposed approach is based on deep learning, and is trained using our new dataset. Sylvain Lobry, Jesse Murray, Diego Marcos, Devis Tuia |
IGARSS | 1 |
| 2019 | Half a Percent of Labels is Enough: Efficient Animal Detection in UAV Imagery Using Deep CNNs and Active LearningabstractWe present an Active Learning (AL) strategy for reusing a deep Convolutional Neural Network (CNN)-based object detector on a new data set. This is of particular interest for wildlife conservation: given a set of images acquired with an Unmanned Aerial Vehicle (UAV) and manually labeled ground truth, our goal is to train an animal detector that can be reused for repeated acquisitions, e.g., in follow-up years. Domain shifts between data sets typically prevent such a direct model application. We thus propose to bridge this gap using AL and introduce a new criterion called Transfer Sampling (TS). TS uses Optimal Transport (OT) to find corresponding regions between the source and the target data sets in the space of CNN activations. The CNN scores in the source data set are used to rank the samples according to their likelihood of being animals, and this ranking is transferred to the target data set. Unlike conventional AL criteria that exploit model uncertainty, TS focuses on very confident samples, thus allowing quick retrieval of true positives in the target data set, where positives are typically extremely rare and difficult to find by visual inspection. We extend TS with a new window cropping strategy that further accelerates sample retrieval. Our experiments show that with both strategies combined, less than half a percent of oracle-provided labels are enough to find almost 80% of the animals in challenging sets of UAV images, beating all baselines by a margin. Benjamin Kellenberger, Diego Marcos, Sylvain Lobry, Devis Tuia |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2018 | Correcting Misaligned Rural Building Annotations in Open Street Map Using Convolutional Neural Networks EvidenceabstractMapping rural buildings in developing countries is crucial to monitor and plan in those vulnerable areas. Despite the existence of some rural building annotations in OpenStreetMap (OSM), those are of insufficient quantity and quality to train models able to map large areas accurately. In particular, these annotations are very often misaligned with respect to the buildings that are present in updated aerial imagery. We propose a Markov Random Field (MRF) method to correct misaligned rural building annotations. To do so, our method uses i) the correlation between candidate aligned OSM annotations and buildings roughly detected on aerial images and ii) the local consistency of the alignment vectors. John E. Vargas-Munoz, Diego Marcos, Sylvain Lobry, Jefersson A. dos Santos, Alexandre X. Falcão, Devis Tuia |
IGARSS | 3 |
| 2017 | Double MRF for water classification in SAR images by joint detection and reflectivity estimationabstractClassification of SAR images is a challenging task as the radiometric properties of a class may not be constant throughout the image. The assumption made in most classification algorithms that a class can be modeled by constant parameters is then not valid. In this paper, we propose a classification algorithm based on two Markov random fields that accounts for local and global variations of the parameters inside the image and produces a regularized classification. This algorithm is applied on airborne TropiSAR and simulated SWOT HR data. Both quantitative and visual results are provided, demonstrating the effectiveness of the proposed method. Sylvain Lobry, Loïc Denis, Florence Tupin, Roger Fjørtoft |
IGARSS | 1 |
| 2017 | Unsupervised detection of thin water surfaces in SWOT images based on segment detection and connectionabstractThe objective of the Surface Water and Ocean Topography (SWOT) mission is to regularly monitor the height of the earth's water surfaces. One of the challenges toward obtaining global measurements of these surfaces is to detect small water areas. In this article we introduce a method for the detection of thin water surfaces, such as rivers, in SWOT images. It combines a low-level step (segment detection) with a high-level regularization of these features. The method is then tested on a simulated SWOT image. Sylvain Lobry, Florence Tupin, Roger Fjørtoft |
IGARSS | 1 |
| 2016 | A decomposition model for scatterers change detection in multi-temporal series of SAR imagesabstractThis paper presents a method for strong scatterers change detection in synthetic aperture radar (SAR) images based on a decomposition for multi-temporal series. The formulated decomposition model jointly estimates the background of the series and the scatterers. The decomposition model retrieves possible changes in scatterers and the date at which they occurred. An exact optimization method of the model is presented and applied to a TerraSAR-X time series. Sylvain Lobry, Florence Tupin, Loïc Denis |
IGARSS | 1 |