EDBT 2026 Demo / reviewers in the wild / expert
Javiera Castillo-Navarro
dblp:247/7727 · also Javiera Castillo Navarro
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2025
0000-0003-4917-5103ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | What to align in multimodal contrastive learning?abstractHumans perceive the world through multisensory integration, blending the information of different modalities to adapt their behavior.
Contrastive learning offers an appealing solution for multimodal self-supervised learning. Indeed, by considering each modality as a different view of the same entity, it learns to align features of different modalities in a shared representation space. However, this approach is intrinsically limited as it only learns shared or redundant information between modalities, while multimodal interactions can arise in other ways. In this work, we introduce CoMM, a Contrastive Multimodal learning strategy that enables the communication between modalities in a single multimodal space. Instead of imposing cross- or intra- modality constraints, we propose to align multimodal representations by maximizing the mutual information between augmented versions of these multimodal features. Our theoretical analysis shows that shared, synergistic and unique terms of information naturally emerge from this formulation, allowing us to estimate multimodal interactions beyond redundancy. We test CoMM both in a controlled and in a series of real-world settings: in the former, we demonstrate that CoMM effectively captures redundant, unique and synergistic information between modalities. In the latter, CoMM learns complex multimodal interactions and achieves state-of-the-art results on seven multimodal tasks. Benoit Dufumier, Javiera Castillo-Navarro, Devis Tuia, Jean-Philippe Thiran |
ICLR | 2 |
| 2024 | ConVQG: Contrastive Visual Question Generation with Multimodal GuidanceabstractAsking questions about visual environments is a crucial way for intelligent agents to understand rich multi-faceted scenes, raising the importance of Visual Question Generation (VQG) systems. Apart from being grounded to the image, existing VQG systems can use textual constraints, such as expected answers or knowledge triplets, to generate focused questions. These constraints allow VQG systems to specify the question content or leverage external commonsense knowledge that can not be obtained from the image content only. However, generating focused questions using textual constraints while enforcing a high relevance to the image content remains a challenge, as VQG systems often ignore one or both forms of grounding. In this work, we propose Contrastive Visual Question Generation (ConVQG), a method using a dual contrastive objective to discriminate questions generated using both modalities from those based on a single one. Experiments on both knowledge-aware and standard VQG benchmarks demonstrate that ConVQG outperforms the state-of-the-art methods and generates image-grounded, text-guided, and knowledge-rich questions. Our human evaluation results also show preference for ConVQG questions compared to non-contrastive baselines. Li Mi, Syrielle Montariol, Javiera Castillo-Navarro, Xianjie Dai, Antoine Bosselut, Devis Tuia |
AAAI | 3 |
| 2024 | ConGeo: Robust Cross-View Geo-Localization Across Ground View Variations
Li Mi, Chang Xu 0027, Javiera Castillo-Navarro, Syrielle Montariol, Wen Yang 0001, Antoine Bosselut, Devis Tuia |
ECCV (14) | 3 |
| 2024 | Knowledge-Aware Visual Question Generation for Remote Sensing ImagesabstractWith the rapid development of remote sensing image archives, asking questions about images has become an effective way of gathering specific information or performing image retrieval. However, automatically generated image-based questions tend to be simplistic and template-based, which hinders the real deployment of question answering or visual dialogue systems. To enrich and diversify the questions, we propose a knowledge-aware remote sensing visual question generation model, KRSVQG, that incorporates external knowledge related to the image content to improve the quality and contextual understanding of the generated questions. The model takes an image and a related knowledge triplet from external knowledge sources as inputs and leverages image captioning as an intermediary representation to enhance the image grounding of the generated questions. To assess the performance of KRSVQG, we utilized two datasets that we manually annotated: NWPU-300 and TextRS-300. Results on these two datasets demonstrate that KRSVQG outperforms existing methods and leads to knowledge-enriched questions, grounded in both image and domain knowledge. Li Mi, Javiera Castillo-Navarro, Devis Tuia |
IGARSS | 3 |
| 2024 | Training Visual Language Models with Object Detection: Grounded Change Descriptions in Satellite ImagesabstractRecently, generalist Vision Language Models (VLMs) have shown exceptional progress in tasks previously dominated by specialized computer vision models. This becomes more prevalent when visual grounding capabilities, such as the ability to reason over input text and image to generate bounding boxes around objects, are required. However, how these capabilities transfer to specialized domains such as remote sensing remains understudied, despite the recent increase in specialized models for Earth observation. In this work, we evaluate how grounding visual entities – by generating bounding-box coordinates – affects VLM performance in satellite imagery. To this end, we create two instruction-following tasks sourced from the xBD dataset, describing changes due to natural disasters observed in satellite images. We fine-tune several instances of MiniGPTv2, an open-source VLM with grounding capabilities, and evaluate their performance under the "grounded" vs. "not grounded" settings. We find that generating bounding boxes to refer to visual entities increases performance in tasks related to objects in the image, but only when the number of entities in the image is limited. João Luis Prado, Syrielle Montariol, Javiera Castillo-Navarro, Devis Tuia, Antoine Bosselut |
IGARSS | 3 |
| 2024 | Knowledge-Aware Text-Image Retrieval for Remote Sensing ImagesabstractImage-based retrieval in large Earth observation archives is challenging because one needs to navigate across thousands of candidate matches only with the query image as a guide. By using text as information supporting the visual query, the retrieval system gains in usability, but at the same time faces difficulties due to the diversity of visual signals that cannot be summarized by a short caption only. For this reason, as a matching-based task, cross-modal text–image retrieval often suffers from information asymmetry between text and images. To address this challenge, we propose a Knowledge-aware Text–Image Retrieval (KTIR) method for remote sensing images. By mining relevant information from an external knowledge graph, KTIR enriches the text scope available in the search query and alleviates the information gaps between text and images for better matching. Moreover, by integrating domain-specific knowledge, KTIR also enhances the adaptation of pretrained vision–language models to remote sensing applications. Experimental results on three commonly used remote sensing text–image retrieval benchmarks show that the proposed knowledge-aware method leads to varied and consistent retrievals, outperforming state-of-the-art retrieval methods. Li Mi, Xianjie Dai, Javiera Castillo-Navarro, Devis Tuia |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Text as a Richer Source of Supervision in Semantic Segmentation TasksabstractThis paper introduces TACOSS a text-image alignment approach that allows explainable land cover semantic segmentation by directly integrating semantic concepts encoded from texts. TACOSS combines convolutional neural networks for visual feature extraction with semantic embeddings provided by a language model. By leveraging contrastive learning approaches, we learn an alignment between the visual and the (fixed) textual representations. In addition to producing standard semantic segmentation outputs, our model enables interactive queries with RS images using natural language prompts. The experimental results obtained on 50cm resolution aerial data from Switzerland show that TACOSS performs similarly to a standard semantic segmentation model while allowing the flexible usage of in- and out-of-vocabulary terms for the interactions with the image. Valérie Zermatten, Javiera Castillo-Navarro, Lloyd Hughes, Tobias Kellenberger, Devis Tuia |
IGARSS | 2 |
| 2022 | Semi-supervised semantic segmentation in Earth Observation: the MiniFrance suite, dataset analysis and multi-task network study
Javiera Castillo-Navarro, Bertrand Le Saux, Alexandre Boulch, Nicolas Audebert, Sébastien Lefèvre |
Mach. Learn. | 1 |
| 2022 | Energy-Based Models in Earth Observation: From Generation to Semisupervised LearningabstractDeep learning, together with the availability of large amounts of data, has transformed the way we process Earth observation (EO) tasks, such as land cover mapping or image registration. Yet, today, new models are needed to push further the revolution and enable new possibilities. This work focuses on a recent framework for generative modeling and explores its applicability to the EO images. The framework learns an energy-based model (EBM) to estimate the underlying joint distribution of the data and the categories, obtaining a neural network that is able to classify and synthesize images. On these two tasks, we show that EBMs reach comparable or better performances than convolutional networks on various public EO datasets and that they are naturally adapted to semisupervised settings, with very few labeled data. Moreover, models of this kind allow us to address high-potential applications, such as out-of-distribution analysis and land cover mapping with confidence estimation. Javiera Castillo-Navarro, Bertrand Le Saux, Alexandre Boulch, Sébastien Lefèvre |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | Classification and Generation of Earth Observation Images Using a Joint Energy-Based ModelabstractDeep learning has changed unbelievably the processing of Earth Observation tasks such as land cover mapping or image registration. Yet, today new models are needed to push further the revolution and enable new possibilities. We propose a new framework for generative modelling of Earth Observation images. It learns an energy-based model to estimate the underlying distribution of the data while jointly training a deep neural network for classification. On the varied image types of the EuroSAT benchmark, we show this model obtains classification results on par with state-of-the-art and moreover allows us to tackle a wide range of high-potential applications: image synthesis, out-of-distribution testing for domain adaptation, and image completion or denoising. Javiera Castillo-Navarro, Bertrand Le Saux, Alexandre Boulch, Sébastien Lefèvre |
IGARSS | 1 |