EDBT 2026 Demo / reviewers in the wild / expert
Philipe A. Dias
dblp:215/4908 · also Philipe Ambrozio Dias
· DBLP profile ↗
16ranked-venue papers
8as first author
13since 2021 · last 2025
0000-0001-9427-7112ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 2 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Active Learning Meets Foundation Models: Fast Remote Sensing Data Annotation for Object Detection
Marvin Burges, Philipe A. Dias, Carson Woody, Sarah Walters, Dalton D. Lunga |
ICCV | 2 |
| 2024 | OReole-FM: successes and challenges toward billion-parameter foundation models for high-resolution satellite imageryabstractWhile the pretraining of Foundation Models (FMs) for remote sensing (RS) imagery is on the rise, models remain restricted to a few hundred million parameters. Scaling models to billions of parameters has been shown to yield unprecedented benefits including emergent abilities, but requires data scaling and computing resources typically not available outside industry R&D labs. In this work, we pair high-performance computing resources including Frontier supercomputer, America's first exascale system, and high-resolution optical RS data to pretrain billion-scale FMs. Our study assesses performance of different pretrained variants of vision Transformers across image classification, semantic segmentation and object detection benchmarks, which highlight the importance of data scaling for effective model scaling. Moreover, we discuss construction of a novel TIU pretraining dataset, model initialization, with data and pretrained models intended for public release. By discussing technical challenges and details often lacking in the related literature, this work is intended to offer best practices to the geospatial community toward efficient training and benchmarking of larger FMs. Philipe A. Dias, Aristeidis Tsaris, Jordan Bowman, Abhishek Potnis, Jacob Arndt, Hsiuhan Lexie Yang, Dalton D. Lunga |
SIGSPATIAL/GIS | 1 |
| 2024 | Towards Diverse and Representative Global Pretraining Datasets for Remote Sensing Foundation ModelsabstractThe design of a pretraining dataset is emerging as a critical component for the generality of foundation models. In the remote sensing realm, large volumes of imagery and benchmark datasets exist that can be leveraged to pretrain foundation models, however using this imagery in absence of a well-crafted sampling strategy is inefficient and has the potential to create biased and less generalizable models. Here, we provide a discussion and vision for the curation and assessment of pretraining datasets for remote sensing geospatial foundation models. We highlight the importance of geographic, temporal, and image acquisition diversity and review possible strategies to enable such diversity at global scale. In addition to these characteristics, support for various spatial-temporal pretext tasks within the dataset is also critical. Ultimately, our primary objective is to place emphasis on and draw attention to the data curation stage of the foundation model development pipeline. By doing so, we think it is possible to reduce biases of geospatial foundation models, as well as enable broader generalization to downstream remote sensing tasks and applications. Jacob Arndt, Philipe A. Dias, Abhishek Potnis, Dalton D. Lunga |
IGARSS | 2 |
| 2024 | Conditional Experts for Improved Building Damage Assessment Across Satellite Imagery View AnglesabstractRapid building damage assessment (BDA) is vital in guiding disaster response missions and estimating population distribution across impacted areas. While commercial satellite imagery providers have enabled near-daily monitoring of the Earth, near-realtime assessment of disaster scenarios frequently requires analysis of off-nadir imagery, as satellites are often far from impacted areas for at-nadir post-event imaging to occur Such scenarios are, however, underrepresented in existing BDA datasets and methodologies. With this motivation, we investigate generalization capabilities of current BDA practices across overhead view-angles and strategies for their improvement. Using a labeled dataset of images capturing conflict-related damages, we first train a baseline BDA architecture using imbalanced and balanced datasets with respect to view-angle. Then, we explore conditional convolutions parameterized on image features, image nadir, and their combination as a mechanism for conditioning on view-angles. Experiments demonstrate the limitations of current practice and the potential of conditional mechanisms to increase model robustness to view-angle variations. Philipe A. Dias, Jacob Arndt, Marie L. Urban, Dalton D. Lunga |
IGARSS | 1 |
| 2024 | Introducing SpaceNet 9 - Cross-Modal Satellite Imagery Registration for Natural Disaster ResponsesabstractComputer vision algorithms are increasingly leveraged to accelerate geospatial analysis for disaster response and recovery. As the diversity of remote sensing imagery grows with optical, SAR, and other modalities, a perquisite for analytics is cross-modal image registration. There is a high potential to harness computer vision for this pre-processing requirement toward enabling downstream analytics such as heterogeneous change detection, automated feature extraction, and data fusion. Advancement in these areas has the potential to simplify data wrangling tasks and further accelerate disaster response timelines. The SpaceNet 9 challenge (launching in mid-2024) focuses on addressing the cross-modal image registration problem and demonstrating the utility of such modules on earthquake impacted scenarios. This paper describes the motivation for the SpaceNet 9 and provides a first overview of the dataset, the baseline algorithm, and implications for seeking cross-modal image registration in Earth observation. Code is available at https://github.com/SpaceNetChallenge/SpaceNet9. Ronny Hänsch, Jacob Arndt, Philipe A. Dias, Abhishek Potnis, Dalton D. Lunga, Desiree Petrie, Todd M. Bacastow |
IGARSS | 3 |
| 2023 | An Agenda for Multimodal Foundation Models for Earth ObservationabstractArchives of remote sensing (RS) data are increasing swiftly as new sensing modalities with enhanced spatiotemporal resolution become operational. While promising new breakthroughs, the sheer volume of RS archives stretches the limits of human analysts and existing AI tools, as most models are: i) limited to single data modalities; ii) task-specific; iii) heavily reliant on labeled data. The emerging Foundation Models (FMs) have the potential to address these limitations. Trained on vast unlabeled datasets through self-supervised learning, FMs enable generic feature extraction that facilitate specialization to a wide variety of downstream tasks. This paper describes a vision towards an FM for multimodal Earth Observation data (FM4EO), discussing key building blocks and open challenges. We put particular emphasis on multimodal reasoning, a topic underexplored in EO. Our ultimate goal is a practical path toward FM4EO with capacity to unlock breakthroughs in few-shot learning scenarios, multimodal geographic knowledge integration, synthesis, and hypothesis generation. Philipe A. Dias, Abhishek Potnis, Sreelekha Guggilam, Hsiuhan Lexie Yang, Aristeidis Tsaris, Henry Medeiros 0001, Dalton D. Lunga |
IGARSS | 1 |
| 2023 | Scaling Automatic Vector Data Alignment to Satellite ImageryabstractGiven the tremendous volume of accessible Earth Observation (EO) data, there is a need to develop scalable Geospatial Artificial Intelligence (GeoAI) solutions for time-sensitive applications. Scalability in this context refers to rapidly processing large-scale EO data using high performance computing resources. Accurate mapping of the built environment from remote sensing (RS) imagery has been one of the crucial components in GeoAI workflows for a wide spectrum of humanitarian applications. Derived vector data of built environment is often leveraged for disaster preparedness and response activities. However, factors such as differences in ortho-rectification, atmospheric conditions and human error, results in spatial misalignment between vector data and the timely available RS imagery. Model training for downstream tasks such as object detection, change analysis, etc., is negatively impacted due to such spatial misalignment. Although there has been progress towards automatic alignment of vector data, the lack of scalability remains an open research challenge. This paper proposes to leverage parallel computing to optimize an automatic vector data alignment workflow. It further employs CPU-level multi-core parallelism for improving the performance of the workflow for scalable built environment mapping. We report observations and discuss findings from the preliminary experiments performed on the Summit Supercomputer. Abhishek Potnis, Dalton D. Lunga, Philipe A. Dias, Hsiuhan Lexie Yang, Jacob Arndt, Jordan Bowman |
IGARSS | 3 |
| 2023 | Towards Geospatial Knowledge Graph Infused Neuro-Symbolic AI for Remote Sensing Scene UnderstandingabstractDeep learning has proven its effectiveness in numerous tasks for remote sensing scene understanding. However there is an increasing interest to explore fusion of domain-specific background information to the deep neural network to further improve its performance. Remote sensing researchers are also working towards developing models that generalize and adapt to multiple applications. Generalization challenges coupled with the scarcity of large corpora of high-quality noise-free labelled data, have together fueled an interest for leveraging background information. Knowledge graphs serve as excellent choice to represent domain-specific information in a structured, standardized and extensible manner. Integrating symbolic knowledge representations in the form of Knowledge Graph Embedding (KGE) to perform neuro-symbolic reasoning is an emerging research direction promising significant impacts. This vision paper seeks to position ideas and provoke early thoughts toward advancing neuro-symbolic artificial intelligence in the context of geospatial challenges. Specifically, it conceptualizes and elaborates on an architecture for infusing geospatial knowledge from knowledge graph in a deep neural network pipeline. As guiding case studies - land-use land-cover classification, object detection and instance segmentation can benefit from infusing spatio-contextual information with remote sensing imagery. The discussion further reflects on and articulates the challenges and explainable AI opportunities anticipated when scaling and maintaining large-scale geospatial knowledge graphs. Abhishek Potnis, Dalton D. Lunga, Alexandre Sorokine, Philipe A. Dias, Hsiuhan Lexie Yang, Jacob Arndt, Jordan Bowman, Jason Wohlgemuth |
IGARSS | 4 |
| 2023 | Towards Rapid Response Updates of Populations at RiskabstractUnderstanding population at risks has been a focus of the LandScan program through its development of population estimates. With advancements in computer vision, deep learning technologies and access to High Performance Computing (HPC) and high resolution imagery, population estimates are now modeled at the building level. However, when those patterns are disrupted, rapid updates to population distribution estimates are needed to support humanitarian aid and response. Oak Ridge National Laboratory (ORNL) recently adapted an existing deep learning building footprint extraction model in development of a scalable approach to Building Damage Assessments (BDA). This new opportunity opens the possibility of automating BDA to support rapid population distribution estimate updates for geographic areas involved in geopolitical conflicts or natural events for humanitarian aid and response or where to focus recovery efforts. In addition, incorporate social surveys to further model human behavior under conflict or other scenarios that disrupt normal patterns of life. Marie L. Urban, Jessica Moehl, Philipe A. Dias, Joseph Tuccillo, Andrew Reith, Kelly M. Sims, Sarah Walters, Jacob Arndt, Abhishek Potnis, Dalton D. Lunga |
IGARSS | 3 |
| 2023 | Uncertainty-Aware Gaze Tracking for Assisted Living EnvironmentsabstractEffective assisted living environments must be able to infer how their occupants interact in a variety of scenarios. Gaze direction provides strong indications of how a person engages with the environment and its occupants. In this paper, we investigate the problem of gaze tracking in multi-camera assisted living environments. We propose a gaze tracking method based on predictions generated by a neural network regressor that relies only on the relative positions of facial keypoints to estimate gaze. For each gaze prediction, our regressor also provides an estimate of its own uncertainty, which is used to weigh the contribution of previously estimated gazes within a tracking framework based on an angular Kalman filter. Our gaze estimation neural network uses confidence gated units to alleviate keypoint prediction uncertainties in scenarios involving partial occlusions or unfavorable views of the subjects. We evaluate our method using videos from the MoDiPro dataset, which we acquired in a real assisted living facility, and on the publicly available MPIIFaceGaze, GazeFollow, and Gaze360 datasets. Experimental results show that our gaze estimation network outperforms sophisticated state-of-the-art methods, while additionally providing uncertainty predictions that are highly correlated with the actual angular error of the corresponding estimates. Finally, an analysis of the temporal integration performance of our method demonstrates that it generates accurate and temporally stable gaze predictions. Paris Her, Logan Manderle, Philipe A. Dias, Henry Medeiros 0001, Francesca Odone |
IEEE Trans. Image Process. | 3 |
| 2022 | Embedding Ethics and Trustworthiness for Sustainable AI in Earth Sciences: Where Do We Begin?abstractAs in many other research domains, Artificial Intelligence (AI) techniques have been increasing their footprint in Earth Sciences to extract meaningful information from the large amount of high-detailed data available from multiple sensor modalities. While on the one hand the existing success cases endorse the great potential of AI to help address open challenges in ES, on the other hand on-going discussions and established lessons from studies on the sustainability, ethics and trustworthiness of AI must be taken into consideration if the community is to ensure that its research efforts move into directions that effectively benefit the society and the environment. In this paper, we discuss insights gathered from a brief literature review on the subtopics of AI Ethics, Sustainable AI, AI Trustworthiness and AI for Earth Sciences in an attempt to identify some of the promising directions and key needs to successfully bring these concepts together. Philipe A. Dias, Dalton D. Lunga |
IGARSS | 1 |
| 2022 | Advancing Data Fusion in Earth SciencesabstractArtificial intelligence (AI) algorithms have proven to be quite effective in Earth observation applications, often, when extensive amounts of representative training data are available. At large, processing large volumes of observation data can be challenging due to a myriad of reasons that include the cost of acquiring labeled samples, computing resources, identifying critical data features for model prototyping, standardization of model building, and deployment. Practical novel tools and approaches are emerging across different communities. In this paper, we discuss several such recent methods from machine learning and share lessons from advanced, scalable workflows that could impact the advancement of multimodal data fusion for Earth Science applications. Dalton D. Lunga, Philipe A. Dias |
IGARSS | 2 |
| 2022 | Model Assumptions and Data Characteristics: Impacts on Domain Adaptation in Building SegmentationabstractStudies on domain adaptation (DA) for remote sensing (RS) imagery analysis lack consistency in selection and description of evaluation scenarios. Without properly characterizing datasets, model assumptions, and evaluation scenarios, it is difficult to objectively compare DA methods and reach conclusions about their suitability across different applications. With this motivation, this work seeks to empirically assess to which extent the interaction between data characteristics and model assumptions influence the effectiveness of DA methods. Using the widely explored task of building footprint segmentation as case study, we perform a large-scale study across over 200 domain adaptation scenarios that include variations across view angles, areas observed, and sensors used for data acquisition. Rather than adopting different model architectures or optimization criteria, we contrast the performances of two DA methods based on adversarial learning that differ only in their assumptions about source and target domains. Informed by metadata and data characteristics unveiled using traditional computer vision techniques as well as pre-trained deep models, we provide a detailed meta-analysis of experiments highlighting the importance of accurately considering data assumptions for DA in RS segmentation tasks. As a “cherry-picking” exercise demonstrates, different claims regarding which model is best could be made by selecting different subsets of evaluation scenarios. While well-calibrated assumptions can be beneficial, mismatching assumptions can lead to negative biases in DA applications. This study intends to motivate the community towards more consistent evaluation protocols, while providing recommendations and insights toward creating novel benchmark datasets, documenting data characteristics, application-specific knowledge, and model assumptions. Philipe A. Dias, Shawn D. Newsam, Aristeidis Tsaris, Jacob D. Hinkle, Dalton D. Lunga |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | Gaze Estimation for Assisted Living EnvironmentsabstractEffective assisted living environments must be able to perform inferences on how their occupants interact with one another as well as with surrounding objects. To accomplish this goal using a vision-based automated approach, multiple tasks such as pose estimation, object segmentation and gaze estimation must be addressed. Gaze direction provides some of the strongest indications of how a person interacts with the environment. In this paper, we propose a simple neural network regressor that estimates the gaze direction of individuals in a multi-camera assisted living scenario, relying only on the relative positions of facial keypoints collected from a single pose estimation model. To handle cases of keypoint occlusion, our model exploits a novel confidence gated unit in its input layer. In addition to the gaze direction, our model also outputs an estimation of its own prediction uncertainty. Experimental results on a public benchmark demonstrate that our approach performs on par with a complex, dataset-specific baseline, while its uncertainty predictions are highly correlated to the actual angular error of corresponding estimations. Finally, experiments on images from a real assisted living environment demonstrate that our model has a higher suitability for its final application. Philipe A. Dias, Damiano Malafronte, Henry Medeiros 0001, Francesca Odone |
WACV | 1 |
| 2019 | FreeLabel: A Publicly Available Annotation Tool Based on Freehand TracesabstractLarge-scale annotation of image segmentation datasets is often prohibitively expensive, as it usually requires a huge number of worker hours to obtain high-quality results. Abundant and reliable data has been, however, crucial for the advances on image understanding tasks recently achieved by deep learning models. In this paper, we introduce FreeLabel, an intuitive open-source web interface that allows users to obtain high-quality segmentation masks with just a few freehand scribbles, in a matter of seconds. The efficacy of FreeLabel is quantitatively demonstrated by experimental results on the PASCAL dataset as well as on a dataset from the agricultural domain. Designed to benefit the computer vision community, FreeLabel can be used for both crowdsourced or private annotation and has a modular structure that can be easily adapted for any image dataset. Philipe A. Dias, Zhou Shen, Amy Tabb, Henry Medeiros 0001 |
WACV | 1 |
| 2018 | Semantic Segmentation Refinement by Monte Carlo Region Growing of High Confidence Detections
Philipe A. Dias, Henry Medeiros 0001 |
ACCV (2) | 1 |