EDBT 2026 Demo / reviewers in the wild / expert
Stefano Ermon
dblp:47/8135
· DBLP profile ↗
10ranked-venue papers in the field
1as first author
6since 2021 · last 2023
0000-0003-0039-2887ORCID · verified
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 7 (1 first)Database Systems & Data Management · 2Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Graph and Geometry Generative Modeling for Drug DiscoveryabstractWith the recent progress in geometric deep learning, generative modeling, and the availability of large-scale biological datasets, molecular graph and geometry generative modeling have emerged as a highly promising direction for scientific discovery such as drug design. These generative methods enable efficient chemical space exploration and potential drug candidate generation. However, by representing molecules as 2D graphs or 3D geometries, there exist many both fundamental and challenging problems for modeling the distribution of these irregular and complex relational data. In this tutorial, we will introduce participants to the latest key developments in this field, covering important topics including 2D molecular graph generation, 3D molecular geometry generation, 2D graph to 3D geometry generation, and conditional 3D molecular geometry generation. We further include antibody generation, where we particularly consider large-size antibody molecules. For each topic, we will outline the underlying problem characteristics, summarize key challenges, present unified views of the representative approaches, and highlight future research direction and potential impacts. We anticipate this lecture-style tutorial would attract a broad audience of researchers and practitioners. Minkai Xu, Meng Liu 0015, Wengong Jin, Shuiwang Ji, Jure Leskovec, Stefano Ermon |
KDD | 6 |
| 2023 | Towards general-purpose representation learning of polygonal geometries
Gengchen Mai, Chiyu Max Jiang, Rui Zhu 0008, Yao Xuan, Ling Cai 0002, Krzysztof Janowicz, Stefano Ermon, Ni Lao |
GeoInformatica | 8 |
| 2022 | Understanding economic development in rural Africa using satellite imagery, building footprints and deep modelsabstractRecent advancements in machine learning enable cost effective methods for understanding societal and economic activities in developing countries using publicly available satellite imagery. However, this progress remains stagnant in rural areas where the largest population under poverty line resides. In this work, we explore deep models' performance in rural areas in Africa and investigate methods that improve the performance. We argue that the geographic displacement noise present in ground surveys for anonymization purposes causes misalignments between input imagery and labels and therefore hampers accuracy, which exacerbates in rural areas. We then propose to incorporate building footprints data and a novel self-attention mechanism to provide more robust and accurate predictions of socioeconomic development. We test our framework against three socioeconomic measures in 21 African countries. Our best models outperform previous baselines in most of these tasks. Amna Elmustafa, Erik Rozi, Gengchen Mai, Stefano Ermon, Marshall Burke, David B. Lobell |
SIGSPATIAL/GIS | 5 |
| 2022 | Towards a foundation model for geospatial artificial intelligence (vision paper)abstractLarge pre-trained models, also known as foundation models (FMs), are trained in a task-agnostic manner on large-scale data and can be adapted to a wide range of downstream tasks by fine tuning, few-shot, or even zero-shot learning. Despite their successes in language and vision tasks, we have yet to see an attempt to develop foundation models for geospatial artificial intelligence (GeoAI). In this work, we explore the promises and challenges for developing multimodal foundation models for GeoAI. We first show the advantages of this idea by testing the performance of existing Large pre-trained Language Models (LLMs) (e.g. GPT-2 and GPT-3) on two geospatial semantics tasks. Results indicate that these task-agnostic LLMs can outperform task-specific fully-supervised models on both tasks with 2--9% improvement in a few-shot learning setting. However, we also show the limitations of these existing foundation models given the multimodality nature of GeoAI, especially when dealing with geometries in conjunction with other modalities. So we discuss the possibility of a multimodal foundation model which can reason over various types of geospatial data through geospatial alignments. We conclude this paper by discussing the unique risks and challenges to develop such model for GeoAI. Gengchen Mai, Chris Cundy, Kristy Choi, Yingjie Hu 0001, Ni Lao, Stefano Ermon |
SIGSPATIAL/GIS | 6 |
| 2021 | Challenges in KDD and ML for Sustainable DevelopmentabstractArtificial Intelligence and machine learning techniques can offer powerful tools for addressing the greatest challenges facing humanity and helping society adapt to a rapidly changing climate, respond to disasters and pandemic crisis, and reach the United Nations (UN) Sustainable Development Goals (SDGs) by 2030. In recent approaches for mitigation and adaptation, data analytics and ML are only one part of the solution that requires interdisciplinary and methodological research and innovations. For example, challenges include multi-modal and multi-source data fusion to combine satellite imagery with other relevant data, handling noisy and missing ground data at various spatio-temporal scales, and ensembling multiple physical and ML models to improve prediction accuracy. Despite recognized successes, there are many areas where ML is not applicable, performs poorly or gives insights that are not actionable. This tutorial will survey the recent and significant contributions in KDD and ML for sustainable development and will highlight current challenges that need to be addressed to transform and equip engaged sustainability science with robust ML-based tools to support actionable decision-making for a more sustainable future. Laure Berti-Équille, David Dao, Stefano Ermon, Bedharta Goswami |
KDD | 3 |
| 2021 | Multi-agent Imitation Learning with Copulas
Hongwei Wang 0004, Lantao Yu, Zhangjie Cao, Stefano Ermon |
ECML/PKDD (1) | 4 |
| 2019 | Predicting Economic Development using Geolocated Wikipedia ArticlesabstractProgress on the UN Sustainable Development Goals (SDGs) is hampered by a persistent lack of data regarding key social, environmental, and economic indicators, particularly in developing countries. For example, data on poverty - the first of seventeen SDGs - is both spatially sparse and infrequently collected in Sub-Saharan Africa due to the high cost of surveys. Here we propose a novel method for estimating socioeconomic indicators using open-source, geolocated textual information from Wikipedia articles. We demonstrate that modern NLP techniques can be used to predict community-level asset wealth and education outcomes using nearby geolocated Wikipedia articles. When paired with nightlights satellite imagery, our method outperforms all previously published benchmarks for this prediction task, indicating the potential of Wikipedia to inform both research in the social sciences and future policy decisions. Evan Sheehan, Chenlin Meng, Matthew Tan, Burak Uzkent, Neal Jean, Marshall Burke, David B. Lobell, Stefano Ermon |
KDD | 8 |
| 2018 | Infrastructure Quality Assessment in Africa using Satellite Imagery and Deep LearningabstractThe UN Sustainable Development Goals allude to the importance of infrastructure quality in three of its seventeen goals. However, monitoring infrastructure quality in developing regions remains prohibitively expensive and impedes efforts to measure progress toward these goals. To this end, we investigate the use of widely available remote sensing data for the prediction of infrastructure quality in Africa. We train a convolutional neural network to predict ground truth labels from the Afrobarometer Round 6 survey using Landsat 8 and Sentinel 1 satellite imagery. Our best models predict infrastructure quality with AUROC scores of 0.881 on Electricity, 0.862 on Sewerage, 0.739 on Piped Water, and 0.786 on Roads using Landsat 8. These performances are significantly better than models that leverage OpenStreetMap or nighttime light intensity on the same tasks. We also demonstrate that our trained model can accurately make predictions in an unseen country after fine-tuning on a small sample of images. Furthermore, the model can be deployed in regions with limited samples to predict infrastructure outcomes with higher performance than nearest neighbor spatial interpolation. Barak Oshri, Annie Hu, Peter Adelson, Xiao Chen 0014, Pascaline Dupas, Jeremy Weinstein, Marshall Burke, David B. Lobell, Stefano Ermon |
KDD | 9 |
| 2012 | Learning Policies for Battery Usage Optimization in Electric Vehicles
Stefano Ermon, Yexiang Xue, Carla P. Gomes, Bart Selman |
ECML/PKDD (2) | 1 |
| 2012 | Feature-Enhanced Probabilistic Models for Diffusion Network Inference
Liaoruo Wang, Stefano Ermon, John E. Hopcroft |
ECML/PKDD (2) | 2 |