VLDB 2026 Research / reviewers in the wild / expert
Juan M. Lavista Ferres
dblp:190/7447
· DBLP profile ↗
15ranked-venue papers
0as first author
12since 2021 · last 2026
0000-0002-9654-3178ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Machine Learning for Sustainable Rice Production: Region-Scale Monitoring of Water-Saving Practices in Punjab, IndiaabstractRice cultivation supplies half the world's population with staple food, while also being a major driver of freshwater depletion--consuming roughly a quarter of global freshwater--and accounting for ~48% of greenhouse gas emissions from croplands. In regions like Punjab, India, where groundwater levels are plummeting at 41.6 cm/year, adopting water-saving rice farming practices is critical. Direct-Seeded Rice (DSR) and Alternate Wetting and Drying (AWD) can cut irrigation water use by 20–40% without hurting yields, yet lack of spatial data on adoption impedes effective adaptation policy and climate action. We present a machine learning framework to bridge this data gap by monitoring sustainable rice farming at scale. In collaboration with agronomy experts and a large-scale farmer training program, we obtained ground-truth data from ~1,400 fields across Punjab. Leveraging this partnership, we developed a novel dimensional classification approach that decouples sowing and irrigation practices, achieving F1 scores of 0.8 and 0.74 respectively, solely employing Sentinel-1 satellite imagery. Explainability analysis reveals that DSR classification is robust while AWD classification depends primarily on planting schedule differences, as Sentinel-1's 12-day revisit frequency cannot capture the higher frequency irrigation cycles characteristic of AWD practices. Applying this model across 3 million fields reveals spatial heterogeneity in adoption at the state level, highlighting gaps and opportunities for policy targeting. Our district-level adoption rates correlate well with government estimates (Spearman's=0.69 and Rank Biased Overlap=0.77). This study provides policymakers and sustainability programs a powerful tool to track practice adoption, inform targeted interventions, and drive data-driven policies for water conservation and climate mitigation at regional scale. Ando Shah, Rajveer Singh, Akram Zaytar, Girmaw Abebe, Caleb Robinson, Negar Tafti, Stephen A. Wood, Rahul Dodhia, Juan M. Lavista Ferres |
AAAI | 9 |
| 2025 | Fields of The World: A Machine Learning Benchmark Dataset for Global Agricultural Field Boundary SegmentationabstractCrop field boundaries are foundational datasets for agricultural monitoring and assessments but are expensive to collect manually. Machine learning (ML) methods for automatically extracting field boundaries from remotely sensed images could help realize the demand for these datasets at a global scale. However, current ML methods for field instance segmentation lack sufficient geographic coverage, accuracy, and generalization capabilities. Further, research on improving ML methods is restricted by the lack of labeled datasets representing the diversity of global agricultural fields. We present Fields of The World (FTW)---a novel ML benchmark dataset for agricultural field instance segmentation spanning 24 countries on four continents (Europe, Africa, Asia, and South America). FTW is an order of magnitude larger than previous datasets with 70,462 samples, each containing instance and semantic segmentation masks paired with multi-date, multi-spectral Sentinel-2 satellite images. We provide results from baseline models for the new FTW benchmark, show that models trained on FTW have better zero-shot and fine-tuning performance in held-out countries than models that aren't pre-trained with diverse datasets, and show positive qualitative zero-shot results of FTW models in a real-world scenario -- running on Sentinel-2 scenes over Ethiopia. Hannah Kerner, Snehal Chaudhari, Aninda Ghosh, Caleb Robinson, Eddie Choi, Nathan Jacobs, Matthias Mohr, Rahul Dodhia, Juan M. Lavista Ferres, Jennifer Marcus |
AAAI | 11 |
| 2025 | PGRID: Power Grid Reconstruction in Informal Developments Using High-Resolution Aerial Imagery
Simone Fobi, Amrita Gupta, Duncan Kebut, Seema Iyer, Luana Marotti, Rahul Dodhia, Juan M. Lavista Ferres, Anthony Ortiz |
WACV | 7 |
| 2025 | Using website referrals to identify unreliable content rabbit holesabstractDoes the URL referral structure of websites lead users into ‘rabbit holes’ of unreliable content? Past work suggests algorithmic recommender systems on sites like YouTube lead users to view more unreliable content. However, websites without algorithmic recommender systems have financial and political motivations to influence the movement of users, potentially creating browsing rabbit holes. We address this gap using browser telemetry that captures referrals to a large sample of domains rated as reliable or unreliable information sources. Our results suggest the incentives for unreliable sites to retain and monetise users create rabbit holes. After landing on an unreliable site, users are very likely to be referred to another page on the site. Further, unreliable sites are better at retaining users than reliable sites. We find less support for political motivations. While reliable and unreliable sites are largely disconnected from one another, the probability of traveling from one unreliable site to another is relatively low. Our findings indicate the need for additional focus on site-level incentives to shape traffic moving through their sites. Kevin T. Greene, Mayana Pereira, Nilima Pisharody, Rahul Dodhia, Juan M. Lavista Ferres, Jacob N. Shapiro |
Behav. Inf. Technol. | 5 |
| 2024 | Weak Labeling for Cropland Mapping in AfricaabstractCropland mapping can play a vital role in addressing environmental, agricultural, and food security challenges. However, in the context of Africa, practical applications are often hindered by the limited availability of high-resolution cropland maps. Such maps typically require extensive human labeling, thereby creating a scalability bottleneck. To address this, we propose an approach that utilizes unsupervised object clustering to refine existing weak labels, such as those obtained from global cropland maps. The refined labels, in conjunction with sparse human annotations, serve as training data for a semantic segmentation network designed to identify cropland areas. We conduct experiments to demonstrate the benefits of the improved weak labels generated by our method. In a scenario where we train our model with only 33 human-annotated labels, the F1score for the cropland category increases from 0.53 to 0.84 when we add the mined negative labels. Gilles Quentin Hacheme, Akram Zaytar, Girmaw Abebe, Caleb Robinson, Rahul Dodhia, Juan M. Lavista Ferres, Stephen A. Wood |
IGARSS | 6 |
| 2024 | Challenges of COVID-19 Case Forecasting in the US, 2020-2021abstractDuring the COVID-19 pandemic, forecasting COVID-19 trends to support planning and response was a priority for scientists and decision makers alike. In the United States, COVID-19 forecasting was coordinated by a large group of universities, companies, and government entities led by the Centers for Disease Control and Prevention and the US COVID-19 Forecast Hub (https://covid19forecasthub.org). We evaluated approximately 9.7 million forecasts of weekly state-level COVID-19 cases for predictions 1-4 weeks into the future submitted by 24 teams from August 2020 to December 2021. We assessed coverage of central prediction intervals and weighted interval scores (WIS), adjusting for missing forecasts relative to a baseline forecast, and used a Gaussian generalized estimating equation (GEE) model to evaluate differences in skill across epidemic phases that were defined by the effective reproduction number. Overall, we found high variation in skill across individual models, with ensemble-based forecasts outperforming other approaches. Forecast skill relative to the baseline was generally higher for larger jurisdictions (e.g., states compared to counties). Over time, forecasts generally performed worst in periods of rapid changes in reported cases (either in increasing or decreasing epidemic phases) with 95% prediction interval coverage dropping below 50% during the growth phases of the winter 2020, Delta, and Omicron waves. Ideally, case forecasts could serve as a leading indicator of changes in transmission dynamics. However, while most COVID-19 case forecasts outperformed a naïve baseline model, even the most accurate case forecasts were unreliable in key phases. Further research could improve forecasts of leading indicators, like COVID-19 cases, by leveraging additional real-time data, addressing performance across phases, improving the characterization of forecast confidence, and ensuring that forecasts were coherent across spatial scales. In the meantime, it is critical for forecast users to appreciate current limitations and use a broad set of indicators to inform pandemic-related decision making. Velma K. Lopez, Estee Y. Cramer, Robert Pagano, John M. Drake, Eamon B. O'Dea, Madeline Adee, Turgay Ayer, Jagpreet Chhatwal, Ozden O. Dalgic, Mary A. Ladd, Benjamin P. Linas, Peter P. Mueller, Jade Xiao, Johannes Bracher, Alvaro J. Castro Rivadeneira, Aaron Gerding, Tilmann Gneiting, Yuxin Huang 0009, Dasuni Jayawardena, Abdul H. Kanji, Khoa Le, Anja Mühlemann, Jarad Niemi, Evan L. Ray, Ariane Stark, Nutcha Wattanachit, Martha W. Zorn, Sen Pei, Jeffrey Shaman, Teresa K. Yamana, Samuel R. Tarasewicz, Daniel J. Wilson 0002, Sid Baccam, Heidi Gurung, Steve Stage, Brad Suchoski, Lei Gao 0011, Zhiling Gu, Myungjin Kim, Guannan Wang, Li Wang 0035, Yueying Wang, Lauren Gardner, Sonia Jindal, Maximilian Marshall, Kristen Nixon, Juan Dent, Alison L. Hill, Joshua Kaminsky, Elizabeth C. Lee, Joseph Chadi Lemaitre, Justin Lessler, Claire P. Smith, Shaun Truelove, Matt Kinsey, Luke C. Mullany, Kaitlin Rainwater-Lovett, Lauren Shin, Katharine Tallaksen, Shelby Wilson, Dean Karlen, Lauren A. Castro, Geoffrey Fairchild, Isaac Michaud, Dave Osthus, Jiang Bian 0002, Wei Cao 0007, Zhifeng Gao, Juan M. Lavista Ferres, Chaozhuo Li, Tie-Yan Liu, Xing Xie 0001, Shun Zheng 0001, Matteo Chinazzi, Jessica T. Davis, Kunpeng Mu, Ana L. Pastore y Piontti, Alessandro Vespignani, Xinyue Xiong, Robert Walraven, Quanquan Gu, Lingxiao Wang 0001, Pan Xu 0002, Difan Zou, Graham Casey Gibson, Daniel Sheldon, Ajitesh Srivastava, Aniruddha Adiga, Benjamin Hurt, Gursharn Kaur, Bryan L. Lewis, Madhav V. Marathe, Akhil Sai Peddireddy, Przemyslaw J. Porebski, Srinivasan Venkatramanan, Lijing Wang 0001, Pragati V. Prasad, Jo W. Walker, Alexander E. Webber, Rachel B. Slayton, Matthew Biggerstaff, Nicholas G. Reich, Michael A. Johansson |
PLoS Comput. Biol. | 72 |
| 2023 | Poverty rate prediction using multi-modal survey and earth observation dataabstractThis work presents an approach for combining household demographic and living standards survey questions with features derived from satellite imagery to predict the poverty rate of a region. Our approach utilizes visual features obtained from a single-step featurization method applied to freely available 10m/px Sentinel-2 surface reflectance satellite imagery. These visual features are combined with ten survey questions in a proxy means test (PMT) to estimate whether a household is below the poverty line. We show that the inclusion of visual features reduces the mean error in poverty rate estimates from 4.09% to 3.88% over a nationally representative out-of-sample test set. In addition to including satellite imagery features in proxy means tests, we propose an approach for selecting a subset of survey questions that are complementary to the visual features extracted from satellite imagery. Specifically, we design a survey variable selection approach guided by the full survey and image features and use the approach to determine the most relevant set of small survey questions to include in a PMT. We validate the choice of small survey questions in a downstream task of predicting the poverty rate using the small set of questions. This approach results in the best performance – errors in poverty rate decrease from 4.09% to 3.71%. We show that extracted visual features encode geographic and urbanization differences between regions. Simone Fobi, Manuel Cardona 0003, Elliott Collins, Caleb Robinson, Anthony Ortiz, Tina Sederholm, Rahul Dodhia, Juan M. Lavista Ferres |
COMPASS | 8 |
| 2022 | TorchGeo: deep learning with geospatial dataabstractRemotely sensed geospatial data are critical for applications including precision agriculture, urban planning, disaster monitoring and response, and climate change research, among others. Deep learning methods are particularly promising for modeling many remote sensing tasks given the success of deep neural networks in similar computer vision tasks and the sheer volume of remotely sensed imagery available. However, the variance in data collection methods and handling of geospatial metadata make the application of deep learning methodology to remotely sensed data nontrivial. For example, satellite imagery often includes additional spectral bands beyond red, green, and blue and must be joined to other geospatial data sources that can have differing coordinate systems, bounds, and resolutions. To help realize the potential of deep learning for remote sensing applications, we introduce TorchGeo, a Python library for integrating geospatial data into the PyTorch deep learning ecosystem. TorchGeo provides data loaders for a variety of benchmark datasets, composable datasets for generic geospatial data sources, samplers for geospatial data, and transforms that work with multispectral imagery. TorchGeo is also the first library to provide pre-trained models for multispectral satellite imagery (e.g., models that use all bands from the Sentinel-2 satellites), allowing for advances in transfer learning on downstream remote sensing tasks with limited labeled data. We use TorchGeo to create reproducible benchmark results on existing datasets and benchmark our proposed method for preprocessing geospatial imagery on the fly. TorchGeo is open source and available on GitHub: https://github.com/microsoft/torchgeo. Adam J. Stewart, Caleb Robinson, Isaac Corley, Anthony Ortiz, Juan M. Lavista Ferres, Arindam Banerjee 0001 |
SIGSPATIAL/GIS | 5 |
| 2021 | Where there's Smoke, there's Fire: Wildfire Risk Predictive Modeling via Historical Climate DataabstractWildfire is a growing global crisis with devastating consequences. Uncontrolled wildfires take away human lives, destroy millions of animals and trees, degrade the air quality, impact the biodiversity of the planet and cause substantial economic costs. It is incredibly challenging to predict the spatio-temporal likelihood of wildfires based on historical data, due to their stochastic nature. Crucially though, the accurate and reliable prediction of wildfires can help the stakeholders and decision-makers take timely, strategic and effective actions to prevent, detect and suppress the wildfires before they become unmanageable. Unfortunately, most previous studies developed predictive models that suffer from some shortcomings: (i) they do not take the temporal aspects into account precisely and they assume the independent and identically distributed random variables in the evaluation phase; (ii) they do not evaluate their approaches comprehensively, thus it is not clear if their proposed predictions and selected models are reliable across different locations and time steps for practical deployment; and (iii) for the supervised learning models, they use predictor features and fire observations from the same time step in the training phase, which makes the inference task infeasible for future fire prediction. In this paper, we revisit the wildfire predictive modeling, explore the inherent challenges from a practical perspective and evaluate our modeling approach comprehensively via historical burned areas, climate and geospatial data from three vast landscapes in India. Shahrzad Gholami, Narendran Kodandapani, Juan M. Lavista Ferres |
AAAI | 4 |
| 2021 | Becoming Good at AI for GoodabstractAI for good (AI4G) projects involve developing and applying artificial intelligence (AI) based solutions to further goals in areas such as sustainability, health, humanitarian aid, and social justice. Developing and deploying such solutions must be done in collaboration with partners who are experts in the domain in question and who already have experience in making progress towards such goals. Based on our experiences, we detail the different aspects of this type of collaboration broken down into four high-level categories: communication, data, modeling, and impact, and distill eleven takeaways to guide such projects in the future. We briefly describe two case studies to illustrate how some of these takeaways were applied in practice during our past collaborations. Meghana Kshirsagar 0001, Caleb Robinson, Shahrzad Gholami, Ivan S. Klyuzhin, Sumit Mukherjee, Md Nasir, Anthony Ortiz, Felipe Oviedo, Darren Tanner, Anusua Trivedi, Yixi Xu, Ming Zhong 0014, Bistra Dilkina, Rahul Dodhia, Juan M. Lavista Ferres |
AIES | 16 |
| 2021 | Temporal Cluster Matching for Change Detection of Structures from Satellite ImageryabstractLongitudinal studies are vital to understanding dynamic changes of the planet, but labels (e.g., buildings, facilities, roads) are often available only for a single point in time. We propose a general model, Temporal Cluster Matching (TCM), for detecting building changes in time series of remotely sensed imagery when footprint labels are observed only once. The intuition behind the model is that the relationship between spectral values inside and outside of building’s footprint will change when a building is constructed (or demolished). For instance, in rural settings, the pre-construction area may look similar to the surrounding environment until the building is constructed. Similarly, in urban settings, the pre-construction areas will look different from the surrounding environment until construction. We further propose a heuristic method for selecting the parameters of our model which allows it to be applied in novel settings without requiring data labeling efforts (to fit the parameters). We apply our model over a dataset of poultry barns from 2016/2017 high-resolution aerial imagery in the Delmarva Peninsula and a dataset of solar farms from a 2020 mosaic of Sentinel 2 imagery in India. Our results show that our model performs as well when fit using the proposed heuristic as it does when fit with labeled data, and further, that supervised versions of our model perform the best among all the baselines we test against. Finally, we show that our proposed approach can act as an effective data augmentation strategy – it enables researchers to augment existing structure footprint labels along the time dimension and thus use imagery from multiple points in time to train deep learning models. We show that this improves the spatial generalization of such models when evaluated on the same change detection task. Caleb Robinson, Anthony Ortiz, Juan M. Lavista Ferres, Brandon R. Anderson, Daniel E. Ho |
COMPASS | 3 |
| 2021 | privGAN: Protecting GANs from membership inference attacks at low cost to utilityabstractAbstract Generative Adversarial Networks (GANs) have made releasing of synthetic images a viable approach to share data without releasing the original dataset. It has been shown that such synthetic data can be used for a variety of downstream tasks such as training classifiers that would otherwise require the original dataset to be shared. However, recent work has shown that the GAN models and their synthetically generated data can be used to infer the training set membership by an adversary who has access to the entire dataset and some auxiliary information. Current approaches to mitigate this problem (such as DPGAN [1]) lead to dramatically poorer generated sample quality than the original non–private GANs. Here we develop a new GAN architecture (privGAN), where the generator is trained not only to cheat the discriminator but also to defend membership inference attacks. The new mechanism is shown to empirically provide protection against this mode of attack while leading to negligible loss in downstream performances. In addition, our algorithm has been shown to explicitly prevent memorization of the training set, which explains why our protection is so effective. The main contributions of this paper are: i) we propose a novel GAN architecture that can generate synthetic data in a privacy preserving manner with minimal hyperparameter tuning and architecture selection, ii) we provide a theoretical understanding of the optimal solution of the privGAN loss function, iii) we empirically demonstrate the effectiveness of our model against several white and black–box attacks on several benchmark datasets, iv) we empirically demonstrate on three common benchmark datasets that synthetic images generated by privGAN lead to negligible loss in downstream performance when compared against non– private GANs. While we have focused on benchmarking privGAN exclusively on image datasets, the architecture of privGAN is not exclusive to image datasets and can be easily extended to other types of datasets. Repository link: https://github.com/microsoft/privGAN . Sumit Mukherjee, Yixi Xu, Anusua Trivedi, Nabajyoti Patowary, Juan M. Lavista Ferres |
Proc. Priv. Enhancing Technol. | 5 |
| 2020 | Using internet search trends to forecast short term drug overdose deaths: A case study on ConnecticutabstractIn the United States, the opioid epidemic is a serious public health crisis which claimed over 130 lives per day in 2018, according to the CDC. While there are many efforts to design effective interventions to prevent drug related deaths, much of them are focused around better prescribing practices. A promising line of inquiry has focused on utilizing machine learning tools to predict addiction or overdose related hospital admission using prior health record information. However, these are strongly reliant on the private health record information of individuals. Here, we propose using publicly available historic death records along with publicly available internet search trends of drug related search terms to predict the number of overdose deaths in the upcoming week. Our model is able to predict both the number of, and spikes in drug overdose deaths with good accuracy compared to several baselines, demonstrating the utility of search data in forecasting overdose deaths. While we demonstrate this approach as a case study in the State of Connecticut, which collects and publishes overdose data, our findings could encourage other state governments to similarly invest in collection, publication and analysis of such data. Sumit Mukherjee, Nicholas Becker, William B. Weeks, Juan M. Lavista Ferres |
ICMLA | 4 |
| 2019 | On Dynamic Network Models and Application to Causal ImpactabstractDynamic extensions of Stochastic block model (SBM) are of importance in several fields that generate temporal interaction data. These models, besides producing compact and interpretable network representations, can be useful in applications such as link prediction or network forecasting. In this paper we present a conditional pseudo-likelihood based extension to dynamic SBM that can be efficiently estimated by optimizing a regularized objective. Our formulation leads to a highly scalable approach that can handle very large networks, even with millions of nodes. We also extend our formalism to causal impact for networks that allows us to quantify the impact of external events on a time dependent sequence of networks. We support our work with extensive results on both synthetic and real networks. Yu-Chia Chen, Avleen Singh Bijral, Juan M. Lavista Ferres |
KDD | 3 |
| 2018 | NonSTOP: A NonSTationary Online Prediction Method for Time SeriesabstractWe present online prediction methods for time series that let us explicitly handle nonstationary artifacts (e.g., trend and seasonality) present in most real time series. Specifically, we show that applying appropriate transformations to such time series before prediction can lead to improved theoretical and empirical prediction performance. Moreover, since these transformations are usually unknown, we employ the learning with experts setting to develop a fully online method (NonSTOP-NonSTationary Online Prediction) for predicting nonstationary time series. This framework allows for seasonality and/or other trends in univariate time series and cointegration in multivariate time series. Our algorithms and regret analysis subsume recent related work while significantly expanding the applicability of such methods. For all the methods, we provide sublinear regret bounds. We support all of our results with experiments on simulated and real data. Christopher Xie, Avleen Singh Bijral, Juan M. Lavista Ferres |
IEEE Signal Process. Lett. | 3 |