David J. Heslop

dblp:310/3406 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2025
0000-0002-1978-770XORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
YearPublicationVenuePosition
2025 EpiScale: Large-Scale Simulation of Infectious Disease Based on Human Mobility
abstract
We present a demonstration of a highly scalable, spatially explicit infectious disease simulation that models the spread of disease across all 220,000+ census block groups in the United States using a compartmental Susceptible-Infectious-Recovered (SIR) epidemiological framework. To achieve this unprecedented scale and resolution, our system leverages efficient sparse matrix and vector operations alongside statistical approximations of large numbers of independent random events via Poisson and Normal distributions. The resulting simulation produces realistic spatiotemporal dynamics that align with empirical patterns observed in major epidemics, including the COVID-19 outbreak. Our live demonstration at the conference will highlight the simulation's computational efficiency and interactive capabilities. Starting from the conference venue in Minneapolis, participants will be able to configure disease parameters and observe the geographic spread of infection in real time, offering both an educational and analytical perspective on pandemic modeling.
Ruochen Kong 0001, Taylor Anderson 0001, David J. Heslop, Matthew Scotch, Flora D. Salim, C. Raina MacIntyre, Andreas Züfle
SIGSPATIAL/GIS3
2025 A Probabilistic Framework for Imputing Genetic Distances in Spatiotemporal Pathogen Models
abstract
Pathogen genome data offers valuable structure for spatial models, but its utility is limited by incomplete sequencing coverage. We propose a probabilistic framework for inferring genetic distances between unsequenced cases and known sequences within defined transmission chains, using time-aware evolutionary distance modeling. The method estimates pairwise divergence from collection dates and observed genetic distances, enabling biologically plausible imputation grounded in observed divergence patterns, without requiring sequence alignment or known transmission chains. Applied to highly pathogenic avian influenza A/H5 cases in wild birds in the United States, this approach supports scalable, uncertainty-aware augmentation of genomic datasets and enhances the integration of evolutionary information into spatiotemporal modeling workflows.
Haley Stone, Jing Du 0003, Hao Xue 0001, Matthew Scotch, David J. Heslop, Andreas Züfle, C. Raina MacIntyre, Flora D. Salim
SIGSPATIAL/GIS5
2025 Simulated Infectious Diseases Datasets with Controlled Data Bias
abstract
Massive datasets related to infectious diseases became available after the COVID-19 pandemic, supporting data-driven approaches in modeling and forecasting infectious diseases. However, these approaches are known to exacerbate data biases present in the training data such as having certain demographic groups being over or underrepresented in the data. Such data collection biases may propagate through the modeling and prediction pipelines to decision-making, and the consequences are relatively unknown. Therefore, efforts are needed to understand how data collection bias affects data-driven infectious disease models. This datasets and benchmarks paper provides a suite of datasets, each corresponding to a simulated disease spread among a population of 5000 simulated agents over 90 days in Atlanta and San Francisco. For each dataset, we provide not only the full (simulated ground truth) of the disease spread in terms of when, where, and by whom the disease spreads, but also information on which cases are observed when different types and degrees of data collection bias are applied. The agents' characteristics, check-ins, and social network data are also available to support downstream tasks. Additionally, we also describe how to use the simulation to re-generate the data and to generate new datasets in different regions and with different parameters. With the provided datasets and the simulation tools, researchers studying the spread of infectious diseases may better understand, account for, and correct the systematic bias caused by the inherent real-world data bias, and hence improve the prediction of infectious diseases.
Ruochen Kong 0001, Taylor Anderson 0001, Matthew Scotch, David J. Heslop, Yonchanok Khaokaew, Hao Xue 0001, Li Xiong 0001, C. Raina MacIntyre, Flora D. Salim, Andreas Züfle
KDD (2)4
2024 An Infectious Disease Spread Simulation to Control Data Bias
abstract
The increased availability of datasets during the COVID-19 pandemic enabled machine-learning approaches for modeling and forecasting infectious diseases. However, such approaches are known to amplify the bias in the data they are trained on. Bias in such input data like clinical case data for COVID-19 is difficult to measure due to disparities in testing availability, reporting standards, and healthcare access among different populations and regions. Furthermore, the way such biases may propagate through the modeling pipeline to decision-making is relatively unknown. Therefore, we present a system that leverages a highly detailed agent-based model (ABM) of infectious disease spread in a city to simulate the collection of biased clinical case data where the bias is known. Our system allows users to load either a pre-selected region or select their own (using OpenStreetMap data for the environment and census data for the population), specify population and infectious disease parameters, and the degree(s) to which different populations will be overrep-resented or underrepresented in the case data. In addition to the system, we provide a large number of benchmark datasets that produce case data at different levels of bias for different regions. We hope that infectious disease modelers will use these datasets to investigate how well their models are robust to data bias or whether their model is overfit to biased data.
Ruochen Kong 0001, Taylor Anderson 0001, David J. Heslop, Andreas Züfle
SIGSPATIAL/GIS3