EDBT 2026 Demo / reviewers in the wild / expert
Ruochen Kong 0001
dblp:289/1537-1
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2025
0009-0006-0329-8019ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | EpiScale: Large-Scale Simulation of Infectious Disease Based on Human MobilityabstractWe present a demonstration of a highly scalable, spatially explicit infectious disease simulation that models the spread of disease across all 220,000+ census block groups in the United States using a compartmental Susceptible-Infectious-Recovered (SIR) epidemiological framework. To achieve this unprecedented scale and resolution, our system leverages efficient sparse matrix and vector operations alongside statistical approximations of large numbers of independent random events via Poisson and Normal distributions. The resulting simulation produces realistic spatiotemporal dynamics that align with empirical patterns observed in major epidemics, including the COVID-19 outbreak. Our live demonstration at the conference will highlight the simulation's computational efficiency and interactive capabilities. Starting from the conference venue in Minneapolis, participants will be able to configure disease parameters and observe the geographic spread of infection in real time, offering both an educational and analytical perspective on pandemic modeling. Ruochen Kong 0001, Taylor Anderson 0001, David J. Heslop, Matthew Scotch, Flora D. Salim, C. Raina MacIntyre, Andreas Züfle |
SIGSPATIAL/GIS | 1 |
| 2025 | Training Machine Learning Models on Human Spatio-temporal Mobility Data: An Experimental Study [Experiment Paper]abstractIndividual-level human mobility prediction has emerged as a significant topic of research. In this paper, we focus on an underexplored problem in human mobility prediction: determining the best practices to train a machine learning model using historical data to forecast an individuals complete trajectory over the next days and weeks. In this experiment paper, we undertake a comprehensive experimental analysis of diverse models, parameter configurations, and training strategies, accompanied by an in-depth examination of the statistical distribution inherent in human mobility patterns. Our empirical evaluations encompass both Long Short-Term Memory and Transformer-based architectures, and further investigate how incorporating individual life patterns can enhance the effectiveness of the prediction. Moreover, since the absence of explicit user information is often missing due to user privacy, we show that the sampling of users may exacerbate data skewness and result in a substantial loss in predictive accuracy. To mitigate data imbalance and preserve diversity, we apply user semantic clustering with stratified sampling to ensure that the sampled dataset remains representative. Our results further show that small-batch stochastic gradient optimization improves model performance, especially when human mobility training data is limited. Lance Kennedy, Ruochen Kong 0001, Joon-Seok Kim 0001, Andreas Züfle |
SIGSPATIAL/GIS | 3 |
| 2025 | Human Mobility Prediction via Sparse Mixture-of-Experts and Hybrid Multi-Scale EncodingabstractThe prediction of individual-level human mobility has emerged as a critical research domain. However, human trajectories are inherently complex and subject to perturbations from exogenous factors such as meteorological conditions, collective social behaviors, and temporal events including public holidays. These external influences introduce substantial variability, thereby complicating predictive modeling. In this study, we propose a novel framework for human mobility trajectory prediction. Unlike existing methods that focus primarily on a single temporal scale, our approach leverages multi-scale modeling through a hybrid Temporal Convolutional Network-Transformer encoder to capture both local and global mobility dynamics. The framework further integrates user profile embeddings to incorporate individual and group-level behavioral patterns, and employs a sparse Mixture-of-Experts architecture with top-1 routing to achieve scalable specialization at fixed inference cost. Ruochen Kong 0001, Andreas Züfle |
SIGSPATIAL/GIS | 2 |
| 2025 | Simulated Infectious Diseases Datasets with Controlled Data BiasabstractMassive datasets related to infectious diseases became available after the COVID-19 pandemic, supporting data-driven approaches in modeling and forecasting infectious diseases. However, these approaches are known to exacerbate data biases present in the training data such as having certain demographic groups being over or underrepresented in the data. Such data collection biases may propagate through the modeling and prediction pipelines to decision-making, and the consequences are relatively unknown. Therefore, efforts are needed to understand how data collection bias affects data-driven infectious disease models. This datasets and benchmarks paper provides a suite of datasets, each corresponding to a simulated disease spread among a population of 5000 simulated agents over 90 days in Atlanta and San Francisco. For each dataset, we provide not only the full (simulated ground truth) of the disease spread in terms of when, where, and by whom the disease spreads, but also information on which cases are observed when different types and degrees of data collection bias are applied. The agents' characteristics, check-ins, and social network data are also available to support downstream tasks. Additionally, we also describe how to use the simulation to re-generate the data and to generate new datasets in different regions and with different parameters. With the provided datasets and the simulation tools, researchers studying the spread of infectious diseases may better understand, account for, and correct the systematic bias caused by the inherent real-world data bias, and hence improve the prediction of infectious diseases. Ruochen Kong 0001, Taylor Anderson 0001, Matthew Scotch, David J. Heslop, Yonchanok Khaokaew, Hao Xue 0001, Li Xiong 0001, C. Raina MacIntyre, Flora D. Salim, Andreas Züfle |
KDD (2) | 1 |
| 2024 | An Infectious Disease Spread Simulation to Control Data BiasabstractThe increased availability of datasets during the COVID-19 pandemic enabled machine-learning approaches for modeling and forecasting infectious diseases. However, such approaches are known to amplify the bias in the data they are trained on. Bias in such input data like clinical case data for COVID-19 is difficult to measure due to disparities in testing availability, reporting standards, and healthcare access among different populations and regions. Furthermore, the way such biases may propagate through the modeling pipeline to decision-making is relatively unknown. Therefore, we present a system that leverages a highly detailed agent-based model (ABM) of infectious disease spread in a city to simulate the collection of biased clinical case data where the bias is known. Our system allows users to load either a pre-selected region or select their own (using OpenStreetMap data for the environment and census data for the population), specify population and infectious disease parameters, and the degree(s) to which different populations will be overrep-resented or underrepresented in the case data. In addition to the system, we provide a large number of benchmark datasets that produce case data at different levels of bias for different regions. We hope that infectious disease modelers will use these datasets to investigate how well their models are robust to data bias or whether their model is overfit to biased data. Ruochen Kong 0001, Taylor Anderson 0001, David J. Heslop, Andreas Züfle |
SIGSPATIAL/GIS | 1 |
| 2022 | Dynamic network embedding survey
Guotong Xue, Ming Zhong 0002, Jianxin Li 0001, Jia Chen 0020, Chengshuai Zhai, Ruochen Kong 0001 |
Neurocomputing | 6 |