VLDB 2026 Research / reviewers in the wild / expert
Sam Abbott
dblp:267/3632
· DBLP profile ↗
11ranked-venue papers
0as first author
10since 2021 · last 2024
0000-0001-8057-8037ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Best practices for estimating and reporting epidemiological delay distributions of infectious diseasesabstractEpidemiological delays are key quantities that inform public health policy and clinical practice. They are used as inputs for mathematical and statistical models, which in turn can guide control strategies. In recent work, we found that censoring, right truncation, and dynamical bias were rarely addressed correctly when estimating delays and that these biases were large enough to have knock-on impacts across a large number of use cases. Here, we formulate a checklist of best practices for estimating and reporting epidemiological delays. We also provide a flowchart to guide practitioners based on their data. Our examples are focused on the incubation period and serial interval due to their importance in outbreak response and modeling, but our recommendations are applicable to other delays. The recommendations, which are based on the literature and our experience estimating epidemiological delay distributions during outbreak responses, can help improve the robustness and utility of reported estimates and provide guidance for the evaluation of estimates for downstream use in transmission models or other analyses. Kelly Charniga, Sang Woo Park, Andrei R. Akhmetzhanov, Anne Cori, Jonathan Dushoff, Sebastian Funk, Katelyn M. Gostic, Natalie M. Linton, Adrian Lison, Christopher E. Overton, Juliet R. C. Pulliam, Simon Cauchemez, Sam Abbott |
PLoS Comput. Biol. | 14 |
| 2024 | EpiFusion: Joint inference of the effective reproduction number by integrating phylodynamic and epidemiological modelling with particle filteringabstractAccurately estimating the effective reproduction number (Rt) of a circulating pathogen is a fundamental challenge in the study of infectious disease. The fields of epidemiology and pathogen phylodynamics both share this goal, but to date, methodologies and data employed by each remain largely distinct. Here we present EpiFusion: a joint approach that can be used to harness the complementary strengths of each field to improve estimation of outbreak dynamics for large and poorly sampled epidemics, such as arboviral or respiratory virus outbreaks, and validate it for retrospective analysis. We propose a model of Rt that estimates outbreak trajectories conditional upon both phylodynamic (time-scaled trees estimated from genetic sequences) and epidemiological (case incidence) data. We simulate stochastic outbreak trajectories that are weighted according to epidemiological and phylodynamic observation models and fit using particle Markov Chain Monte Carlo. To assess performance, we test EpiFusion on simulated outbreaks in which transmission and/or surveillance rapidly changes and find that using EpiFusion to combine epidemiological and phylodynamic data maintains accuracy and increases certainty in trajectory and Rt estimates, compared to when each data type is used alone. We benchmark EpiFusion's performance against existing methods to estimate Rt and demonstrate advances in speed and accuracy. Importantly, our approach scales efficiently with dataset size. Finally, we apply our model to estimate Rt during the 2014 Ebola outbreak in Sierra Leone. EpiFusion is designed to accommodate future extensions that will improve its utility, such as explicitly modelling population structure, accommodations for phylogenetic uncertainty, and the ability to weight the contributions of genomic or case incidence to the inference. Ciara E. Judge, Timothy G. Vaughan, Timothy Russell, Sam Abbott, Louis du Plessis, Tanja Stadler, Oliver Brady, Sarah Hill |
PLoS Comput. Biol. | 4 |
| 2024 | Generative Bayesian modeling to nowcast the effective reproduction number from line list data with missing symptom onset datesabstractThe time-varying effective reproduction number Rt is a widely used indicator of transmission dynamics during infectious disease outbreaks. Timely estimates of Rt can be obtained from reported cases counted by their date of symptom onset, which is generally closer to the time of infection than the date of report. Case counts by date of symptom onset are typically obtained from line list data, however these data can have missing information and are subject to right truncation. Previous methods have addressed these problems independently by first imputing missing onset dates, then adjusting truncated case counts, and finally estimating the effective reproduction number. This stepwise approach makes it difficult to propagate uncertainty and can introduce subtle biases during real-time estimation due to the continued impact of assumptions made in previous steps. In this work, we integrate imputation, truncation adjustment, and Rt estimation into a single generative Bayesian model, allowing direct joint inference of case counts and Rt from line list data with missing symptom onset dates. We then use this framework to compare the performance of nowcasting approaches with different stepwise and generative components on synthetic line list data for multiple outbreak scenarios and across different epidemic phases. We find that under reporting delays realistic for hospitalization data (50% of reports delayed by more than a week), intermediate smoothing, as is common practice in stepwise approaches, can bias nowcasts of case counts and Rt, which is avoided in a joint generative approach due to shared regularization of all model components. On incomplete line list data, a fully generative approach enables the quantification of uncertainty due to missing onset dates without the need for an initial multiple imputation step. In a real-world comparison using hospitalization line list data from the COVID-19 pandemic in Switzerland, we observe the same qualitative differences between approaches. The generative modeling components developed in this work have been integrated and further extended in the R package epinowcast, providing a flexible and interpretable tool for real-time surveillance. Adrian Lison, Sam Abbott, Jana S. Huisman, Tanja Stadler |
PLoS Comput. Biol. | 2 |
| 2023 | Scoring epidemiological forecasts on transformed scalesabstractForecast evaluation is essential for the development of predictive epidemic models and can inform their use for public health decision-making. Common scores to evaluate epidemiological forecasts are the Continuous Ranked Probability Score (CRPS) and the Weighted Interval Score (WIS), which can be seen as measures of the absolute distance between the forecast distribution and the observation. However, applying these scores directly to predicted and observed incidence counts may not be the most appropriate due to the exponential nature of epidemic processes and the varying magnitudes of observed values across space and time. In this paper, we argue that transforming counts before applying scores such as the CRPS or WIS can effectively mitigate these difficulties and yield epidemiologically meaningful and easily interpretable results. Using the CRPS on log-transformed values as an example, we list three attractive properties: Firstly, it can be interpreted as a probabilistic version of a relative error. Secondly, it reflects how well models predicted the time-varying epidemic growth rate. And lastly, using arguments on variance-stabilizing transformations, it can be shown that under the assumption of a quadratic mean-variance relationship, the logarithmic transformation leads to expected CRPS values which are independent of the order of magnitude of the predicted quantity. Applying a transformation of log(x + 1) to data and forecasts from the European COVID-19 Forecast Hub, we find that it changes model rankings regardless of stratification by forecast date, location or target types. Situations in which models missed the beginning of upward swings are more strongly emphasised while failing to predict a downturn following a peak is less severely penalised when scoring transformed forecasts as opposed to untransformed ones. We conclude that appropriate transformations, of which the natural logarithm is only one particularly attractive option, should be considered when assessing the performance of different models in the context of infectious disease incidence. Nikos I. Bosse, Sam Abbott, Anne Cori, Edwin van Leeuwen, Johannes Bracher, Sebastian Funk |
PLoS Comput. Biol. | 2 |
| 2023 | Why are different estimates of the effective reproductive number so different? A case study on COVID-19 in GermanyabstractThe effective reproductive number Rt has taken a central role in the scientific, political, and public discussion during the COVID-19 pandemic, with numerous real-time estimates of this quantity routinely published. Disagreement between estimates can be substantial and may lead to confusion among decision-makers and the general public. In this work, we compare different estimates of the national-level effective reproductive number of COVID-19 in Germany in 2020 and 2021. We consider the agreement between estimates from the same method but published at different time points (within-method agreement) as well as retrospective agreement across eight different approaches (between-method agreement). Concerning the former, estimates from some methods are very stable over time and hardly subject to revisions, while others display considerable fluctuations. To evaluate between-method agreement, we reproduce the estimates generated by different groups using a variety of statistical approaches, standardizing analytical choices to assess how they contribute to the observed disagreement. These analytical choices include the data source, data pre-processing, assumed generation time distribution, statistical tuning parameters, and various delay distributions. We find that in practice, these auxiliary choices in the estimation of Rt may affect results at least as strongly as the selection of the statistical approach. They should thus be communicated transparently along with the estimates. Elisabeth K. Brockhaus, Daniel Wolffram, Tanja Stadler, Michael Osthege, Tanmay Mitra, Jonas M. Littek, Ekaterina A. Krymova, Anna J. Klesen, Jana S. Huisman, Stefan Heyder, Laura Marie Helleckes, Matthias an der Heiden, Sebastian Funk, Sam Abbott, Johannes Bracher |
PLoS Comput. Biol. | 14 |
| 2023 | Evaluating the use of social contact data to produce age-specific short-term forecasts of SARS-CoV-2 incidence in EnglandabstractMathematical and statistical models can be used to make predictions of how epidemics may progress in the near future and form a central part of outbreak mitigation and control. Renewal equation based models allow inference of epidemiological parameters from historical data and forecast future epidemic dynamics without requiring complex mechanistic assumptions. However, these models typically ignore interaction between age groups, partly due to challenges in parameterising a time varying interaction matrix. Social contact data collected regularly during the COVID-19 epidemic provide a means to inform interaction between age groups in real-time. We developed an age-specific forecasting framework and applied it to two age-stratified time-series: incidence of SARS-CoV-2 infection, estimated from a national infection and antibody prevalence survey; and, reported cases according to the UK national COVID-19 dashboard. Jointly fitting our model to social contact data from the CoMix study, we inferred a time-varying next generation matrix which we used to project infections and cases in the four weeks following each of 29 forecast dates between October 2020 and November 2021. We evaluated the forecasts using proper scoring rules and compared performance with three other models with alternative data and specifications alongside two naive baseline models. Overall, incorporating age interaction improved forecasts of infections and the CoMix-data-informed model was the best performing model at time horizons between two and four weeks. However, this was not true when forecasting cases. We found that age group interaction was most important for predicting cases in children and older adults. The contact-data-informed models performed best during the winter months of 2020-2021, but performed comparatively poorly in other periods. We highlight challenges regarding the incorporation of contact data in forecasting and offer proposals as to how to extend and adapt our approach, which may lead to more successful forecasts in future. James D. Munday, Sam Abbott, Sophie Meakin, Sebastian Funk |
PLoS Comput. Biol. | 2 |
| 2023 | Nowcasting the 2022 mpox outbreak in EnglandabstractIn May 2022, a cluster of mpox cases were detected in the UK that could not be traced to recent travel history from an endemic region. Over the coming months, the outbreak grew, with over 3000 total cases reported in the UK, and similar outbreaks occurring worldwide. These outbreaks appeared linked to sexual contact networks between gay, bisexual and other men who have sex with men. Following the COVID-19 pandemic, local health systems were strained, and therefore effective surveillance for mpox was essential for managing public health policy. However, the mpox outbreak in the UK was characterised by substantial delays in the reporting of the symptom onset date and specimen collection date for confirmed positive cases. These delays led to substantial backfilling in the epidemic curve, making it challenging to interpret the epidemic trajectory in real-time. Many nowcasting models exist to tackle this challenge in epidemiological data, but these lacked sufficient flexibility. We have developed a nowcasting model using generalised additive models that makes novel use of individual-level patient data to correct the mpox epidemic curve in England. The aim of this model is to correct for backfilling in the epidemic curve and provide real-time characteristics of the state of the epidemic, including the real-time growth rate. This model benefited from close collaboration with individuals involved in collecting and processing the data, enabling temporal changes in the reporting structure to be built into the model, which improved the robustness of the nowcasts generated. The resulting model accurately captured the true shape of the epidemic curve in real time. Christopher E. Overton, Sam Abbott, Rachel Christie, Fergus Cumming, Julie Day, Owen Jones, Rob Paton, Charlie Turner |
PLoS Comput. Biol. | 2 |
| 2023 | Collaborative nowcasting of COVID-19 hospitalization incidences in GermanyabstractReal-time surveillance is a crucial element in the response to infectious disease outbreaks. However, the interpretation of incidence data is often hampered by delays occurring at various stages of data gathering and reporting. As a result, recent values are biased downward, which obscures current trends. Statistical nowcasting techniques can be employed to correct these biases, allowing for accurate characterization of recent developments and thus enhancing situational awareness. In this paper, we present a preregistered real-time assessment of eight nowcasting approaches, applied by independent research teams to German 7-day hospitalization incidences during the COVID-19 pandemic. This indicator played an important role in the management of the outbreak in Germany and was linked to levels of non-pharmaceutical interventions via certain thresholds. Due to its definition, in which hospitalization counts are aggregated by the date of case report rather than admission, German hospitalization incidences are particularly affected by delays and can take several weeks or months to fully stabilize. For this study, all methods were applied from 22 November 2021 to 29 April 2022, with probabilistic nowcasts produced each day for the current and 28 preceding days. Nowcasts at the national, state, and age-group levels were collected in the form of quantiles in a public repository and displayed in a dashboard. Moreover, a mean and a median ensemble nowcast were generated. We find that overall, the compared methods were able to remove a large part of the biases introduced by delays. Most participating teams underestimated the importance of very long delays, though, resulting in nowcasts with a slight downward bias. The accompanying prediction intervals were also too narrow for almost all methods. Averaged over all nowcast horizons, the best performance was achieved by a model using case incidences as a covariate and taking into account longer delays than the other approaches. For the most recent days, which are often considered the most relevant in practice, a mean ensemble of the submitted nowcasts performed best. We conclude by providing some lessons learned on the definition of nowcasting targets and practical challenges. Daniel Wolffram, Sam Abbott, Matthias an der Heiden, Sebastian Funk, Felix Günther 0003, Davide Hailer, Stefan Heyder, Thomas Hotz, Jan van de Kassteele, Helmut Küchenhoff, Sören Müller-Hansen, Diellë Syliqi, Alexander Ullrich, Maximilian Weigert, Melanie Schienle, Johannes Bracher |
PLoS Comput. Biol. | 2 |
| 2022 | Comparing human and model-based forecasts of COVID-19 in Germany and PolandabstractForecasts based on epidemiological modelling have played an important role in shaping public policy throughout the COVID-19 pandemic. This modelling combines knowledge about infectious disease dynamics with the subjective opinion of the researcher who develops and refines the model and often also adjusts model outputs. Developing a forecast model is difficult, resource- and time-consuming. It is therefore worth asking what modelling is able to add beyond the subjective opinion of the researcher alone. To investigate this, we analysed different real-time forecasts of cases of and deaths from COVID-19 in Germany and Poland over a 1-4 week horizon submitted to the German and Polish Forecast Hub. We compared crowd forecasts elicited from researchers and volunteers, against a) forecasts from two semi-mechanistic models based on common epidemiological assumptions and b) the ensemble of all other models submitted to the Forecast Hub. We found crowd forecasts, despite being overconfident, to outperform all other methods across all forecast horizons when forecasting cases (weighted interval score relative to the Hub ensemble 2 weeks ahead: 0.89). Forecasts based on computational models performed comparably better when predicting deaths (rel. WIS 1.26), suggesting that epidemiological modelling and human judgement can complement each other in important ways. Nikos I. Bosse, Sam Abbott, Johannes Bracher, Habakuk Hain, Billy J. Quilty, Mark Jit, Edwin van Leeuwen, Anne Cori, Sebastian Funk |
PLoS Comput. Biol. | 2 |
| 2021 | Correction: Practical considerations for measuring the effective reproductive number, RtabstractThere is an error in the caption for Fig 1, "Instantaneous reproductive number as estimated by the method of Cori et al. vs. cohort reproductive number estimated by Wallinga and Teunis" sentence four. Katelyn M. Gostic, Lauren McGough, Edward B. Baskerville, Sam Abbott, Keya Joshi, Christine Tedijanto, Rebecca Kahn, Rene Niehus, James A. Hay, Pablo M. De Salazar, Joel Hellewell, Sophie Meakin, James D. Munday, Nikos I. Bosse, Katharine Sherratt, Robin N. Thompson, Laura F. White, Jana S. Huisman, Jérémie Scire, Sebastian Bonhoeffer, Tanja Stadler, Jacco Wallinga, Sebastian Funk, Marc Lipsitch, Sarah Cobey |
PLoS Comput. Biol. | 4 |
| 2020 | Practical considerations for measuring the effective reproductive number, RtabstractEstimation of the effective reproductive number Rt is important for detecting changes in disease transmission over time. During the Coronavirus Disease 2019 (COVID-19) pandemic, policy makers and public health officials are using Rt to assess the effectiveness of interventions and to inform policy. However, estimation of Rt from available data presents several challenges, with critical implications for the interpretation of the course of the pandemic. The purpose of this document is to summarize these challenges, illustrate them with examples from synthetic data, and, where possible, make recommendations. For near real-time estimation of Rt, we recommend the approach of Cori and colleagues, which uses data from before time t and empirical estimates of the distribution of time between infections. Methods that require data from after time t, such as Wallinga and Teunis, are conceptually and methodologically less suited for near real-time estimation, but may be appropriate for retrospective analyses of how individuals infected at different time points contributed to the spread. We advise caution when using methods derived from the approach of Bettencourt and Ribeiro, as the resulting Rt estimates may be biased if the underlying structural assumptions are not met. Two key challenges common to all approaches are accurate specification of the generation interval and reconstruction of the time series of new infections from observations occurring long after the moment of transmission. Naive approaches for dealing with observation delays, such as subtracting delays sampled from a distribution, can introduce bias. We provide suggestions for how to mitigate this and other technical challenges and highlight open problems in Rt estimation. Katelyn M. Gostic, Lauren McGough, Edward B. Baskerville, Sam Abbott, Keya Joshi, Christine Tedijanto, Rebecca Kahn, Rene Niehus, James A. Hay, Pablo M. De Salazar, Joel Hellewell, Sophie Meakin, James D. Munday, Nikos I. Bosse, Katharine Sherratt, Robin N. Thompson, Laura F. White, Jana S. Huisman, Jérémie Scire, Sebastian Bonhoeffer, Tanja Stadler, Jacco Wallinga, Sebastian Funk, Marc Lipsitch, Sarah Cobey |
PLoS Comput. Biol. | 4 |