Jeffrey Shaman

dblp:155/3664 · DBLP profile ↗
← Back
25ranked-venue papers
1as first author
6since 2021 · last 2024
0000-0002-7216-7809ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 25 · 1 first-author · 6 since 2021
YearPublicationVenuePosition
2024 Challenges of COVID-19 Case Forecasting in the US, 2020-2021
abstract
During the COVID-19 pandemic, forecasting COVID-19 trends to support planning and response was a priority for scientists and decision makers alike. In the United States, COVID-19 forecasting was coordinated by a large group of universities, companies, and government entities led by the Centers for Disease Control and Prevention and the US COVID-19 Forecast Hub (https://covid19forecasthub.org). We evaluated approximately 9.7 million forecasts of weekly state-level COVID-19 cases for predictions 1-4 weeks into the future submitted by 24 teams from August 2020 to December 2021. We assessed coverage of central prediction intervals and weighted interval scores (WIS), adjusting for missing forecasts relative to a baseline forecast, and used a Gaussian generalized estimating equation (GEE) model to evaluate differences in skill across epidemic phases that were defined by the effective reproduction number. Overall, we found high variation in skill across individual models, with ensemble-based forecasts outperforming other approaches. Forecast skill relative to the baseline was generally higher for larger jurisdictions (e.g., states compared to counties). Over time, forecasts generally performed worst in periods of rapid changes in reported cases (either in increasing or decreasing epidemic phases) with 95% prediction interval coverage dropping below 50% during the growth phases of the winter 2020, Delta, and Omicron waves. Ideally, case forecasts could serve as a leading indicator of changes in transmission dynamics. However, while most COVID-19 case forecasts outperformed a naïve baseline model, even the most accurate case forecasts were unreliable in key phases. Further research could improve forecasts of leading indicators, like COVID-19 cases, by leveraging additional real-time data, addressing performance across phases, improving the characterization of forecast confidence, and ensuring that forecasts were coherent across spatial scales. In the meantime, it is critical for forecast users to appreciate current limitations and use a broad set of indicators to inform pandemic-related decision making.
Velma K. Lopez, Estee Y. Cramer, Robert Pagano, John M. Drake, Eamon B. O'Dea, Madeline Adee, Turgay Ayer, Jagpreet Chhatwal, Ozden O. Dalgic, Mary A. Ladd, Benjamin P. Linas, Peter P. Mueller, Jade Xiao, Johannes Bracher, Alvaro J. Castro Rivadeneira, Aaron Gerding, Tilmann Gneiting, Yuxin Huang 0009, Dasuni Jayawardena, Abdul H. Kanji, Khoa Le, Anja Mühlemann, Jarad Niemi, Evan L. Ray, Ariane Stark, Nutcha Wattanachit, Martha W. Zorn, Sen Pei, Jeffrey Shaman, Teresa K. Yamana, Samuel R. Tarasewicz, Daniel J. Wilson 0002, Sid Baccam, Heidi Gurung, Steve Stage, Brad Suchoski, Lei Gao 0011, Zhiling Gu, Myungjin Kim, Guannan Wang, Li Wang 0035, Yueying Wang, Lauren Gardner, Sonia Jindal, Maximilian Marshall, Kristen Nixon, Juan Dent, Alison L. Hill, Joshua Kaminsky, Elizabeth C. Lee, Joseph Chadi Lemaitre, Justin Lessler, Claire P. Smith, Shaun Truelove, Matt Kinsey, Luke C. Mullany, Kaitlin Rainwater-Lovett, Lauren Shin, Katharine Tallaksen, Shelby Wilson, Dean Karlen, Lauren A. Castro, Geoffrey Fairchild, Isaac Michaud, Dave Osthus, Jiang Bian 0002, Wei Cao 0007, Zhifeng Gao, Juan M. Lavista Ferres, Chaozhuo Li, Tie-Yan Liu, Xing Xie 0001, Shun Zheng 0001, Matteo Chinazzi, Jessica T. Davis, Kunpeng Mu, Ana L. Pastore y Piontti, Alessandro Vespignani, Xinyue Xiong, Robert Walraven, Quanquan Gu, Lingxiao Wang 0001, Pan Xu 0002, Difan Zou, Graham Casey Gibson, Daniel Sheldon, Ajitesh Srivastava, Aniruddha Adiga, Benjamin Hurt, Gursharn Kaur, Bryan L. Lewis, Madhav V. Marathe, Akhil Sai Peddireddy, Przemyslaw J. Porebski, Srinivasan Venkatramanan, Lijing Wang 0001, Pragati V. Prasad, Jo W. Walker, Alexander E. Webber, Rachel B. Slayton, Matthew Biggerstaff, Nicholas G. Reich, Michael A. Johansson
PLoS Comput. Biol.30
2023 Inference of transmission dynamics and retrospective forecast of invasive meningococcal disease
abstract
The pathogenic bacteria Neisseria meningitidis, which causes invasive meningococcal disease (IMD), predominantly colonizes humans asymptomatically; however, invasive disease occurs in a small proportion of the population. Here, we explore the seasonality of IMD and develop and validate a suite of models for simulating and forecasting disease outcomes in the United States. We combine the models into multi-model ensembles (MME) based on the past performance of the individual models, as well as a naive equally weighted aggregation, and compare the retrospective forecast performance over a six-month forecast horizon. Deployment of the complete vaccination regimen, introduced in 2011, coincided with a change in the periodicity of IMD, suggesting altered transmission dynamics. We found that a model forced with the period obtained by local power wavelet decomposition best fit and forecast observations. In addition, the MME performed the best across the entire study period. Finally, our study included US-level data until 2022, allowing study of a possible IMD rebound after relaxation of non-pharmaceutical interventions imposed in response to the COVID-19 pandemic; however, no evidence of a rebound was found. Our findings demonstrate the ability of process-based models to retrospectively forecast IMD and provide a first analysis of the seasonality of IMD before and after the complete vaccination regimen.
Jaime Cascante-Vega, Marta Galanti, Katharina Schley, Sen Pei, Jeffrey Shaman
PLoS Comput. Biol.5
2023 Hindcasts and forecasts of suicide mortality in US: A modeling study
abstract
Deaths by suicide, as well as suicidal ideations, plans and attempts, have been increasing in the US for the past two decades. Deployment of effective interventions would require timely, geographically well-resolved estimates of suicide activity. In this study, we evaluated the feasibility of a two-step process for predicting suicide mortality: a) generation of hindcasts, mortality estimates for past months for which observational data would not have been available if forecasts were generated in real-time; and b) generation of forecasts with observational data augmented with hindcasts. Calls to crisis hotline services and online queries to the Google search engine for suicide-related terms were used as proxy data sources to generate hindcasts. The primary hindcast model (auto) is an Autoregressive Integrated Moving average model (ARIMA), trained on suicide mortality rates alone. Three regression models augment hindcast estimates from auto with call rates (calls), GHT search rates (ght) and both datasets together (calls_ght). The 4 forecast models used are ARIMA models trained with corresponding hindcast estimates. All models were evaluated against a baseline random walk with drift model. Rolling monthly 6-month ahead forecasts for all 50 states between 2012 and 2020 were generated. Quantile score (QS) was used to assess the quality of the forecast distributions. Median QS for auto was better than baseline (0.114 vs. 0.21. Median QS of augmented models were lower than auto, but not significantly different from each other (Wilcoxon signed-rank test, p > .05). Forecasts from augmented models were also better calibrated. Together, these results provide evidence that proxy data can address delays in release of suicide mortality data and improve forecast quality. An operational forecast system of state-level suicide risk may be feasible with sustained engagement between modelers and public health departments to appraise data sources and methods as well as to continuously evaluate forecast accuracy.
Sasikiran Kandula, Mark Olfson, Madelyn S. Gould, Katherine M. Keyes, Jeffrey Shaman
PLoS Comput. Biol.5
2023 Development of Accurate Long-lead COVID-19 Forecast
abstract
Coronavirus disease 2019 (COVID-19) will likely remain a major public health burden; accurate forecast of COVID-19 epidemic outcomes several months into the future is needed to support more proactive planning. Here, we propose strategies to address three major forecast challenges, i.e., error growth, the emergence of new variants, and infection seasonality. Using these strategies in combination we generate retrospective predictions of COVID-19 cases and deaths 6 months in the future for 10 representative US states. Tallied over >25,000 retrospective predictions through September 2022, the forecast approach using all three strategies consistently outperformed a baseline forecast approach without these strategies across different variant waves and locations, for all forecast targets. Overall, probabilistic forecast accuracy improved by 64% and 38% and point prediction accuracy by 133% and 87% for cases and deaths, respectively. Real-time 6-month lead predictions made in early October 2022 suggested large attack rates in most states but a lower burden of deaths than previous waves during October 2022 -March 2023; these predictions are in general accurate compared to reported data. The superior skill of the forecast methods developed here demonstrate means for generating more accurate long-lead forecast of COVID-19 and possibly other infectious diseases.
Wan Yang, Jeffrey Shaman
PLoS Comput. Biol.2
2022 Epidemic management and control through risk-dependent individual contact interventions
abstract
Testing, contact tracing, and isolation (TTI) is an epidemic management and control approach that is difficult to implement at scale because it relies on manual tracing of contacts. Exposure notification apps have been developed to digitally scale up TTI by harnessing contact data obtained from mobile devices; however, exposure notification apps provide users only with limited binary information when they have been directly exposed to a known infection source. Here we demonstrate a scalable improvement to TTI and exposure notification apps that uses data assimilation (DA) on a contact network. Network DA exploits diverse sources of health data together with the proximity data from mobile devices that exposure notification apps rely upon. It provides users with continuously assessed individual risks of exposure and infection, which can form the basis for targeting individual contact interventions. Simulations of the early COVID-19 epidemic in New York City are used to establish proof-of-concept. In the simulations, network DA identifies up to a factor 2 more infections than contact tracing when both harness the same contact data and diagnostic test data. This remains true even when only a relatively small fraction of the population uses network DA. When a sufficiently large fraction of the population (≳ 75%) uses network DA and complies with individual contact interventions, targeting contact interventions with network DA reduces deaths by up to a factor 4 relative to TTI. Network DA can be implemented by expanding the computational backend of existing exposure notification apps, thus greatly enhancing their capabilities. Implemented at scale, it has the potential to precisely and effectively control future epidemics while minimizing economic disruption.
Tapio Schneider, Oliver R. A. Dunbar, Lucas Böttcher, Dmitry Burov, Alfredo Garbuno-Inigo, Gregory LeClaire Wagner, Sen Pei, Chiara Daraio, Raffaele Ferrari, Jeffrey Shaman
PLoS Comput. Biol.11
2022 Inference and dynamic simulation of malaria using a simple climate-driven entomological model of malaria transmission
abstract
Given the crucial role of climate in malaria transmission, many mechanistic models of malaria represent vector biology and the parasite lifecycle as functions of climate variables in order to accurately capture malaria transmission dynamics. Lower dimension mechanistic models that utilize implicit vector dynamics have relied on indirect climate modulation of transmission processes, which compromises investigation of the ecological role played by climate in malaria transmission. In this study, we develop an implicit process-based malaria model with direct climate-mediated modulation of transmission pressure borne through the Entomological Inoculation Rate (EIR). The EIR, a measure of the number of infectious bites per person per unit time, includes the effects of vector dynamics, resulting from mosquito development, survivorship, feeding activity and parasite development, all of which are moderated by climate. We combine this EIR-model framework, which is driven by rainfall and temperature, with Bayesian inference methods, and evaluate the model's ability to simulate local transmission across 42 regions in Rwanda over four years. Our findings indicate that the biologically-motivated, EIR-model framework is capable of accurately simulating seasonal malaria dynamics and capturing of some of the inter-annual variation in malaria incidence. However, the model unsurprisingly failed to reproduce large declines in malaria transmission during 2018 and 2019 due to elevated anti-malaria measures, which were not accounted for in the model structure. The climate-driven transmission model also captured regional variation in malaria incidence across Rwanda's diverse climate, while identifying key entomological and epidemiological parameters important to seasonal malaria dynamics. In general, this new model construct advances the capabilities of implicitly-forced lower dimension dynamical malaria models by leveraging climate drivers of malaria ecology and transmission.
Israel Ukawuba, Jeffrey Shaman
PLoS Comput. Biol.2
2020 arcasHLA: high-resolution HLA typing from RNAseq
abstract
MOTIVATION: The human leukocyte antigen (HLA) locus plays a critical role in tissue compatibility and regulates the host response to many diseases, including cancers and autoimmune di3orders. Recent improvements in the quality and accessibility of next-generation sequencing have made HLA typing from standard short-read data practical. However, this task remains challenging given the high level of polymorphism and homology between HLA genes. HLA typing from RNA sequencing is further complicated by post-transcriptional modifications and bias due to amplification. RESULTS: Here, we present arcasHLA: a fast and accurate in silico tool that infers HLA genotypes from RNA-sequencing data. Our tool outperforms established tools on the gold-standard benchmark dataset for HLA typing in terms of both accuracy and speed, with an accuracy rate of 100% at two-field resolution for Class I genes, and over 99.7% for Class II. Furthermore, we evaluate the performance of our tool on a new biological dataset of 447 single-end total RNA samples from nasopharyngeal swabs, and establish the applicability of arcasHLA in metatranscriptome studies. AVAILABILITY AND IMPLEMENTATION: arcasHLA is available at https://github.com/RabadanLab/arcasHLA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Rose Orenbuch, Ioan Filip, Devon Comito, Jeffrey Shaman, Itsik Pe'er, Raul Rabadan
Bioinform.4
2020 Forecasting influenza in Europe using a metapopulation model incorporating cross-border commuting and air travel
abstract
Past work has shown that models incorporating human travel can improve the quality of influenza forecasts. Here, we develop and validate a metapopulation model of twelve European countries, in which international translocation of virus is driven by observed commuting and air travel flows, and use this model to generate influenza forecasts in conjunction with incidence data from the World Health Organization. We find that, although the metapopulation model fits the data well, it offers no improvement over isolated models in forecast quality. We discuss several potential reasons for these results. In particular, we note the need for data that are more comparable from country to country, and offer suggestions as to how surveillance systems might be improved to achieve this goal.
Sarah C. Kramer, Sen Pei, Jeffrey Shaman
PLoS Comput. Biol.3
2020 Aggregating forecasts of multiple respiratory pathogens supports more accurate forecasting of influenza-like illness
abstract
Influenza-like illness (ILI) is a commonly measured syndromic signal representative of a range of acute respiratory infections. Reliable forecasts of ILI can support better preparation for patient surges in healthcare systems. Although ILI is an amalgamation of multiple pathogens with variable seasonal phasing and attack rates, most existing process-based forecasting systems treat ILI as a single infectious agent. Here, using ILI records and virologic surveillance data, we show that ILI signal can be disaggregated into distinct viral components. We generate separate predictions for six contributing pathogens (influenza A/H1, A/H3, B, respiratory syncytial virus, and human parainfluenza virus types 1-2 and 3), and develop a method to forecast ILI by aggregating these predictions. The relative contribution of each pathogen to the total ILI signal is estimated using a Markov Chain Monte Carlo (MCMC) method upon forecast aggregation. We find highly variable overall contributions from influenza type A viruses across seasons, but relatively stable contributions for the other pathogens. Using historical data from 1997 to 2014 at US national and regional levels, the proposed forecasting system generates improved predictions of both seasonal and near-term targets relative to a baseline method that simulates ILI as a single pathogen. The hierarchical forecasting system can generate predictions for each viral component, as well as infer and predict their contributions to ILI, which may additionally help physicians determine the etiological causes of ILI in clinical settings.
Sen Pei, Jeffrey Shaman
PLoS Comput. Biol.2
2019 Reappraising the utility of Google Flu Trends
abstract
Estimation of influenza-like illness (ILI) using search trends activity was intended to supplement traditional surveillance systems, and was a motivation behind the development of Google Flu Trends (GFT). However, several studies have previously reported large errors in GFT estimates of ILI in the US. Following recent release of time-stamped surveillance data, which better reflects real-time operational scenarios, we reanalyzed GFT errors. Using three data sources-GFT: an archive of weekly ILI estimates from Google Flu Trends; ILIf: fully-observed ILI rates from ILINet; and, ILIp: ILI rates available in real-time based on partial reporting-five influenza seasons were analyzed and mean square errors (MSE) of GFT and ILIp as estimates of ILIf were computed. To correct GFT errors, a random forest regression model was built with ILI and GFT rates from the previous three weeks as predictors. An overall reduction in error of 44% was observed and the errors of the corrected GFT are lower than those of ILIp. An 80% reduction in error during 2012/13, when GFT had large errors, shows that extreme failures of GFT could have been avoided. Using autoregressive integrated moving average (ARIMA) models, one- to four-week ahead forecasts were generated with two separate data streams: ILIp alone, and with both ILIp and corrected GFT. At all forecast targets and seasons, and for all but two regions, inclusion of GFT lowered MSE. Results from two alternative error measures, mean absolute error and mean absolute proportional error, were largely consistent with results from MSE. Taken together these findings provide an error profile of GFT in the US, establish strong evidence for the adoption of search trends based 'nowcasts' in influenza forecast systems, and encourage reevaluation of the utility of this data source in diverse domains.
Sasikiran Kandula, Jeffrey Shaman
PLoS Comput. Biol.2
2019 Development and validation of influenza forecasting for 64 temperate and tropical countries
abstract
Accurate forecasts of influenza incidence can be used to inform medical and public health decision-making and response efforts.However, forecasting systems are uncommon in most countries, with a few notable exceptions.Here we use publicly available data from the World Health Organization to generate retrospective forecasts of influenza peak timing and peak intensity for 64 countries, including 18 tropical and subtropical countries.We find that accurate and well-calibrated forecasts can be generated for countries in temperate regions, with peak timing and intensity accuracy exceeding 50% at four and two weeks prior to the predicted epidemic peak, respectively.Forecasts are significantly less accurate in the tropics and subtropics for both peak timing and intensity.This work indicates that, in temperate regions around the world, forecasts can be generated with sufficient lead time to prepare for upcoming outbreak peak incidence. Author summaryInfluenza is responsible for an estimated 3-5 million cases and 300-650,000 deaths each year worldwide.If produced early enough, accurate forecasts of influenza activity could guide public health practitioners and medical professionals in preparing for an outbreak, reducing the subsequent morbidity and mortality.For example, hospitals could use these forecasts to determine how many beds will be needed when an outbreak is most intense.Despite this potential impact, influenza forecasts are primarily generated for the United States, with forecasts for other countries being comparatively rare.Here, we use publically available influenza data to forecast influenza activity in 64 countries.We find that accurate forecasts can be produced several weeks before the outbreak's peak in temperate countries, where influenza outbreaks occur regularly during the winter.Forecast accuracy is lower in the tropics and subtropics, where outbreaks occur more sporadically.Overall, our results suggest that forecasts have potential as an important public health tool in many countries, not only in the US.
Sarah C. Kramer, Jeffrey Shaman
PLoS Comput. Biol.2
2019 Predictability in process-based ensemble forecast of influenza
abstract
Process-based models have been used to simulate and forecast a number of nonlinear dynamical systems, including influenza and other infectious diseases. In this work, we evaluate the effects of model initial condition error and stochastic fluctuation on forecast accuracy in a compartmental model of influenza transmission. These two types of errors are found to have qualitatively similar growth patterns during model integration, indicating that dynamic error growth, regardless of source, is a dominant component of forecast inaccuracy. We therefore examine the nonlinear growth of model initial error and compute the fastest growing directions using singular vector analysis. Using this information, we generate perturbations in an ensemble forecast system of influenza to obtain more optimal ensemble spread. In retrospective forecasts of historical outbreaks for 95 US cities from 2003 to 2014, this approach improves short-term forecast of incidence over the next one to four weeks.
Sen Pei, Mark A. Cane, Jeffrey Shaman
PLoS Comput. Biol.3
2019 Accuracy of real-time multi-model ensemble forecasts for seasonal influenza in the U.S
abstract
Seasonal influenza results in substantial annual morbidity and mortality in the United States and worldwide. Accurate forecasts of key features of influenza epidemics, such as the timing and severity of the peak incidence in a given season, can inform public health response to outbreaks. As part of ongoing efforts to incorporate data and advanced analytical methods into public health decision-making, the United States Centers for Disease Control and Prevention (CDC) has organized seasonal influenza forecasting challenges since the 2013/2014 season. In the 2017/2018 season, 22 teams participated. A subset of four teams created a research consortium called the FluSight Network in early 2017. During the 2017/2018 season they worked together to produce a collaborative multi-model ensemble that combined 21 separate component models into a single model using a machine learning technique called stacking. This approach creates a weighted average of predictive densities where the weight for each component is determined by maximizing overall ensemble accuracy over past seasons. In the 2017/2018 influenza season, one of the largest seasonal outbreaks in the last 15 years, this multi-model ensemble performed better on average than all individual component models and placed second overall in the CDC challenge. It also outperformed the baseline multi-model ensemble created by the CDC that took a simple average of all models submitted to the forecasting challenge. This project shows that collaborative efforts between research teams to develop ensemble forecasting approaches can bring measurable improvements in forecast accuracy and important reductions in the variability of performance from year to year. Efforts such as this, that emphasize real-time testing and evaluation of forecasting models and facilitate the close collaboration between public health officials and modeling researchers, are essential to improving our understanding of how best to use forecasts to improve public health response to seasonal and emerging epidemic threats.
Nicholas G. Reich, Craig J. McGowan, Teresa K. Yamana, Abhinav Tushar, Evan L. Ray, Dave Osthus, Sasikiran Kandula, Logan C. Brooks, Willow Crawford-Crudell, Graham Casey Gibson, Evan Moore, Rebecca Silva, Matthew Biggerstaff, Michael A. Johansson, Ronald Rosenfeld, Jeffrey Shaman
PLoS Comput. Biol.16
2019 Characteristics of measles epidemics in China (1951-2004) and implications for elimination: A case study of three key locations
abstract
Measles is a highly infectious, severe viral disease. The disease is targeted for global eradication; however, this result has proven challenging. In China, where countrywide vaccination coverage for the last decade has been above 95% (the threshold for measles elimination), measles continues to cause large epidemics. To diagnose factors contributing to the persistency of measles, here we develop a model-inference system to infer measles transmission dynamics in China. The model-inference system uses demographic and vaccination data for each year as model inputs to directly account for changing population dynamics (including births, deaths, migrations, and vaccination). In addition, it simultaneously estimates unobserved model variables and parameters based on incidence data. When fitted to yearly incidence data for the entire population, it is able to accurately estimate independent, out-of-sample age-specific incidence. Using this validated model-inference system, we are thus able to estimate epidemiological and demographical characteristics key to measles transmission during 1951-2004 for three key locations in China, including its capital Beijing. These characteristics include age-specific population susceptibility and incidence rates, the basic reproductive number (R0), reporting rate, population mixing intensity, and amplitude of seasonality. Key differences among the three sites reveal population and epidemiological characteristics crucial for understanding the current persistence of measles epidemics in China. We also discuss the implications our findings have for future elimination strategies.
Wan Yang, Jeffrey Shaman
PLoS Comput. Biol.3
2019 Correction: Geospatial characteristics of measles transmission in China during 2005-2014
abstract
Measles is a highly contagious and severe disease.Despite mass vaccination, it remains a leading cause of death in children in developing regions, killing 114,900 globally in 2014.In 2006, China committed to eliminating measles by 2012; to this end, the country enhanced its mandatory vaccination programs and achieved vaccination rates reported above 95% by 2008.However, in spite of these efforts, during the last 3 years (2013-2015) China documented 27,695, 52,656, and 42,874 confirmed measles cases.How measles manages to spread in China-the world's largest population-in the mass vaccination era remains poorly understood.To address this conundrum and provide insights for future public health efforts, we analyze the geospatial pattern of measles transmission across China during 2005-2014.We map measles incidence and incidence rates for each of the 344 cities in mainland China, identify the key socioeconomic and demographic features associated with high disease burden, and identify transmission clusters based on the synchrony of outbreak cycles.Using hierarchical cluster analysis, we identify 21 epidemic clusters, of which 12 were cross-regional.The cross-regional clusters included more underdeveloped cities with large numbers of emigrants than would be expected by chance (p = 0.011; bootstrap sampling), indicating that cities in these clusters were likely linked by internal worker migration in response to uneven economic development.In contrast, cities in regional clusters were more likely to have high rates of minorities and high natural growth rates than would be expected by chance (p = 0.074; bootstrap sampling).Our findings suggest that multiple highly connected foci of measles transmission coexist in China and that migrant workers likely facilitate the transmission of measles across regions.This complex connection renders eradication of measles challenging in China despite its high overall vaccination coverage.Future immunization programs should therefore target these transmission foci simultaneously.
Wan Yang, Liang Wen, Shen-Long Li, Wen-Yi Zhang, Jeffrey Shaman
PLoS Comput. Biol.6
2018 Use of temperature to improve West Nile virus forecasts
abstract
Ecological and laboratory studies have demonstrated that temperature modulates West Nile virus (WNV) transmission dynamics and spillover infection to humans. Here we explore whether inclusion of temperature forcing in a model depicting WNV transmission improves WNV forecast accuracy relative to a baseline model depicting WNV transmission without temperature forcing. Both models are optimized using a data assimilation method and two observed data streams: mosquito infection rates and reported human WNV cases. Each coupled model-inference framework is then used to generate retrospective ensemble forecasts of WNV for 110 outbreak years from among 12 geographically diverse United States counties. The temperature-forced model improves forecast accuracy for much of the outbreak season. From the end of July until the beginning of October, a timespan during which 70% of human cases are reported, the temperature-forced model generated forecasts of the total number of human cases over the next 3 weeks, total number of human cases over the season, the week with the highest percentage of infectious mosquitoes, and the peak percentage of infectious mosquitoes that on average increased absolute forecast accuracy 5%, 10%, 12%, and 6%, respectively, over the non-temperature forced baseline model. These results indicate that use of temperature forcing improves WNV forecast accuracy and provide further evidence that temperature influences rates of WNV transmission. The findings provide a foundation for implementation of a statistically rigorous system for real-time forecast of seasonal WNV outbreaks and their use as a quantitative decision support tool for public health officials and mosquito control programs.
Nicholas B. DeFelice, Zachary D. Schneider, Eliza Little, Christopher Barker, Kevin A. Caillouët, Scott R. Campbell, Dan Damian, Patrick Irwin, Herff M. P. Jones, John Townsend, Jeffrey Shaman
PLoS Comput. Biol.11
2017 The use of ambient humidity conditions to improve influenza forecast
abstract
Laboratory and epidemiological evidence indicate that ambient humidity modulates the survival and transmission of influenza. Here we explore whether the inclusion of humidity forcing in mathematical models describing influenza transmission improves the accuracy of forecasts generated with those models. We generate retrospective forecasts for 95 cities over 10 seasons in the United States and assess both forecast accuracy and error. Overall, we find that humidity forcing improves forecast performance (at 1-4 lead weeks, 3.8% more peak week and 4.4% more peak intensity forecasts are accurate than with no forcing) and that forecasts generated using daily climatological humidity forcing generally outperform forecasts that utilize daily observed humidity forcing (4.4% and 2.6% respectively). These findings hold for predictions of outbreak peak intensity, peak timing, and incidence over 2- and 4-week horizons. The results indicate that use of climatological humidity forcing is warranted for current operational influenza forecast.
Jeffrey Shaman, Sasikiran Kandula, Wan Yang, Alicia Karspeck
PLoS Comput. Biol.1
2017 Individual versus superensemble forecasts of seasonal influenza outbreaks in the United States
abstract
Recent research has produced a number of methods for forecasting seasonal influenza outbreaks. However, differences among the predicted outcomes of competing forecast methods can limit their use in decision-making. Here, we present a method for reconciling these differences using Bayesian model averaging. We generated retrospective forecasts of peak timing, peak incidence, and total incidence for seasonal influenza outbreaks in 48 states and 95 cities using 21 distinct forecast methods, and combined these individual forecasts to create weighted-average superensemble forecasts. We compared the relative performance of these individual and superensemble forecast methods by geographic location, timing of forecast, and influenza season. We find that, overall, the superensemble forecasts are more accurate than any individual forecast method and less prone to producing a poor forecast. Furthermore, we find that these advantages increase when the superensemble weights are stratified according to the characteristics of the forecast or geographic location. These findings indicate that different competing influenza prediction systems can be combined into a single more accurate forecast product for operational delivery in real time.
Teresa K. Yamana, Sasikiran Kandula, Jeffrey Shaman
PLoS Comput. Biol.3
2017 Geospatial characteristics of measles transmission in China during 2005-2014
abstract
Measles is a highly contagious and severe disease. Despite mass vaccination, it remains a leading cause of death in children in developing regions, killing 114,900 globally in 2014. In 2006, China committed to eliminating measles by 2012; to this end, the country enhanced its mandatory vaccination programs and achieved vaccination rates reported above 95% by 2008. However, in spite of these efforts, during the last 3 years (2013-2015) China documented 27,695, 52,656, and 42,874 confirmed measles cases. How measles manages to spread in China-the world's largest population-in the mass vaccination era remains poorly understood. To address this conundrum and provide insights for future public health efforts, we analyze the geospatial pattern of measles transmission across China during 2005-2014. We map measles incidence and incidence rates for each of the 344 cities in mainland China, identify the key socioeconomic and demographic features associated with high disease burden, and identify transmission clusters based on the synchrony of outbreak cycles. Using hierarchical cluster analysis, we identify 21 epidemic clusters, of which 12 were cross-regional. The cross-regional clusters included more underdeveloped cities with large numbers of emigrants than would be expected by chance (p = 0.011; bootstrap sampling), indicating that cities in these clusters were likely linked by internal worker migration in response to uneven economic development. In contrast, cities in regional clusters were more likely to have high rates of minorities and high natural growth rates than would be expected by chance (p = 0.074; bootstrap sampling). Our findings suggest that multiple highly connected foci of measles transmission coexist in China and that migrant workers likely facilitate the transmission of measles across regions. This complex connection renders eradication of measles challenging in China despite its high overall vaccination coverage. Future immunization programs should therefore target these transmission foci simultaneously.
Wan Yang, Liang Wen, Shen-Long Li, Wen-Yi Zhang, Jeffrey Shaman
PLoS Comput. Biol.6
2016 Retrospective Parameter Estimation and Forecast of Respiratory Syncytial Virus in the United States
abstract
Recent studies have shown that systems combining mathematical modeling and Bayesian inference methods can be used to generate real-time forecasts of future infectious disease incidence. Here we develop such a system to study and forecast respiratory syncytial virus (RSV). RSV is the most common cause of acute lower respiratory infection and bronchiolitis. Advanced warning of the epidemic timing and volume of RSV patient surges has the potential to reduce well-documented delays of treatment in emergency departments. We use a susceptible-infectious-recovered (SIR) model in conjunction with an ensemble adjustment Kalman filter (EAKF) and ten years of regional U.S. specimen data provided by the Centers for Disease Control and Prevention. The data and EAKF are used to optimize the SIR model and i) estimate critical epidemiological parameters over the course of each outbreak and ii) generate retrospective forecasts. The basic reproductive number, R0, is estimated at 3.0 (standard deviation 0.6) across all seasons and locations. The peak magnitude of RSV outbreaks is forecast with nearly 70% accuracy (i.e. nearly 70% of forecasts within 25% of the actual peak), four weeks before the predicted peak. This work represents a first step in the development of a real-time RSV prediction system.
Julia Reis, Jeffrey Shaman
PLoS Comput. Biol.2
2016 Forecasting Influenza Outbreaks in Boroughs and Neighborhoods of New York City
abstract
The ideal spatial scale, or granularity, at which infectious disease incidence should be monitored and forecast has been little explored.By identifying the optimal granularity for a given disease and host population, and matching surveillance and prediction efforts to this scale, response to emergent and recurrent outbreaks can be improved.Here we explore how granularity and representation of spatial structure affect influenza forecast accuracy within New York City.We develop network models at the borough and neighborhood levels, and use them in conjunction with surveillance data and a data assimilation method to forecast influenza activity.These forecasts are compared to an alternate system that predicts influenza for each borough or neighborhood in isolation.At the borough scale, influenza epidemics are highly synchronous despite substantial differences in intensity, and inclusion of network connectivity among boroughs generally improves forecast accuracy.At the neighborhood scale, we observe much greater spatial heterogeneity among influenza outbreaks including substantial differences in local outbreak timing and structure; however, inclusion of the network model structure generally degrades forecast accuracy.One notable exception is that local outbreak onset, particularly when signal is modest, is better predicted with the network model.These findings suggest that observation and forecast at sub-municipal scales within New York City provides richer, more discriminant information on influenza incidence, particularly at the neighborhood scale where greater heterogeneity exists, and that the spatial spread of influenza among localities can be forecast. Author SummaryInfluenza, or the flu, causes significant morbidity and mortality during both seasonal and pandemic outbreaks.Recently developed influenza forecast systems have the potential to aid public health planning for and mitigation of the burden of this disease.However, current forecasts are often generated at spatial scales (e.g.national level) that are coarser than the scales at which public health measures and interventions are implemented (e.g.
Wan Yang, Donald R. Olson, Jeffrey Shaman
PLoS Comput. Biol.3
2015 Impact of School Cycles and Environmental Forcing on the Timing of Pandemic Influenza Activity in Mexican States, May-December 2009
abstract
While a relationship between environmental forcing and influenza transmission has been established in inter-pandemic seasons, the drivers of pandemic influenza remain debated. In particular, school effects may predominate in pandemic seasons marked by an atypical concentration of cases among children. For the 2009 A/H1N1 pandemic, Mexico is a particularly interesting case study due to its broad geographic extent encompassing temperate and tropical regions, well-documented regional variation in the occurrence of pandemic outbreaks, and coincidence of several school breaks during the pandemic period. Here we fit a series of transmission models to daily laboratory-confirmed influenza data in 32 Mexican states using MCMC approaches, considering a meta-population framework or the absence of spatial coupling between states. We use these models to explore the effect of environmental, school-related and travel factors on the generation of spatially-heterogeneous pandemic waves. We find that the spatial structure of the pandemic is best understood by the interplay between regional differences in specific humidity (explaining the occurrence of pandemic activity towards the end of the school term in late May-June 2009 in more humid southeastern states), school vacations (preventing influenza transmission during July-August in all states), and regional differences in residual susceptibility (resulting in large outbreaks in early fall 2009 in central and northern Mexico that had yet to experience fully-developed outbreaks). Our results are in line with the concept that very high levels of specific humidity, as present during summer in southeastern Mexico, favor influenza transmission, and that school cycles are a strong determinant of pandemic wave timing.
James Tamerius, Cécile Viboud, Jeffrey Shaman, Gerardo Chowell
PLoS Comput. Biol.3
2015 Forecasting Influenza Epidemics in Hong Kong
abstract
Recent advances in mathematical modeling and inference methodologies have enabled development of systems capable of forecasting seasonal influenza epidemics in temperate regions in real-time. However, in subtropical and tropical regions, influenza epidemics can occur throughout the year, making routine forecast of influenza more challenging. Here we develop and report forecast systems that are able to predict irregular non-seasonal influenza epidemics, using either the ensemble adjustment Kalman filter or a modified particle filter in conjunction with a susceptible-infected-recovered (SIR) model. We applied these model-filter systems to retrospectively forecast influenza epidemics in Hong Kong from January 1998 to December 2013, including the 2009 pandemic. The forecast systems were able to forecast both the peak timing and peak magnitude for 44 epidemics in 16 years caused by individual influenza strains (i.e., seasonal influenza A(H1N1), pandemic A(H1N1), A(H3N2), and B), as well as 19 aggregate epidemics caused by one or more of these influenza strains. Average forecast accuracies were 37% (for both peak timing and magnitude) at 1-3 week leads, and 51% (peak timing) and 50% (peak magnitude) at 0 lead. Forecast accuracy increased as the spread of a given forecast ensemble decreased; the forecast accuracy for peak timing (peak magnitude) increased up to 43% (45%) for H1N1, 93% (89%) for H3N2, and 53% (68%) for influenza B at 1-3 week leads. These findings suggest that accurate forecasts can be made at least 3 weeks in advance for subtropical and tropical regions.
Wan Yang, Benjamin J. Cowling, Eric H. Y. Lau, Jeffrey Shaman
PLoS Comput. Biol.4
2014 Spatial Transmission of 2009 Pandemic Influenza in the US
abstract
The 2009 H1N1 influenza pandemic provides a unique opportunity for detailed examination of the spatial dynamics of an emerging pathogen. In the US, the pandemic was characterized by substantial geographical heterogeneity: the 2009 spring wave was limited mainly to northeastern cities while the larger fall wave affected the whole country. Here we use finely resolved spatial and temporal influenza disease data based on electronic medical claims to explore the spread of the fall pandemic wave across 271 US cities and associated suburban areas. We document a clear spatial pattern in the timing of onset of the fall wave, starting in southeastern cities and spreading outwards over a period of three months. We use mechanistic models to tease apart the external factors associated with the timing of the fall wave arrival: differential seeding events linked to demographic factors, school opening dates, absolute humidity, prior immunity from the spring wave, spatial diffusion, and their interactions. Although the onset of the fall wave was correlated with school openings as previously reported, models including spatial spread alone resulted in better fit. The best model had a combination of the two. Absolute humidity or prior exposure during the spring wave did not improve the fit and population size only played a weak role. In conclusion, the protracted spread of pandemic influenza in fall 2009 in the US was dominated by short-distance spatial spread partially catalysed by school openings rather than long-distance transmission events. This is in contrast to the rapid hierarchical transmission patterns previously described for seasonal influenza. The findings underline the critical role that school-age children play in facilitating the geographic spread of pandemic influenza and highlight the need for further information on the movement and mixing patterns of this age group.
Julia R. Gog, Sébastien Ballesteros, Cécile Viboud, Lone Simonsen, Ottar N. Bjørnstad, Jeffrey Shaman, Dennis L. Chao, Farid Khan, Bryan T. Grenfell
PLoS Comput. Biol.6
2014 Comparison of Filtering Methods for the Modeling and Retrospective Forecasting of Influenza Epidemics
abstract
A variety of filtering methods enable the recursive estimation of system state variables and inference of model parameters. These methods have found application in a range of disciplines and settings, including engineering design and forecasting, and, over the last two decades, have been applied to infectious disease epidemiology. For any system of interest, the ideal filter depends on the nonlinearity and complexity of the model to which it is applied, the quality and abundance of observations being entrained, and the ultimate application (e.g. forecast, parameter estimation, etc.). Here, we compare the performance of six state-of-the-art filter methods when used to model and forecast influenza activity. Three particle filters--a basic particle filter (PF) with resampling and regularization, maximum likelihood estimation via iterated filtering (MIF), and particle Markov chain Monte Carlo (pMCMC)--and three ensemble filters--the ensemble Kalman filter (EnKF), the ensemble adjustment Kalman filter (EAKF), and the rank histogram filter (RHF)--were used in conjunction with a humidity-forced susceptible-infectious-recovered-susceptible (SIRS) model and weekly estimates of influenza incidence. The modeling frameworks, first validated with synthetic influenza epidemic data, were then applied to fit and retrospectively forecast the historical incidence time series of seven influenza epidemics during 2003-2012, for 115 cities in the United States. Results suggest that when using the SIRS model the ensemble filters and the basic PF are more capable of faithfully recreating historical influenza incidence time series, while the MIF and pMCMC do not perform as well for multimodal outbreaks. For forecast of the week with the highest influenza activity, the accuracies of the six model-filter frameworks are comparable; the three particle filters perform slightly better predicting peaks 1-5 weeks in the future; the ensemble filters are more accurate predicting peaks in the past.
Wan Yang, Alicia Karspeck, Jeffrey Shaman
PLoS Comput. Biol.3