David L. Buckeridge

dblp:74/3031 · DBLP profile ↗
← Back
49ranked-venue papers
7as first author
9since 2021 · last 2026
0000-0003-1817-5047ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 43 · 7 first-author · 6 since 2021Artificial intelligence and machine learning · 7 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TrajGPT: Irregular Time-Series Representation Learning of Health Trajectory
abstract
In the healthcare domain, time-series data are often irregularly sampled with varying intervals through outpatient visits, posing challenges for existing models designed for equally spaced sequential data. To address this, we propose Trajectory Generative Pre-trained Transformer (TrajGPT) for representation learning on irregularly-sampled healthcare time series. TrajGPT introduces a novel Selective Recurrent Attention (SRA) module that leverages a data-dependent decay to adaptively filter irrelevant past information. As a discretized ordinary differential equation (ODE) framework, TrajGPT captures underlying continuous dynamics and enables a time-specific inference for forecasting arbitrary target timesteps without auto-regressive prediction. Experimental results based on the longitudinal EHR data PopHR from Montreal health system and eICU from PhysioNet showcase TrajGPT's superior zero-shot performance in disease forecasting, drug usage prediction, and sepsis detection. The inferred trajectories of diabetic and cardiac patients reveal meaningful comorbidity conditions, underscoring TrajGPT as a useful tool for forecasting patient health evolution.
Qincheng Lu, David L. Buckeridge, Yue Li 0017
IEEE J. Biomed. Health Informatics4
2024 BAND: Biomedical Alert News Dataset
abstract
Infectious disease outbreaks continue to pose a significant threat to human health and well-being. To improve disease surveillance and understanding of disease spread, several surveillance systems have been developed to monitor daily news alerts and social media. However, existing systems lack thorough epidemiological analysis in relation to corresponding alerts or news, largely due to the scarcity of well-annotated reports data. To address this gap, we introduce the Biomedical Alert News Dataset (BAND), which includes 1,508 samples from existing reported news articles, open emails, and alerts, as well as 30 epidemiology-related questions. These questions necessitate the model's expert reasoning abilities, thereby offering valuable insights into the outbreak of the disease. The BAND dataset brings new challenges to the NLP world, requiring better inference capability of the content and the ability to infer important information. We provide several benchmark tasks, including Named Entity Recognition (NER), Question Answering (QA), and Event Extraction (EE), to demonstrate existing models' capabilities and limitations in handling epidemiology-specific tasks. It is worth noting that some models may lack the human-like inference capability required to fully utilize the corpus. To the best of our knowledge, the BAND corpus is the largest corpus of well-annotated biomedical outbreak alert news with elaborately designed questions, making it a valuable resource for epidemiologists and NLP researchers alike.
Meiru Zhang, Zaiqiao Meng, Yannan Shen, David L. Buckeridge, Nigel Collier
AAAI5
2024 CODA: an open-source platform for federated analysis and machine learning on distributed healthcare data
abstract
OBJECTIVES: Distributed computations facilitate multi-institutional data analysis while avoiding the costs and complexity of data pooling. Existing approaches lack crucial features, such as built-in medical standards and terminologies, no-code data visualizations, explicit disclosure control mechanisms, and support for basic statistical computations, in addition to gradient-based optimization capabilities. MATERIALS AND METHODS: We describe the development of the Collaborative Data Analysis (CODA) platform, and the design choices undertaken to address the key needs identified during our survey of stakeholders. We use a public dataset (MIMIC-IV) to demonstrate end-to-end multi-modal FL using CODA. We assessed the technical feasibility of deploying the CODA platform at 9 hospitals in Canada, describe implementation challenges, and evaluate its scalability on large patient populations. RESULTS: The CODA platform was designed, developed, and deployed between January 2020 and January 2023. Software code, documentation, and technical documents were released under an open-source license. Multi-modal federated averaging is illustrated using the MIMIC-IV and MIMIC-CXR datasets. To date, 8 out of the 9 participating sites have successfully deployed the platform, with a total enrolment of >1M patients. Mapping data from legacy systems to FHIR was the biggest barrier to implementation. DISCUSSION AND CONCLUSION: The CODA platform was developed and successfully deployed in a public healthcare setting in Canada, with heterogeneous information technology systems and capabilities. Ongoing efforts will use the platform to develop and prospectively validate models for risk assessment, proactive monitoring, and resource usage. Further work will also make tools available to facilitate migration from legacy formats to FHIR and DICOM.
Louis Mullie, Jonathan Afilalo, Patrick M. Archambault, Rima Bouchakri, Kip Brown, David L. Buckeridge, Yiorgos Alexandros Cavayas, Alexis F. Turgeon, Denis Martineau, François Lamontagne, Martine Lebrasseur, Renald Lemieux, Jeffrey Li, Michaël Sauthier, Pascal St-Onge, An Tang, William Witteman, Michael Chassé
J. Am. Medical Informatics Assoc.6
2022 Automatic Phenotyping by a Seed-guided Topic Model
abstract
Electronic health records (EHRs) provide rich clinical information and the opportunities to extract epidemiological patterns to understand and predict patient disease risks with suitable machine learning methods such as topic models. However, existing topic models do not generate identifiable topics each predicting a unique phenotype. One promising direction is to use known phenotype concepts to guide topic inference. We present a seed-guided Bayesian topic model called MixEHR-Seed with 3 contributions: (1) for each phenotype, we infer a dual-form of topic distribution: a seed-topic distribution over a small set of key EHR codes and a regular topic distribution over the entire EHR vocabulary; (2) we model age-dependent disease progression as Markovian dynamic topic priors; (3) we infer seed-guided multi-modal topics over distinct EHR data types. For inference, we developed a variational inference algorithm. Using MixEHR-Seed, we inferred 1569 PheCode-guided phenotype topics from an EHR database in Quebec, Canada covering 1.3 million patients for up to 20-year follow-up with 122 million records for 8539 and 1126 unique diagnostic and drug codes, respectively. We observed (1) accurate phenotype prediction by the guided topics, (2) clinically relevant PheCode-guided disease topics, (3) meaningful age-dependent disease prevalence. Source code is available at GitHub: https://github.com/li-lab-mcgill/MixEHR-Seed.
Yuanyi Hu, Aman Verma, David L. Buckeridge, Yue Li 0017
KDD4
2022 BioCaster in 2021: automatic disease outbreaks detection from global news media
abstract
SUMMARY: BioCaster was launched in 2008 to provide an ontology-based text mining system for early disease detection from open news sources. Following a 6-year break, we have re-launched the system in 2021. Our goal is to systematically upgrade the methodology using state-of-the-art neural network language models, whilst retaining the original benefits that the system provided in terms of logical reasoning and automated early detection of infectious disease outbreaks. Here, we present recent extensions such as neural machine translation in 10 languages, neural classification of disease outbreak reports and a new cloud-based visualization dashboard. Furthermore, we discuss our vision for further improvements, including combining risk assessment with event semantics and assessing the risk of outbreaks with multi-granularity. We hope that these efforts will benefit the global public health community. AVAILABILITY AND IMPLEMENTATION: BioCaster web-portal is freely accessible at http://biocaster.org.
Zaiqiao Meng, Anya Okhmatovskaia, Maxime Polleri, Yannan Shen, Guido Powell, Iris Ganser, Meiru Zhang, Nicholas B. King, David L. Buckeridge, Nigel Collier
Bioinform.10
2022 MixEHR-Guided: A guided multi-modal topic modeling approach for large-scale automatic phenotyping using the electronic health record
abstract
Electronic Health Records (EHRs) contain rich clinical data collected at the point of the care, and their increasing adoption offers exciting opportunities for clinical informatics, disease risk prediction, and personalized treatment recommendation. However, effective use of EHR data for research and clinical decision support is often hampered by a lack of reliable disease labels. To compile gold-standard labels, researchers often rely on clinical experts to develop rule-based phenotyping algorithms from billing codes and other surrogate features. This process is tedious and error-prone due to recall and observer biases in how codes and measures are selected, and some phenotypes are incompletely captured by a handful of surrogate features. To address this challenge, we present a novel automatic phenotyping model called MixEHR-Guided (MixEHR-G), a multimodal hierarchical Bayesian topic model that efficiently models the EHR generative process by identifying latent phenotype structure in the data. Unlike existing topic modeling algorithms wherein the inferred topics are not identifiable, MixEHR-G uses prior information from informative surrogate features to align topics with known phenotypes. We applied MixEHR-G to an openly-available EHR dataset of 38,597 intensive care patients (MIMIC-III) in Boston, USA and to administrative claims data for a population-based cohort (PopHR) of 1.3 million people in Quebec, Canada. Qualitatively, we demonstrate that MixEHR-G learns interpretable phenotypes and yields meaningful insights about phenotype similarities, comorbidities, and epidemiological associations. Quantitatively, MixEHR-G outperforms existing unsupervised phenotyping methods on a phenotype label annotation task, and it can accurately estimate relative phenotype prevalence functions without gold-standard phenotype information. Altogether, MixEHR-G is an important step towards building an interpretable and automated phenotyping system using EHR data.
Yuri Ahuja, Yuesong Zou, Aman Verma, David L. Buckeridge, Yue Li 0017
J. Biomed. Informatics4
2022 Novel informatics approaches to COVID-19 Research: From methods to applications
Hua Xu 0001, David L. Buckeridge, Fei Wang 0001, Peter Tarczy-Hornoch
J. Biomed. Informatics2
2021 Predicting Infectiousness for Proactive Contact Tracing
Yoshua Bengio, Prateek Gupta, Tegan Maharaj, Nasim Rahaman, Martin Weiss, Tristan Deleu, Eilif B. Muller, Meng Qu, Victor Schmidt, Pierre-Luc St-Charles, Hannah Alsdurf, Olexa Bilaniuk, David L. Buckeridge, Gaétan Marceau-Caron, Pierre Luc Carrier, Joumana Ghosn, Satya Ortiz-Gagne, Christopher Joseph Pal, Irina Rish, Bernhard Schölkopf, Jian Tang 0005, Andrew Robert Williams
ICLR13
2021 Guest Editorial Explainable AI: Towards Fairness, Accountability, Transparency and Trust in Healthcare
abstract
The papers in this special section focus on explainable artificial intelligence (AI) in healthcare services. Recent advances in AI, precision health, and medicine have paved the way for the accelerated adaptation and use of intelligent tools and systems in decision-making processes across the healthcare spectrum. Insights and knowledge derived from complex analytics are used to implement diagnostic and therapeutic solutions and targeted interventions in individuals and communities across the globe. Given the complexity of the current multi-dimensional clinical and public health data landscape, providing explainability in the context of socio-environmental and technical systems is a key to revealing pathways from socio-economic disadvantages to health disparities and implementing equitable interventions. As the complexity of the underlying data sets and AI-based algorithms increases, the explainability and justifiability of the insights generated decrease. Humans need to understand the underlying mechanism behind these insights to know whether they are sound, correct, trustable, and justifiable to make informed decisions. Lack of understandability and explainability in the biomedical domain often leads to poor transparency and accountability and ultimately lower quality of care and suboptimal and unfair health policies. Explainability is considered one of the prerequisites for deep medicine, where AI is meant to provide composite, panoramic views of individuals’ medical data.
Arash Shaban-Nejad, Martin Michalowski, John S. Brownstein, David L. Buckeridge
IEEE J. Biomed. Health Informatics4
2020 Seven pillars of precision digital health and medicine
Arash Shaban-Nejad, Martin Michalowski, Niels Peek, John S. Brownstein, David L. Buckeridge
Artif. Intell. Medicine5
2019 A systematic review of aberration detection algorithms used in public health surveillance
Mengru Yuan, Nikita Boston-Fisher, Yu T. Luo, Aman Verma, David L. Buckeridge
J. Biomed. Informatics5
2018 Usage and accuracy of medication data from nationwide health information exchange in Quebec, Canada
abstract
Objective: (1) To describe the usage of medication data from the Health Information Exchange (HIE) at the health care system level in the province of Quebec; (2) To assess the accuracy of the medication list obtained from the HIE. Methods: A descriptive study was conducted utilizing usage data obtained from the Ministry of Health at the individual provider level from January 1 to December 31, 2015. Usage patterns by role, type of site, and tool used to access the HIE were investigated. The list of medications of 111 high risk patients arriving at the emergency department of an academic healthcare center was obtained from the HIE and compared with the list obtained through the medication reconciliation process. Results: There were 31 022 distinct users accessing the HIE 11 085 653 times in 2015. The vast majority of pharmacists and general practitioners accessed it, compared to a minority of specialists and nurses. The top 1% of users was responsible of 19% of access. Also, 63% of the access was made using the Viewer application, while using a certified electronic medical record application seemed to facilitate usage. Among 111 patients, 71 (64%) had at least one discrepancy between the medication list obtained from the HIE and the reference list. Conclusions: Early adopters were mostly in primary care settings, and were accessing it more frequently when using a certified electronic medical record. Further work is needed to investigate how to resolve accuracy issues with the medication list and how certain tools provide different features.
Aude Motulsky, Daniala L. Weir, Isabelle Couture, Claude Sicotte, Marie-Pierre Gagnon, David L. Buckeridge, Robyn Tamblyn
J. Am. Medical Informatics Assoc.6
2018 Improving patient safety and efficiency of medication reconciliation through the development and adoption of a computer-assisted tool with automated electronic integration of population-based community drug data: the RightRx project
abstract
Background and Objective: Many countries require hospitals to implement medication reconciliation for accreditation, but the process is resource-intensive, thus adherence is poor. We report on the impact of prepopulating and aligning community and hospital drug lists with data from population-based and hospital-based drug information systems to reduce workload and enhance adoption and use of an e-medication reconciliation application, RightRx. Methods: The prototype e-medical reconciliation web-based software was developed for a cluster-randomized trial at the McGill University Health Centre. User-centered design and agile development processes were used to develop features intended to enhance adoption, safety, and efficiency. RightRx was implemented in medical and surgical wards, with support and training provided by unit champions and field staff. The time spent per professional using RightRx was measured, as well as the medication reconciliation completion rates in the intervention and control units during the first 20 months of the trial. Results: Users identified required modifications to the application, including the need for dose-based prescribing, the role of the discharge physician in prescribing community-based medication, and access to the rationale for medication decisions made during hospitalization. In the intervention units, both physicians and pharmacists were involved in discharge reconciliation, for 96.1% and 71.9% of patients, respectively. Medication reconciliation was completed for 80.7% (surgery) to 96.0% (medicine) of patients in the intervention units, and 0.7% (surgery) to 82.7% of patients in the control units. The odds of completing medication reconciliation were 9 times greater in the intervention compared to control units (odds ratio: 9.0, 95% confidence interval, 7.4-10.9, P < .0001) after adjusting for differences in patient characteristics. Conclusion: High rates of medication reconciliation completion were achieved with automated prepopulation and alignment of community and hospital medication lists.
Robyn Tamblyn, Nancy Winslade, Todd C. Lee, Aude Motulsky, Ari Meguerditchian, Melissa Bustillo, Sarah Elsayed, David L. Buckeridge, Isabelle Couture, Christina J. Qian, Teresa Moraga, Allen Huang
J. Am. Medical Informatics Assoc.8
2017 A novel application of point-of-sales grocery transaction data to enhance community nutrition monitoring
Hiroshi Mamiya, Erica E. Moodie, David L. Buckeridge
AMIA3
2017 Initial Usability Evaluation of a Knowledge-Based Population Health Information System: The Population Health Record (PopHR)
Mengru Yuan, David L. Buckeridge, Guido Powell, Maxime Lavigne, Anya Okhmatovskaia
AMIA2
2016 Enteric disease episodes and the risk of acquiring a future sexually transmitted infection: a prediction model in Montreal residents
abstract
OBJECTIVE: The sexual transmission of enteric diseases poses an important public health challenge. We aimed to build a prediction model capable of identifying individuals with a reported enteric disease who could be at risk of acquiring future sexually transmitted infections (STIs). MATERIALS AND METHODS: Passive surveillance data on Montreal residents with at least 1 enteric disease report was used to construct the prediction model. Cases were defined as all subjects with at least 1 STI report following their initial enteric disease episode. A final logistic regression prediction model was chosen using forward stepwise selection. RESULTS: The prediction model with the greatest validity included age, sex, residential location, number of STI episodes experienced prior to the first enteric disease episode, type of enteric disease acquired, and an interaction term between age and male sex. This model had an area under the curve of 0.77 and had acceptable calibration. DISCUSSION: A coordinated public health response to the sexual transmission of enteric diseases requires that a distinction be made between cases of enteric diseases transmitted through sexual activity from those transmitted through contaminated food or water. A prediction model can aid public health officials in identifying individuals who may have a higher risk of sexually acquiring a reportable disease. Once identified, these individuals could receive specialized intervention to prevent future infection. CONCLUSION: The information produced from a prediction model capable of identifying higher risk individuals can be used to guide efforts in investigating and controlling reported cases of enteric diseases and STIs.
Melissa Caron, Robert Allard, Lucie Bédard, Jérôme Latreille, David L. Buckeridge
J. Am. Medical Informatics Assoc.5
2015 The SDIDS System for Integrating Global Health Surveillance Data: An Example Application to Malaria Surveillance in Uganda
Kate Zinszer, Anya Okhmatovskaia, Arash Shaban-Nejad, Lauren N. Carroll, Neil F. Abernethy, David L. Buckeridge
AMIA6
2015 A novel method of adverse event detection can accurately identify venous thromboembolisms (VTEs) from narrative electronic health record data
abstract
BACKGROUND: Venous thromboembolisms (VTEs), which include deep vein thrombosis (DVT) and pulmonary embolism (PE), are associated with significant mortality, morbidity, and cost in hospitalized patients. To evaluate the success of preventive measures, accurate and efficient methods for monitoring VTE rates are needed. Therefore, we sought to determine the accuracy of statistical natural language processing (NLP) for identifying DVT and PE from electronic health record data. METHODS: We randomly sampled 2000 narrative radiology reports from patients with a suspected DVT/PE in Montreal (Canada) between 2008 and 2012. We manually identified DVT/PE within each report, which served as our reference standard. Using a bag-of-words approach, we trained 10 alternative support vector machine (SVM) models predicting DVT, and 10 predicting PE. SVM training and testing was performed with nested 10-fold cross-validation, and the average accuracy of each model was measured and compared. RESULTS: On manual review, 324 (16.2%) reports were DVT-positive and 154 (7.7%) were PE-positive. The best DVT model achieved an average sensitivity of 0.80 (95% CI 0.76 to 0.85), specificity of 0.98 (98% CI 0.97 to 0.99), positive predictive value (PPV) of 0.89 (95% CI 0.85 to 0.93), and an area under the curve (AUC) of 0.98 (95% CI 0.97 to 0.99). The best PE model achieved sensitivity of 0.79 (95% CI 0.73 to 0.85), specificity of 0.99 (95% CI 0.98 to 0.99), PPV of 0.84 (95% CI 0.75 to 0.92), and AUC of 0.99 (95% CI 0.98 to 1.00). CONCLUSIONS: Statistical NLP can accurately identify VTE from narrative radiology reports.
Christian M. Rochefort, Aman Verma, Tewodros Eguale, Todd C. Lee, David L. Buckeridge
J. Am. Medical Informatics Assoc.5
2015 Using age, triage score, and disposition data from emergency department electronic records to improve Influenza-like illness surveillance
abstract
OBJECTIVE: Markers of illness severity are increasingly captured in emergency department (ED) electronic systems, but their value for surveillance is not known. We assessed the value of age, triage score, and disposition data from ED electronic records for predicting influenza-related hospitalizations. MATERIALS AND METHODS: From June 2006 to January 2011, weekly counts of pneumonia and influenza (P&I) hospitalizations from five Montreal hospitals were modeled using negative binomial regression. Over lead times of 0-5 weeks, we assessed the predictive ability of weekly counts of 1) total ED visits, 2) ED visits with influenza-like illness (ILI), and 3) ED visits with ILI stratified by age, triage score, or disposition. Models were adjusted for secular trends, seasonality, and autocorrelation. Model fit was assessed using Akaike information criterion, and predictive accuracy using the mean absolute scaled error (MASE). RESULTS: Predictive accuracy for P&I hospitalizations during non-pandemic years was improved when models included visits from patients ≥65 years old and visits resulting in admission/transfer/death (MASE of 0.64, 95% confidence interval (95% CI) 0.54-0.80) compared to overall ILI visits (0.89, 95% CI 0.69-1.10). During the H1N1 pandemic year, including visits from patients <18 years old, visits with high priority triage scores, or visits resulting in admission/transfer/death resulted in the best model fit. DISCUSSION: Age and disposition data improved model fit and moderately reduced the prediction error for P&I hospitalizations; triage score improved model fit only during the pandemic year. CONCLUSION: Incorporation of age and severity measures available in ED records can improve ILI surveillance algorithms.
Noémie Savard, Lucie Bédard, Robert Allard, David L. Buckeridge
J. Am. Medical Informatics Assoc.4
2015 Quantifying the determinants of outbreak detection performance through simulation and machine learning
Nastaran Jafarpour, Masoumeh T. Izadi, Doina Precup, David L. Buckeridge
J. Biomed. Informatics4
2015 Towards probabilistic decision support in public health practice: Predicting recent transmission of tuberculosis from patient attributes
Hiroshi Mamiya, Kevin Schwartzman, Aman Verma, Christian Jauvin, Marcel Behr, David L. Buckeridge
J. Biomed. Informatics6
2014 Implementation of a Population Health Record in Montreal, Canada
Maxime Lavigne, Lam Dang-Duy, Catherine Ghassemian, Alexis Hamel, Anya Okhmatovskaia, Mojtaba Peyvandy, Arash Shaban-Nejad, David L. Buckeridge
AMIA8
2013 Using Hierarchical Mixture of Experts Model for Fusion of Outbreak Detection Methods
Nastaran Jafarpour, Doina Precup, Masoumeh T. Izadi, David L. Buckeridge
AMIA4
2013 Assessing the Predictability of Hospital Readmission Using Machine Learning
abstract
Unplanned hospital readmissions raise health care costs and cause significant distress to patients. Hence, predicting which patients are at risk to be readmitted is of great interest. In this paper, we mine large amounts of administrative information from claim data, including patients demographics, dispensed drugs, medical or surgical procedures performed, and medical diagnosis, in order to predict readmission using supervised learning methods. Our objective is to gain knowledge about the predictive power of the available information. Our preliminary results on data from the provincial hospital system in Quebec illustrate the potential for this approach to reveal important information on factors that trigger hospital readmission. Our findings suggest that a substantial portion of readmissions is inherently hard to predict. Consequently, the use of the raw readmission rate as an indicator of the quality of provided care might not be appropriate.
Arian Hosseinzadeh, Masoumeh T. Izadi, Aman Verma, Doina Precup, David L. Buckeridge
IAAI5
2012 Self-reported fever and measured temperature in emergency department records used for syndromic surveillance
abstract
Many public health agencies monitor population health using syndromic surveillance, generally employing information from emergency department (ED) visit records. When combined with other information, objective evidence of fever may enhance the accuracy with which surveillance systems detect syndromes of interest, such as influenza-like illness. This study found that patient chief complaint of self-reported fever was more readily available in ED records than measured temperature and that the majority of patients with an elevated temperature recorded also self-reported fever. Due to its currently limited availability, we conclude that measured temperature is likely to add little value to self-reported fever in syndromic surveillance for febrile illness using ED records.
Taha A. Kass-Hout, David L. Buckeridge, John S. Brownstein, Zhiheng Xu, Paul McMurray, Charles K. T. Ishikawa, Julia E. Gunn, Barbara L. Massoudi
J. Am. Medical Informatics Assoc.2
2012 Application of change point analysis to daily influenza-like illness emergency department visits
abstract
BACKGROUND: The utility of healthcare utilization data from US emergency departments (EDs) for rapid monitoring of changes in influenza-like illness (ILI) activity was highlighted during the recent influenza A (H1N1) pandemic. Monitoring has tended to rely on detection algorithms, such as the Early Aberration Reporting System (EARS), which are limited in their ability to detect subtle changes and identify disease trends. OBJECTIVE: To evaluate a complementary approach, change point analysis (CPA), for detecting changes in the incidence of ED visits due to ILI. METHODOLOGY AND PRINCIPAL FINDINGS: Data collected through the Distribute project (isdsdistribute.org), which aggregates data on ED visits for ILI from over 50 syndromic surveillance systems operated by state or local public health departments were used. The performance was compared of the cumulative sum (CUSUM) CPA method in combination with EARS and the performance of three CPA methods (CUSUM, structural change model and Bayesian) in detecting change points in daily time-series data from four contiguous US states participating in the Distribute network. Simulation data were generated to assess the impact of autocorrelation inherent in these time-series data on CPA performance. The CUSUM CPA method was robust in detecting change points with respect to autocorrelation in time-series data (coverage rates at 90% when -0.2≤ρ≤0.2 and 80% when -0.5≤ρ≤0.5). During the 2008-9 season, 21 change points were detected and ILI trends increased significantly after 12 of these change points and decreased nine times. In the 2009-10 flu season, we detected 11 change points and ILI trends increased significantly after two of these change points and decreased nine times. Using CPA combined with EARS to analyze automatically daily ED-based ILI data, a significant increase was detected of 3% in ILI on April 27, 2009, followed by multiple anomalies in the ensuing days, suggesting the onset of the H1N1 pandemic in the four contiguous states. CONCLUSIONS AND SIGNIFICANCE: As a complementary approach to EARS and other aberration detection methods, the CPA method can be used as a tool to detect subtle changes in time-series data more effectively and determine the moving direction (ie, up, down, or stable) in ILI trends between change points. The combined use of EARS and CPA might greatly improve the accuracy of outbreak detection in syndromic surveillance systems.
Taha A. Kass-Hout, Zhiheng Xu, Paul McMurray, Soyoun Park, David L. Buckeridge, John S. Brownstein, Lyn Finelli, Samuel L. Groseclose
J. Am. Medical Informatics Assoc.5
2012 Adjusting outbreak detection algorithms for surveillance during epidemic and non-epidemic periods
abstract
Many aberration detection algorithms are used in infectious disease surveillance systems to assist in the early detection of potential outbreaks. In this study, we explored a novel approach to adjusting aberration detection algorithms to account for the impact of seasonality inherent in some surveillance data. By using surveillance data for hand-foot-and-mouth disease in Shandong province, China, we evaluated the use of seasonally-adjusted alerting thresholds with three aberration detection methods (C1, C2, and C3). We found that the optimal thresholds of C1, C2, and C3 varied between the epidemic and non-epidemic seasons of hand-foot-and-mouth disease, and the application of seasonally adjusted thresholds improved the performance of outbreak detection by maintaining the same sensitivity and timeliness while decreasing by nearly half the false alert rate during the non-epidemic season. Our preliminary findings suggest a general approach to improving aberration detection for outbreaks of infectious disease with seasonally variable incidence.
Shengjie Lai, David L. Buckeridge, Honglong Zhang, Yajia Lan, Weizhong Yang
J. Am. Medical Informatics Assoc.3
2012 The effectiveness of a new generation of computerized drug alerts in reducing the risk of injury from drug side effects: a cluster randomized trial
abstract
CONTEXT: Computerized drug alerts for psychotropic drugs are expected to reduce fall-related injuries in older adults. However, physicians over-ride most alerts because they believe the benefit of the drugs exceeds the risk. OBJECTIVE: To determine whether computerized prescribing decision support with patient-specific risk estimates would increase physician response to psychotropic drug alerts and reduce injury risk in older people. DESIGN: Cluster randomized controlled trial of 81 family physicians and 5628 of their patients aged 65 and older who were prescribed psychotropic medication. INTERVENTION: Intervention physicians received information about patient-specific risk of injury computed at the time of each visit using statistical models of non-modifiable risk factors and psychotropic drug doses. Risk thermometers presented changes in absolute and relative risk with each change in drug treatment. Control physicians received commercial drug alerts. MAIN OUTCOME MEASURES: Injury risk at the end of follow-up based on psychotropic drug doses and non-modifiable risk factors. Electronic health records and provincial insurance administrative data were used to measure outcomes. RESULTS: Mean patient age was 75.2 years. Baseline risk of injury was 3.94 per 100 patients per year. Intermediate-acting benzodiazepines (56.2%) were the most common psychotropic drug. Intervention physicians reviewed therapy in 83.3% of visits and modified therapy in 24.6%. The intervention reduced the risk of injury by 1.7 injuries per 1000 patients (95% CI 0.2/1000 to 3.2/1000; p=0.02). The effect of the intervention was greater for patients with higher baseline risks of injury (p<0.03). CONCLUSION: Patient-specific risk estimates provide an effective method of reducing the risk of injury for high-risk older people. TRIAL REGISTRATION NUMBER: clinicaltrials.gov Identifier: NCT00818285.
Robyn Tamblyn, Tewodros Eguale, David L. Buckeridge, Allen Huang, James A. Hanley, Kristen Reidel, Sherry Shi, Nancy Winslade
J. Am. Medical Informatics Assoc.3
2012 Infectious Disease Informatics: Syndromic Surveillance for Public Health and Bio-Defense, Hsinchun Chen, Daniel Zeng, Ping Yan. Springer (2010)
David L. Buckeridge
J. Biomed. Informatics1
2011 COPE: Childhood Obesity Prevention [Knowledge] Enterprise
Arash Shaban-Nejad, David L. Buckeridge, Laurette Dubé
AIME2
2011 A secure protocol for protecting the identity of providers when disclosing data for disease surveillance
abstract
BACKGROUND: Providers have been reluctant to disclose patient data for public-health purposes. Even if patient privacy is ensured, the desire to protect provider confidentiality has been an important driver of this reluctance. METHODS: Six requirements for a surveillance protocol were defined that satisfy the confidentiality needs of providers and ensure utility to public health. The authors developed a secure multi-party computation protocol using the Paillier cryptosystem to allow the disclosure of stratified case counts and denominators to meet these requirements. The authors evaluated the protocol in a simulated environment on its computation performance and ability to detect disease outbreak clusters. RESULTS: Theoretical and empirical assessments demonstrate that all requirements are met by the protocol. A system implementing the protocol scales linearly in terms of computation time as the number of providers is increased. The absolute time to perform the computations was 12.5 s for data from 3000 practices. This is acceptable performance, given that the reporting would normally be done at 24 h intervals. The accuracy of detection disease outbreak cluster was unchanged compared with a non-secure distributed surveillance protocol, with an F-score higher than 0.92 for outbreaks involving 500 or more cases. CONCLUSION: The protocol and associated software provide a practical method for providers to disclose patient data for sentinel, syndromic or other indicator-based surveillance while protecting patient privacy and the identity of individual providers.
Khaled El Emam, Jay Mercer, Liam Peyton, Murat Kantarcioglu, Bradley A. Malin, David L. Buckeridge, Saeed Samet, Craig Earle
J. Am. Medical Informatics Assoc.7
2011 Outpatient physician billing data for age and setting specific syndromic surveillance of influenza-like illnesses
Emily H. Chan, Robyn Tamblyn, Katia M. L. Charland, David L. Buckeridge
J. Biomed. Informatics4
2010 Developing syndrome definitions based on consensus and current use
abstract
OBJECTIVE: Standardized surveillance syndromes do not exist but would facilitate sharing data among surveillance systems and comparing the accuracy of existing systems. The objective of this study was to create reference syndrome definitions from a consensus of investigators who currently have or are building syndromic surveillance systems. DESIGN: Clinical condition-syndrome pairs were catalogued for 10 surveillance systems across the United States and the representatives of these systems were brought together for a workshop to discuss consensus syndrome definitions. RESULTS: Consensus syndrome definitions were generated for the four syndromes monitored by the majority of the 10 participating surveillance systems: Respiratory, gastrointestinal, constitutional, and influenza-like illness (ILI). An important element in coming to consensus quickly was the development of a sensitive and specific definition for respiratory and gastrointestinal syndromes. After the workshop, the definitions were refined and supplemented with keywords and regular expressions, the keywords were mapped to standard vocabularies, and a web ontology language (OWL) ontology was created. LIMITATIONS: The consensus definitions have not yet been validated through implementation. CONCLUSION: The consensus definitions provide an explicit description of the current state-of-the-art syndromes used in automated surveillance, which can subsequently be systematically evaluated against real data to improve the definitions. The method for creating consensus definitions could be applied to other domains that have diverse existing definitions.
Wendy W. Chapman, John N. Dowling, Atar Baer, David L. Buckeridge, Dennis Cochrane, Michael A. Conway, Peter L. Elkin, Jeremy U. Espino, Julia E. Gunn, Craig M. Hales, Lori Hutwagner, Mikaela Keller, Catherine Larson, Rebecca Noe, Anya Okhmatovskaia, Karen Olson, Marc Paladini, Matthew Scholer, Carol Sniegoski, William B. Lober
J. Am. Medical Informatics Assoc.4
2009 A Bayesian Network Model for Analysis of Detection Performance in Surveillance Systems
Masoumeh T. Izadi, David L. Buckeridge, Anya Okhmatovskaia, Samson W. Tu, Martin J. O'Connor, Csongor Nyulas, Mark A. Musen
AMIA2
2009 Sensitivity Analysis of POMDP Value Functions
abstract
In sequential decision making under uncertainty, as in many other modeling endeavors, researchers observe a dynamical system and collect data measuring its behavior over time. These data are often used to build models that explain relationships between the measured variables, and are eventually used for planning and control purposes. However, these measurements cannot always be exact, systems can change over time, and discovering these facts or fixing these problems is not always feasible. Therefore it is important to formally describe the degree to which the model can tolerate noise, in order to keep near optimal behavior. The problem of finding tolerance bounds has been the focus of many studies for Markov Decision Processes (MDPs) due to their usefulness in practical applications. In this paper, we consider Partially Observable MDPs (POMDPs), which is a more realistic extension of MDPs with a wider scope of applications. We address two types of perturbations in POMDP model parameters, namely additive and multiplicative, and provide theoretical bounds for the impact of these changes in the value function. Experimental results are provided to illustrate our POMDP perturbation analysis in practice.
Stéphane Ross, Masoumeh T. Izadi, Mark Mercer, David L. Buckeridge
ICMLA4
2008 Predicting Outbreak Detection in Public Health Surveillance: Quantitative Analysis to Enable Evidence-Based Method Selection
David L. Buckeridge, Anya Okhmatovskaia, Samson W. Tu, Martin J. O'Connor, Csongor Nyulas, Mark A. Musen
AMIA1
2008 Model Formulation: Understanding Detection Performance in Public Health Surveillance: Modeling Aberrancy-detection Algorithms
abstract
OBJECTIVE: Statistical aberrancy-detection algorithms play a central role in automated public health systems, analyzing large volumes of clinical and administrative data in real-time with the goal of detecting disease outbreaks rapidly and accurately. Not all algorithms perform equally well in terms of sensitivity, specificity, and timeliness in detecting disease outbreaks and the evidence describing the relative performance of different methods is fragmented and mainly qualitative. DESIGN: We developed and evaluated a unified model of aberrancy-detection algorithms and a software infrastructure that uses this model to conduct studies to evaluate detection performance. We used a task-analytic methodology to identify the common features and meaningful distinctions among different algorithms and to provide an extensible framework for gathering evidence about the relative performance of these algorithms using a number of evaluation metrics. We implemented our model as part of a modular software infrastructure (Biological Space-Time Outbreak Reasoning Module, or BioSTORM) that allows configuration, deployment, and evaluation of aberrancy-detection algorithms in a systematic manner. MEASUREMENT: We assessed the ability of our model to encode the commonly used EARS algorithms and the ability of the BioSTORM software to reproduce an existing evaluation study of these algorithms. RESULTS: Using our unified model of aberrancy-detection algorithms, we successfully encoded the EARS algorithms, deployed these algorithms using BioSTORM, and were able to reproduce and extend previously published evaluation results. CONCLUSION: The validated model of aberrancy-detection algorithms and its software implementation will enable principled comparison of algorithms, synthesis of results from evaluation studies, and identification of surveillance algorithms for use in specific public health settings.
David L. Buckeridge, Anya Okhmatovskaia, Samson W. Tu, Martin J. O'Connor, Csongor Nyulas, Mark A. Musen
J. Am. Medical Informatics Assoc.1
2007 Optimizing Anthrax Outbreak Detection Using Reinforcement Learning
Masoumeh T. Izadi, David L. Buckeridge
AAAI2
2007 Decision Theoretic Analysis of Improving Epidemic Detection
Masoumeh T. Izadi, David L. Buckeridge
AMIA2
2007 Methods Paper: Finding Leading Indicators for Disease Outbreaks: Filtering, Cross-correlation, and Caveats
abstract
Bioterrorism and emerging infectious diseases such as influenza have spurred research into rapid outbreak detection. One primary thrust of this research has been to identify data sources that provide early indication of a disease outbreak by being leading indicators relative to other established data sources. Researchers tend to rely on the sample cross-correlation function (CCF) to quantify the association between two data sources. There has been, however, little consideration by medical informatics researchers of the influence of methodological choices on the ability of the CCF to identify a lead-lag relationship between time series. We draw on experience from the econometric and environmental health communities, and we use simulation to demonstrate that the sample CCF is highly prone to bias. Specifically, long-scale phenomena tend to overwhelm the CCF, obscuring phenomena at shorter wave lengths. Researchers seeking lead-lag relationships in surveillance data must therefore stipulate the scale length of the features of interest (e.g., short-scale spikes versus long-scale seasonal fluctuations) and then filter the data appropriately--to diminish the influence of other features, which may mask the features of interest. Otherwise, conclusions drawn from the sample CCF of bi-variate time-series data will inevitably be ambiguous and often altogether misleading.
Ronald M. Bloom, David L. Buckeridge, Karen E. Cheng
J. Am. Medical Informatics Assoc.2
2007 Outbreak detection through automated surveillance: A review of the determinants of detection
David L. Buckeridge
J. Biomed. Informatics1
2005 Decision Support for Community-Based Empirical Antibiotic Prescribing
Dana Teltsch, David Pinelle, Nancy Winslade, James A. Hanley, David L. Buckeridge, Robyn Tamblyn
AMIA5
2005 Algorithms for rapid outbreak detection: a research synthesis
David L. Buckeridge, Howard S. Burkom, Murray Campbell, William R. Hogan, Andrew W. Moore 0001
J. Biomed. Informatics1
2004 Review Paper: Implementing Syndromic Surveillance: A Practical Guide Informed by the Early Experience
abstract
Syndromic surveillance refers to methods relying on detection of individual and population health indicators that are discernible before confirmed diagnoses are made. In particular, prior to the laboratory confirmation of an infectious disease, ill persons may exhibit behavioral patterns, symptoms, signs, or laboratory findings that can be tracked through a variety of data sources. Syndromic surveillance systems are being developed locally, regionally, and nationally. The efforts have been largely directed at facilitating the early detection of a covert bioterrorist attack, but the technology may also be useful for general public health, clinical medicine, quality improvement, patient safety, and research. This report, authored by developers and methodologists involved in the design and deployment of the first wave of syndromic surveillance systems, is intended to serve as a guide for informaticians, public health managers, and practitioners who are currently planning deployment of such systems in their regions.
Kenneth D. Mandl, J. Marc Overhage, Michael M. Wagner 0001, William B. Lober, Paola Sebastiani, Farzad Mostashari, Julie A. Pavlin, Per H. Gesteland, Tracee Treadwell, Eileen Koski, Lori Hutwagner, David L. Buckeridge, Raymond D. Aller, Shaun J. Grannis
J. Am. Medical Informatics Assoc.12
2003 An Analytic Framework for Space-Time Aberrancy Detection in Public Health Surveillance Data
David L. Buckeridge, Mark A. Musen, Paul Switzer, Monica Crubézy
AMIA1
2003 BioSTORM: A System for Automated Surveillance of Diverse Data Sources
Martin J. O'Connor, David L. Buckeridge, Michael Choy, Monica Crubézy, Zachary Pincus, Mark A. Musen
AMIA2
2002 Knowledge-based bioterrorism surveillance
David L. Buckeridge, Justin Graham, Martin J. O'Connor, Michael Choy, Samson W. Tu, Mark A. Musen
AMIA1
2002 Conceptual Heterogeneity Complicates Automated Syndromic Surveillance for Bioterrorism
Justin Graham, David L. Buckeridge, Michael Choy, Mark A. Musen
AMIA2
2000 What Is Happenning In Canadian Health Informatics Research and Education?
Daniel Gordon, Kevin J. Leonard, Michael Carter, David L. Buckeridge
AMIA4