VLDB 2026 Research / reviewers in the wild / expert
Ben Y. Reis
dblp:92/4699
· DBLP profile ↗
15ranked-venue papers
5as first author
4since 2021 · last 2023
0000-0001-9908-5523ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 15 · 5 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | The value of parental medical records for the prediction of diabetes and cardiovascular disease: a novel method for generating and incorporating family historiesabstractOBJECTIVE: To determine whether data-driven family histories (DDFH) derived from linked EHRs of patients and their parents can improve prediction of patients' 10-year risk of diabetes and atherosclerotic cardiovascular disease (ASCVD). MATERIALS AND METHODS: A retrospective cohort study using data from Israel's largest healthcare organization. A random sample of 200 000 subjects aged 40-60 years on the index date (January 1, 2010) was included. Subjects with insufficient history (<1 year) or insufficient follow-up (<10 years) were excluded. Two separate XGBoost models were developed-1 for diabetes and 1 for ASCVD-to predict the 10-year risk for each outcome based on data available prior to the index date of January 1, 2010. RESULTS: Overall, the study included 110 734 subject-father-mother triplets. There were 22 153 cases of diabetes (20%) and 11 715 cases of ASCVD (10.6%). The addition of parental information significantly improved prediction of diabetes risk (P < .001), but not ASCVD risk. For both outcomes, maternal medical history was more predictive than paternal medical history. A binary variable summarizing parental disease state delivered similar predictive results to the full parental EHR. DISCUSSION: The increasing availability of EHRs for multiple family generations makes DDFH possible and can assist in delivering more personalized and precise medicine to patients. Consent frameworks must be established to enable sharing of information across generations, and the results suggest that sharing the full records may not be necessary. CONCLUSION: DDFH can address limitations of patient self-reported family history, and it improves clinical predictions for some conditions, but not for all, and particularly among younger adults. Yuval Barak-Corren, David Tsurel, Daphna Keidar, Ilan Gofer, Dafna Shahaf, Maya Leventer-Roberts, Noam Barda, Ben Y. Reis |
J. Am. Medical Informatics Assoc. | 8 |
| 2022 | A SMART on FHIR Application to Improve Suicide Risk Prediction and Management
William J. Gordon, Kate Bentley, Amy Fitzpatrick, Brian Kaney, Ronald Kessler, Matthew K. Nock, Ben Y. Reis, Mark Schechter, Vicki Strateman, Dewar Tan, Sarah Young, Jordan W. Smoller |
AMIA | 7 |
| 2021 | Prediction of patient disposition: comparison of computer and human approaches and a proposed synthesisabstractOBJECTIVE: To compare the accuracy of computer versus physician predictions of hospitalization and to explore the potential synergies of hybrid physician-computer models. MATERIALS AND METHODS: A single-center prospective observational study in a tertiary pediatric hospital in Boston, Massachusetts, United States. Nine emergency department (ED) attending physicians participated in the study. Physicians predicted the likelihood of admission for patients in the ED whose hospitalization disposition had not yet been decided. In parallel, a random-forest computer model was developed to predict hospitalizations from the ED, based on data available within the first hour of the ED encounter. The model was tested on the same cohort of patients evaluated by the participating physicians. RESULTS: 198 pediatric patients were considered for inclusion. Six patients were excluded due to incomplete or erroneous physician forms. Of the 192 included patients, 54 (28%) were admitted and 138 (72%) were discharged. The positive predictive value for the prediction of admission was 66% for the clinicians, 73% for the computer model, and 86% for a hybrid model combining the two. To predict admission, physicians relied more heavily on the clinical appearance of the patient, while the computer model relied more heavily on technical data-driven features, such as the rate of prior admissions or distance traveled to hospital. DISCUSSION: Computer-generated predictions of patient disposition were more accurate than clinician-generated predictions. A hybrid prediction model improved accuracy over both individual predictions, highlighting the complementary and synergistic effects of both approaches. CONCLUSION: The integration of computer and clinician predictions can yield improved predictive performance. Yuval Barak-Corren, Isha Agarwal, Kenneth A. Michelson, Todd W. Lyons, Mark I Neuman, Susan C. Lipsett, Amir A. Kimia, Matthew A. Eisenberg, Andrew J. Capraro, Jason A. Levy, Joel D. Hudgins, Ben Y. Reis, Andrew M. Fine |
J. Am. Medical Informatics Assoc. | 12 |
| 2021 | Temporally informed random forests for suicide risk predictionabstractOBJECTIVE: Suicide is one of the leading causes of death worldwide, yet clinicians find it difficult to reliably identify individuals at high risk for suicide. Algorithmic approaches for suicide risk detection have been developed in recent years, mostly based on data from electronic health records (EHRs). Significant room for improvement remains in the way these models take advantage of temporal information to improve predictions. MATERIALS AND METHODS: We propose a temporally enhanced variant of the random forest (RF) model-Omni-Temporal Balanced Random Forests (OT-BRFs)-that incorporates temporal information in every tree within the forest. We develop and validate this model using longitudinal EHRs and clinician notes from the Mass General Brigham Health System recorded between 1998 and 2018, and compare its performance to a baseline Naive Bayes Classifier and 2 standard versions of balanced RFs. RESULTS: Temporal variables were found to be associated with suicide risk: Elevated suicide risk was observed in individuals with a higher total number of visits as well as those with a low rate of visits over time, while lower suicide risk was observed in individuals with a longer period of EHR coverage. RF models were more accurate than Naive Bayesian classifiers at predicting suicide risk in advance (area under the receiver operating curve = 0.824 vs. 0.754, respectively). The proposed OT-BRF model performed best among all RF approaches, yielding a sensitivity of 0.339 at 95% specificity, compared to 0.290 and 0.286 for the other 2 RF models. Temporal variables were assigned high importance by the models that incorporated them. DISCUSSION: We demonstrate that temporal variables have an important role to play in suicide risk detection and that requiring their inclusion in all RF trees leads to increased predictive performance. Integrating temporal information into risk prediction models helps the models interpret patient data in temporal context, improving predictive performance. Ilkin Bayramli, Victor M. Castro, Yuval Barak-Corren, Emily M. Madsen, Matthew K. Nock, Jordan W. Smoller, Ben Y. Reis |
J. Am. Medical Informatics Assoc. | 7 |
| 2013 | Concordance and Predictive Value of Two Adverse Drug Event Data Sets
Aurel Cami, Ben Y. Reis |
AMIA | 2 |
| 2010 | Research paper: Use of population health data to refine diagnostic decision-making for pertussisabstractOBJECTIVE: To improve identification of pertussis cases by developing a decision model that incorporates recent, local, population-level disease incidence. DESIGN: Retrospective cohort analysis of 443 infants tested for pertussis (2003-7). MEASUREMENTS: Three models (based on clinical data only, local disease incidence only, and a combination of clinical data and local disease incidence) to predict pertussis positivity were created with demographic, historical, physical exam, and state-wide pertussis data. Models were compared using sensitivity, specificity, area under the receiver-operating characteristics (ROC) curve (AUC), and related metrics. RESULTS: The model using only clinical data included cyanosis, cough for 1 week, and absence of fever, and was 89% sensitive (95% CI 79 to 99), 27% specific (95% CI 22 to 32) with an area under the ROC curve of 0.80. The model using only local incidence data performed best when the proportion positive of pertussis cultures in the region exceeded 10% in the 8-14 days prior to the infant's associated visit, achieving 13% sensitivity, 53% specificity, and AUC 0.65. The combined model, built with patient-derived variables and local incidence data, included cyanosis, cough for 1 week, and the variable indicating that the proportion positive of pertussis cultures in the region exceeded 10% 8-14 days prior to the infant's associated visit. This model was 100% sensitive (p<0.04, 95% CI 92 to 100), 38% specific (p<0.001, 95% CI 33 to 43), with AUC 0.82. CONCLUSIONS: Incorporating recent, local population-level disease incidence improved the ability of a decision model to correctly identify infants with pertussis. Our findings support fostering bidirectional exchange between public health and clinical practice, and validate a method for integrating large-scale public health datasets with rich clinical data to improve decision-making and public health. Andrew M. Fine, Ben Y. Reis, Lise E. Nigrovic, Donald A. Goldmann, Tracy N. LaPorte, Karen L. Olson, Kenneth D. Mandl |
J. Am. Medical Informatics Assoc. | 2 |
| 2008 | Model Formulation: HealthMap: Global Infectious Disease Monitoring through Automated Classification and Visualization of Internet Media ReportsabstractOBJECTIVE: Unstructured electronic information sources, such as news reports, are proving to be valuable inputs for public health surveillance. However, staying abreast of current disease outbreaks requires scouring a continually growing number of disparate news sources and alert services, resulting in information overload. Our objective is to address this challenge through the HealthMap.org Web application, an automated system for querying, filtering, integrating and visualizing unstructured reports on disease outbreaks. DESIGN: This report describes the design principles, software architecture and implementation of HealthMap and discusses key challenges and future plans. MEASUREMENTS: We describe the process by which HealthMap collects and integrates outbreak data from a variety of sources, including news media (e.g., Google News), expert-curated accounts (e.g., ProMED Mail), and validated official alerts. Through the use of text processing algorithms, the system classifies alerts by location and disease and then overlays them on an interactive geographic map. We measure the accuracy of the classification algorithms based on the level of human curation necessary to correct misclassifications, and examine geographic coverage. RESULTS: As part of the evaluation of the system, we analyzed 778 reports with HealthMap, representing 87 disease categories and 89 countries. The automated classifier performed with 84% accuracy, demonstrating significant usefulness in managing the large volume of information processed by the system. Accuracy for ProMED alerts is 91% compared to Google News reports at 81%, as ProMED messages follow a more regular structure. CONCLUSION: HealthMap is a useful free and open resource employing text-processing algorithms to identify important disease outbreak information through a user-friendly interface. Clark C. Freifeld, Kenneth D. Mandl, Ben Y. Reis, John S. Brownstein |
J. Am. Medical Informatics Assoc. | 3 |
| 2007 | Research paper: Linking Surveillance to Action: Incorporation of Real-time Regional Data into a Medical Decision RuleabstractOBJECTIVE: Broadly, to create a bidirectional communication link between public health surveillance and clinical practice. Specifically, to measure the impact of integrating public health surveillance data into an existing clinical prediction rule. We incorporate data about recent local trends in meningitis epidemiology into a prediction model differentiating aseptic from bacterial meningitis. DESIGN AND MEASUREMENTS: Retrospective analysis of a cohort of all 696 children with meningitis admitted to a large urban pediatric hospital from 1992 to 2000. We modified a published bacterial meningitis score by adding a new epidemiological context adjustor variable. We examined 540 possible rules for this adjustor, varying both the number of aseptic meningitis cases that needed to be seen, and the recent time window in which they were seen. We performed sensitivity analyses with each of 540 possibilities in order to identify the optimal rule--namely, the one that included the most cases of aseptic meningitis without missing additional cases of bacterial meningitis, as compared with the published prediction model. We used bootstrap methods to validate this new score. RESULTS: The optimal rule was found to be: "at least four cases of aseptic meningitis in the previous 10 days." The epidemiological context adjustor based on surveillance of recent cases of meningitis allowed the correct identification of an additional 47 cases (7%) of aseptic meningitis without missing any additional cases of bacterial meningitis. The epidemiological context adjustor was validated, showing significance in 84% of 1,000 bootstrap samples. CONCLUSION: Epidemiological contextual information can improve the performance of a clinical prediction rule. We provide a methodological framework for leveraging regional surveillance data to improve medical decision-making. Andrew M. Fine, Lise E. Nigrovic, Ben Y. Reis, E. Francis Cook, Kenneth D. Mandl |
J. Am. Medical Informatics Assoc. | 3 |
| 2007 | Model Formulation: A Self-scaling, Distributed Information Architecture for Public Health, Research, and Clinical CareabstractOBJECTIVE: This study sought to define a scalable architecture to support the National Health Information Network (NHIN). This architecture must concurrently support a wide range of public health, research, and clinical care activities. STUDY DESIGN: The architecture fulfils five desiderata: (1) adopt a distributed approach to data storage to protect privacy, (2) enable strong institutional autonomy to engender participation, (3) provide oversight and transparency to ensure patient trust, (4) allow variable levels of access according to investigator needs and institutional policies, (5) define a self-scaling architecture that encourages voluntary regional collaborations that coalesce to form a nationwide network. RESULTS: Our model has been validated by a large-scale, multi-institution study involving seven medical centers for cancer research. It is the basis of one of four open architectures developed under funding from the Office of the National Coordinator of Health Information Technology, fulfilling the biosurveillance use case defined by the American Health Information Community. The model supports broad applicability for regional and national clinical information exchanges. CONCLUSIONS: This model shows the feasibility of an architecture wherein the requirements of care providers, investigators, and public health authorities are served by a distributed model that grants autonomy, protects privacy, and promotes participation. Andrew J. McMurry, Clint A. Gilbert, Ben Y. Reis, Henry C. Chueh, Isaac S. Kohane, Kenneth D. Mandl |
J. Am. Medical Informatics Assoc. | 3 |
| 2007 | Application of Information Technology: AEGIS: A Robust and Scalable Real-time Public Health Surveillance SystemabstractIn this report, we describe the Automated Epidemiological Geotemporal Integrated Surveillance system (AEGIS), developed for real-time population health monitoring in the state of Massachusetts. AEGIS provides public health personnel with automated near-real-time situational awareness of utilization patterns at participating healthcare institutions, supporting surveillance of bioterrorism and naturally occurring outbreaks. As real-time public health surveillance systems become integrated into regional and national surveillance initiatives, the challenges of scalability, robustness, and data security become increasingly prominent. A modular and fault tolerant design helps AEGIS achieve scalability and robustness, while a distributed storage model with local autonomy helps to minimize risk of unauthorized disclosure. The report includes a description of the evolution of the design over time in response to the challenges of a regional and national integration environment. Ben Y. Reis, Chaim Kirby, Lucy E. Hadden, Karen L. Olson, Andrew J. McMurry, James B. Daniel, Kenneth D. Mandl |
J. Am. Medical Informatics Assoc. | 1 |
| 2003 | Integrating Syndromic Surveillance Data across Multiple Locations: Effects on Outbreak Detection Performance
Ben Y. Reis, Kenneth D. Mandl |
AMIA | 1 |
| 2002 | Defining Expected Daily Emergency Department Utilization Rates for Detection of Bioterrorist Attacks
Ben Y. Reis, Kenneth D. Mandl |
AMIA | 1 |
| 2001 | Comparing the Similarity of Time-Series Gene Expression Using Signal Processing Metrics
Atul J. Butte, Ling Bao, Ben Y. Reis, Timothy W. Watkins, Isaac S. Kohane |
J. Biomed. Informatics | 3 |
| 2001 | Extracting Knowledge from Dynamics in Gene Expression
Ben Y. Reis, Atul J. Butte, Isaac S. Kohane |
J. Biomed. Informatics | 1 |
| 2001 | Reply
Ben Y. Reis, Atul J. Butte, Isaac S. Kohane |
J. Biomed. Informatics | 1 |