VLDB 2026 Research / reviewers in the wild / expert
David J. Albers
dblp:63/10358
· DBLP profile ↗
52ranked-venue papers
17as first author
20since 2021 · last 2026
0000-0002-5369-526XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 47 · 17 first-author · 18 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 since 2021Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Interdisciplinary development and application of computational methods in informatics for clinical applicationsabstractThis focus issue serves to highlight the challenging but highly valuable work of interdisciplinary teams collaborating across traditional scientific silos.1–25 There were several points of origin that motivated our choice to highlight interdisciplinary work. One point of origin that initially motivated us as guest associate editors of this focus issue was our shared personal experiences on high-impact clinical informatics projects.26–30 A second point of origin comes from talking to others who are engaged in similarly scoped efforts like the ICU Cockpit.31,32 For example, Dr Keller who led the ICU Cockpit work has spoken of similar roadblocks, shared experiences, need for time spent communicating and listening, and of how few people understand how difficult and time-consuming these efforts are—often 10-15 years from beginning to deployment. The projects we have been part of, and projects of similar scope whose leaders we have commiserated with, took years of collaborative effort from large interdisciplinary teams drawing members across the research and deployment pipelines. The collaborative clinical informatics projects that motivated this focus issue included highly engaged experts spanning a range of diverse teams of practicing clinicians, computational scientists, human–computer interaction and implementation science researchers, informaticians, and operational engineers who run the day-to-day electronic health record (EHR) systems. Interestingly, we observed that experts from diverse teams who presented components of these collaborative projects outside of their direct field of application were met with misunderstandings which manifested in dismissal of ideas, underestimation of the difficulty of another field’s problems, underestimation of the deep innovation required to translate and assemble the science, and general undervaluation of translating research to operations within the clinical informatics space. It is well known that open communication across highly distinct scientific domains is required to solve complex real-world problems and is the reason why, for instance, Oppenheimer fought so hard for open dialog between all scientists working on the Manhattan project.33 A third point of origin was our belief that clinical informatics communities could increase their engagement in interdisciplinary work and that missed opportunities for impact abound when collaborative interdisciplinary expertise is lacking. The goal of increasing dialog between distinct fields to drive increased interdisciplinary work served as a motivation for the Banff International Research Station titled: Dynamics and Data Assimilation, Physiology and Bioinformatics: Mathematics at the Interface of Theory and Clinical Applications in 2022 comprising researchers from a wide variety of interdisciplinary fields including several of the focus issue guest associate editors who were motivated to move these ideas forward through a journal focus issue. Despite some high-profile examples of interdisciplinary work, our anecdotal experience has been that examples of cross-field communication and interdisciplinary teams are hard to identify in peer-reviewed literature. We would like to see this change. David J. Albers, Kenrick Cato, Anita Layton, Sarah Collins Rossetti |
J. Am. Medical Informatics Assoc. | 1 |
| 2026 | Predicting intracranial pressure monitor placement in children with traumatic brain injury: a prospective cohort study to develop a clinical decision support toolabstractOBJECTIVE: Clinicians currently make decisions about placing an intracranial pressure (ICP) monitor in children with traumatic brain injury (TBI) without the benefit of an accurate clinical decision support tool. The goal of this study was to develop and validate a model that predicts placement of an ICP monitor and updates as new information becomes available. MATERIALS AND METHODS: A prospective observational cohort study was conducted from September 2014 to January 2024. The setting included one US hospital designated as an American College of Surgeons Level 1 Pediatric Trauma Center. Participants were 389 children with acute TBI admitted to the ICU who had at least one Glasgow Coma Scale (GCS) score ≤ 8 or intubation with at least one GCS-Motor ≤ 5. We excluded children who received ICP monitors prior to arrival, those with GCS = 3 and bilateral fixed, dilated pupils, and those with a do not resuscitate order. RESULTS: Of the 389 participants, 138 received ICP monitoring. Several machine learning models, including a recurrent neural network (RNN), were developed and validated using 4 combinations of input data. The best performing model, an RNN, achieved an F1 of 0.71 within 720 minutes of hospital arrival. The cumulative F1 of the RNN from minute 0 to 720 was 0.61. The best performing non-neural network model, standard logistic regression, achieved an F1 of 0.36 within 720 minutes of hospital arrival. CONCLUSIONS: These findings will contribute to design and implementation of a multidisciplinary clinical decision support tool for ICP monitor placement in children with TBI. Seth Russell, Peter E. Dewitt, Laura J. Helmkamp, Kathryn Colborn, Charlotte Gray, Margaret Rebull, Yamila L. Sierra, Rachel Greer, Lexi Petruccelli, Sara Shankman, Todd C. Hankinson, Fuyong Xing, David J. Albers, Tellen D. Bennett |
J. Am. Medical Informatics Assoc. | 13 |
| 2026 | Navigating the landscape of personalized oncology: overcoming challenges and expanding horizons with computational modelingabstractOBJECTIVES: We discuss challenges using computational modeling approaches for personalized prediction in clinical practice to predict treatment response for rare diseases treated by novel therapies using clinical oncology as an example context. Several challenges are discussed, including data scarcity, data sparsity, and difficulties in establishing interdisciplinary teams. Machine learning (ML), mechanistic modeling (MM), and hybrid modeling (HM) are discussed in the context of these challenges. MATERIALS AND METHODS: We present an HM approach, combining ML and MM techniques for improved personalized model estimation in the context of chimeric antigen receptor T-cell therapy for aggressive lymphoma. RESULTS: The HM approach improved the root mean squared error by 61.27±23.21% compared to using MM alone (MM: 2.36*105∓1.68*105and HM: 9.57*104∓8.37*104, where the units are in cells), computed from 13 patients included in this study. DISCUSSION: By exploiting the complementary strengths of ML and MM approaches, the developed HM method addresses common limitations such as data scarcity and sparsity in medical settings, especially common for rare diseases. CONCLUSION: The HM techniques are likely required to overcome data scarcity and sparsity issues in broad medical settings. Developing these techniques requires dedicated interdisciplinary teams. Melike Sirlanci, David J. Albers, Jennifer K. Briggs, Clayton Smith, Tellen D. Bennett, Steven M. Bair |
J. Am. Medical Informatics Assoc. | 2 |
| 2025 | T2 Coach: A Qualitative Study of an Automated Health Coach for Diabetes Self-ManagementabstractComputational intelligence is increasingly common in interactive systems in many domains, including health. Health coaching with conversational agents (CA) can reach wide populations, but the level of computational intelligence needed for a positive coaching experience is unclear. We conducted a study with sixteen individuals with diabetes and prediabetes who used a CA for health coaching, T2 Coach. Qualitative interviews revealed that participants saw T2 Coach as reliable in helping them stay on track with self-management, appreciated the flexibility in choosing personally meaningful goals and engaging on their own terms, and felt it provided encouragement and even compared it favorably with human coaches. However, they also noted that coaching experience could be improved with more fluid conversations, more tailoring to their personal preferences and lifestyles, and more sensitivity to specific contexts, all of which require more computational intelligence. We discuss implications and design directions for more intelligent coaching CA in health. Elliot G. Mitchell, Pooja M. Desai, Arlene M. Smaldone, Andrea Cassells, Jonathan N. Tobin, David J. Albers, Matthew E. Levine, Lena Mamykina |
CHI | 6 |
| 2023 | Interpretable physiological forecasting in the ICU using constrained data assimilation and electronic health record dataabstractOBJECTIVE: Prediction of physiological mechanics are important in medical practice because interventions are guided by predicted impacts of interventions. But prediction is difficult in medicine because medicine is complex and difficult to understand from data alone, and the data are sparse relative to the complexity of the generating processes. Computational methods can increase prediction accuracy, but prediction with clinical data is difficult because the data are sparse, noisy and nonstationary. This paper focuses on predicting physiological processes given sparse, non-stationary, electronic health record data in the intensive care unit using data assimilation (DA), a broad collection of methods that pair mechanistic models with inference methods. METHODS: A methodological pipeline embedding a glucose-insulin model into a new DA framework, the constrained ensemble Kalman filter (CEnKF) to forecast blood glucose was developed. The data include tube-fed patients whose nutrition, blood glucose, administered insulins and medications were extracted by hand due to their complexity and to ensure accuracy. The model was estimated using an individual's data as if they arrived in real-time, and the estimated model was run forward producing a forecast. Both constrained and unconstrained ensemble Kalman filters were estimated to compare the impact of constraints. Constraint boundaries, model parameter sets estimated, and data used to estimate the models were varied to investigate their influence on forecasting accuracy. Forecasting accuracy was evaluated according to mean squared error between the model-forecasted glucose and the measurements and by comparing distributions of measured glucose and forecast ensemble means. RESULTS: The novel CEnKF produced substantial gains in robustness and accuracy while minimizing the data requirements compared to the unconstrained ensemble Kalman filters. Administered insulin and tube-nutrition were important for accurate forecasting, but including glucose in IV medication delivery did not increase forecast accuracy. Model flexibility, controlled by constraint boundaries and estimated parameters, did influence forecasting accuracy. CONCLUSION: Accurate and robust physiological forecasting with sparse clinical data is possible with DA. Introducing constrained inference, particularly on unmeasured states and parameters, reduced forecast error and data requirements. The results are not particularly sensitive to model flexibility such as constraint boundaries, but over or under constraining increased forecasting errors. David J. Albers, Melike Sirlanci, Matthew E. Levine, Jan Claassen, Caroline Der Nigoghossian, George Hripcsak |
J. Biomed. Informatics | 1 |
| 2023 | Who needs what (features) when? Personalizing engagement with data-driven self-management to improve health equity
Marissa Burgermaster, Pooja M. Desai, Elizabeth M. Heitkemper, Filippa Juul, Elliot G. Mitchell, Meghan Reading Turchioe, David J. Albers, Matthew E. Levine, Dagny Larson, Lena Mamykina |
J. Biomed. Informatics | 7 |
| 2023 | Hypothesis-driven modeling of the human lung-ventilator system: A characterization tool for Acute Respiratory Distress Syndrome research
J. N. Stroh, Bradford J. Smith, Peter D. Sottile, George Hripcsak, David J. Albers |
J. Biomed. Informatics | 5 |
| 2023 | A methodology of phenotyping ICU patients from EHR data: High-fidelity, personalized, and interpretable phenotypes estimationabstractOBJECTIVE: Computing phenotypes that provide high-fidelity, time-dependent characterizations and yield personalized interpretations is challenging, especially given the complexity of physiological and healthcare systems and clinical data quality. This paper develops a methodological pipeline to estimate unmeasured physiological parameters and produce high-fidelity, personalized phenotypes anchored to physiological mechanics from electronic health record (EHR). METHODS: A methodological phenotyping pipeline is developed that computes new phenotypes defined with unmeasurable computational biomarkers quantifying specific physiological properties in real time. Working within the inverse problem framework, this pipeline is applied to the glucose-insulin system for ICU patients using data assimilation to estimate an established mathematical physiological model with stochastic optimization. This produces physiological model parameter vectors of clinically unmeasured endocrine properties, here insulin secretion, clearance, and resistance, estimated for individual patient. These physiological parameter vectors are used as inputs to unsupervised machine learning methods to produce phenotypic labels and discrete physiological phenotypes. These phenotypes are inherently interpretable because they are based on parametric physiological descriptors. To establish potential clinical utility, the computed phenotypes are evaluated with external EHR data for consistency and reliability and with clinician face validation. RESULTS: The phenotype computation was performed on a cohort of 109 ICU patients who received no or short-acting insulin therapy, rendering continuous and discrete physiological phenotypes as specific computational biomarkers of unmeasured insulin secretion, clearance, and resistance on time windows of three days. Six, six, and five discrete phenotypes were found in the first, middle, and last three-day periods of ICU stays, respectively. Computed phenotypic labels were predictive with an average accuracy of 89%. External validation of discrete phenotypes showed coherence and consistency in clinically observable differences based on laboratory measurements and ICD 9/10 codes and clinical concordance from face validity. A particularly clinically impactful parameter, insulin secretion, had a concordance accuracy of 83%±27%. CONCLUSION: The new physiological phenotypes computed with individual patient ICU data and defined by estimates of mechanistic model parameters have high physiological fidelity, are continuous, time-specific, personalized, interpretable, and predictive. This methodology is generalizable to other clinical and physiological settings and opens the door for discovering deeper physiological information to personalize medical care. J. N. Stroh, George Hripcsak, Cecilia C. Low Wang, Tellen D. Bennett, Julia Wrobel, Caroline Der Nigoghossian, Scott W. Mueller, Jan Claassen, David J. Albers |
J. Biomed. Informatics | 10 |
| 2022 | Optimizing Strategies of Pressure Reactivity Index and Optimal Cerebral Perfusion Pressure Identification for Cerebral Autoregulatory-Guided Clinical Decision Support
Jennifer K. Briggs, J. N. Stroh, Tellen D. Bennett, Soojin Park, David J. Albers, Brandon Foreman |
AMIA | 5 |
| 2022 | Informatics Research and Implementation During COVID: Challenges, Opportunities and Recommendations for Building a Sustainable Infrastructure
Patricia C. Dykes, Sarah Collins Rossetti, Patricia Sengstack, Guilherme Del Fiol, David J. Albers |
AMIA | 5 |
| 2022 | Using Data Assimilation to Predict Post-Operative Bariatric Surgery Glycemic Status in Adolescents
Lauren R. Richter, Benjamin Albert, Linying Zhang, Ilene Fennoy, David J. Albers, George Hripcsak |
AMIA | 5 |
| 2022 | Gaining Purchase on Ventilator-Induced Lung Injury: A Interpretable Approach to Describing Complex System Data via Informed Modeling
J. N. Stroh, Bradford J. Smith, Peter D. Sottile, George Hripcsak, David J. Albers |
AMIA | 5 |
| 2022 | A methodology of phenotyping ICU patients: high-fidelity, personalized, and interpretable phenotypes estimation
J. N. Stroh, George Hripcsak, Cecilia C. Low Wang, Julia Wrobel, Caroline Der Nigoghossian, Tellen D. Bennett, David J. Albers |
AMIA | 8 |
| 2021 | Toward phenotyping of ventilator-induced lung injury with a damage-informed pulmonary model of lung-ventilator interaction
David J. Albers, Deepak K. Agrawal, Bradford J. Smith, Peter D. Sottile, Tellen D. Bennett, J. N. Stroh, George Hripcsak |
AMIA | 1 |
| 2021 | From Reflection to Action: Combining Machine Learning with Expert Knowledge for Nutrition Goal RecommendationsabstractSelf-tracking can help personalize self-management interventions for chronic conditions like type 2 diabetes (T2D), but reflecting on personal data requires motivation and literacy. Machine learning (ML) methods can identify patterns, but a key challenge is making actionable suggestions based on personal health data. We introduce GlucoGoalie, which combines ML with an expert system to translate ML output into personalized nutrition goal suggestions for individuals with T2D. In a controlled experiment, participants with T2D found that goal suggestions were understandable and actionable. A 4-week in-the-wild deployment study showed that receiving goal suggestions augmented participants' self-discovery, choosing goals highlighted the multifaceted nature of personal preferences, and the experience of following goals demonstrated the importance of feedback and context. However, we identified tensions between abstract goals and concrete eating experiences and found static text too ambiguous for complex concepts. We discuss implications for ML-based interventions and the need for systems that offer more interactivity, feedback, and negotiation. Elliot G. Mitchell, Elizabeth M. Heitkemper, Marissa Burgermaster, Matthew E. Levine, Yishen Miao, Maria L. Hwang, Pooja M. Desai, Andrea Cassells, Jonathan N. Tobin, Esteban G. Tabak, David J. Albers, Arlene M. Smaldone, Lena Mamykina |
CHI | 11 |
| 2021 | Utilizing timestamps of longitudinal electronic health record data to classify clinical deterioration eventsabstractOBJECTIVE: To propose an algorithm that utilizes only timestamps of longitudinal electronic health record data to classify clinical deterioration events. MATERIALS AND METHODS: This retrospective study explores the efficacy of machine learning algorithms in classifying clinical deterioration events among patients in intensive care units using sequences of timestamps of vital sign measurements, flowsheets comments, order entries, and nursing notes. We design a data pipeline to partition events into discrete, regular time bins that we refer to as timesteps. Logistic regressions, random forest classifiers, and recurrent neural networks are trained on datasets of different length of timesteps, respectively, against a composite outcome of death, cardiac arrest, and Rapid Response Team calls. Then these models are validated on a holdout dataset. RESULTS: A total of 6720 intensive care unit encounters meet the criteria and the final dataset includes 830 578 timestamps. The gated recurrent unit model utilizes timestamps of vital signs, order entries, flowsheet comments, and nursing notes to achieve the best performance on the time-to-outcome dataset, with an area under the precision-recall curve of 0.101 (0.06, 0.137), a sensitivity of 0.443, and a positive predictive value of 0. 092 at the threshold of 0.6. DISCUSSION AND CONCLUSION: This study demonstrates that our recurrent neural network models using only timestamps of longitudinal electronic health record data that reflect healthcare processes achieve well-performing discriminative power. Li-heng Fu, Christopher Knaplund, Kenrick Cato, Adler J. Perotte, Min-Jeoung Kang, Patricia C. Dykes, David J. Albers, Sarah Collins Rossetti |
J. Am. Medical Informatics Assoc. | 7 |
| 2021 | Healthcare Process Modeling to Phenotype Clinician Behaviors for Exploiting the Signal Gain of Clinical Expertise (HPM-ExpertSignals): Development and evaluation of a conceptual frameworkabstractOBJECTIVE: There are signals of clinicians' expert and knowledge-driven behaviors within clinical information systems (CIS) that can be exploited to support clinical prediction. Describe development of the Healthcare Process Modeling Framework to Phenotype Clinician Behaviors for Exploiting the Signal Gain of Clinical Expertise (HPM-ExpertSignals). MATERIALS AND METHODS: We employed an iterative framework development approach that combined data-driven modeling and simulation testing to define and refine a process for phenotyping clinician behaviors. Our framework was developed and evaluated based on the Communicating Narrative Concerns Entered by Registered Nurses (CONCERN) predictive model to detect and leverage signals of clinician expertise for prediction of patient trajectories. RESULTS: Seven themes-identified during development and simulation testing of the CONCERN model-informed framework development. The HPM-ExpertSignals conceptual framework includes a 3-step modeling technique: (1) identify patterns of clinical behaviors from user interaction with CIS; (2) interpret patterns as proxies of an individual's decisions, knowledge, and expertise; and (3) use patterns in predictive models for associations with outcomes. The CONCERN model differentiated at risk patients earlier than other early warning scores, lending confidence to the HPM-ExpertSignals framework. DISCUSSION: The HPM-ExpertSignals framework moves beyond transactional data analytics to model clinical knowledge, decision making, and CIS interactions, which can support predictive modeling with a focus on the rapid and frequent patient surveillance cycle. CONCLUSIONS: We propose this framework as an approach to embed clinicians' knowledge-driven behaviors in predictions and inferences to facilitate capture of healthcare processes that are activated independently, and sometimes well before, physiological changes are apparent. Sarah Collins Rossetti, Christopher Knaplund, David J. Albers, Patricia C. Dykes, Min-Jeoung Kang, Zfania Tom Korach, Li Zhou 0007, Kumiko Schnock, Jose P. Garcia, Jessica Schwartz-Dillard, Li-heng Fu, Jeffrey G. Klann, Graham Lowenthal, Kenrick Cato |
J. Am. Medical Informatics Assoc. | 3 |
| 2021 | Real-time electronic health record mortality prediction during the COVID-19 pandemic: a prospective cohort studyabstractOBJECTIVE: To rapidly develop, validate, and implement a novel real-time mortality score for the COVID-19 pandemic that improves upon sequential organ failure assessment (SOFA) for decision support for a Crisis Standards of Care team. MATERIALS AND METHODS: We developed, verified, and deployed a stacked generalization model to predict mortality using data available in the electronic health record (EHR) by combining 5 previously validated scores and additional novel variables reported to be associated with COVID-19-specific mortality. We verified the model with prospectively collected data from 12 hospitals in Colorado between March 2020 and July 2020. We compared the area under the receiver operator curve (AUROC) for the new model to the SOFA score and the Charlson Comorbidity Index. RESULTS: The prospective cohort included 27 296 encounters, of which 1358 (5.0%) were positive for SARS-CoV-2, 4494 (16.5%) required intensive care unit care, 1480 (5.4%) required mechanical ventilation, and 717 (2.6%) ended in death. The Charlson Comorbidity Index and SOFA scores predicted mortality with an AUROC of 0.72 and 0.90, respectively. Our novel score predicted mortality with AUROC 0.94. In the subset of patients with COVID-19, the stacked model predicted mortality with AUROC 0.90, whereas SOFA had AUROC of 0.85. DISCUSSION: Stacked regression allows a flexible, updatable, live-implementable, ethically defensible predictive analytics tool for decision support that begins with validated models and includes only novel information that improves prediction. CONCLUSION: We developed and validated an accurate in-hospital mortality prediction score in a live EHR for automatic and continuous calculation using a novel model that improved upon SOFA. Peter D. Sottile, David J. Albers, Peter E. Dewitt, Seth Russell, J. N. Stroh, David P. Kao, Bonnie Adrian, Matthew E. Levine, Ryan Mooney, Lenny Larchick, Jean S. Kutner, Matthew K. Wynia, Jeffrey J. Glasheen, Tellen D. Bennett |
J. Am. Medical Informatics Assoc. | 2 |
| 2021 | Enabling personalized decision support with patient-generated data and attributable components
Elliot G. Mitchell, Esteban G. Tabak, Matthew E. Levine, Lena Mamykina, David J. Albers |
J. Biomed. Informatics | 5 |
| 2021 | Correction: Personalized glucose forecasting for type 2 diabetes using data assimilationabstract[This corrects the article DOI: 10.1371/journal.pcbi.1005232.]. David J. Albers, Matthew E. Levine, Bruce J. Gluckman, Henry N. Ginsberg, George Hripcsak, Lena Mamykina |
PLoS Comput. Biol. | 1 |
| 2020 | Lessons learned from assimilating knowledge into machine learning to forecast and control glucose in a critical care setting
David J. Albers, Melike Sirlanci Tuysuzoglu, Matthew E. Levine, Caroline Der Nigoghossian, Andrew M. Stuart, Jan Claassen, Bruce J. Gluckman, George Hripcsak |
AMIA | 1 |
| 2020 | Utilizing Timestamps of Longitudinal Data from Electronic Health Record to Predict Clinical Deterioration Events
Li-heng Fu, Christopher Knaplund, Kenrick Cato, David J. Albers, Sarah Collins Rossetti |
AMIA | 4 |
| 2020 | Development and validation of early warning score system: A systematic literature review
Li-heng Fu, Jessica Schwartz-Dillard, Amanda J. Moy, Christopher Knaplund, Min-Jeoung Kang, Kumiko Schnock, Jose P. Garcia, Haomiao Jia, Patricia C. Dykes, Kenrick Cato, David J. Albers, Sarah Collins Rossetti |
J. Biomed. Informatics | 11 |
| 2019 | Feasibility of a machine learning based method to generate personalized nutrition goals for diabetes self-management
Elliot G. Mitchell, Marissa Burgermaster, Elizabeth M. Heitkemper, Matthew E. Levine, Yishen Miao, Esteban G. Tabak, Arlene M. Smaldone, David J. Albers, Lena Mamykina |
AMIA | 8 |
| 2019 | Machine learning for personalized decision support with patient-generated health data
Elliot G. Mitchell, Lena Mamykina, Matthew E. Levine, Esteban G. Tabak, David J. Albers |
AMIA | 5 |
| 2019 | Leveraging Clinical Expertise as a Feature - not an Outcome - of Predictive Models: Evaluation of an Early Warning System Use Case
Sarah Collins Rossetti, Christopher Knaplund, David J. Albers, Abdul A. Tariq, Kui Tang, David K. Vawdrey, Natalie Yip, Patricia C. Dykes, Jeffrey G. Klann, Min-Jeoung Kang, Jose P. Garcia, Li-heng Fu, Kumiko Schnock, Kenrick Cato |
AMIA | 3 |
| 2019 | Personal Health Oracle: Explorations of Personalized Predictions in Diabetes Self-ManagementabstractThe increasing availability of health data and knowledge about computationally modeling human physiology opens new opportunities for personalized predictions in health. Yet little is known about how individuals interact and reason with personalized predictions. To explore these questions, we developed a smartphone app, GlucOracle, that uses self-tracking data of individuals with type 2 diabetes to generate personalized forecasts for post-meal blood glucose levels. We pilot-tested GlucOracle with two populations: members of an online diabetes community, knowledgeable about diabetes and technologically savvy; and individuals from a low socio-economic status community, characterized by high prevalence of diabetes, low literacy and limited experience with mobile apps. Individuals in both communities engaged with personal glucose forecasts and found them useful for adjusting immediate meal options, and planning future meals. However, the study raised new questions as to appropriate time, form, and focus of forecasts and suggested new research directions for personalized predictions in health. Pooja M. Desai, Elliot G. Mitchell, Maria L. Hwang, Matthew E. Levine, David J. Albers, Lena Mamykina |
CHI | 5 |
| 2019 | Data-driven modeling and prediction of blood glucose dynamics: Machine learning applications in type 1 diabetesabstractBACKGROUND: Diabetes mellitus (DM) is a metabolic disorder that causes abnormal blood glucose (BG) regulation that might result in short and long-term health complications and even death if not properly managed. Currently, there is no cure for diabetes. However, self-management of the disease, especially keeping BG in the recommended range, is central to the treatment. This includes actively tracking BG levels and managing physical activity, diet, and insulin intake. The recent advancements in diabetes technologies and self-management applications have made it easier for patients to have more access to relevant data. In this regard, the development of an artificial pancreas (a closed-loop system), personalized decision systems, and BG event alarms are becoming more apparent than ever. Techniques such as predicting BG (modeling of a personalized profile), and modeling BG dynamics are central to the development of these diabetes management technologies. The increased availability of sufficient patient historical data has paved the way for the introduction of machine learning and its application for intelligent and improved systems for diabetes management. The capability of machine learning to solve complex tasks with dynamic environment and knowledge has contributed to its success in diabetes research. MOTIVATION: Recently, machine learning and data mining have become popular, with their expanding application in diabetes research and within BG prediction services in particular. Despite the increasing and expanding popularity of machine learning applications in BG prediction services, updated reviews that map and materialize the current trends in modeling options and strategies are lacking within the context of BG prediction (modeling of personalized profile) in type 1 diabetes. OBJECTIVE: The objective of this review is to develop a compact guide regarding modeling options and strategies of machine learning and a hybrid system focusing on the prediction of BG dynamics in type 1 diabetes. The review covers machine learning approaches pertinent to the controller of an artificial pancreas (closed-loop systems), modeling of personalized profiles, personalized decision support systems, and BG alarm event applications. Generally, the review will identify, assess, analyze, and discuss the current trends of machine learning applications within these contexts. METHOD: A rigorous literature review was conducted between August 2017 and February 2018 through various online databases, including Google Scholar, PubMed, ScienceDirect, and others. Additionally, peer-reviewed journals and articles were considered. Relevant studies were first identified by reviewing the title, keywords, and abstracts as preliminary filters with our selection criteria, and then we reviewed the full texts of the articles that were found relevant. Information from the selected literature was extracted based on predefined categories, which were based on previous research and further elaborated through brainstorming among the authors. RESULTS: The initial search was done by analyzing the title, abstract, and keywords. A total of 624 papers were retrieved from DBLP Computer Science (25), Diabetes Technology and Therapeutics (31), Google Scholar (193), IEEE (267), Journal of Diabetes Science and Technology (31), PubMed/Medline (27), and ScienceDirect (50). After removing duplicates from the list, 417 records remained. Then, we independently assessed and screened the articles based on the inclusion and exclusion criteria, which eliminated another 204 papers, leaving 213 relevant papers. After a full-text assessment, 55 articles were left, which were critically analyzed. The inter-rater agreement was measured using a Cohen Kappa test, and disagreements were resolved through discussion. CONCLUSION: Due to the complexity of BG dynamics, it remains difficult to achieve a universal model that produces an accurate prediction in every circumstance (i.e., hypo/eu/hyperglycemia events). Recently, machine learning techniques have received wider attention and increased popularity in diabetes research in general and BG prediction in particular, coupled with the ever-growing availability of a self-collected health data. The state-of-the-art demonstrates that various machine learning techniques have been tested to predict BG, such as recurrent neural networks, feed-forward neural networks, support vector machines, self-organizing maps, the Gaussian process, genetic algorithm and programs, deep neural networks, and others, using various group of input parameters and training algorithms. The main limitation of the current approaches is the lack of a well-defined approach to estimate carbohydrate intake, which is mainly done manually by individual users and is prone to an error that can severely affect the predictive performance. Moreover, a universal approach has not been established to estimate and quantify the approximate effect of physical activities, stress, and infections on the BG level. No researchers have assessed model predictive performance during stress and infection incidences in a free-living condition, which should be considered in future studies. Furthermore, a little has been done regarding model portability that can capture inter- and intra-variability among patients. It seems that the effect of time lags between the CGM readings and the actual BG levels is not well covered. However, in general, we foresee that these developments might foster the advancement of next-generation BG prediction algorithms, which will make a great contribution in the effort to develop the long-awaited, so-called artificial pancreas (a closed-loop system). Ashenafi Zebene Woldaregay, Eirik Årsand, Ståle Walderhaug, David J. Albers, Lena Mamykina, Taxiarchis Botsis, Gunnar Hartvigsen |
Artif. Intell. Medicine | 4 |
| 2018 | Using mechanistic machine learning to forecast glucose and infer physiologic phenotypes in the ICU: what is possible and what are the challenges
David J. Albers, Matthew E. Levine, Andrew M. Stuart, Jan Claassen, Bruce J. Gluckman, George Hripcsak |
AMIA | 1 |
| 2018 | Pictures Worth a Thousand Words: Reflections on Visualizing Personal Blood Glucose Forecasts for Individuals with Type 2 DiabetesabstractType 2 Diabetes Mellitus (T2DM) is a common chronic condition that requires management of one's lifestyle, including nutrition. Critically, patients often lack a clear understanding of how everyday meals impact their blood glucose. New predictive analytics approaches can provide personalized mealtime blood glucose forecasts. While communicating forecasts can be challenging, effective strategies for doing so remain little explored. In this study, we conducted focus groups with 13 participants to identify approaches to visualizing personalized blood glucose forecasts that can promote diabetes self-management and understand key styles and visual features that resonate with individuals with diabetes. Focus groups demonstrated that individuals rely on simple heuristics and tend to take a reactive approach to their health and nutrition management. Further, the study highlighted the need for simple and explicit, yet information-rich design. Effective visualizations were found to utilize common metaphors alongside words, numbers, and colors to convey a sense of authority and encourage action and learning. Pooja M. Desai, Matthew E. Levine, David J. Albers, Lena Mamykina |
CHI | 3 |
| 2018 | Mechanistic machine learning: how data assimilation leverages physiologic knowledge using Bayesian inference to forecast the future, infer the present, and phenotypeabstractWe introduce data assimilation as a computational method that uses machine learning to combine data with human knowledge in the form of mechanistic models in order to forecast future states, to impute missing data from the past by smoothing, and to infer measurable and unmeasurable quantities that represent clinically and scientifically important phenotypes. We demonstrate the advantages it affords in the context of type 2 diabetes by showing how data assimilation can be used to forecast future glucose values, to impute previously missing glucose values, and to infer type 2 diabetes phenotypes. At the heart of data assimilation is the mechanistic model, here an endocrine model. Such models can vary in complexity, contain testable hypotheses about important mechanics that govern the system (eg, nutrition's effect on glucose), and, as such, constrain the model space, allowing for accurate estimation using very little data. David J. Albers, Matthew E. Levine, Andrew M. Stuart, Lena Mamykina, Bruce J. Gluckman, George Hripcsak |
J. Am. Medical Informatics Assoc. | 1 |
| 2018 | A visual analytics approach for pattern-recognition in patient-generated dataabstractObjective: To develop and test a visual analytics tool to help clinicians identify systematic and clinically meaningful patterns in patient-generated data (PGD) while decreasing perceived information overload. Methods: Participatory design was used to develop Glucolyzer, an interactive tool featuring hierarchical clustering and a heatmap visualization to help registered dietitians (RDs) identify associative patterns between blood glucose levels and per-meal macronutrient composition for individuals with type 2 diabetes (T2DM). Ten RDs participated in a within-subjects experiment to compare Glucolyzer to a static logbook format. For each representation, participants had 25 minutes to examine 1 month of diabetes self-monitoring data captured by an individual with T2DM and identify clinically meaningful patterns. We compared the quality and accuracy of the observations generated using each representation. Results: Participants generated 50% more observations when using Glucolyzer (98) than when using the logbook format (64) without any loss in accuracy (69% accuracy vs 62%, respectively, p = .17). Participants identified more observations that included ingredients other than carbohydrates using Glucolyzer (36% vs 16%, p = .027). Fewer RDs reported feelings of information overload using Glucolyzer compared to the logbook format. Study participants displayed variable acceptance of hierarchical clustering. Conclusions: Visual analytics have the potential to mitigate provider concerns about the volume of self-monitoring data. Glucolyzer helped dietitians identify meaningful patterns in self-monitoring data without incurring perceived information overload. Future studies should assess whether similar tools can support clinicians in personalizing behavioral interventions that improve patient outcomes. Daniel J. Feller, Marissa Burgermaster, Matthew E. Levine, Arlene M. Smaldone, Patricia G. Davidson, David J. Albers, Lena Mamykina |
J. Am. Medical Informatics Assoc. | 6 |
| 2018 | High-fidelity phenotyping: richness and freedom from biasabstractElectronic health record phenotyping is the use of raw electronic health record data to assert characterizations about patients. Researchers have been doing it since the beginning of biomedical informatics, under different names. Phenotyping will benefit from an increasing focus on fidelity, both in the sense of increasing richness, such as measured levels, degree or severity, timing, probability, or conceptual relationships, and in the sense of reducing bias. Research agendas should shift from merely improving binary assignment to studying and improving richer representations. The field is actively researching new temporal directions and abstract representations, including deep learning. The field would benefit from research in nonlinear dynamics, in combining mechanistic models with empirical data, including data assimilation, and in topology. The health care process produces substantial bias, and studying that bias explicitly rather than treating it as merely another source of noise would facilitate addressing it. George Hripcsak, David J. Albers |
J. Am. Medical Informatics Assoc. | 2 |
| 2018 | Estimating summary statistics for electronic health record laboratory data for use in high-throughput phenotyping algorithmsabstractWe study the question of how to represent or summarize raw laboratory data taken from an electronic health record (EHR) using parametric model selection to reduce or cope with biases induced through clinical care. It has been previously demonstrated that the health care process (Hripcsak and Albers, 2012, 2013), as defined by measurement context (Hripcsak and Albers, 2013; Albers et al., 2012) and measurement patterns (Albers and Hripcsak, 2010, 2012), can influence how EHR data are distributed statistically (Kohane and Weber, 2013; Pivovarov et al., 2014). We construct an algorithm, PopKLD, which is based on information criterion model selection (Burnham and Anderson, 2002; Claeskens and Hjort, 2008), is intended to reduce and cope with health care process biases and to produce an intuitively understandable continuous summary. The PopKLD algorithm can be automated and is designed to be applicable in high-throughput settings; for example, the output of the PopKLD algorithm can be used as input for phenotyping algorithms. Moreover, we develop the PopKLD-CAT algorithm that transforms the continuous PopKLD summary into a categorical summary useful for applications that require categorical data such as topic modeling. We evaluate our methodology in two ways. First, we apply the method to laboratory data collected in two different health care contexts, primary versus intensive care. We show that the PopKLD preserves known physiologic features in the data that are lost when summarizing the data using more common laboratory data summaries such as mean and standard deviation. Second, for three disease-laboratory measurement pairs, we perform a phenotyping task: we use the PopKLD and PopKLD-CAT algorithms to define high and low values of the laboratory variable that are used for defining a disease state. We then compare the relationship between the PopKLD-CAT summary disease predictions and the same predictions using empirically estimated mean and standard deviation to a gold standard generated by clinical review of patient records. We find that the PopKLD laboratory data summary is substantially better at predicting disease state. The PopKLD or PopKLD-CAT algorithms are not meant to be used as phenotyping algorithms, but we use the phenotyping task to show what information can be gained when using a more informative laboratory data summary. In the process of evaluation our method we show that the different clinical contexts and laboratory measurements necessitate different statistical summaries. Similarly, leveraging the principle of maximum entropy we argue that while some laboratory data only have sufficient information to estimate a mean and standard deviation, other laboratory data captured in an EHR contain substantially more information than can be captured in higher-parameter models. David J. Albers, Noémie Elhadad, Jan Claassen, Rimma Perotte, Andrew Goldstein, George Hripcsak |
J. Biomed. Informatics | 1 |
| 2018 | Methodological variations in lagged regression for detecting physiologic drug effects in EHR data
Matthew E. Levine, David J. Albers, George Hripcsak |
J. Biomed. Informatics | 2 |
| 2017 | Why predicting postprandial glucose using self-monitoring data is difficult
David J. Albers, Matthew E. Levine, Andrew M. Stuart, Bruce J. Gluckman, George Hripcsak |
AMIA | 1 |
| 2017 | Reflecting on Diabetes Self-Management Logs with Simulated, Continuous Blood Glucose Curves: A Pilot Study
Elliot G. Mitchell, Matthew E. Levine, David J. Albers, Lena Mamykina |
AMIA | 3 |
| 2017 | Personalized glucose forecasting for type 2 diabetes using data assimilationabstractType 2 diabetes leads to premature death and reduced quality of life for 8% of Americans. Nutrition management is critical to maintaining glycemic control, yet it is difficult to achieve due to the high individual differences in glycemic response to nutrition. Anticipating glycemic impact of different meals can be challenging not only for individuals with diabetes, but also for expert diabetes educators. Personalized computational models that can accurately forecast an impact of a given meal on an individual's blood glucose levels can serve as the engine for a new generation of decision support tools for individuals with diabetes. However, to be useful in practice, these computational engines need to generate accurate forecasts based on limited datasets consistent with typical self-monitoring practices of individuals with type 2 diabetes. This paper uses three forecasting machines: (i) data assimilation, a technique borrowed from atmospheric physics and engineering that uses Bayesian modeling to infuse data with human knowledge represented in a mechanistic model, to generate real-time, personalized, adaptable glucose forecasts; (ii) model averaging of data assimilation output; and (iii) dynamical Gaussian process model regression. The proposed data assimilation machine, the primary focus of the paper, uses a modified dual unscented Kalman filter to estimate states and parameters, personalizing the mechanistic models. Model selection is used to make a personalized model selection for the individual and their measurement characteristics. The data assimilation forecasts are empirically evaluated against actual postprandial glucose measurements captured by individuals with type 2 diabetes, and against predictions generated by experienced diabetes educators after reviewing a set of historical nutritional records and glucose measurements for the same individual. The evaluation suggests that the data assimilation forecasts compare well with specific glucose measurements and match or exceed in accuracy expert forecasts. We conclude by examining ways to present predictions as forecast-derived range quantities and evaluate the comparative advantages of these ranges. David J. Albers, Matthew E. Levine, Bruce J. Gluckman, Henry N. Ginsberg, George Hripcsak, Lena Mamykina |
PLoS Comput. Biol. | 1 |
| 2016 | Using data assimilation to forecast post-meal glucose for patients with type 2 diabetes
David J. Albers, Matthew E. Levine, Andrew M. Stuart, George Hripcsak, Lena Mamykina |
AMIA | 1 |
| 2016 | Approaches for using temporal and other filters for next generation phenotype discovery
David J. Albers, Adler J. Perotte, George Hripcsak |
AMIA | 1 |
| 2016 | Comparing Lagged Linear Correlation, Lagged Regression, Granger Causality, and Vector Autoregression for Uncovering Associations in EHR Data
Matthew E. Levine, David J. Albers, George Hripcsak |
AMIA | 2 |
| 2016 | Data-driven health management: reasoning about personally generated data in diabetes with information technologiesabstractOBJECTIVE: To investigate how individuals with diabetes and diabetes educators reason about data collected through self-monitoring and to draw implications for the design of data-driven self-management technologies. MATERIALS AND METHODS: Ten individuals with diabetes (six type 1 and four type 2) and 2 experienced diabetes educators were presented with a set of self-monitoring data captured by an individual with type 2 diabetes. The set included digital images of meals and their textual descriptions, and blood glucose (BG) readings captured before and after these meals. The participants were asked to review a set of meals and associated BG readings, explain differences in postprandial BG levels for these meals, and predict postprandial BG levels for the same individual for a different set of meals. Researchers compared conclusions and predictions reached by the participants with those arrived at by quantitative analysis of the collected data. RESULTS: The participants used both macronutrient composition of meals, most notably the inclusion of carbohydrates, and names of dishes and ingredients to reason about changes in postprandial BG levels. Both individuals with diabetes and diabetes educators reported difficulties in generating predictions of postprandial BG; their predictions varied in their correlations with the actual captured readings from r = 0.008 to r = 0.75. CONCLUSION: Overall, the study showed that identifying trends in the data collected with self-monitoring is a complex process, and that conclusions reached by both individuals with diabetes and diabetes educators are not always reliable. This suggests the need for new ways to facilitate individuals' reasoning with informatics interventions. Lena Mamykina, Matthew E. Levine, Patricia G. Davidson, Arlene M. Smaldone, Noémie Elhadad, David J. Albers |
J. Am. Medical Informatics Assoc. | 6 |
| 2015 | Personalized medicine beyond genetics: using personalized model-based forecasting to help type 2 diabetics understand and predict their post-meal glucose
David J. Albers, Matthew E. Levine, Bruce J. Gluckman, George Hripcsak, Lena Mamykina |
AMIA | 1 |
| 2015 | Model Selection For EHR Laboratory Tests Preserving Healthcare Context and Underlying Physiology
David J. Albers, Rimma Perotte, J. Michael Schmidt, Noémie Elhadad, George Hripcsak |
AMIA | 1 |
| 2015 | Parameterizing time in electronic health record studiesabstractBACKGROUND: Fields like nonlinear physics offer methods for analyzing time series, but many methods require that the time series be stationary-no change in properties over time.Objective Medicine is far from stationary, but the challenge may be able to be ameliorated by reparameterizing time because clinicians tend to measure patients more frequently when they are ill and are more likely to vary. METHODS: We compared time parameterizations, measuring variability of rate of change and magnitude of change, and looking for homogeneity of bins of temporal separation between pairs of time points. We studied four common laboratory tests drawn from 25 years of electronic health records on 4 million patients. RESULTS: We found that sequence time-that is, simply counting the number of measurements from some start-produced more stationary time series, better explained the variation in values, and had more homogeneous bins than either traditional clock time or a recently proposed intermediate parameterization. Sequence time produced more accurate predictions in a single Gaussian process model experiment. CONCLUSIONS: Of the three parameterizations, sequence time appeared to produce the most stationary series, possibly because clinicians adjust their sampling to the acuity of the patient. Parameterizing by sequence time may be applicable to association and clustering experiments on electronic health record data. A limitation of this study is that laboratory data were derived from only one institution. Sequence time appears to be an important potential parameterization. George Hripcsak, David J. Albers, Adler J. Perotte |
J. Am. Medical Informatics Assoc. | 2 |
| 2014 | Model selection for EHR laboratory variables: how physiology and the health care process can influence EHR laboratory data and their model representations
David J. Albers, Rimma Perotte, Noémie Elhadad, George Hripcsak |
AMIA | 1 |
| 2014 | Temporal trends of hemoglobin A1c testingabstractOBJECTIVE: The study of utilization patterns can quantify potential overuse of laboratory tests and find new ways to reduce healthcare costs. We demonstrate the use of distributional analytics for comparing electronic health record (EHR) laboratory test orders across time to diagnose and quantify overutilization. MATERIALS AND METHODS: We looked at hemoglobin A1c (HbA1c) testing across 119,000 patients and 15 years of hospital records. We examined the patterns of HbA1c ordering before and after the publication of the 2002 American Diabetes Association guidelines for HbA1c testing. We conducted analyses to answer three questions. What are the patterns of HbA1c ordering? Do HbA1c orders follow the guidelines with respect to frequency of measurement? If not, how and why do they depart from the guidelines? RESULTS: The raw number of HbA1c orderings has steadily increased over time, with a specific increase in low-measurement orderings (<6.5%). There is a change in ordering pattern following the 2002 guideline (p<0.001). However, by comparing ordering distributions, we found that the changes do not reflect the guidelines and rather exhibit a new practice of rapid-repeat testing. The rapid-retesting phenomenon does not follow the 2009 guidelines for diabetes diagnosis either, illustrated by a stratified HbA1c value analysis. DISCUSSION: Results suggest HbA1c test overutilization, and contributing factors include lack of care coordination, unexpected values prompting retesting, and point-of-care tests followed by confirmatory laboratory tests. CONCLUSIONS: We present a method of comparing ordering distributions in an EHR across time as a useful diagnostic approach for identifying and assessing the trend of inappropriate use over time. Rimma Perotte, David J. Albers, George Hripcsak, Jorge L. Sepulveda, Noémie Elhadad |
J. Am. Medical Informatics Assoc. | 2 |
| 2014 | Identifying and mitigating biases in EHR laboratory tests
Rimma Perotte, David J. Albers, Jorge L. Sepulveda, Noémie Elhadad |
J. Biomed. Informatics | 2 |
| 2013 | Using patient laboratory measurement values and dynamics to deconvolve EHR bias and define acuity-based phenotypes
David J. Albers, Rimma Perotte, George Hripcsak, Noémie Elhadad |
AMIA | 1 |
| 2013 | Next-generation phenotyping of electronic health recordsabstractThe national adoption of electronic health records (EHR) promises to make an unprecedented amount of data available for clinical research, but the data are complex, inaccurate, and frequently missing, and the record reflects complex processes aside from the patient's physiological state. We believe that the path forward requires studying the EHR as an object of interest in itself, and that new models, learning from data, and collaboration will lead to efficient use of the valuable information currently locked in health records. George Hripcsak, David J. Albers |
J. Am. Medical Informatics Assoc. | 2 |
| 2012 | Using Empirical orthogonal functions to identify temporally important variables to understand time-dependent pathophysiologic and phenotypic differences in patients
David J. Albers, Jan Claassen, George Hripcsak |
AMIA | 1 |
| 2011 | Exploiting time in electronic health record correlationsabstractOBJECTIVE: To demonstrate that a large, heterogeneous clinical database can reveal fine temporal patterns in clinical associations; to illustrate several types of associations; and to ascertain the value of exploiting time. MATERIALS AND METHODS: Lagged linear correlation was calculated between seven clinical laboratory values and 30 clinical concepts extracted from resident signout notes from a 22-year, 3-million-patient database of electronic health records. Time points were interpolated, and patients were normalized to reduce inter-patient effects. RESULTS: The method revealed several types of associations with detailed temporal patterns. Definitional associations included low blood potassium preceding 'hypokalemia.' Low potassium preceding the drug spironolactone with high potassium following spironolactone exemplified intentional and physiologic associations, respectively. Counterintuitive results such as the fact that diseases appeared to follow their effects may be due to the workflow of healthcare, in which clinical findings precede the clinician's diagnosis of a disease even though the disease actually preceded the findings. Fully exploiting time by interpolating time points produced less noisy results. DISCUSSION: Electronic health records are not direct reflections of the patient state, but rather reflections of the healthcare process and the recording process. With proper techniques and understanding, and with proper incorporation of time, interpretable associations can be derived from a large clinical database. CONCLUSION: A large, heterogeneous clinical database can reveal clinical associations, time is an important feature, and care must be taken to interpret the results. George Hripcsak, David J. Albers, Adler J. Perotte |
J. Am. Medical Informatics Assoc. | 2 |