VLDB 2026 Research / reviewers in the wild / expert
Michael E. Matheny
dblp:47/6669 · also Michael Edwin Matheny
· DBLP profile ↗
84ranked-venue papers
12as first author
18since 2021 · last 2025
0000-0003-3217-4147ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 81 · 12 first-author · 17 since 2021Security and privacy · 2 · 1 since 2021Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Emerging algorithmic bias: fairness drift as the next dimension of model maintenance and sustainabilityabstractOBJECTIVES: While performance drift of clinical prediction models is well-documented, the potential for algorithmic biases to emerge post-deployment has had limited characterization. A better understanding of how temporal model performance may shift across subpopulations is required to incorporate fairness drift into model maintenance strategies. MATERIALS AND METHODS: We explore fairness drift in a national population over 11 years, with and without model maintenance aimed at sustaining population-level performance. We trained random forest models predicting 30-day post-surgical readmission, mortality, and pneumonia using 2013 data from US Department of Veterans Affairs facilities. We evaluated performance quarterly from 2014 to 2023 by self-reported race and sex. We estimated discrimination, calibration, and accuracy, and operationalized fairness using metric parity measured as the gap between disadvantaged and advantaged groups. RESULTS: Our cohort included 1 739 666 surgical cases. We observed fairness drift in both the original and temporally updated models. Model updating had a larger impact on overall performance than fairness gaps. During periods of stable fairness, updating models at the population level increased, decreased, or did not impact fairness gaps. During periods of fairness drift, updating models restored fairness in some cases and exacerbated fairness gaps in others. DISCUSSION: This exploratory study highlights that algorithmic fairness cannot be assured through one-time assessments during model development. Temporal changes in fairness may take multiple forms and interact with model updating strategies in unanticipated ways. CONCLUSION: Equitable and sustainable clinical artificial intelligence deployments will require novel methods to monitor algorithmic fairness, detect emerging bias, and adopt model updates that promote fairness. Sharon E. Davis, Chad Dorn, Daniel J. Park, Michael E. Matheny |
J. Am. Medical Informatics Assoc. | 4 |
| 2025 | A machine learning framework to adjust for learning effects in medical device safety evaluationabstractOBJECTIVES: Traditional methods for medical device post-market surveillance often fail to accurately account for operator learning effects, leading to biased assessments of device safety. These methods struggle with non-linearity, complex learning curves, and time-varying covariates, such as physician experience. To address these limitations, we sought to develop a machine learning (ML) framework to detect and adjust for operator learning effects. MATERIALS AND METHODS: A gradient-boosted decision tree ML method was used to analyze synthetic datasets that replicate the complexity of clinical scenarios involving high-risk medical devices. We designed this process to detect learning effects using a risk-adjusted cumulative sum method, quantify the excess adverse event rate attributable to operator inexperience, and adjust for these alongside patient factors in evaluating device safety signals. To maintain integrity, we employed blinding between data generation and analysis teams. Synthetic data used underlying distributions and patient feature correlations based on clinical data from the Department of Veterans Affairs between 2005 and 2012. We generated 2494 synthetic datasets with widely varying characteristics including number of patient features, operators and institutions, and the operator learning form. Each dataset contained a hypothetical study device, Device B, and a reference device, Device A. We evaluated accuracy in identifying learning effects and identifying and estimating the strength of the device safety signal. Our approach also evaluated different clinically relevant thresholds for safety signal detection. RESULTS: Our framework accurately identified the presence or absence of learning effects in 93.6% of datasets and correctly determined device safety signals in 93.4% of cases. The estimated device odds ratios' 95% confidence intervals were accurately aligned with the specified ratios in 94.7% of datasets. In contrast, a comparative model excluding operator learning effects significantly underperformed in detecting device signals and in accuracy. Notably, our framework achieved 100% specificity for clinically relevant safety signal thresholds, although sensitivity varied with the threshold applied. DISCUSSION: A machine learning framework, tailored for the complexities of post-market device evaluation, may provide superior performance compared to standard parametric techniques when operator learning is present. CONCLUSION: Demonstrating the capacity of ML to overcome complex evaluative challenges, our framework addresses the limitations of traditional statistical methods in current post-market surveillance processes. By offering a reliable means to detect and adjust for learning effects, it may significantly improve medical device safety evaluation. Jejo Koola, Karthik Ramesh, Jialin Mao, Minyoung Ahn, Sharon E. Davis, Usha Govindarajulu, Amy Perkins, Dax M. Westerman, Henry Ssemaganda, Theodore Speroff, Lucila Ohno-Machado, Craig Ramsay, Art Sedrakyan, Frederic S. Resnic, Michael E. Matheny |
J. Am. Medical Informatics Assoc. | 15 |
| 2025 | Detecting Opioid Use Disorder in Health Claims Data With Positive Unlabeled LearningabstractAccurate detection and prevalence estimation of behavioral health conditions, such as opioid use disorder (OUD), are crucial for identifying at-risk individuals, determining treatment needs, monitoring prevention and intervention efforts, and recruiting treatment-naive participants for clinical trials. The availability of extensive health data, combined with advancements in machine learning (ML) frameworks, has enabled researchers to employ various ML techniques to predict or identify OUD within patient health data. Ideally, we could directly estimate the prevalence, or the proportion of a population with a condition over time. However, underdiagnosis and undercoding of conditions in patient health records make it challenging to determine the true prevalence of these conditions and to identify at-risk patients with less severe conditions who are more likely to be missed. Consequently, patients without diagnoses may comprise positive and negative examples for a given condition. Treating all undiagnosed (uncoded) patients as negative when applying ML methods can introduce bias into models, affecting their predictive power. To address this issue, we employed Positive Unlabeled Learning Selected Not At Random (PULSNAR), a Positive and Unlabeled (PU) learning technique, to estimate the probability of a given patient having OUD during a time window and the overall population prevalence of OUD. In a sample of 3,342,044 commercially insured US patients with at least one opioid prescription filled, PULSNAR estimated that 5.08% of patients have a cumulative prevalence of OUD over a 2-5 a observation period, compared to the 1.35% with a recorded OUD diagnosis, with 73.5% of cases not diagnosed/coded. The prevalence estimates provided by PULSNAR are consistent with those reported in other studies. Fariha Moomtaheen, Scott A. Malec, Jeremy J. Yang, Cristian Bologa, Kristan Alexander Schneider, Yiliang Zhu 0001, Mauricio Tohen, Gerardo Villarreal, Douglas J. Perkins, Elliot M. Fielstein, Sharon E. Davis, Michael E. Matheny, Christophe G. Lambert |
IEEE J. Biomed. Health Informatics | 13 |
| 2024 | Robin Hood: A De-identification Method to Preserve Minority Representation for Disparities Research
J. Thomas Brown, Ellen Wright Clayton, Michael E. Matheny, Murat Kantarcioglu, Yevgeniy Vorobeychik, Bradley A. Malin |
PSD | 3 |
| 2024 | Sustainable deployment of clinical prediction tools - a 360° approach to model maintenanceabstractBACKGROUND: As the enthusiasm for integrating artificial intelligence (AI) into clinical care grows, so has our understanding of the challenges associated with deploying impactful and sustainable clinical AI models. Complex dataset shifts resulting from evolving clinical environments strain the longevity of AI models as predictive accuracy and associated utility deteriorate over time. OBJECTIVE: Responsible practice thus necessitates the lifecycle of AI models be extended to include ongoing monitoring and maintenance strategies within health system algorithmovigilance programs. We describe a framework encompassing a 360° continuum of preventive, preemptive, responsive, and reactive approaches to address model monitoring and maintenance from critically different angles. DISCUSSION: We describe the complementary advantages and limitations of these four approaches and highlight the importance of such a coordinated strategy to help ensure the promise of clinical AI is not short-lived. Sharon E. Davis, Peter J. Embí, Michael E. Matheny |
J. Am. Medical Informatics Assoc. | 3 |
| 2023 | A framework for understanding label leakage in machine learning for health careabstractINTRODUCTION: The pitfalls of label leakage, contamination of model input features with outcome information, are well established. Unfortunately, avoiding label leakage in clinical prediction models requires more nuance than the common advice of applying "no time machine rule." FRAMEWORK: We provide a framework for contemplating whether and when model features pose leakage concerns by considering the cadence, perspective, and applicability of predictions. To ground these concepts, we use real-world clinical models to highlight examples of appropriate and inappropriate label leakage in practice. RECOMMENDATIONS: Finally, we provide recommendations to support clinical and technical stakeholders as they evaluate the leakage tradeoffs associated with model design, development, and implementation decisions. By providing common language and dimensions to consider when designing models, we hope the clinical prediction community will be better prepared to develop statistically valid and clinically useful machine learning models. Sharon E. Davis, Michael E. Matheny, Suresh Balu, Mark P. Sendak |
J. Am. Medical Informatics Assoc. | 2 |
| 2023 | Inpatient nurses' preferences and decisions with risk information visualizationabstractOBJECTIVE: We examined the influence of 4 different risk information formats on inpatient nurses' preferences and decisions with an acute clinical deterioration decision-support system. MATERIALS AND METHODS: We conducted a comparative usability evaluation in which participants provided responses to multiple user interface options in a simulated setting. We collected qualitative data using think aloud methods. We collected quantitative data by asking participants which action they would perform after each time point in 3 different patient scenarios. RESULTS: More participants (n = 6) preferred the probability format over relative risk ratios (n = 2), absolute differences (n = 2), and number of persons out of 100 (n = 0). Participants liked average lines, having a trend graph to supplement the risk estimate, and consistent colors between trend graphs and possible actions. Participants did not like too much text information or the presence of confidence intervals. From a decision-making perspective, use of the probability format was associated with greater concordance in actions taken by participants compared to the other 3 risk information formats. DISCUSSION: By focusing on nurses' preferences and decisions with several risk information display formats and collecting both qualitative and quantitative data, we have provided meaningful insights for the design of clinical decision-support systems containing complex quantitative information. CONCLUSION: This study adds to our knowledge of presenting risk information to nurses within clinical decision-support systems. We encourage those developing risk-based systems for inpatient nurses to consider expressing risk in a probability format and include a graph (with average line) to display the patient's recent trends. Alvin D. Jeffery, Carrie Reale, Janelle Faiman, Vera Borkowski, Russ Beebe, Michael E. Matheny, Shilo Anders |
J. Am. Medical Informatics Assoc. | 6 |
| 2023 | Blockchain-enabled immutable, distributed, and highly available clinical research activity logging system for federated COVID-19 data analysis from multiple institutionsabstractOBJECTIVE: We aimed to develop a distributed, immutable, and highly available cross-cloud blockchain system to facilitate federated data analysis activities among multiple institutions. MATERIALS AND METHODS: We preprocessed 9166 COVID-19 Structured Query Language (SQL) code, summary statistics, and user activity logs, from the GitHub repository of the Reliable Response Data Discovery for COVID-19 (R2D2) Consortium. The repository collected local summary statistics from participating institutions and aggregated the global result to a COVID-19-related clinical query, previously posted by clinicians on a website. We developed both on-chain and off-chain components to store/query these activity logs and their associated queries/results on a blockchain for immutability, transparency, and high availability of research communication. We measured run-time efficiency of contract deployment, network transactions, and confirmed the accuracy of recorded logs compared to a centralized baseline solution. RESULTS: The smart contract deployment took 4.5 s on an average. The time to record an activity log on blockchain was slightly over 2 s, versus 5-9 s for baseline. For querying, each query took on an average less than 0.4 s on blockchain, versus around 2.1 s for baseline. DISCUSSION: The low deployment, recording, and querying times confirm the feasibility of our cross-cloud, blockchain-based federated data analysis system. We have yet to evaluate the system on a larger network with multiple nodes per cloud, to consider how to accommodate a surge in activities, and to investigate methods to lower querying time as the blockchain grows. CONCLUSION: Blockchain technology can be used to support federated data analysis among multiple institutions. Tsung-Ting Kuo, Anh Pham, Maxim E. Edelson, Jihoon Kim 0001, Yash Gupta, Lucila Ohno-Machado, David M. Anderson, Chandrasekar Balacha, Tyler Bath, Sally L. Baxter, Andrea Becker-Pennrich, Douglas S. Bell, Elmer V. Bernstam, Ngan Chau, Michele E. Day, Jason N. Doctor, Scott L. DuVall, Robert El-Kareh, Renato Florian, Robert W. Follett, Benjamin P. Geisler, Alessandro Ghigi, Assaf Gottlieb, Christian Hinske, Zhaoxian Hu, Diana Ir, Xiaoqian Jiang, Katherine K. Kim, Tara K. Knight, Jejo Koola, Ulrich Mansmann, Michael E. Matheny, Daniella Meeker, Zongyang Mou, Larissa Neumann, Nghia H. Nguyen, Nicholas R. Anderson 0001, Eunice Park, Paulina Paul, Mark J. Pletcher, Kai W. Post, Clemens Rieder, Clemens Scherer, Lisa M. Schilling, Andrey Soares, Spencer L. SooHoo, Ekin Soysal, Steven Covington, Brian Tep, Brian Toy, Baocheng Wang, Zhen R. Wu, Hua Xu 0001, Yong K. Choi, Kai Zheng 0002, Yujia Zhou 0003, Rachel A Zucker |
J. Am. Medical Informatics Assoc. | 34 |
| 2022 | A Framework for Generating Synthetic Clinical Datasets with Learning Effects to Support Methods Development and Validation
Sharon E. Davis, Henry Ssemaganda, Jejo Koola, Jialin Mao, Dax M. Westerman, Theodore Speroff, Usha Govindarajulu, Craig Ramsay, Lucila Ohno-Machado, Frederic S. Resnic, Michael E. Matheny |
AMIA | 11 |
| 2022 | A Framework for Detecting Medical Device Safety Signals Confounded by Learning Effects Using Machine Learning
Jejo Koola, Jialin Mao, Sharon E. Davis, Henry Ssemaganda, Dax M. Westerman, Lucila Ohno-Machado, Frederic S. Resnic, Michael E. Matheny |
AMIA | 8 |
| 2022 | Can Data Science Move Medical Device Surveillance Forward? Highlights of Solutions and Remaining Challenges in Real-World Evidence Generation
Michael E. Matheny, Art Sedrakyan, Omar Badawi, Danica Marinac-Dabic, Frederic S. Resnic |
AMIA | 1 |
| 2022 | Disentangling and Characterizing Device Safety Signals and Learning Effects
Henry Ssemaganda, Frederic S. Resnic, Sharon E. Davis, Usha Govindarajulu, Jejo Koola, Jialin Mao, Dax M. Westerman, Theodore Speroff, Craig Ramsay, Art Sedrakyan, Lucila Ohno-Machado, Michael E. Matheny |
AMIA | 12 |
| 2022 | AUTO-PILOT: Development and Validation of a Tool to Extract Structured Data in Interfacility Transfer Paperwork
Jesse O. Wrenn, Sara Lin, Ruth Reeves, Dax M. Westerman, Michael E. Matheny, Michael J. Ward |
AMIA | 6 |
| 2021 | Disparities in Coded and Imputed Post-Traumatic Stress Disorder and Self-Harm Among US Veterans
Sharon E. Davis, Nicolas R. Lauve, Sharidan K. Parr, Daniel Park, Michael E. Matheny, Gerardo Villarreal, George Uhl, Yiliang Zhu 0001, Mauricio Tohen, Douglas J. Perkins, Christophe G. Lambert |
AMIA | 6 |
| 2021 | Intrinsic Evaluation of Contextual and Non-contextual Word Embeddings using Radiology Reports
Mirza S. Khan, Bennett A. Landman, Steve Deppen, Michael E. Matheny |
AMIA | 4 |
| 2021 | Privacy-protecting, reliable response data discovery using COVID-19 patient observationsabstractOBJECTIVE: To utilize, in an individual and institutional privacy-preserving manner, electronic health record (EHR) data from 202 hospitals by analyzing answers to COVID-19-related questions and posting these answers online. MATERIALS AND METHODS: We developed a distributed, federated network of 12 health systems that harmonized their EHRs and submitted aggregate answers to consortia questions posted at https://www.covid19questions.org. Our consortium developed processes and implemented distributed algorithms to produce answers to a variety of questions. We were able to generate counts, descriptive statistics, and build a multivariate, iterative regression model without centralizing individual-level data. RESULTS: Our public website contains answers to various clinical questions, a web form for users to ask questions in natural language, and a list of items that are currently pending responses. The results show, for example, that patients who were taking angiotensin-converting enzyme inhibitors and angiotensin II receptor blockers, within the year before admission, had lower unadjusted in-hospital mortality rates. We also showed that, when adjusted for, age, sex, and ethnicity were not significantly associated with mortality. We demonstrated that it is possible to answer questions about COVID-19 using EHR data from systems that have different policies and must follow various regulations, without moving data out of their health systems. DISCUSSION AND CONCLUSIONS: We present an alternative or a complement to centralized COVID-19 registries of EHR data. We can use multivariate distributed logistic regression on observations recorded in the process of care to generate results without transferring individual-level data outside the health systems. Jihoon Kim 0001, Larissa Neumann, Paulina Paul, Michele E. Day, Michael Aratow, Douglas S. Bell, Jason N. Doctor, Christian Hinske, Xiaoqian Jiang, Katherine K. Kim, Michael E. Matheny, Daniella Meeker, Mark J. Pletcher, Lisa M. Schilling, Spencer L. SooHoo, Hua Xu 0001, Kai Zheng 0002, Lucila Ohno-Machado |
J. Am. Medical Informatics Assoc. | 11 |
| 2021 | Enhancing trust in AI through industry self-governanceabstractArtificial intelligence (AI) is critical to harnessing value from exponentially growing health and healthcare data. Expectations are high for AI solutions to effectively address current health challenges. However, there have been prior periods of enthusiasm for AI followed by periods of disillusionment, reduced investments, and progress, known as "AI Winters." We are now at risk of another AI Winter in health/healthcare due to increasing publicity of AI solutions that are not representing touted breakthroughs, and thereby decreasing trust of users in AI. In this article, we first highlight recently published literature on AI risks and mitigation strategies that would be relevant for groups considering designing, implementing, and promoting self-governance. We then describe a process for how a diverse group of stakeholders could develop and define standards for promoting trust, as well as AI risk-mitigating practices through greater industry self-governance. We also describe how adherence to such standards could be verified, specifically through certification/accreditation. Self-governance could be encouraged by governments to complement existing regulatory schema or legislative efforts to mitigate AI risks. Greater adoption of industry self-governance could fill a critical gap to construct a more comprehensive approach to the governance of AI solutions than US legislation/regulations currently encompass. In this more comprehensive approach, AI developers, AI users, and government/legislators all have critical roles to play to advance practices that maintain trust in AI and prevent another AI Winter. Joachim Roski, Ezekiel J. Maier, Kevin Vigilante, Elizabeth A. Kane, Michael E. Matheny |
J. Am. Medical Informatics Assoc. | 5 |
| 2021 | Adaptation of an NLP system to a new healthcare environment to identify social determinants of health
Ruth M. Reeves, Lee M. Christensen, Jeremiah R. Brown, Michael Conway, Maxwell Levis, Glenn T. Gobbel, Rashmee U. Shah, Christine Goodrich, Iben Ricket, Freneka F. Minter, Andrew Bohm, Bruce E. Bray, Michael E. Matheny, Wendy W. Chapman |
J. Biomed. Informatics | 13 |
| 2020 | A Surveillance Framework for Monitoring and Updating Clinical Prediction Models
Sharon E. Davis, Robert A. Greevy Jr., Thomas A. Lasko, Colin G. Walsh, Michael E. Matheny |
AMIA | 5 |
| 2020 | Pain Management Collaboratory Coordinating Center (PMC3): Insights from Multiple Work Groups for Clinical Research Informatics Support
Michael E. Matheny, Stephen Luther, William T. Roddy, Cynthia Brandt, Joseph Erdos |
AMIA | 1 |
| 2020 | Efficient determination of equivalence for encrypted data
Jason N. Doctor, Jaideep Vaidya, Xiaoqian Jiang, Shuang Wang 0002, Lisa M. Schilling, Toan Ong, Michael E. Matheny, Lucila Ohno-Machado, Daniella Meeker |
Comput. Secur. | 7 |
| 2020 | COVID-19 TestNorm: A tool to normalize COVID-19 testing names to LOINC codesabstractLarge observational data networks that leverage routine clinical practice data in electronic health records (EHRs) are critical resources for research on coronavirus disease 2019 (COVID-19). Data normalization is a key challenge for the secondary use of EHRs for COVID-19 research across institutions. In this study, we addressed the challenge of automating the normalization of COVID-19 diagnostic tests, which are critical data elements, but for which controlled terminology terms were published after clinical implementation. We developed a simple but effective rule-based tool called COVID-19 TestNorm to automatically normalize local COVID-19 testing names to standard LOINC (Logical Observation Identifiers Names and Codes) codes. COVID-19 TestNorm was developed and evaluated using 568 test names collected from 8 healthcare systems. Our results show that it could achieve an accuracy of 97.4% on an independent test set. COVID-19 TestNorm is available as an open-source package for developers and as an online Web application for end users (https://clamp.uth.edu/covid/loinc.php). We believe that it will be a useful tool to support secondary use of EHRs for research on COVID-19. Jianfu Li, Ekin Soysal, Jiang Bian 0001, Scott L. DuVall, Elizabeth Hanchrow, Kristine E. Lynch, Michael E. Matheny, Karthik Natarajan, Lucila Ohno-Machado, Serguei V. S. Pakhomov, Ruth M. Reeves, Amy M. Sitapati, Swapna Abhyankar, Theresa A. Cullen, Jami Deckard, Xiaoqian Jiang, Robert Murphy, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 9 |
| 2020 | Detection of calibration drift in clinical prediction models to inform model updating
Sharon E. Davis, Robert A. Greevy Jr., Thomas A. Lasko, Colin G. Walsh, Michael E. Matheny |
J. Biomed. Informatics | 5 |
| 2019 | Developing Synthetic VA Healthcare Data in OMOP CDM Model
Jiantao Bian, Hamid Saoudian, Brett R. South, Kristine E. Lynch, Benjamin Viernes, Michael E. Matheny, Scott L. DuVall |
AMIA | 6 |
| 2019 | Comparison of Prediction Model Performance Updating Protocols: Using a Data-Driven Testing Procedure to Guide Updating
Sharon E. Davis, Robert A. Greevy Jr., Thomas A. Lasko, Colin G. Walsh, Michael E. Matheny |
AMIA | 5 |
| 2019 | Mapping Medications to Standards with Machine Learning and Noisy Labels
Alvin D. Jeffery, Sharidan K. Parr, Michael E. Matheny |
AMIA | 3 |
| 2019 | Longitudinal modeling of prescription refill records to predict medication use
Kimberley Kondratieff, Candace D. McNaughton, Michael E. Matheny, Thomas A. Lasko |
AMIA | 3 |
| 2019 | Value Set Development for Studies Involving Historical Data
Kristin Lyman, Melanie Canterberry, Elizabeth Crull, Jason Denton, Fern FitzHenry, Qiaohong Hu, Sharidan K. Parr, Amy Perkins, Dina Zein, Michael E. Matheny, Jason N. Doctor, Daniella Meeker |
AMIA | 10 |
| 2019 | A nonparametric updating method to correct clinical prediction model driftabstractOBJECTIVE: Clinical prediction models require updating as performance deteriorates over time. We developed a testing procedure to select updating methods that minimizes overfitting, incorporates uncertainty associated with updating sample sizes, and is applicable to both parametric and nonparametric models. MATERIALS AND METHODS: We describe a procedure to select an updating method for dichotomous outcome models by balancing simplicity against accuracy. We illustrate the test's properties on simulated scenarios of population shift and 2 models based on Department of Veterans Affairs inpatient admissions. RESULTS: In simulations, the test generally recommended no update under no population shift, no update or modest recalibration under case mix shifts, intercept correction under changing outcome rates, and refitting under shifted predictor-outcome associations. The recommended updates provided superior or similar calibration to that achieved with more complex updating. In the case study, however, small update sets lead the test to recommend simpler updates than may have been ideal based on subsequent performance. DISCUSSION: Our test's recommendations highlighted the benefits of simple updating as opposed to systematic refitting in response to performance drift. The complexity of recommended updating methods reflected sample size and magnitude of performance drift, as anticipated. The case study highlights the conservative nature of our test. CONCLUSIONS: This new test supports data-driven updating of models developed with both biostatistical and machine learning approaches, promoting the transportability and maintenance of a wide array of clinical prediction models and, in turn, a variety of applications relying on modern prediction tools. Sharon E. Davis, Robert A. Greevy Jr., Christopher Fonnesbeck, Thomas A. Lasko, Colin G. Walsh, Michael E. Matheny |
J. Am. Medical Informatics Assoc. | 6 |
| 2019 | Prognostic models will be victims of their own success, unlessabstractPredictive analytics have begun to change the workflows of healthcare by giving insight into our future health. Deploying prognostic models into clinical workflows should change behavior and motivate interventions that affect outcomes. As users respond to model predictions, downstream characteristics of the data, including the distribution of the outcome, may change. The ever-changing nature of healthcare necessitates maintenance of prognostic models to ensure their longevity. The more effective a model and intervention(s) are at improving outcomes, the faster a model will appear to degrade. Improving outcomes can disrupt the association between the model's predictors and the outcome. Model refitting may not always be the most effective response to these challenges. These problems will need to be mitigated by systematically incorporating interventions into prognostic models and by maintaining robust performance surveillance of models in clinical use. Holistically modeling the outcome and intervention(s) can lead to resilience to future compromises in performance. Matthew C. Lenert, Michael E. Matheny, Colin G. Walsh |
J. Am. Medical Informatics Assoc. | 2 |
| 2019 | Explicit causal reasoning is preferred, but not necessary for pragmatic valueabstractIn their response, researchers Sperrin, Jenkins, Martin, and Peek discuss some of the benefits of applying causal inference frameworks (CIFs) to predict treatment naïve risk in the domain of risk modeling. We agree that causality-based models using diagrams are a powerful tool and that these models can avoid the pitfalls of model-mediated changes to the outcome process.1 CIFs have also demonstrated robustness to unobserved confounders.2 There are many reasons why explicitly considering causality and estimating baseline risk in the absence of treatments are important when deploying and maintaining prognostic models in clinical operations. While these models have many desirable properties, they are not without their challenges, as Sperrin et al note. CIFs demonstrate a firm understanding of the processes one wishes to improve. Getting to the requisite level of insight to build such a diagram is a long and arduous scientific process. This is not to say many processes cannot be diagramed using current knowledge. We feel that incorporating causality where it is well understood is useful, but there are circumstances in which CIFs are likely to be incorrect and have the potential to cause error. Furthermore, causal models require data elements that reflect how a process works. Current bulwark data streams (revenue cycle-focused electronic health records) are not likely to include data relevant or sufficient for CIFs for a number (if not most) use-cases. Matthew C. Lenert, Michael E. Matheny, Colin G. Walsh |
J. Am. Medical Informatics Assoc. | 2 |
| 2018 | Developing a Testing Procedure to Select Model Updating Methods
Sharon E. Davis, Robert A. Greevy Jr., Michael E. Matheny |
AMIA | 3 |
| 2018 | The Department of Defense (DoD) and Department of Veterans Affairs (VA) Infrastructure for Clinical Intelligence (DaVINCI)
Scott L. DuVall, Michael E. Matheny, Ildar R. Ibragimov, Trey D. Oats, Jay N. Tucker, Brett R. South, Augie Turano, Hamid Saoudian, Casey Kangas, Keith D. Hofmann, Wendy Funk, Chris Nichols, Albert Bonnema, Louis Ferrucci, Jonathan R. Nebeker |
AMIA | 2 |
| 2018 | Automated Mapping of Laboratory Tests to LOINC Codes using Noisy Labels in a National Electronic Health Record System Database
Sharidan K. Parr, Thomas A. Lasko, Alvin D. Jeffery, Matthew S. Shotwell, Michael E. Matheny |
AMIA | 5 |
| 2018 | Annotating Social Determinants of Health and Functional Status Information Using Publicly Accessible Corpora
Ruth M. Reeves, Brett R. South, Glenn T. Gobbel, Lee M. Christensen, Wendy W. Chapman, Michael E. Matheny, Jeremiah R. Brown |
AMIA | 7 |
| 2018 | Flipping the Model for Biomedical Informatics Research
Brett R. South, Kristine W. Lynch, Michael E. Matheny, Julie A. Lynch, Olga Efimova, Catherine Chanfreau-Coffinier, Olga V. Patterson, Benjamin Viernes, Scott L. DuVall |
AMIA | 3 |
| 2018 | Integrating Cancer Genomic Data into Decision Support at the VA: Hematologic Oncology as a Use Case
Dax M. Westerman, Fern FitzHenry, Michael E. Matheny, Michele L. LeNoue-Newton, Claudio A. Mosse, Ian Maurer, Matthew Stachowiak, Justin A. Yeakley, James Cole, Regina M. Thomas, Mia A. Levy |
AMIA | 3 |
| 2018 | Automated mapping of laboratory tests to LOINC codes using noisy labels in a national electronic health record system databaseabstractObjective: Standards such as the Logical Observation Identifiers Names and Codes (LOINC®) are critical for interoperability and integrating data into common data models, but are inconsistently used. Without consistent mapping to standards, clinical data cannot be harmonized, shared, or interpreted in a meaningful context. We sought to develop an automated machine learning pipeline that leverages noisy labels to map laboratory data to LOINC codes. Materials and Methods: Across 130 sites in the Department of Veterans Affairs Corporate Data Warehouse, we selected the 150 most commonly used laboratory tests with numeric results per site from 2000 through 2016. Using source data text and numeric fields, we developed a machine learning model and manually validated random samples from both labeled and unlabeled datasets. Results: The raw laboratory data consisted of >6.5 billion test results, with 2215 distinct LOINC codes. The model predicted the correct LOINC code in 85% of the unlabeled data and 96% of the labeled data by test frequency. In the subset of labeled data where the original and model-predicted LOINC codes disagreed, the model-predicted LOINC code was correct in 83% of the data by test frequency. Conclusion: Using a completely automated process, we are able to assign LOINC codes to unlabeled data with high accuracy. When the model-predicted LOINC code differed from the original LOINC code, the model prediction was correct in the vast majority of cases. This scalable, automated algorithm may improve data quality and interoperability, while substantially reducing the manual effort currently needed to accurately map laboratory data. Sharidan K. Parr, Matthew S. Shotwell, Alvin D. Jeffery, Thomas A. Lasko, Michael E. Matheny |
J. Am. Medical Informatics Assoc. | 5 |
| 2018 | Development of an automated phenotyping algorithm for hepatorenal syndrome
Jejo Koola, Sharon E. Davis, Omar Al-Nimri, Sharidan K. Parr, Daniel Fabbri, Bradley A. Malin, Samuel B. Ho, Michael E. Matheny |
J. Biomed. Informatics | 8 |
| 2017 | Calibration Drift Among Regression and Machine Learning Models for Hospital Mortality
Sharon E. Davis, Thomas A. Lasko, Guanhua Chen 0002, Michael E. Matheny |
AMIA | 4 |
| 2017 | Machine Learning Models to Predict Readmission for Patients with Cirrhosis
Jejo Koola, Aize Cao, Guanhua Chen 0002, Amy Perkins, Samuel B. Ho, Sharon E. Davis, Michael E. Matheny |
AMIA | 7 |
| 2017 | Natural Language Data Sampling: Principles & Strategies in Corpus Selection
Ruth M. Reeves, Nancy Gentry, Fern FitzHenry, Glenn T. Gobbel, Michael E. Matheny |
AMIA | 5 |
| 2017 | Identification of adverse drug-drug interactions through causal association rule discovery from spontaneous adverse event reports
Ruichu Cai, Yong Hu 0002, Brittany Melton, Michael E. Matheny, Hua Xu 0001, Lemuel R. Waitman |
Artif. Intell. Medicine | 5 |
| 2017 | Calibration drift in regression and machine learning models for acute kidney injuryabstractOBJECTIVE: Predictive analytics create opportunities to incorporate personalized risk estimates into clinical decision support. Models must be well calibrated to support decision-making, yet calibration deteriorates over time. This study explored the influence of modeling methods on performance drift and connected observed drift with data shifts in the patient population. MATERIALS AND METHODS: Using 2003 admissions to Department of Veterans Affairs hospitals nationwide, we developed 7 parallel models for hospital-acquired acute kidney injury using common regression and machine learning methods, validating each over 9 subsequent years. RESULTS: Discrimination was maintained for all models. Calibration declined as all models increasingly overpredicted risk. However, the random forest and neural network models maintained calibration across ranges of probability, capturing more admissions than did the regression models. The magnitude of overprediction increased over time for the regression models while remaining stable and small for the machine learning models. Changes in the rate of acute kidney injury were strongly linked to increasing overprediction, while changes in predictor-outcome associations corresponded with diverging patterns of calibration drift across methods. CONCLUSIONS: Efficient and effective updating protocols will be essential for maintaining accuracy of, user confidence in, and safety of personalized risk predictions to support decision-making. Model updating protocols should be tailored to account for variations in calibration drift across methods and respond to periods of rapid performance drift rather than be limited to regularly scheduled annual or biannual intervals. Sharon E. Davis, Thomas A. Lasko, Guanhua Chen 0002, Edward D. Siew, Michael E. Matheny |
J. Am. Medical Informatics Assoc. | 5 |
| 2017 | Congestive heart failure information extraction framework for automated treatment performance measures assessmentabstractOBJECTIVE: This paper describes a new congestive heart failure (CHF) treatment performance measure information extraction system - CHIEF - developed as part of the Automated Data Acquisition for Heart Failure project, a Veterans Health Administration project aiming at improving the detection of patients not receiving recommended care for CHF. DESIGN: CHIEF is based on the Apache Unstructured Information Management Architecture framework, and uses a combination of rules, dictionaries, and machine learning methods to extract left ventricular function mentions and values, CHF medications, and documented reasons for a patient not receiving these medications. MEASUREMENTS: The training and evaluation of CHIEF were based on subsets of a reference standard of various clinical notes from 1083 Veterans Health Administration patients. Domain experts manually annotated these notes to create our reference standard. Metrics used included recall, precision, and the F 1 -measure. RESULTS: In general, CHIEF extracted CHF medications with high recall (>0.990) and good precision (0.960-0.978). Mentions of Left Ventricular Ejection Fraction were also extracted with high recall (0.978-0.986) and precision (0.986-0.994), and quantitative values of Left Ventricular Ejection Fraction were found with 0.910-0.945 recall and with high precision (0.939-0.976). Reasons for not prescribing CHF medications were more difficult to extract, only reaching fair accuracy with about 0.310-0.400 recall and 0.250-0.320 precision. CONCLUSION: This study demonstrated that applying natural language processing to unlock the rich and detailed clinical information found in clinical narrative text notes makes fast and scalable quality improvement approaches possible, eventually improving management and outpatient treatment of patients suffering from CHF. Stéphane M. Meystre, Glenn T. Gobbel, Michael E. Matheny, Andrew Redd, Bruce E. Bray, Jennifer H. Garvin |
J. Am. Medical Informatics Assoc. | 4 |
| 2016 | Informatics Challenges in Working with "Big Data" - A Use Case in Identifying Predictors of Constipation
Fern FitzHenry, Svetlana K. Eden, Jason N. Denton, Robert J. LoCasale, Aize Cao, Ruth M. Reeves, Nancy L. Wells, Michael E. Matheny |
AMIA | 9 |
| 2016 | Workflow-guided development of a clinical decision support tool for patients with advanced liver disease
Samuel B. Ho, Julie Ducom, Jennifer H. Garvin, Jejo Koola, Russ Beebe, Jason Slagle, Dax M. Westerman, Carrie Reale, Matthew B. Weinger, Erik Groessl, Michael E. Matheny |
AMIA | 12 |
| 2016 | An Electronic Health Record Phenotyping Algorithm for Identifying Patients with Hepatorenal Syndrome
Jejo Koola, Samuel B. Ho, Michael E. Matheny |
AMIA | 3 |
| 2016 | Using Results of Statistical Text Mining in Big Data Analysis
Stephen Luther, James A. McCart, Dezon Finch, Lina Bouayad, William Lapcevic, Michael E. Matheny |
AMIA | 6 |
| 2016 | An Integrated Privacy Preserving Collaborative Analytics Platform: The PCORnet pSCANNER-PopMedNet TM Software Suite
Michael E. Matheny, Dax M. Westerman, Laura Pearlman, Josh Gieringer, Xiaoqian Jiang, Claudiu Farcas, Tara K. Knight, Shuang Wang 0002, Amy Perkins, Lucila Ohno-Machado, Bill Clarke, Daniella Meeker |
AMIA | 1 |
| 2016 | Calibration of Predictive Models for Clinical Decision Making: Personalizing Prevention, Treatment, and Disease Progression
Lucila Ohno-Machado, George Hripcsak, Michael E. Matheny, Yuan Wu 0003, Xiaoqian Jiang |
AMIA | 3 |
| 2016 | Event Coreference in Support of Temporal Reasoning in Mental Health Notes
Ruth M. Reeves, Marcus Verhagen, Cynthia Brandt, Wendy W. Chapman, Michael E. Matheny, Steven H. Brown, Brian Marx, Theodore Speroff |
AMIA | 5 |
| 2016 | The Discriminative Power of Non-Specific Laboratory Results
Jacob P. VanHouten, Christopher Fonnesbeck, Michael E. Matheny, Thomas A. Lasko |
AMIA | 3 |
| 2015 | Transforming the National Department of Veterans Affairs Data Warehouse to the OMOP Common Data Model
Fern FitzHenry, Jesse Brannen, Jason N. Denton, Jonathan R. Nebeker, Scott L. DuVall, Freneka F. Minter, Jeffrey Scehnet, Brian C. Sauer, Lucila Ohno-Machado, Michael E. Matheny |
AMIA | 10 |
| 2015 | RapTAT: A Tool for Assisted Annotation and Reviewer Training via Online Machine Learning
Glenn T. Gobbel, Ruth M. Reeves, Brett R. South, Wendy W. Chapman, Jay Jarman, Steven Lay, Michael E. Matheny |
AMIA | 7 |
| 2015 | A Novel Visualization for Rapid Summarization of Patient History: Application to Cirrhosis
Jejo Koola, Samuel B. Ho, Michael E. Matheny |
AMIA | 3 |
| 2015 | National Veterans Health Administration inpatient risk stratification models for hospital-acquired acute kidney injuryabstractOBJECTIVE: Hospital-acquired acute kidney injury (HA-AKI) is a potentially preventable cause of morbidity and mortality. Identifying high-risk patients prior to the onset of kidney injury is a key step towards AKI prevention. MATERIALS AND METHODS: A national retrospective cohort of 1,620,898 patient hospitalizations from 116 Veterans Affairs hospitals was assembled from electronic health record (EHR) data collected from 2003 to 2012. HA-AKI was defined at stage 1+, stage 2+, and dialysis. EHR-based predictors were identified through logistic regression, least absolute shrinkage and selection operator (lasso) regression, and random forests, and pair-wise comparisons between each were made. Calibration and discrimination metrics were calculated using 50 bootstrap iterations. In the final models, we report odds ratios, 95% confidence intervals, and importance rankings for predictor variables to evaluate their significance. RESULTS: The area under the receiver operating characteristic curve (AUC) for the different model outcomes ranged from 0.746 to 0.758 in stage 1+, 0.714 to 0.720 in stage 2+, and 0.823 to 0.825 in dialysis. Logistic regression had the best AUC in stage 1+ and dialysis. Random forests had the best AUC in stage 2+ but the least favorable calibration plots. Multiple risk factors were significant in our models, including some nonsteroidal anti-inflammatory drugs, blood pressure medications, antibiotics, and intravenous fluids given during the first 48 h of admission. CONCLUSIONS: This study demonstrated that, although all the models tested had good discrimination, performance characteristics varied between methods, and the random forests models did not calibrate as well as the lasso or logistic regression models. In addition, novel modifiable risk factors were explored and found to be significant. Robert M. Cronin, Jacob P. VanHouten, Edward D. Siew, Svetlana K. Eden, Stephan D. Fihn, Christopher D. Nielson, Josh F. Peterson, Clifton R. Baker, T. Alp Ikizler, Theodore Speroff, Michael E. Matheny |
J. Am. Medical Informatics Assoc. | 11 |
| 2015 | A system to build distributed multivariate models and manage disparate data sharing policies: implementation in the scalable national network for effectiveness researchabstractBACKGROUND: Centralized and federated models for sharing data in research networks currently exist. To build multivariate data analysis for centralized networks, transfer of patient-level data to a central computation resource is necessary. The authors implemented distributed multivariate models for federated networks in which patient-level data is kept at each site and data exchange policies are managed in a study-centric manner. OBJECTIVE: The objective was to implement infrastructure that supports the functionality of some existing research networks (e.g., cohort discovery, workflow management, and estimation of multivariate analytic models on centralized data) while adding additional important new features, such as algorithms for distributed iterative multivariate models, a graphical interface for multivariate model specification, synchronous and asynchronous response to network queries, investigator-initiated studies, and study-based control of staff, protocols, and data sharing policies. MATERIALS AND METHODS: Based on the requirements gathered from statisticians, administrators, and investigators from multiple institutions, the authors developed infrastructure and tools to support multisite comparative effectiveness studies using web services for multivariate statistical estimation in the SCANNER federated network. RESULTS: The authors implemented massively parallel (map-reduce) computation methods and a new policy management system to enable each study initiated by network participants to define the ways in which data may be processed, managed, queried, and shared. The authors illustrated the use of these systems among institutions with highly different policies and operating under different state laws. DISCUSSION AND CONCLUSION: Federated research networks need not limit distributed query functionality to count queries, cohort discovery, or independently estimated analytic models. Multivariate analyses can be efficiently and securely conducted without patient-level data transport, allowing institutions with strict local data storage requirements to participate in sophisticated analyses based on federated research networks. Daniella Meeker, Xiaoqian Jiang, Michael E. Matheny, Claudiu Farcas, Mike D'Arcy, Laura Pearlman, Lavanya Nookala, Michele E. Day, Katherine K. Kim, Hyeon-Eui Kim, Aziz A. Boxwala, Robert El-Kareh, Grace Kuo, Frederic S. Resnic, Carl Kesselman, Lucila Ohno-Machado |
J. Am. Medical Informatics Assoc. | 3 |
| 2014 | Development of a Cirrhosis-Associated Symptom/Finding Detection Tool
Jejo Koola, Robert M. Cronin, Ruth M. Reeves, Jason N. Denton, Samuel B. Ho, Michael E. Matheny |
AMIA | 6 |
| 2014 | Research and applications: Assisted annotation of medical free text using RapTATabstractOBJECTIVE: To determine whether assisted annotation using interactive training can reduce the time required to annotate a clinical document corpus without introducing bias. MATERIALS AND METHODS: A tool, RapTAT, was designed to assist annotation by iteratively pre-annotating probable phrases of interest within a document, presenting the annotations to a reviewer for correction, and then using the corrected annotations for further machine learning-based training before pre-annotating subsequent documents. Annotators reviewed 404 clinical notes either manually or using RapTAT assistance for concepts related to quality of care during heart failure treatment. Notes were divided into 20 batches of 19-21 documents for iterative annotation and training. RESULTS: The number of correct RapTAT pre-annotations increased significantly and annotation time per batch decreased by ~50% over the course of annotation. Annotation rate increased from batch to batch for assisted but not manual reviewers. Pre-annotation F-measure increased from 0.5 to 0.6 to >0.80 (relative to both assisted reviewer and reference annotations) over the first three batches and more slowly thereafter. Overall inter-annotator agreement was significantly higher between RapTAT-assisted reviewers (0.89) than between manual reviewers (0.85). DISCUSSION: The tool reduced workload by decreasing the number of annotations needing to be added and helping reviewers to annotate at an increased rate. Agreement between the pre-annotations and reference standard, and agreement between the pre-annotations and assisted annotations, were similar throughout the annotation process, which suggests that pre-annotation did not introduce bias. CONCLUSIONS: Pre-annotations generated by a tool capable of interactive training can reduce the time required to create an annotated document corpus by up to 50%. Glenn T. Gobbel, Jennifer H. Garvin, Ruth M. Reeves, Robert M. Cronin, Julia Heavirland, Jenifer Williams, Allison Weaver, Shrimalini Jayaramaraja, Dario A. Giuse, Theodore Speroff, Steven H. Brown, Hua Xu 0001, Michael E. Matheny |
J. Am. Medical Informatics Assoc. | 13 |
| 2014 | Determining molecular predictors of adverse drug reactions with causality analysis based on structure learningabstractOBJECTIVE: Adverse drug reaction (ADR) can have dire consequences. However, our current understanding of the causes of drug-induced toxicity is still limited. Hence it is of paramount importance to determine molecular factors of adverse drug responses so that safer therapies can be designed. METHODS: We propose a causality analysis model based on structure learning (CASTLE) for identifying factors that contribute significantly to ADRs from an integration of chemical and biological properties of drugs. This study aims to address two major limitations of the existing ADR prediction studies. First, ADR prediction is mostly performed by assessing the correlations between the input features and ADRs, and the identified associations may not indicate causal relations. Second, most predictive models lack biological interpretability. RESULTS: CASTLE was evaluated in terms of prediction accuracy on 12 organ-specific ADRs using 830 approved drugs. The prediction was carried out by first extracting causal features with structure learning and then applying them to a support vector machine (SVM) for classification. Through rigorous experimental analyses, we observed significant increases in both macro and micro F1 scores compared with the traditional SVM classifier, from 0.88 to 0.89 and 0.74 to 0.81, respectively. Most importantly, identified links between the biological factors and organ-specific drug toxicities were partially supported by evidence in Online Mendelian Inheritance in Man. CONCLUSIONS: The proposed CASTLE model not only performed better in prediction than the baseline SVM but also produced more interpretable results (ie, biological factors responsible for ADRs), which is critical to discovering molecular activators of ADRs. Ruichu Cai, Yong Hu 0002, Michael E. Matheny, Jingchun Sun, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 4 |
| 2014 | Brief communication: pSCANNER: patient-centered Scalable National Network for Effectiveness ResearchabstractThis article describes the patient-centered Scalable National Network for Effectiveness Research (pSCANNER), which is part of the recently formed PCORnet, a national network composed of learning healthcare systems and patient-powered research networks funded by the Patient Centered Outcomes Research Institute (PCORI). It is designed to be a stakeholder-governed federated network that uses a distributed architecture to integrate data from three existing networks covering over 21 million patients in all 50 states: (1) VA Informatics and Computing Infrastructure (VINCI), with data from Veteran Health Administration's 151 inpatient and 909 ambulatory care and community-based outpatient clinics; (2) the University of California Research exchange (UC-ReX) network, with data from UC Davis, Irvine, Los Angeles, San Francisco, and San Diego; and (3) SCANNER, a consortium of UCSD, Tennessee VA, and three federally qualified health systems in the Los Angeles area supplemented with claims and health information exchange data, led by the University of Southern California. Initial use cases will focus on three conditions: (1) congestive heart failure; (2) Kawasaki disease; (3) obesity. Stakeholders, such as patients, clinicians, and health service researchers, will be engaged to prioritize research questions to be answered through the network. We will use a privacy-preserving distributed computation model with synchronous and asynchronous modes. The distributed system will be based on a common data model that allows the construction and evaluation of distributed multivariate models for a variety of statistical analyses. Lucila Ohno-Machado, Zia Agha, Douglas S. Bell, Lisa Dahm, Michele E. Day, Jason N. Doctor, Davera Gabriel, Maninder K. Kahlon, Katherine K. Kim, Michael A. Hogarth, Michael E. Matheny, Daniella Meeker, Jonathan R. Nebeker |
J. Am. Medical Informatics Assoc. | 11 |
| 2014 | Development and evaluation of RapTAT: A machine learning system for concept mapping of phrases from medical narratives
Glenn T. Gobbel, Ruth M. Reeves, Shrimalini Jayaramaraja, Dario A. Giuse, Theodore Speroff, Steven H. Brown, Peter L. Elkin, Michael E. Matheny |
J. Biomed. Informatics | 8 |
| 2013 | Use and Evaluation of RapTAT-Assisted Annotation for Extraction of Acute Kidney Injury Concepts from Clinical Free Text
Glenn T. Gobbel, Ruth M. Reeves, Fern FitzHenry, Diane Montella, Robert M. Cronin, Shrimalini Jayaramaraja, Theodore Speroff, Steven H. Brown, Dario A. Giuse, Michael E. Matheny |
AMIA | 10 |
| 2013 | Methods for Detection of Kidney Disease Risk Factors in Clinical Reports
Ruth M. Reeves, Glenn T. Gobbel, Shrimalini Jayaramaraja, Fern FitzHenry, Robert M. Cronin, Theodore Speroff, Michael E. Matheny |
AMIA | 7 |
| 2013 | Comparative analysis of pharmacovigilance methods in the detection of adverse drug reactions using electronic medical recordsabstractOBJECTIVE: Medication safety requires that each drug be monitored throughout its market life as early detection of adverse drug reactions (ADRs) can lead to alerts that prevent patient harm. Recently, electronic medical records (EMRs) have emerged as a valuable resource for pharmacovigilance. This study examines the use of retrospective medication orders and inpatient laboratory results documented in the EMR to identify ADRs. METHODS: Using 12 years of EMR data from Vanderbilt University Medical Center (VUMC), we designed a study to correlate abnormal laboratory results with specific drug administrations by comparing the outcomes of a drug-exposed group and a matched unexposed group. We assessed the relative merits of six pharmacovigilance measures used in spontaneous reporting systems (SRSs): proportional reporting ratio (PRR), reporting OR (ROR), Yule's Q (YULE), the χ(2) test (CHI), Bayesian confidence propagation neural networks (BCPNN), and a gamma Poisson shrinker (GPS). RESULTS: We systematically evaluated the methods on two independently constructed reference standard datasets of drug-event pairs. The dataset of Yoon et al contained 470 drug-event pairs (10 drugs and 47 laboratory abnormalities). Using VUMC's EMR, we created another dataset of 378 drug-event pairs (nine drugs and 42 laboratory abnormalities). Evaluation on our reference standard showed that CHI, ROR, PRR, and YULE all had the same F score (62%). When the reference standard of Yoon et al was used, ROR had the best F score of 68%, with 77% precision and 61% recall. CONCLUSIONS: Results suggest that EMR-derived laboratory measurements and medication orders can help to validate previously reported ADRs, and detect new ADRs. Eugenia R. McPeek Hinz, Michael E. Matheny, Joshua C. Denny, Jonathan S. Schildcrout, Randolph A. Miller, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 3 |
| 2012 | Data Harmonization Barriers for Post Marketing Surveillance of Medications
Fern FitzHenry, Frederic S. Resnic, Arijit Basu, Susan Robbins, Kenneth Nunes, Robert El-Kareh, Grace Kuo, Michael E. Matheny |
AMIA | 8 |
| 2012 | Evaluation of the RapTAT Automated Annotation System for Identifying Post-Traumatic Stress Disorder-Related Concepts in Clinical Notes
Glenn T. Gobbel, Ruth M. Reeves, Shrimalini Jayaramaraja, Svetlana K. Eden, Fern FitzHenry, Theodore Speroff, Steven H. Brown, Stephen Luther, Maryan Zirkle, Dezon Finch, James A. McCart, Michael E. Matheny |
AMIA | 12 |
| 2012 | National Post-Marketing Surveillance of Embolic Protection Devices in the Veterans Administration CART Program
Michael E. Matheny, Fern FitzHenry, Arijit Basu, Susan Robbins, Richard Cope, Mary Plomondon, Thomas Maddox, Thomas Tsai, John Rumsfeld, Frederic S. Resnic |
AMIA | 1 |
| 2012 | Developing the Veterans Health Administration ETL for the OMOP Common Data Model Version 3
Lalit Nookala, Michael E. Matheny, Frederic S. Resnic, Susan Robbins, Fern FitzHenry |
AMIA | 2 |
| 2012 | Who Said It? Establishing Professional Attribution among Authors of Veterans' Electronic Health Records
Ruth M. Reeves, Fern FitzHenry, Steven H. Brown, Kristen Kotter, Glenn T. Gobbel, Diane Montella, Harvey J. Murff, Theodore Speroff, Michael E. Matheny |
AMIA | 9 |
| 2012 | Developing an Integrated Temporal Reasoning System with TimeML and UMLS Concepts
Ruth M. Reeves, Zhao Zou, Glenn T. Gobbel, Shrimalini Jayaramaraja, Steven H. Brown, Michael E. Matheny, Theodore Speroff |
AMIA | 6 |
| 2012 | Large-scale prediction of adverse drug reactions using chemical, biological, and phenotypic properties of drugsabstractOBJECTIVE: Adverse drug reaction (ADR) is one of the major causes of failure in drug development. Severe ADRs that go undetected until the post-marketing phase of a drug often lead to patient morbidity. Accurate prediction of potential ADRs is required in the entire life cycle of a drug, including early stages of drug design, different phases of clinical trials, and post-marketing surveillance. METHODS: Many studies have utilized either chemical structures or molecular pathways of the drugs to predict ADRs. Here, the authors propose a machine-learning-based approach for ADR prediction by integrating the phenotypic characteristics of a drug, including indications and other known ADRs, with the drug's chemical structures and biological properties, including protein targets and pathway information. A large-scale study was conducted to predict 1385 known ADRs of 832 approved drugs, and five machine-learning algorithms for this task were compared. RESULTS: This evaluation, based on a fivefold cross-validation, showed that the support vector machine algorithm outperformed the others. Of the three types of information, phenotypic data were the most informative for ADR prediction. When biological and phenotypic features were added to the baseline chemical information, the ADR prediction model achieved significant improvements in area under the curve (from 0.9054 to 0.9524), precision (from 43.37% to 66.17%), and recall (from 49.25% to 63.06%). Most importantly, the proposed model successfully predicted the ADRs associated with withdrawal of rofecoxib and cerivastatin. CONCLUSION: The results suggest that phenotypic information on drugs is valuable for ADR prediction. Moreover, they demonstrate that different models that combine chemical, biological, or phenotypic information can be built from approved drugs, and they have the potential to detect clinically important ADRs in both preclinical and post-marketing phases. Yonghui Wu 0001, Yukun Chen 0001, Jingchun Sun, Zhongming Zhao, Xue-wen Chen 0001, Michael E. Matheny, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 7 |
| 2012 | iDASH: integrating data for analysis, anonymization, and sharingabstractiDASH (integrating data for analysis, anonymization, and sharing) is the newest National Center for Biomedical Computing funded by the NIH. It focuses on algorithms and tools for sharing data in a privacy-preserving manner. Foundational privacy technology research performed within iDASH is coupled with innovative engineering for collaborative tool development and data-sharing capabilities in a private Health Insurance Portability and Accountability Act (HIPAA)-certified cloud. Driving Biological Projects, which span different biological levels (from molecules to individuals to populations) and focus on various health conditions, help guide research and development within this Center. Furthermore, training and dissemination efforts connect the Center with its stakeholders and educate data owners and data consumers on how to share and use clinical and biological data. Through these various mechanisms, iDASH implements its goal of providing biomedical and behavioral researchers with access to data, software, and a high-performance computing environment, thus enabling them to generate and test new hypotheses. Lucila Ohno-Machado, Vineet Bafna, Aziz A. Boxwala, Brian E. Chapman, Wendy W. Chapman, Kamalika Chaudhuri, Michele E. Day, Claudiu Farcas, Nathaniel D. Heintzman, Xiaoqian Jiang, Hyeon-Eui Kim, Jihoon Kim 0001, Michael E. Matheny, Frederic S. Resnic, Staal Amund Vinterbo |
J. Am. Medical Informatics Assoc. | 13 |
| 2009 | Detection of Blood Culture Bacterial Contamination using Natural Language Processing
Michael E. Matheny, Fern FitzHenry, Theodore Speroff, Jacob Hathaway, Harvey J. Murff, Steven H. Brown, Elliot M. Fielstein, Robert S. Dittus, Peter L. Elkin |
AMIA | 1 |
| 2009 | Research Paper: Impact of Non-interruptive Medication Laboratory Monitoring Alerts in Ambulatory CareabstractOBJECTIVE: Interruptive alerts within electronic applications can cause "alert fatigue" if they fire too frequently or are clinically reasonable only some of the time. We assessed the impact of non-interruptive, real-time medication laboratory alerts on provider lab test ordering. DESIGN: We enrolled 22 outpatient practices into a prospective, randomized, controlled trial. Clinics either used the existing system or received on-screen recommendations for baseline laboratory tests when prescribing new medications. Since the warnings were non-interruptive, providers did not have to act upon or acknowledge the notification to complete a medication request. MEASUREMENTS: Data were collected each time providers performed suggested laboratory testing within 14 days of a new prescription order. Findings were adjusted for patient and provider characteristics as well as patient clustering within clinics. RESULTS: Among 12 clinics with 191 providers in the control group and 10 clinics with 175 providers in the intervention group, there were 3673 total events where baseline lab tests would have been advised: 1988 events in the control group and 1685 in the intervention group. In the control group, baseline labs were requested for 771 (39%) of the medications. In the intervention group, baseline labs were ordered by clinicians in 689 (41%) of the cases. Overall, no significant association existed between the intervention and the rate of ordering appropriate baseline laboratory tests. CONCLUSION: We found that non-interruptive medication laboratory monitoring alerts were not effective in improving receipt of recommended baseline laboratory test monitoring for medications. Further work is necessary to optimize compliance with non-critical recommendations. Helen G. Lo, Michael E. Matheny, Diane L. Seger, David W. Bates, Tejal K. Gandhi |
J. Am. Medical Informatics Assoc. | 2 |
| 2009 | A comparison of methods for assessing penetrating trauma on retrospective multi-center data
Bilal A. Ahmed, Michael E. Matheny, Phillip L. Rice, John R. Clarke, Omolola Ogunyemi |
J. Biomed. Informatics | 2 |
| 2008 | Research Paper: A Randomized Trial of Electronic Clinical Reminders to Improve Medication Laboratory MonitoringabstractOBJECTIVE: Recommendations for routine laboratory monitoring to reduce the risk of adverse medication events are not consistently followed. We evaluated the impact of electronic reminders delivered to primary care physicians on rates of appropriate routine medication laboratory monitoring. DESIGN: We enrolled 303 primary care physicians caring for 1,922 patients across 20 ambulatory clinics that had at least one overdue routine laboratory test for a given medication between January and June 2004. Clinics were randomized so that physicians received either usual care or electronic reminders at the time of office visits focused on potassium, creatinine, liver function, thyroid function, and therapeutic drug levels. MEASUREMENTS: Primary outcomes were the receipt of recommended laboratory monitoring within 14 days following an outpatient clinic visit. The effect of the intervention was assessed for each reminder after adjusting for clustering within clinics, as well as patient and provider characteristics. RESULTS: Medication-laboratory monitoring non-compliance ranged from 1.6% (potassium monitoring with potassium-supplement use) to 6.3% (liver function monitoring with HMG CoA Reductase Inhibitor use). Rates of appropriate laboratory monitoring following an outpatient visit ranged from 14% (therapeutic drug levels) to 64% (potassium monitoring with potassium-sparing diuretic use). Reminders for appropriate laboratory monitoring had no impact on rates of receiving appropriate testing for creatinine, potassium, liver function, renal function, or therapeutic drug level monitoring. CONCLUSION: We identified high rates of appropriate laboratory monitoring, and electronic reminders did not significantly improve these monitoring rates. Future studies should focus on settings with lower baseline adherence rates and alternate drug-laboratory combinations. Michael E. Matheny, Thomas D. Sequist, Andrew C. Seger, Julie M. Fiskio, Michael Sperling, Donald Bugbee, David W. Bates, Tejal K. Gandhi |
J. Am. Medical Informatics Assoc. | 1 |
| 2007 | Rare Adverse Event Monitoring of Medical Devices with the Use of an Automated Surveillance Tool
Michael E. Matheny, Nipun Arora, Lucila Ohno-Machado, Frederic S. Resnic |
AMIA | 1 |
| 2007 | Effects of SVM parameter optimization on discrimination and calibration for post-procedural PCI mortality
Michael E. Matheny, Frederic S. Resnic, Nipun Arora, Lucila Ohno-Machado |
J. Biomed. Informatics | 1 |
| 2006 | Research Paper: Monitoring Device Safety in Interventional CardiologyabstractOBJECTIVE: A variety of postmarketing surveillance strategies to monitor the safety of medical devices have been supported by the U.S. Food and Drug Administration, but there are few systems to automate surveillance. Our objective was to develop a system to perform real-time monitoring of safety data using a variety of process control techniques. DESIGN: The Web-based Data Extraction and Longitudinal Time Analysis (DELTA) system imports clinical data in real-time from an electronic database and generates alerts for potentially unsafe devices or procedures. The statistical techniques used are statistical process control (SPC), logistic regression (LR), and Bayesian updating statistics (BUS). MEASUREMENTS: We selected in-patient mortality following implantation of the Cypher drug-eluting coronary stent to evaluate our system. Data from the University of Michigan Consortium Bare-Metal Stent Study was used to calculate the event rate alerting boundaries. Data analysis was performed on local catheterization data from Brigham and Women's Hospital from July 1, 2003, shortly after the Cypher release, to December 31, 2004, including 2,270 cases with 27 observed deaths. RESULTS: The single-stratum SPC had alerts in months 4 and 10. The multistrata SPC had alerts in months 5, 10, and 18 in the moderate-risk stratum, and months 1, 4, 7, and 10 in the high-risk stratum. The only cumulative alerts were in the first month for the high-risk stratum of the multistrata SPC. The LR method showed no monthly or cumulative alerts. The BUS method showed an alert in the first month for the high-risk stratum. CONCLUSION: The system performed adequately within the Brigham and Women's Hospital Intranet environment based on the design goals. All three cumulative methods agreed that the overall observed event rates were not significantly higher for the new medical device than for a closely related medical device and were consistent with the observation that the initial concerns about this device dissipated as more data accumulated. Michael E. Matheny, Lucila Ohno-Machado, Frederic S. Resnic |
J. Am. Medical Informatics Assoc. | 1 |
| 2005 | Exploration of a Bayesian Updating Tool to Provide Real-Time Safety Monitoring for New Medical Devices
Michael E. Matheny, Lucila Ohno-Machado, Frederic S. Resnic |
AMIA | 1 |
| 2005 | Evaluating the Discriminatory Power of a Computer-based System for Assessing Penetrating Trauma on Retrospective Multi-Center Data
Michael E. Matheny, Omolola Ogunyemi, Phillip L. Rice, John R. Clarke |
AMIA | 1 |
| 2005 | Discrimination and calibration of mortality risk prediction models in interventional cardiology
Michael E. Matheny, Lucila Ohno-Machado, Frederic S. Resnic |
J. Biomed. Informatics | 1 |