VLDB 2026 Research / reviewers in the wild / expert
Patrick B. Ryan
dblp:67/9194
· DBLP profile ↗
54ranked-venue papers
0as first author
18since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 54 · 18 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Evaluation of the impact of defining observable time in real-world data on outcome incidenceabstractOBJECTIVE: In real-world data (RWD), defining the observation period-the time during which a patient is considered observable-is critical for estimating incidence rates (IRs) and other outcomes. Yet, in the absence of explicit enrollment information, this period must often be inferred, introducing potential bias. MATERIALS AND METHODS: This study evaluates methods for defining observation periods and their impact on IR estimates across multiple database types. We applied 3 methods for defining observation periods: (1) a persistence + surveillance window approach, (2) an age- and gender-adjusted method based on time between healthcare events, and (3) the min/max method. These were tested across 11 RWD databases, including both enrollment-based and encounter-based sources. Enrollment time was used as the reference standard in eligible databases. To assess the impact on epidemiologic results, we replicated a prior study of adverse event incidence, comparing IRs and calculating mean squared error between methods. RESULTS: Incidence rates decreased as observation periods lengthened, driven by increases in the person-time denominator. The persistence + surveillance method produced estimates closest to enrollment-based rates when appropriately balanced. The min/max approach yielded inconsistent results, particularly in encounter-based databases, with greater error observed in databases with longer time spans. DISCUSSION: These findings suggest that assumptions about data completeness and population observability significantly affect incidence estimates. Observation period definitions substantially influence outcome measurement in RWD studies. CONCLUSION: Standardized, transparent approaches are necessary to ensure valid, reproducible results-especially in databases lacking defined enrollment. Clair Blacketer, Frank J. DeFalco, Mitchell Conover, Patrick B. Ryan, Martijn J. Schuemie, Peter R. Rijnbeek |
J. Am. Medical Informatics Assoc. | 4 |
| 2025 | Objective study validity diagnostics: a framework requiring pre-specified, empirical verification to increase trust in the reliability of real-world evidenceabstractOBJECTIVE: Propose a framework to empirically evaluate and report validity of findings from observational studies using pre-specified objective diagnostics, increasing trust in real-world evidence (RWE). MATERIALS AND METHODS: The framework employs objective diagnostic measures to assess the appropriateness of study designs, analytic assumptions, and threats to validity in generating reliable evidence addressing causal questions. Diagnostic evaluations should be interpreted before the unblinding of study results or, alternatively, only unblind results from analyses that pass pre-specified thresholds. We provide a conceptual overview of objective diagnostic measures and demonstrate their impact on the validity of RWE from a large-scale comparative new-user study of various antihypertensive medications. We evaluated expected absolute systematic error (EASE) before and after applying diagnostic thresholds, using a large set of negative control outcomes. RESULTS: Applying objective diagnostics reduces bias and improves evidence reliability in observational studies. Among 11 716 analyses (EASE = 0.38), 13.9% met pre-specified diagnostic thresholds which reduced EASE to zero. Objective diagnostics provide a comprehensive and empirical set of tests that increase confidence when passed and raise doubts when failed. DISCUSSION: The increasing use of real-world data presents a scientific opportunity; however, the complexity of the evidence generation process poses challenges for understanding study validity and trusting RWE. Deploying objective diagnostics is crucial to reducing bias and improving reliability in RWE generation. Under ideal conditions, multiple study designs pass diagnostics and generate consistent results, deepening understanding of causal relationships. Open-source, standardized programs can facilitate implementation of diagnostic analyses. CONCLUSION: Objective diagnostics are a valuable addition to the RWE generation process. Mitchell Conover, Patrick B. Ryan, Yong Chen 0016, Marc A. Suchard, George Hripcsak, Martijn J. Schuemie |
J. Am. Medical Informatics Assoc. | 2 |
| 2025 | CLEAR: A vision to support clinical evidence lifecycle with continuous learningabstractHuman knowledge of diseases, treatments, and prevention techniques is constantly evolving. The generation of clinical evidence using randomized controlled trials on human subjects occurs notably slowly and inefficiently. The Learning Health System (LHS) has been proposed to facilitate the continuous improvement of individual and population health through a cycle of knowledge, practice, and data. However, the gap between the demand for high-quality evidence to support clinical decisions and the available evidence continues to enlarge. While the current LHS vision articulates the integration of Real-World Data (RWD), the rapid generation of RWD often outpaces the rate of effective evidence synthesis and implementation. Considering this, we propose a new framework that more effectively leverages RWD to support the entire clinical evidence lifecycle through a continuous learning mechanism. This framework, powered by modern data science and informatics, offers enhanced scalability and efficiency. In this vision, specifically, RWD is integrated into the clinical evidence lifecycle via four closed feedback loops: 1) guiding research prioritization and study design, 2) facilitating clinical guideline development, 3) assisting guideline evaluation, and 4) supporting shared decision-making. Our framework enables rapid responsiveness to emerging health data and evolving healthcare needs, timely development of clinical guidelines to optimize clinical recommendations, and sustained improvements in clinical practice and patient outcomes. This vision calls for informatics support for an efficient, scalable, and stakeholder-aware clinical evidence lifecycle. Yilu Fang, Fangyi Chen, George Hripcsak, Yifan Peng 0002, Patrick B. Ryan, Chunhua Weng |
J. Biomed. Informatics | 6 |
| 2025 | Evaluating the Bias, type I error and statistical power of the prior Knowledge-Guided integrated likelihood estimation (PIE) for bias reduction in EHR based association studiesabstract• Question: How does PIE perform in various types of real-world scenarios, in terms of estimation and hypothesis testing? • Findings: Under non-differential misclassification, PIE had a smaller bias in estimated associations compared to the naïve method, but it had similar type I error and power. • The bias reduction of PIE was superior when the prior distribution of sensitivity and specificity of the phenotyping algorithm is more accurate (i.e., close to the true operating characteristics of the phenotyping algorithm). The impact of prior is relatively small when the outcome has low prevalence and is larger when the outcome is common. • PIE can effectively reduce the bias due to phenotyping error under a wide spectrum of real-world settings. However, its main advantage is in the reduction of bias in estimation but not in hypothesis testing. Binary outcomes in electronic health records (EHR) derived using automated phenotype algorithms may suffer from phenotyping error, resulting in bias in association estimation. Huang et al. [1] proposed the Prior Knowledge-Guided Integrated Likelihood Estimation (PIE) method to mitigate the estimation bias, however, their investigation focused on point estimation without statistical inference, and the evaluation of PIE therein using simulation was a proof-of-concept with only a limited scope of scenarios. This study aims to comprehensively assess PIE’s performance including (1) how well PIE performs under a wide spectrum of operating characteristics of phenotyping algorithms under real-world scenarios (e. g., low prevalence, low sensitivity, high specificity); (2) beyond point estimation, how much variation of the PIE estimator was introduced by the prior distribution; and (3) from a hypothesis testing point of view, if PIE improves type I error and statistical power relative to the naïve method (i.e., ignoring the phenotyping error). Synthetic data and use-case analysis were utilized to evaluate PIE. The synthetic data were generated under diverse outcome prevalence, phenotyping algorithm sensitivity, and association effect sizes. Simulation studies compared PIE under different prior distributions with the naïve method, assessing bias, variance, type I error, and power. Use-case analysis compared the performance of PIE and the naïve method in estimating the association of multiple predictors with COVID-19 infection. PIE exhibited reduced bias compared to the naïve method across varied simulation settings, with comparable type I error and power. As the effect size became larger, the bias reduced by PIE was larger. PIE has superior performance when prior distributions aligned closely with true phenotyping algorithm characteristics. Impact of prior quality was minor for low-prevalence outcomes but large for common outcomes. In use-case analysis, PIE maintains a relatively accurate estimation across different scenarios, particularly outperforming the naïve approach under large effect sizes. PIE effectively mitigates estimation bias in a wide spectrum of real-world settings, particularly with accurate prior information. Its main benefit lies in bias reduction rather than hypothesis testing. The impact of the prior is small for low-prevalence outcomes. Naimin Jing, Jiayi Tong, James Weaver, Patrick B. Ryan, Hua Xu 0001, Yong Chen 0016 |
J. Biomed. Informatics | 5 |
| 2024 | OHDSI Standardized Vocabularies - a large-scale centralized reference ontology for international data harmonizationabstractIMPORTANCE: The Observational Health Data Sciences and Informatics (OHDSI) is the largest distributed data network in the world encompassing more than 331 data sources with 2.1 billion patient records across 34 countries. It enables large-scale observational research through standardizing the data into a common data model (CDM) (Observational Medical Outcomes Partnership [OMOP] CDM) and requires a comprehensive, efficient, and reliable ontology system to support data harmonization. MATERIALS AND METHODS: We created the OHDSI Standardized Vocabularies-a common reference ontology mandatory to all data sites in the network. It comprises imported and de novo-generated ontologies containing concepts and relationships between them, and the praxis of converting the source data to the OMOP CDM based on these. It enables harmonization through assigned domains according to clinical categories, comprehensive coverage of entities within each domain, support for commonly used international coding schemes, and standardization of semantically equivalent concepts. RESULTS: The OHDSI Standardized Vocabularies comprise over 10 million concepts from 136 vocabularies. They are used by hundreds of groups and several large data networks. More than 8600 users have performed 50 000 downloads of the system. This open-source resource has proven to address an impediment of large-scale observational research-the dependence on the context of source data representation. With that, it has enabled efficient phenotyping, covariate construction, patient-level prediction, population-level estimation, and standard reporting. DISCUSSION AND CONCLUSION: OHDSI has made available a comprehensive, open vocabulary system that is unmatched in its ability to support global observational research. We encourage researchers to exploit it and contribute their use cases to this dynamic resource. Christian G. Reich, Anna Ostropolets, Patrick B. Ryan, Peter R. Rijnbeek, Martijn J. Schuemie, Dmitry Dymshyts, George Hripcsak |
J. Am. Medical Informatics Assoc. | 3 |
| 2023 | Reproducible variability: assessing investigator discordance across 9 research teams attempting to reproduce the same observational studyabstractOBJECTIVE: Observational studies can impact patient care but must be robust and reproducible. Nonreproducibility is primarily caused by unclear reporting of design choices and analytic procedures. This study aimed to: (1) assess how the study logic described in an observational study could be interpreted by independent researchers and (2) quantify the impact of interpretations' variability on patient characteristics. MATERIALS AND METHODS: Nine teams of highly qualified researchers reproduced a cohort from a study by Albogami et al. The teams were provided the clinical codes and access to the tools to create cohort definitions such that the only variable part was their logic choices. We executed teams' cohort definitions against the database and compared the number of subjects, patient overlap, and patient characteristics. RESULTS: On average, the teams' interpretations fully aligned with the master implementation in 4 out of 10 inclusion criteria with at least 4 deviations per team. Cohorts' size varied from one-third of the master cohort size to 10 times the cohort size (2159-63 619 subjects compared to 6196 subjects). Median agreement was 9.4% (interquartile range 15.3-16.2%). The teams' cohorts significantly differed from the master implementation by at least 2 baseline characteristics, and most of the teams differed by at least 5. CONCLUSIONS: Independent research teams attempting to reproduce the study based on its free-text description alone produce different implementations that vary in the population size and composition. Sharing analytical code supported by a common data model and open-source tools allows reproducing a study unambiguously thereby preserving initial design choices. Anna Ostropolets, Yasser Albogami, Mitchell Conover, Juan M. Banda, William A. Baumgartner Jr., Clair Blacketer, Priyamvada Desai, Scott L. DuVall, Stephen P. Fortin, James P. Gilbert, Asieh Golozar, Joshua Ide, Andrew S. Kanter, David M. Kern, Chungsoo Kim, Lana Y. H. Lai, Kristine E. Lynch, Evan P. Minty, Maria Inês Neves, Ding Quan Ng, Tontel Obene, Victor Pera, Nicole Pratt, Gowtham Rao, Nadav Rappoport, Ines Reinecke, Paola Saroufim, Azza Shoaibi, Katherine Simon, Marc A. Suchard, Joel N. Swerdel, Erica A. Voss, James Weaver, Linying Zhang, George Hripcsak, Patrick B. Ryan |
J. Am. Medical Informatics Assoc. | 38 |
| 2023 | Scalable and interpretable alternative to chart review for phenotype evaluation using standardized structured data from electronic health recordsabstractOBJECTIVES: Chart review as the current gold standard for phenotype evaluation cannot support observational research on electronic health records and claims data sources at scale. We aimed to evaluate the ability of structured data to support efficient and interpretable phenotype evaluation as an alternative to chart review. MATERIALS AND METHODS: We developed Knowledge-Enhanced Electronic Profile Review (KEEPER) as a phenotype evaluation tool that extracts patient's structured data elements relevant to a phenotype and presents them in a standardized fashion following clinical reasoning principles. We evaluated its performance (interrater agreement, intermethod agreement, accuracy, and review time) compared to manual chart review for 4 conditions using randomized 2-period, 2-sequence crossover design. RESULTS: Case ascertainment with KEEPER was twice as fast compared to manual chart review. 88.1% of the patients were classified concordantly using charts and KEEPER, but agreement varied depending on the condition. Missing data and differences in interpretation accounted for most of the discrepancies. Pairs of clinicians agreed in case ascertainment in 91.2% of the cases when using KEEPER compared to 76.3% when using charts. Patient classification aligned with the gold standard in 88.1% and 86.9% of the cases respectively. CONCLUSION: Structured data can be used for efficient and interpretable phenotype evaluation if they are limited to relevant subset and organized according to the clinical reasoning principles. A system that implements these principles can achieve noninferior performance compared to chart review at a fraction of time. Anna Ostropolets, George Hripcsak, Syed A. Husain, Lauren R. Richter, Matthew E. Spotnitz, Ahmed Elhussein, Patrick B. Ryan |
J. Am. Medical Informatics Assoc. | 7 |
| 2023 | Representing and utilizing clinical textual data for real world studies: An OHDSI approach
Vipina Kuttichi Keloth, Juan M. Banda, Michael J. Gurley, Paul M. Heider, Georgina Kennedy, Timothy A. Miller, Karthik Natarajan, Olga V. Patterson, Yifan Peng 0002, Kalpana Raja, Ruth M. Reeves, Masoud Rouhizadeh, Jianlin Shi, Yanshan Wang, Wei-Qi Wei, Andrew E. Williams, Rui Zhang 0028, Rimma Belenkaya, Christian G. Reich, Clair Blacketer, Patrick B. Ryan, George Hripcsak, Noémie Elhadad, Hua Xu 0001 |
J. Biomed. Informatics | 24 |
| 2023 | Padé approximant meets federated learning: A nearly lossless, one-shot algorithm for evidence synthesis in distributed research networks with rare outcomes
Martijn J. Schuemie, Marc A. Suchard, Patrick B. Ryan, George Hripcsak, Charles A. Rohde, Yong Chen 0016 |
J. Biomed. Informatics | 4 |
| 2022 | Phenotyping in distributed data networks: selecting the right codes for the right patients
Anna Ostropolets, Patrick B. Ryan, George Hripcsak |
AMIA | 2 |
| 2022 | PheValuator 2.0: Methodological improvements for the PheValuator approach to semi-automated phenotype algorithm evaluationabstractPURPOSE: Phenotype algorithms are central to performing analyses using observational data. These algorithms translate the clinical idea of a health condition into an executable set of rules allowing for queries of data elements from a database. PheValuator, a software package in the Observational Health Data Sciences and Informatics (OHDSI) tool stack, provides a method to assess the performance characteristics of these algorithms, namely, sensitivity, specificity, and positive and negative predictive value. It uses machine learning to develop predictive models for determining a probabilistic gold standard of subjects for assessment of cases and non-cases of health conditions. PheValuator was developed to complement or even replace the traditional approach of algorithm validation, i.e., by expert assessment of subject records through chart review. Results in our first PheValuator paper suggest a systematic underestimation of the PPV compared to previous results using chart review. In this paper we evaluate modifications made to the method designed to improve its performance. METHODS: The major changes to PheValuator included allowing all diagnostic conditions, clinical observations, drug prescriptions, and laboratory measurements to be included as predictors within the modeling process whereas in the prior version there were significant restrictions on the included predictors. We also have allowed for the inclusion of the temporal relationships of the predictors in the model. To evaluate the performance of the new method, we compared the results from the new and original methods against results found from the literature using traditional validation of algorithms for 19 phenotypes. We performed these tests using data from five commercial databases. RESULTS: In the assessment aggregating all phenotype algorithms, the median difference between the PheValuator estimate and the gold standard estimate for PPV was reduced from -21 (IQR -34, -3) in Version 1.0 to 4 (IQR -3, 15) using Version 2.0. We found a median difference in specificity of 3 (IQR 1, 4.25) for Version 1.0 and 3 (IQR 1, 4) for Version 2.0. The median difference between the two versions of PheValuator and the gold standard for estimates of sensitivity was reduced from -39 (-51, -20) to -16 (-34, -6). CONCLUSION: PheValuator 2.0 produces estimates for the performance characteristics for phenotype algorithms that are significantly closer to estimates from traditional validation through chart review compared to version 1.0. With this tool in researcher's toolkits, methods, such as quantitative bias analysis, may now be used to improve the reliability and reproducibility of research studies using observational data. Joel N. Swerdel, Martijn J. Schuemie, Gayle Murray, Patrick B. Ryan |
J. Biomed. Informatics | 4 |
| 2021 | PHenotype Observed Entity Baseline Endorsements (PHOEBE) - recommender system for concept selection in phenotype algorithm development
Anna Ostropolets, Patrick B. Ryan, George Hripcsak |
AMIA | 2 |
| 2021 | Increasing trust in real-world evidence through evaluation of observational data qualityabstractOBJECTIVE: Advances in standardization of observational healthcare data have enabled methodological breakthroughs, rapid global collaboration, and generation of real-world evidence to improve patient outcomes. Standardizations in data structure, such as use of common data models, need to be coupled with standardized approaches for data quality assessment. To ensure confidence in real-world evidence generated from the analysis of real-world data, one must first have confidence in the data itself. MATERIALS AND METHODS: We describe the implementation of check types across a data quality framework of conformance, completeness, plausibility, with both verification and validation. We illustrate how data quality checks, paired with decision thresholds, can be configured to customize data quality reporting across a range of observational health data sources. We discuss how data quality reporting can become part of the overall real-world evidence generation and dissemination process to promote transparency and build confidence in the resulting output. RESULTS: The Data Quality Dashboard is an open-source R package that reports potential quality issues in an OMOP CDM instance through the systematic execution and summarization of over 3300 configurable data quality checks. DISCUSSION: Transparently communicating how well common data model-standardized databases adhere to a set of quality measures adds a crucial piece that is currently missing from observational research. CONCLUSION: Assessing and improving the quality of our data will inherently improve the quality of the evidence we generate. Clair Blacketer, Frank J. DeFalco, Patrick B. Ryan, Peter R. Rijnbeek |
J. Am. Medical Informatics Assoc. | 3 |
| 2021 | Towards clinical data-driven eligibility criteria optimization for interventional COVID-19 clinical trialsabstractOBJECTIVE: This research aims to evaluate the impact of eligibility criteria on recruitment and observable clinical outcomes of COVID-19 clinical trials using electronic health record (EHR) data. MATERIALS AND METHODS: On June 18, 2020, we identified frequently used eligibility criteria from all the interventional COVID-19 trials in ClinicalTrials.gov (n = 288), including age, pregnancy, oxygen saturation, alanine/aspartate aminotransferase, platelets, and estimated glomerular filtration rate. We applied the frequently used criteria to the EHR data of COVID-19 patients in Columbia University Irving Medical Center (CUIMC) (March 2020-June 2020) and evaluated their impact on patient accrual and the occurrence of a composite endpoint of mechanical ventilation, tracheostomy, and in-hospital death. RESULTS: There were 3251 patients diagnosed with COVID-19 from the CUIMC EHR included in the analysis. The median follow-up period was 10 days (interquartile range 4-28 days). The composite events occurred in 18.1% (n = 587) of the COVID-19 cohort during the follow-up. In a hypothetical trial with common eligibility criteria, 33.6% (690/2051) were eligible among patients with evaluable data and 22.2% (153/690) had the composite event. DISCUSSION: By adjusting the thresholds of common eligibility criteria based on the characteristics of COVID-19 patients, we could observe more composite events from fewer patients. CONCLUSIONS: This research demonstrated the potential of using the EHR data of COVID-19 patients to inform the selection of eligibility criteria and their thresholds, supporting data-driven optimization of participant selection towards improved statistical power of COVID-19 trials. Jae Hyun Kim, Casey N. Ta, Cong Liu 0020, Cynthia Sung 0002, Alex M. Butler, Latoya A. Stewart, Lyudmila Ena, James R. Rogers, Anna Ostropolets, Patrick B. Ryan, Hao Liu 0054, Shing M. Lee, Mitchell S. V. Elkind, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 11 |
| 2021 | Data Consult Service: Can we use observational data to address immediate clinical needs?abstractOBJECTIVE: A number of clinical decision support tools aim to use observational data to address immediate clinical needs, but few of them address challenges and biases inherent in such data. The goal of this article is to describe the experience of running a data consult service that generates clinical evidence in real time and characterize the challenges related to its use of observational data. MATERIALS AND METHODS: In 2019, we launched the Data Consult Service pilot with clinicians affiliated with Columbia University Irving Medical Center. We created and implemented a pipeline (question gathering, data exploration, iterative patient phenotyping, study execution, and assessing validity of results) for generating new evidence in real time. We collected user feedback and assessed issues related to producing reliable evidence. RESULTS: We collected 29 questions from 22 clinicians through clinical rounds, emails, and in-person communication. We used validated practices to ensure reliability of evidence and answered 24 of them. Questions differed depending on the collection method, with clinical rounds supporting proactive team involvement and gathering more patient characterization questions and questions related to a current patient. The main challenges we encountered included missing and incomplete data, underreported conditions, and nonspecific coding and accurate identification of drug regimens. CONCLUSIONS: While the Data Consult Service has the potential to generate evidence and facilitate decision making, only a portion of questions can be answered in real time. Recognizing challenges in patient phenotyping and designing studies along with using validated practices for observational research are mandatory to produce reliable evidence. Anna Ostropolets, Philip Zachariah, Patrick B. Ryan, George Hripcsak |
J. Am. Medical Informatics Assoc. | 3 |
| 2021 | Erratum to: Large-Scale Evidence Generation and Evaluation across a Network of Databases (LEGEND): Assessing Validity Using Hypertension as a Case StudyabstractJournal of the American Medical Informatics Association, 27(8), 2020, 1268–1277; doi: 10.1093/jamia/ocaa124 Upon the original publication of this article, several author corrections to the reference section were inadvertently left out, making literature referencing in the article inaccurate. This error has now been corrected online. The publisher apologises for the error. Martijn J. Schuemie, Patrick B. Ryan, Nicole Pratt, Seng Chan You, Harlan M. Krumholz, David Madigan, George Hripcsak, Marc A. Suchard |
J. Am. Medical Informatics Assoc. | 2 |
| 2021 | From clinical trials to clinical practice: How long are drugs tested and then used by patients?abstractOBJECTIVE: Evidence is scarce regarding the safety of long-term drug use, especially for drugs treating chronic diseases. To bridge this knowledge gap, this research investigated the differences in drug exposure between clinical trials and clinical practice. MATERIALS AND METHODS: We extracted drug follow-up times from clinical trials in ClinicalTrials.gov and compared the difference between clinical trials and real-world usage data for 914 drugs taken by 96 645 927 patients. RESULTS: A total of 17.5% of drugs had longer median exposure in practice than in trials, 6% of patients had extended exposure to at least 1 drug, and drugs treating nervous system disorders and cardiovascular diseases were the most common among drugs with high rates of extended exposure. CONCLUSIONS: For most of patients, the drug use length is shorter than the tested length in clinical trials. Still, a remarkable number of patients experienced extended drug exposure, particularly for drugs treating nervous system disorders or cardiovascular disorders. Chi Yuan, Patrick B. Ryan, Casey N. Ta, Jae Hyun Kim, Ziran Li, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 2 |
| 2021 | A conceptual framework for external validity
Amelia J. Averitt, Patrick B. Ryan, Chunhua Weng, Adler J. Perotte |
J. Biomed. Informatics | 2 |
| 2020 | Evaluation of Large-scale Propensity Score Modeling and Covariate Balance on Potential Unmeasured Confounding in Observational Research
Martijn J. Schuemie, Marc A. Suchard, Anna Ostropolets, Linying Zhang, Patrick B. Ryan, George Hripcsak |
AMIA | 6 |
| 2020 | Characterizing database granularity using SNOMED-CT hierarchy
Anna Ostropolets, Christian G. Reich, Patrick B. Ryan, Chunhua Weng, Anthony Molinaro, Frank J. DeFalco, Jitendra Jonnagaddala, Siaw-Teng Liaw, Hokyun Jeon, Rae Woong Park, Matthew E. Spotnitz, Karthik Natarajan, Kristin Kostka, George Argyriou, Robert T. Miller, Andrew E. Williams, Evan P. Minty, José D. Posada, George Hripcsak |
AMIA | 3 |
| 2020 | Characterization and Comparison of Embedding Algorithms for Phenotyping across a Network of Observational Databases
Harry Reyes Nieva, Krishna Kalluri, Tony Sun, Xinzhuo Jiang, Victor Alfonso Rodriguez, Patrick B. Ryan, Karthik Natarajan |
AMIA | 8 |
| 2020 | Phenotype Concept Set Construction from Concept Pair Likelihoods
Victor Alfonso Rodriguez, Tony Sun, Phyllis Thangaraj, Krishna Kalluri, Xinzhuo Jiang, Karthik Natarajan, Patrick B. Ryan, Anna Ostropolets |
AMIA | 8 |
| 2020 | Contemporary Use of Real World Data for Clinical Trial Conduct
James R. Rogers, Patrick B. Ryan, George Hripcsak, Chunhua Weng |
AMIA | 4 |
| 2020 | Large-scale evidence generation and evaluation across a network of databases (LEGEND): assessing validity using hypertension as a case studyabstractOBJECTIVES: To demonstrate the application of the Large-scale Evidence Generation and Evaluation across a Network of Databases (LEGEND) principles described in our companion article to hypertension treatments and assess internal and external validity of the generated evidence. MATERIALS AND METHODS: LEGEND defines a process for high-quality observational research based on 10 guiding principles. We demonstrate how this process, here implemented through large-scale propensity score modeling, negative and positive control questions, empirical calibration, and full transparency, can be applied to compare antihypertensive drug therapies. We assess internal validity through covariate balance, confidence-interval coverage, between-database heterogeneity, and transitivity of results. We assess external validity through comparison to direct meta-analyses of randomized controlled trials (RCTs). RESULTS: From 21.6 million unique antihypertensive new users, we generate 6 076 775 effect size estimates for 699 872 research questions on 12 946 treatment comparisons. Through propensity score matching, we achieve balance on all baseline patient characteristics for 75% of estimates, observe 95.7% coverage in our effect-estimate 95% confidence intervals, find high between-database consistency, and achieve transitivity in 84.8% of triplet hypotheses. Compared with meta-analyses of RCTs, our results are consistent with 28 of 30 comparisons while providing narrower confidence intervals. CONCLUSION: We find that these LEGEND results show high internal validity and are congruent with meta-analyses of RCTs. For these reasons we believe that evidence generated by LEGEND is of high quality and can inform medical decision-making where evidence is currently lacking. Subsequent publications will explore the clinical interpretations of this evidence. Martijn J. Schuemie, Patrick B. Ryan, Nicole Pratt, Seng Chan You, Harlan M. Krumholz, David Madigan, George Hripcsak, Marc A. Suchard |
J. Am. Medical Informatics Assoc. | 2 |
| 2020 | Principles of Large-scale Evidence Generation and Evaluation across a Network of Databases (LEGEND)abstractEvidence derived from existing health-care data, such as administrative claims and electronic health records, can fill evidence gaps in medicine. However, many claim such data cannot be used to estimate causal treatment effects because of the potential for observational study bias; for example, due to residual confounding. Other concerns include P hacking and publication bias. In response, the Observational Health Data Sciences and Informatics international collaborative launched the Large-scale Evidence Generation and Evaluation across a Network of Databases (LEGEND) research initiative. Its mission is to generate evidence on the effects of medical interventions using observational health-care databases while addressing the aforementioned concerns by following a recently proposed paradigm. We define 10 principles of LEGEND that enshrine this new paradigm, prescribing the generation and dissemination of evidence on many research questions at once; for example, comparing all treatments for a disease for many outcomes, thus preventing publication bias. These questions are answered using a prespecified and systematic approach, avoiding P hacking. Best-practice statistical methods address measured confounding, and control questions (research questions where the answer is known) quantify potential residual bias. Finally, the evidence is generated in a network of databases to assess consistency by sharing open-source analytics code to enhance transparency and reproducibility, but without sharing patient-level information. Here we detail the LEGEND principles and provide a generic overview of a LEGEND study. Our companion paper highlights an example study on the effects of hypertension treatments, and evaluates the internal and external validity of the evidence we generate. Martijn J. Schuemie, Patrick B. Ryan, Nicole Pratt, Seng Chan You, Harlan M. Krumholz, David Madigan, George Hripcsak, Marc A. Suchard |
J. Am. Medical Informatics Assoc. | 2 |
| 2020 | Adapting electronic health records-derived phenotypes to claims data: Lessons learned in using limited clinical data for phenotyping
Anna Ostropolets, Christian G. Reich, Patrick B. Ryan, Ning Shang 0004, George Hripcsak, Chunhua Weng |
J. Biomed. Informatics | 3 |
| 2019 | Learning Across a Healthcare Data Network to Improve Model Robustness and Evidence Reliability
Noémie Elhadad, Iñigo Urteaga, Alison Callahan, Jenna Reps, Patrick B. Ryan |
AMIA | 5 |
| 2019 | Challenges with quality of race and ethnicity data in observational databasesabstractOBJECTIVE: We sought to assess the quality of race and ethnicity information in observational health databases, including electronic health records (EHRs), and to propose patient self-recording as an improvement strategy. MATERIALS AND METHODS: We assessed completeness of race and ethnicity information in large observational health databases in the United States (Healthcare Cost and Utilization Project and Optum Labs), and at a single healthcare system in New York City serving a racially and ethnically diverse population. We compared race and ethnicity data collected via administrative processes with data recorded directly by respondents via paper surveys (National Health and Nutrition Examination Survey and Hospital Consumer Assessment of Healthcare Providers and Systems). Respondent-recorded data were considered the gold standard for the collection of race and ethnicity information. RESULTS: Among the 160 million patients from the Healthcare Cost and Utilization Project and Optum Labs datasets, race or ethnicity was unknown for 25%. Among the 2.4 million patients in the single New York City healthcare system's EHR, race or ethnicity was unknown for 57%. However, when patients directly recorded their race and ethnicity, 86% provided clinically meaningful information, and 66% of patients reported information that was discrepant with the EHR. DISCUSSION: Race and ethnicity data are critical to support precision medicine initiatives and to determine healthcare disparities; however, the quality of this information in observational databases is concerning. Patient self-recording through the use of patient-facing tools can substantially increase the quality of the information while engaging patients in their health. CONCLUSIONS: Patient self-recording may improve the completeness of race and ethnicity information. Fernanda Polubriaginof, Patrick B. Ryan, Hojjat Salmasian, Andrea W. Shapiro, Adler J. Perotte, Monika M. Safford, George Hripcsak, Shaun Smith, Nicholas P. Tatonetti, David K. Vawdrey |
J. Am. Medical Informatics Assoc. | 2 |
| 2019 | Criteria2Query: a natural language interface to clinical databases for cohort definitionabstractOBJECTIVE: Cohort definition is a bottleneck for conducting clinical research and depends on subjective decisions by domain experts. Data-driven cohort definition is appealing but requires substantial knowledge of terminologies and clinical data models. Criteria2Query is a natural language interface that facilitates human-computer collaboration for cohort definition and execution using clinical databases. MATERIALS AND METHODS: Criteria2Query uses a hybrid information extraction pipeline combining machine learning and rule-based methods to systematically parse eligibility criteria text, transforms it first into a structured criteria representation and next into sharable and executable clinical data queries represented as SQL queries conforming to the OMOP Common Data Model. Users can interactively review, refine, and execute queries in the ATLAS web application. To test effectiveness, we evaluated 125 criteria across different disease domains from ClinicalTrials.gov and 52 user-entered criteria. We evaluated F1 score and accuracy against 2 domain experts and calculated the average computation time for fully automated query formulation. We conducted an anonymous survey evaluating usability. RESULTS: Criteria2Query achieved 0.795 and 0.805 F1 score for entity recognition and relation extraction, respectively. Accuracies for negation detection, logic detection, entity normalization, and attribute normalization were 0.984, 0.864, 0.514 and 0.793, respectively. Fully automatic query formulation took 1.22 seconds/criterion. More than 80% (11+ of 13) of users would use Criteria2Query in their future cohort definition tasks. CONCLUSIONS: We contribute a novel natural language interface to clinical databases. It is open source and supports fully automated and interactive modes for autonomous data-driven cohort definition by researchers with minimal human effort. We demonstrate its promising user friendliness and usability. Chi Yuan, Patrick B. Ryan, Casey N. Ta, Yixuan Guo, Ziran Li, Jill Hardin, Rupa Makadia, Ning Shang 0004, Tian Kang, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 2 |
| 2019 | Supplementing claims data analysis using self-reported data to develop a probabilistic phenotype model for current smoking status
Jenna Reps, Peter R. Rijnbeek, Patrick B. Ryan |
J. Biomed. Informatics | 3 |
| 2019 | PheValuator: Development and evaluation of a phenotype algorithm evaluator
Joel N. Swerdel, George Hripcsak, Patrick B. Ryan |
J. Biomed. Informatics | 3 |
| 2018 | Treatment Pathways in Patients with Cancer Using a Large-scale Observational Data Network
Patrick B. Ryan, Karthik Natarajan, Thomas Falconer, Christian G. Reich, Rohit Vashisht, Nigam H. Shah, George Hripcsak |
AMIA | 2 |
| 2018 | Uncovering exposures responsible for birth season - disease effects: a global studyabstractOBJECTIVE: Birth month and climate impact lifetime disease risk, while the underlying exposures remain largely elusive. We seek to uncover distal risk factors underlying these relationships by probing the relationship between global exposure variance and disease risk variance by birth season. MATERIAL AND METHODS: This study utilizes electronic health record data from 6 sites representing 10.5 million individuals in 3 countries (United States, South Korea, and Taiwan). We obtained birth month-disease risk curves from each site in a case-control manner. Next, we correlated each birth month-disease risk curve with each exposure. A meta-analysis was then performed of correlations across sites. This allowed us to identify the most significant birth month-exposure relationships supported by all 6 sites while adjusting for multiplicity. We also successfully distinguish relative age effects (a cultural effect) from environmental exposures. RESULTS: Attention deficit hyperactivity disorder was the only identified relative age association. Our methods identified several culprit exposures that correspond well with the literature in the field. These include a link between first-trimester exposure to carbon monoxide and increased risk of depressive disorder (R = 0.725, confidence interval [95% CI], 0.529-0.847), first-trimester exposure to fine air particulates and increased risk of atrial fibrillation (R = 0.564, 95% CI, 0.363-0.715), and decreased exposure to sunlight during the third trimester and increased risk of type 2 diabetes mellitus (R = -0.816, 95% CI, -0.5767, -0.929). CONCLUSION: A global study of birth month-disease relationships reveals distal risk factors involved in causal biological pathways that underlie them. Mary Regina Boland, Pradipta Parhi, Li Li 0062, Riccardo Miotto, Robert J. Carroll, Usman Iqbal, Phung Anh Nguyen, Martijn J. Schuemie, Seng Chan You, Donahue Smith, Sean D. Mooney, Patrick B. Ryan, Yu-Chuan Li, Rae Woong Park, Joshua C. Denny, Joel Dudley, George Hripcsak, Pierre Gentine, Nicholas P. Tatonetti |
J. Am. Medical Informatics Assoc. | 12 |
| 2018 | Decentralized and reproducible geocoding and characterization of community and environmental exposures for multisite studiesabstractOBJECTIVE: Geocoding and characterizing geographic, community, and environmental characteristics of study participants is frequently done in epidemiological studies. However, participant addresses are identifiable protected health information (PHI) and geocoding must be conducted in a Health Insurance Portability and Accountability Act-compliant manner. Our objective was to create a software application for this process that addresses limitations in current approaches. MATERIALS AND METHODS: We used a containerization platform to create DeGAUSS (Decentralized Geomarker Assessment for Multi-Site Studies), a software application that facilitates reproducible geocoding and geomarker assessment while maintaining the confidentiality of PHI. To validate the software, 215 350 addresses in Hamilton County, Ohio, were geocoded using DeGAUSS, ArcGIS, Google, and SAS and compared to a gold-standard approach. We distributed the DeGAUSS software to sites in an ongoing multisite study (Electronic Medical Records and Genomics, or eMERGE), and individual sites independently geocoded and assigned median census tract-level income and distance to nearest major roadway to their participants' addresses, removed associated PHI, and returned deidentified data. RESULTS: Within a multisite study, 52 244 study participants' addresses across 5 sites were geocoded with a median distance to roadway of 10 022m and a median census tract income of $57 266, demonstrating the feasibility of DeGAUSS within a multisite study. Compared to other commonly used geocoding platforms, DeGAUSS had similar geocoding and geomarker assessment accuracies. CONCLUSION: The open source DeGAUSS software overcomes multiple challenges in the use of address data in multisite studies and also serves as a more general reproducible research tool for geocoding and geomarker assessment. Cole Brokamp, Chris Wolfe, Todd Lingren, John Harley, Patrick B. Ryan |
J. Am. Medical Informatics Assoc. | 5 |
| 2018 | Effect of vocabulary mapping for conditions on phenotype cohortsabstractObjective: To study the effect on patient cohorts of mapping condition (diagnosis) codes from source billing vocabularies to a clinical vocabulary. Materials and Methods: Nine International Classification of Diseases, Ninth Revision, Clinical Modification (ICD9-CM) concept sets were extracted from eMERGE network phenotypes, translated to Systematized Nomenclature of Medicine - Clinical Terms concept sets, and applied to patient data that were mapped from source ICD9-CM and ICD10-CM codes to Systematized Nomenclature of Medicine - Clinical Terms codes using Observational Health Data Sciences and Informatics (OHDSI) Observational Medical Outcomes Partnership (OMOP) vocabulary mappings. The original ICD9-CM concept set and a concept set extended to ICD10-CM were used to create patient cohorts that served as gold standards. Results: Four phenotype concept sets were able to be translated to Systematized Nomenclature of Medicine - Clinical Terms without ambiguities and were able to perform perfectly with respect to the gold standards. The other 5 lost performance when 2 or more ICD9-CM or ICD10-CM codes mapped to the same Systematized Nomenclature of Medicine - Clinical Terms code. The patient cohorts had a total error (false positive and false negative) of up to 0.15% compared to querying ICD9-CM source data and up to 0.26% compared to querying ICD9-CM and ICD10-CM data. Knowledge engineering was required to produce that performance; simple automated methods to generate concept sets had errors up to 10% (one outlier at 250%). Discussion: The translation of data from source vocabularies to Systematized Nomenclature of Medicine - Clinical Terms (SNOMED CT) resulted in very small error rates that were an order of magnitude smaller than other error sources. Conclusion: It appears possible to map diagnoses from disparate vocabularies to a single clinical vocabulary and carry out research using a single set of definitions, thus improving efficiency and transportability of research. George Hripcsak, Matthew E. Levine, Ning Shang 0004, Patrick B. Ryan |
J. Am. Medical Informatics Assoc. | 4 |
| 2018 | Design and implementation of a standardized framework to generate and evaluate patient-level prediction models using observational healthcare dataabstractObjective: To develop a conceptual prediction model framework containing standardized steps and describe the corresponding open-source software developed to consistently implement the framework across computational environments and observational healthcare databases to enable model sharing and reproducibility. Methods: Based on existing best practices we propose a 5 step standardized framework for: (1) transparently defining the problem; (2) selecting suitable datasets; (3) constructing variables from the observational data; (4) learning the predictive model; and (5) validating the model performance. We implemented this framework as open-source software utilizing the Observational Medical Outcomes Partnership Common Data Model to enable convenient sharing of models and reproduction of model evaluation across multiple observational datasets. The software implementation contains default covariates and classifiers but the framework enables customization and extension. Results: As a proof-of-concept, demonstrating the transparency and ease of model dissemination using the software, we developed prediction models for 21 different outcomes within a target population of people suffering from depression across 4 observational databases. All 84 models are available in an accessible online repository to be implemented by anyone with access to an observational database in the Common Data Model format. Conclusions: The proof-of-concept study illustrates the framework's ability to develop reproducible models that can be readily shared and offers the potential to perform extensive external validation of models, and improve their likelihood of clinical uptake. In future work the framework will be applied to perform an "all-by-all" prediction analysis to assess the observational data prediction domain across numerous target populations, outcomes and time, and risk settings. Jenna Reps, Martijn J. Schuemie, Marc A. Suchard, Patrick B. Ryan, Peter R. Rijnbeek |
J. Am. Medical Informatics Assoc. | 4 |
| 2018 | The representativeness of eligible patients in type 2 diabetes trials: a case study using GIST 2.0abstractOBJECTIVE: The population representativeness of a clinical study is influenced by how real-world patients qualify for the study. We analyze the representativeness of eligible patients for multiple type 2 diabetes trials and the relationship between representativeness and other trial characteristics. METHODS: Sixty-nine study traits available in the electronic health record data for 2034 patients with type 2 diabetes were used to profile the target patients for type 2 diabetes trials. A set of 1691 type 2 diabetes trials was identified from ClinicalTrials.gov, and their population representativeness was calculated using the published Generalizability Index of Study Traits 2.0 metric. The relationships between population representativeness and number of traits and between trial duration and trial metadata were statistically analyzed. A focused analysis with only phase 2 and 3 interventional trials was also conducted. RESULTS: A total of 869 of 1691 trials (51.4%) and 412 of 776 phase 2 and 3 interventional trials (53.1%) had a population representativeness of <5%. The overall representativeness was significantly correlated with the representativeness of the Hba1c criterion. The greater the number of criteria or the shorter the trial, the less the representativeness. Among the trial metadata, phase, recruitment status, and start year were found to have a statistically significant effect on population representativeness. For phase 2 and 3 interventional trials, only start year was significantly associated with representativeness. CONCLUSIONS: Our study quantified the representativeness of multiple type 2 diabetes trials. The common low representativeness of type 2 diabetes trials could be attributed to specific study design requirements of trials or safety concerns. Rather than criticizing the low representativeness, we contribute a method for increasing the transparency of the representativeness of clinical trials. Anando Sen, Andrew Goldstein, Shreya Chakrabarti, Ning Shang 0004, Tian Kang, Anil Yaman, Patrick B. Ryan, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 7 |
| 2017 | From Large-Scale Network Analytics to Clinical Solutions in OHDSI
Jon D. Duke, George Hripcsak, Patrick B. Ryan, Nigam H. Shah |
AMIA | 3 |
| 2017 | Criteria2Query: Automatically Transforming Clinical Research Eligibility Criteria Text to OMOP Common Data Model (CDM)-based Cohort Queries
Chi Yuan, Patrick B. Ryan, Yixuan Guo, Tian Kang, Chunhua Weng |
AMIA | 2 |
| 2017 | Procedure prediction from symbolic Electronic Health Records via time intervals analytics
Robert Moskovitch, Fernanda Polubriaginof, Aviram Weiss, Patrick B. Ryan, Nicholas P. Tatonetti |
J. Biomed. Informatics | 4 |
| 2017 | Accuracy of an automated knowledge base for identifying drug adverse reactions
Erica A. Voss, Richard D. Boyce, Patrick B. Ryan, Johan van der Lei, Peter R. Rijnbeek, Martijn J. Schuemie |
J. Biomed. Informatics | 3 |
| 2016 | Ensuring Reproducibility in Observational Research: Building and Sharing Knowledge Resources in the OHDSI Network
Jon D. Duke, Nigam H. Shah, George Hripcsak, Patrick B. Ryan |
AMIA | 4 |
| 2016 | Big Data for Healthcare and Life Sciences: Learning Useful Insights from Imperfect Data
Jianying Hu, Nigam H. Shah, Bradley A. Malin, Patrick B. Ryan |
AMIA | 4 |
| 2016 | Multivariate analysis of the population representativeness of related clinical studies
Zhe He 0001, Patrick B. Ryan, Julia Hoxha, Simona Carini, Ida Sim, Chunhua Weng |
J. Biomed. Informatics | 2 |
| 2016 | GIST 2.0: A scalable multi-trait metric for quantifying population representativeness of individual clinical studies
Anando Sen, Shreya Chakrabarti, Andrew Goldstein, Patrick B. Ryan, Chunhua Weng |
J. Biomed. Informatics | 5 |
| 2015 | Simulation-based Evaluation of the Generalizability Index for Study Traits
Zhe He 0001, Praveen Chandar Ravichandran, Patrick B. Ryan, Chunhua Weng |
AMIA | 3 |
| 2015 | OHDSI: An Open-Source Platform for Observational Data Analytics and Collaborative Research
Jon D. Duke, Frank J. DeFalco, Chris Knoll, Vojtech Huser, Richard D. Boyce, Patrick B. Ryan |
AMIA | 6 |
| 2015 | The Value of an Open-Source Observational Research Collaboratory: Results from the OHDSI Initiative
Jon D. Duke, George Hripcsak, Nigam H. Shah, Patrick B. Ryan |
AMIA | 4 |
| 2015 | Feasibility and utility of applications of the common data model to multiple, disparate observational health databasesabstractOBJECTIVES: To evaluate the utility of applying the Observational Medical Outcomes Partnership (OMOP) Common Data Model (CDM) across multiple observational databases within an organization and to apply standardized analytics tools for conducting observational research. MATERIALS AND METHODS: Six deidentified patient-level datasets were transformed to the OMOP CDM. We evaluated the extent of information loss that occurred through the standardization process. We developed a standardized analytic tool to replicate the cohort construction process from a published epidemiology protocol and applied the analysis to all 6 databases to assess time-to-execution and comparability of results. RESULTS: Transformation to the CDM resulted in minimal information loss across all 6 databases. Patients and observations excluded were due to identified data quality issues in the source system, 96% to 99% of condition records and 90% to 99% of drug records were successfully mapped into the CDM using the standard vocabulary. The full cohort replication and descriptive baseline summary was executed for 2 cohorts in 6 databases in less than 1 hour. DISCUSSION: The standardization process improved data quality, increased efficiency, and facilitated cross-database comparisons to support a more systematic approach to observational research. Comparisons across data sources showed consistency in the impact of inclusion criteria, using the protocol and identified differences in patient characteristics and coding practices across databases. CONCLUSION: Standardizing data structure (through a CDM), content (through a standard vocabulary with source code mappings), and analytics can enable an institution to apply a network-based approach to observational research across multiple, disparate observational health databases. Erica A. Voss, Rupa Makadia, Amy Matcho, Chris Knoll, Martijn J. Schuemie, Frank J. DeFalco, Ajit Londhe, Vivienne J. Zhu, Patrick B. Ryan |
J. Am. Medical Informatics Assoc. | 10 |
| 2014 | Transforming NHANES Database to the OMOP Common Data Model
Vivienne J. Zhu, Rupa Makadia, Amy Matcho, Martijn J. Schuemie, Patrick B. Ryan |
AMIA | 5 |
| 2013 | Developing a Systematic Approach to Measure Medication Adherence
Vivienne J. Zhu, Patrick B. Ryan, J. Marc Overhage, Paul E. Stang, Jesse Berlin |
AMIA | 2 |
| 2012 | Validation of a common data model for active safety surveillance researchabstractOBJECTIVE: Systematic analysis of observational medical databases for active safety surveillance is hindered by the variation in data models and coding systems. Data analysts often find robust clinical data models difficult to understand and ill suited to support their analytic approaches. Further, some models do not facilitate the computations required for systematic analysis across many interventions and outcomes for large datasets. Translating the data from these idiosyncratic data models to a common data model (CDM) could facilitate both the analysts' understanding and the suitability for large-scale systematic analysis. In addition to facilitating analysis, a suitable CDM has to faithfully represent the source observational database. Before beginning to use the Observational Medical Outcomes Partnership (OMOP) CDM and a related dictionary of standardized terminologies for a study of large-scale systematic active safety surveillance, the authors validated the model's suitability for this use by example. VALIDATION BY EXAMPLE: To validate the OMOP CDM, the model was instantiated into a relational database, data from 10 different observational healthcare databases were loaded into separate instances, a comprehensive array of analytic methods that operate on the data model was created, and these methods were executed against the databases to measure performance. CONCLUSION: There was acceptable representation of the data from 10 observational databases in the OMOP CDM using the standardized terminologies selected, and a range of analytic methods was developed and executed with sufficient performance to be useful for active safety surveillance. J. Marc Overhage, Patrick B. Ryan, Christian G. Reich, Abraham G. Hartzema, Paul E. Stang |
J. Am. Medical Informatics Assoc. | 2 |
| 2012 | Evaluation of alternative standardized terminologies for medical conditions within a network of observational healthcare databases
Christian G. Reich, Patrick B. Ryan, Paul E. Stang, Mitra Rocca |
J. Biomed. Informatics | 2 |
| 2010 | Development and evaluation of a common data model enabling active drug safety surveillance using disparate healthcare databasesabstractOBJECTIVE: Active drug safety surveillance may be enhanced by analysis of multiple observational healthcare databases, including administrative claims and electronic health records. The objective of this study was to develop and evaluate a common data model (CDM) enabling rapid, comparable, systematic analyses across disparate observational data sources to identify and evaluate the effects of medicines. DESIGN: The CDM uses a person-centric design, with attributes for demographics, drug exposures, and condition occurrence. Drug eras, constructed to represent periods of persistent drug use, are derived from available elements from pharmacy dispensings, prescriptions written, and other medication history. Condition eras aggregate diagnoses that occur within a single episode of care. Drugs and conditions from source data are mapped to biomedical ontologies to standardize terminologies and enable analyses of higher-order effects. MEASUREMENTS: The CDM was applied to two source types: an administrative claims and an electronic medical record database. Descriptive statistics were used to evaluate transformation rules. Two case studies demonstrate the ability of the CDM to enable standard analyses across disparate sources: analyses of persons exposed to rofecoxib and persons with an acute myocardial infarction. RESULTS: Over 43 million persons, with nearly 1 billion drug exposures and 3.7 billion condition occurrences from both databases were successfully transformed into the CDM. An analysis routine applied to transformed data from each database produced consistent, comparable results. CONCLUSION: A CDM can normalize the structure and content of disparate observational data, enabling standardized analyses that are meaningfully comparable when assessing the effects of medicines. Stephanie J. Reisinger, Patrick B. Ryan, Donald J. O'Hara, Gregory E. Powell, Jeffery L. Painter, Edward N. Pattishall, Jonathan A. Morris |
J. Am. Medical Informatics Assoc. | 2 |