EDBT 2026 Demo / reviewers in the wild / expert
Nicholas P. Tatonetti
dblp:41/8767
· DBLP profile ↗
36ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0002-2700-2597ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 35 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Probing Large Language Model Hidden States for Adverse Drug Reaction Knowledge
Jacob S. Berkowitz, Davy Weissenbacher, Apoorva Srinivasan, Nadine A. Friedrich, Jose Miguel Acitores Cortina, Sophia Kivelson, Graciela Gonzalez-Hernandez, Nicholas P. Tatonetti |
AIME (1) | 8 |
| 2025 | Evaluating the Reporting Quality of 21,041 Randomized Controlled Trial Articles with Large Language Models: A Large-Scale Transparency Analysis
Apoorva Srinivasan, Sophia Kivelson, Nadine A. Friedrich, Jacob S. Berkowitz, Nicholas P. Tatonetti |
AIME (1) | 5 |
| 2025 | Biomedical text normalization through generative modelingabstractOBJECTIVE: A large proportion of electronic health record (EHR) data consists of unstructured medical language text. The formatting of this text is often flexible and inconsistent, making it challenging to use for predictive modeling, clinical decision support, and data mining. Large language models' (LLMs) ability to understand context and semantic variations makes them promising tools for standardizing medical text. In this study, we develop and assess clinical text normalization pipelines built using large-language models. METHODS: We implemented four LLM-based normalization strategies (Zero-Shot Recall, Prompt Recall, Semantic Search, and Retrieval-Augmented Generation based normalization [RAGnorm]) and one baseline approach using TF-IDF based String Matching. We evaluated performance across three datasets of SNOMED-mapped condition terms: [1] an oncology-specific dataset, [2] a representative sample of institutional medical conditions, and [3] a dataset of commonly occurring condition codes (>1000 uses) from our institution. We measured performance by recording the mean shortest path length between predicted and true SNOMED CT terms. Additionally, we benchmarked our models against the TAC 2017 drug label annotations, which normalizes terms to the Medical Dictionary for Regulatory Activities (MedDRA) Preferred Terms. RESULTS: We found that RAGnorm was the most effective throughout each dataset, achieving a mean shortest path length of 0.21 for the domain-specific dataset, 0.58 for the sampled dataset, and 0.90 for the top terms dataset. It achieved a micro F1 score of 88.01 on task 4 of the TAC2017 conference, surpassing all other models without viewing the provided training data. CONCLUSION: We find that retrieval-focused approaches overcome traditional LLM limitations for this task. RAGnorm and related retrieval techniques should be explored further for the normalization of biomedical free text. Jacob S. Berkowitz, Apoorva Srinivasan, Jose Miguel Acitores Cortina, Yasaman Fatapour, Nicholas P. Tatonetti |
J. Biomed. Informatics | 5 |
| 2024 | KG-LIME: predicting individualized risk of adverse drug events for multiple sclerosis disease-modifying therapyabstractOBJECTIVE: The aim of this project was to create time-aware, individual-level risk score models for adverse drug events related to multiple sclerosis disease-modifying therapy and to provide interpretable explanations for model prediction behavior. MATERIALS AND METHODS: We used temporal sequences of observational medical outcomes partnership common data model (OMOP CDM) concepts derived from an electronic health record as model features. Each concept was assigned an embedding representation that was learned from a graph convolution network trained on a knowledge graph (KG) of OMOP concept relationships. Concept embeddings were fed into long short-term memory networks for 1-year adverse event prediction following drug exposure. Finally, we implemented a novel extension of the local interpretable model agnostic explanation (LIME) method, knowledge graph LIME (KG-LIME) to leverage the KG and explain individual predictions of each model. RESULTS: For a set of 4859 patients, we found that our model was effective at predicting 32 out of 56 adverse event types (P < .05) when compared to demographics and past diagnosis as variables. We also assessed discrimination in the form of area under the curve (AUC = 0.77 ± 0.15) and area under the precision-recall curve (AUC-PR = 0.31 ± 0.27) and assessed calibration in the form of Brier score (BS = 0.04 ± 0.04). Additionally, KG-LIME generated interpretable literature-validated lists of relevant medical concepts used for prediction. DISCUSSION AND CONCLUSION: Many of our risk models demonstrated high calibration and discrimination for adverse event prediction. Furthermore, our novel KG-LIME method was able to utilize the knowledge graph to highlight concepts that were important to prediction. Future work will be required to further explore the temporal window of adverse event occurrence beyond the generic 1-year window used here, particularly for short-term inpatient adverse events and long-term severe adverse events. Jason Patterson, Nicholas P. Tatonetti |
J. Am. Medical Informatics Assoc. | 2 |
| 2022 | Using Time Series Clustering to Segment and Infer Emergency Department Nursing Shifts from Electronic Health Record Log Files
Amanda J. Moy, Kenrick Cato, Jennifer Withall, Nicholas P. Tatonetti, Eugene Y. Kim, Sarah Collins Rossetti |
AMIA | 4 |
| 2022 | Deep learning on time series laboratory test results from electronic health records for early detection of pancreatic cancer
Jiheum Park, Michael G. Artin, Kate E. Lee, Yoanna S. Pumpalova, Myles A. Ingram, Benjamin L. May, Michael Park, Chin Hur, Nicholas P. Tatonetti |
J. Biomed. Informatics | 9 |
| 2021 | E-Pedigrees: a large-scale automatic family pedigree prediction applicationabstractMOTIVATION: The use and functionality of Electronic Health Records (EHR) have increased rapidly in the past few decades. EHRs are becoming an important depository of patient health information and can capture family data. Pedigree analysis is a longstanding and powerful approach that can gain insight into the underlying genetic and environmental factors in human health, but traditional approaches to identifying and recruiting families are low-throughput and labor-intensive. Therefore, high-throughput methods to automatically construct family pedigrees are needed. RESULTS: We developed a stand-alone application: Electronic Pedigrees, or E-Pedigrees, which combines two validated family prediction algorithms into a single software package for high throughput pedigrees construction. The convenient platform considers patients' basic demographic information and/or emergency contact data to infer high-accuracy parent-child relationship. Importantly, E-Pedigrees allows users to layer in additional pedigree data when available and provides options for applying different logical rules to improve accuracy of inferred family relationships. This software is fast and easy to use, is compatible with different EHR data sources, and its output is a standard PED file appropriate for multiple downstream analyses. AVAILABILITY AND IMPLEMENTATION: The Python 3.3+ version E-Pedigrees application is freely available on: https://github.com/xiayuan-huang/E-pedigrees. Xiayuan Huang, Nicholas P. Tatonetti, Katie Larow, Brooke Delgoffe, John Mayer, David Page, Scott J. Hebbring |
Bioinform. | 2 |
| 2019 | Chronic Condition Symptom Enrichment from Electronic Health Records
Theresa A. Koleck, Suzanne Bakken, Nicholas P. Tatonetti |
AMIA | 3 |
| 2019 | Breast Cancer Screening Rates Among Family Members of Affected Relatives
Fernanda Polubriaginof, Silis Y. Jiang, Nicholas P. Tatonetti, David K. Vawdrey |
AMIA | 3 |
| 2019 | Translational medicine in the Age of Big DataabstractThe ability to collect, store and analyze massive amounts of molecular and clinical data is fundamentally transforming the scientific method and its application in translational medicine. Collecting observations has always been a prerequisite for discovery, and great leaps in scientific understanding are accompanied by an expansion of this ability. Particle physics, astronomy and climate science, for example, have all greatly benefited from the development of new technologies enabling the collection of larger and more diverse data. Unlike medicine, however, each of these fields also has a mature theoretical framework on which new data can be evaluated and incorporated-to say it another way, there are no 'first principals' from which a healthy human could be analytically derived. The worry, and it is a valid concern, is that, without a strong theoretical underpinning, the inundation of data will cause medical research to devolve into a haphazard enterprise without discipline or rigor. The Age of Big Data harbors tremendous opportunity for biomedical advances, but will also be treacherous and demanding on future scientists. Nicholas P. Tatonetti |
Briefings Bioinform. | 1 |
| 2019 | PatientExploreR: an extensible application for dynamic visualization of patient clinical history from electronic health records in the OMOP common data modelabstractMOTIVATION: Electronic health records (EHRs) are quickly becoming omnipresent in healthcare, but interoperability issues and technical demands limit their use for biomedical and clinical research. Interactive and flexible software that interfaces directly with EHR data structured around a common data model (CDM) could accelerate more EHR-based research by making the data more accessible to researchers who lack computational expertise and/or domain knowledge. RESULTS: We present PatientExploreR, an extensible application built on the R/Shiny framework that interfaces with a relational database of EHR data in the Observational Medical Outcomes Partnership CDM format. PatientExploreR produces patient-level interactive and dynamic reports and facilitates visualization of clinical data without any programming required. It allows researchers to easily construct and export patient cohorts from the EHR for analysis with other software. This application could enable easier exploration of patient-level data for physicians and researchers. PatientExploreR can incorporate EHR data from any institution that employs the CDM for users with approved access. The software code is free and open source under the MIT license, enabling institutions to install and users to expand and modify the application for their own purposes. AVAILABILITY AND IMPLEMENTATION: PatientExploreR can be freely obtained from GitHub: https://github.com/BenGlicksberg/PatientExploreR. We provide instructions for how researchers with approved access to their institutional EHR can use this package. We also release an open sandbox server of synthesized patient data for users without EHR access to explore: http://patientexplorer.ucsf.edu. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Benjamin S. Glicksberg, Boris Oskotsky, Phyllis Thangaraj, Nicholas Giangreco, Marcus A. Badgeley, Kipp W. Johnson, Debajyoti Datta, Vivek A. Rudrapatna, Nadav Rappoport, Mark M. Shervey, Riccardo Miotto, Theodore C. Goldstein, Eugenia Rutenberg, Remi Frazier, Sharat Israni, Rick Larsen, Bethany Percha, Li Li 0062, Joel Dudley, Nicholas P. Tatonetti, Atul J. Butte |
Bioinform. | 21 |
| 2019 | Challenges with quality of race and ethnicity data in observational databasesabstractOBJECTIVE: We sought to assess the quality of race and ethnicity information in observational health databases, including electronic health records (EHRs), and to propose patient self-recording as an improvement strategy. MATERIALS AND METHODS: We assessed completeness of race and ethnicity information in large observational health databases in the United States (Healthcare Cost and Utilization Project and Optum Labs), and at a single healthcare system in New York City serving a racially and ethnically diverse population. We compared race and ethnicity data collected via administrative processes with data recorded directly by respondents via paper surveys (National Health and Nutrition Examination Survey and Hospital Consumer Assessment of Healthcare Providers and Systems). Respondent-recorded data were considered the gold standard for the collection of race and ethnicity information. RESULTS: Among the 160 million patients from the Healthcare Cost and Utilization Project and Optum Labs datasets, race or ethnicity was unknown for 25%. Among the 2.4 million patients in the single New York City healthcare system's EHR, race or ethnicity was unknown for 57%. However, when patients directly recorded their race and ethnicity, 86% provided clinically meaningful information, and 66% of patients reported information that was discrepant with the EHR. DISCUSSION: Race and ethnicity data are critical to support precision medicine initiatives and to determine healthcare disparities; however, the quality of this information in observational databases is concerning. Patient self-recording through the use of patient-facing tools can substantially increase the quality of the information while engaging patients in their health. CONCLUSIONS: Patient self-recording may improve the completeness of race and ethnicity information. Fernanda Polubriaginof, Patrick B. Ryan, Hojjat Salmasian, Andrea W. Shapiro, Adler J. Perotte, Monika M. Safford, George Hripcsak, Shaun Smith, Nicholas P. Tatonetti, David K. Vawdrey |
J. Am. Medical Informatics Assoc. | 9 |
| 2019 | Pathway analysis of genomic pathology tests for prognostic cancer subtyping
Olga Lyudovyk, Yufeng Shen, Nicholas P. Tatonetti, Susan J. Hsiao, Mahesh M. Mansukhani, Chunhua Weng |
J. Biomed. Informatics | 3 |
| 2018 | Comparison of Electronic Medication Orders Versus Administration Records for Identifying Prevalence of Postoperative Nausea and Vomiting
Theresa A. Koleck, Suzanne Bakken, Nicholas P. Tatonetti |
AMIA | 3 |
| 2018 | Low Screening Rates for Diabetes Mellitus Among Family Members of Affected Relatives
Fernanda Polubriaginof, Ning Shang 0004, George Hripcsak, Nicholas P. Tatonetti, David K. Vawdrey |
AMIA | 4 |
| 2018 | Using model classifications as quantitative traits to estimate stroke heritability
Phyllis Thangaraj, Fernanda Polubriaginof, Benjamin Kummer, Rami Vanguri, Tal Lorberbaum, Joseph D. Romano, Kayla Quinnies, Alexandre Yahi, Mitchell S. V. Elkind, Nicholas P. Tatonetti |
AMIA | 10 |
| 2018 | Uncovering exposures responsible for birth season - disease effects: a global studyabstractOBJECTIVE: Birth month and climate impact lifetime disease risk, while the underlying exposures remain largely elusive. We seek to uncover distal risk factors underlying these relationships by probing the relationship between global exposure variance and disease risk variance by birth season. MATERIAL AND METHODS: This study utilizes electronic health record data from 6 sites representing 10.5 million individuals in 3 countries (United States, South Korea, and Taiwan). We obtained birth month-disease risk curves from each site in a case-control manner. Next, we correlated each birth month-disease risk curve with each exposure. A meta-analysis was then performed of correlations across sites. This allowed us to identify the most significant birth month-exposure relationships supported by all 6 sites while adjusting for multiplicity. We also successfully distinguish relative age effects (a cultural effect) from environmental exposures. RESULTS: Attention deficit hyperactivity disorder was the only identified relative age association. Our methods identified several culprit exposures that correspond well with the literature in the field. These include a link between first-trimester exposure to carbon monoxide and increased risk of depressive disorder (R = 0.725, confidence interval [95% CI], 0.529-0.847), first-trimester exposure to fine air particulates and increased risk of atrial fibrillation (R = 0.564, 95% CI, 0.363-0.715), and decreased exposure to sunlight during the third trimester and increased risk of type 2 diabetes mellitus (R = -0.816, 95% CI, -0.5767, -0.929). CONCLUSION: A global study of birth month-disease relationships reveals distal risk factors involved in causal biological pathways that underlie them. Mary Regina Boland, Pradipta Parhi, Li Li 0062, Riccardo Miotto, Robert J. Carroll, Usman Iqbal, Phung Anh Nguyen, Martijn J. Schuemie, Seng Chan You, Donahue Smith, Sean D. Mooney, Patrick B. Ryan, Yu-Chuan Li, Rae Woong Park, Joshua C. Denny, Joel Dudley, George Hripcsak, Pierre Gentine, Nicholas P. Tatonetti |
J. Am. Medical Informatics Assoc. | 19 |
| 2017 | The Use of Informatics to Reduce Disparities in Transgender Health
Kenrick Cato, Joseph D. Romano, Rami Vanguri, Nicholas P. Tatonetti |
AMIA | 4 |
| 2017 | Deep recurrent neural networks identify transgender patients
Joseph D. Romano, Kenrick Cato, Rami Vanguri, Nicholas P. Tatonetti |
AMIA | 4 |
| 2017 | Procedure prediction from symbolic Electronic Health Records via time intervals analytics
Robert Moskovitch, Fernanda Polubriaginof, Aviram Weiss, Patrick B. Ryan, Nicholas P. Tatonetti |
J. Biomed. Informatics | 5 |
| 2017 | Ten Simple Rules to Enable Multi-site Collaborations through Data SharingabstractOpen access, open data, and software are critical for advancing science and enabling collaboration across multiple institutions and throughout the world.Despite near universal recognition of its importance, major barriers still exist to sharing raw data, software, and research products throughout the scientific community.Many of these barriers vary by specialty [1], increasing the difficulties for interdisciplinary and/or translational researchers to engage in collaborative research.Multi-site collaborations are vital for increasing both the impact and the generalizability of research results.However, they often present unique data sharing challenges.We discuss enabling multi-site collaborations through enhanced data sharing in this set of Ten Simple Rules.Collaboration is an essential component of research [2] that takes many forms, including internal (across departments within a single institution) and external collaborations (across institutions).However, multi-site collaborations with more than two institutions encounter more complex challenges because of institutional-specific restrictions and guidelines [3].Vicens and Bourne focus on collaborators working together on a shared research grant [4].They do not discuss the specific complexities of multi-site collaborations and the vital need for enhanced data sharing in the multi-site and large-scale collaboration context, in which participants may or may not have the same funding source and/or research grant.While challenging, multi-site collaborations are equally rewarding and result in increased research productivity [5,6].One highly successful multi-site and translational collaboration is the Electronic Medical Records and Genomics (eMERGE) network (URL: https://emerge.mc. vanderbilt.edu/)initiated in 2007 [7].The eMERGE network links biorepository data with clinical information from Electronic Health Records (EHRs).They were able to find novel associations and replicate many known associations between genetic variants and clinical phenotypes that would have been more difficult without the collaboration [8].eMERGE members also collaborated with other consortiums and networks, including the Alzheimer's Disease Genetics Consortium [9] and the NINDS Stroke Genetics Network [10], to name a few.Other successful collaborations include OHDSI: Observational Health Data Sciences and Informatics (http://www.ohdsi.org/),which builds off of the methodology from the Observational Medical Outcomes Partnership (OMOP) [11], and CIRCLE: Clinical Informatics Research Collaborative (http://circleinformatics.org/).In genetics, there are many consortiums, including ExAC: The Exome Aggregation Consortium (http://exac.broadinstitute.org/), the 1000 Genomes Project Consortium (http://www.1000genomes.org/),the Australian BioGRID Mary Regina Boland, Konrad J. Karczewski, Nicholas P. Tatonetti |
PLoS Comput. Biol. | 3 |
| 2017 | Prognosis of Clinical Outcomes with Temporal Patterns and Experiences with One Class Feature SelectionabstractAccurate prognosis of outcome events, such as clinical procedures or disease diagnosis, is central in medicine. The emergence of longitudinal clinical data, like the Electronic Health Records (EHR), represents an opportunity to develop automated methods for predicting patient outcomes. However, these data are highly dimensional and very sparse, complicating the application of predictive modeling techniques. Further, their temporal nature is not fully exploited by current methods, and temporal abstraction was recently used which results in symbolic time intervals representation. We present Maitreya, a framework for the prediction of outcome events that leverages these symbolic time intervals. Using Maitreya, learn predictive models based on the temporal patterns in the clinical records that are prognostic markers and use these markers to train predictive models for eight clinical procedures. In order to decrease the number of patterns that are used as features, we propose the use of three one class feature selection methods. We evaluate the performance of Maitreya under several parameter settings, including the one-class feature selection, and compare our results to that of atemporal approaches. In general, we found that the use of temporal patterns outperformed the atemporal methods, when representing the number of pattern occurrences. Robert Moskovitch, Hyunmi Choi, George Hripcsak, Nicholas P. Tatonetti |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2016 | Patient-provided Data Improves Race and Ethnicity Data Quality in Electronic Health Records
Fernanda Polubriaginof, Hojjat Salmasian, Andrea W. Shapiro, Jennifer E. Prey, George Hripcsak, Adler J. Perotte, Nicholas P. Tatonetti, David K. Vawdrey |
AMIA | 7 |
| 2016 | Predicting G protein-coupled receptor downstream signaling by tissue expressionabstractMOTIVATION: G protein-coupled receptors (GPCRs) are central to how cells respond to their environment and a major class of pharmacological targets. However, comprehensive knowledge of which pathways are activated and deactivated by these essential sensors is largely unknown. To better understand the mechanism of GPCR signaling system, we integrated five independent genome-wide expression datasets, representing 275 human tissues and cell lines, with protein-protein interactions and functional pathway data. RESULTS: We found that tissue-specificity plays a crucial part in the function of GPCR signaling system. Only a few GPCRs are expressed in each tissue, which are coupled by different combinations of G-proteins or β-arrestins to trigger specific downstream pathways. Based on this finding, we predicted the downstream pathways of GPCR in human tissues and validated our results with L1000 knockdown data. In total, we identified 154,988 connections between 294 GPCRs and 690 pathways in 240 tissues and cell types. AVAILABILITY AND IMPLEMENTATION: The source code and results supporting the conclusions of this article are available at http://tatonettilab.org/resources/GOTE/source_code/ CONTACT: [email protected] information: Supplementary data are available at Bioinformatics online. Nicholas P. Tatonetti |
Bioinform. | 2 |
| 2016 | Improving condition severity classification with an efficient active learning based framework
Nir Nissim, Mary Regina Boland, Nicholas P. Tatonetti, Yuval Elovici, George Hripcsak, Yuval Shahar, Robert Moskovitch |
J. Biomed. Informatics | 3 |
| 2015 | An Active Learning Framework for Efficient Condition Severity Classification
Nir Nissim, Mary Regina Boland, Robert Moskovitch, Nicholas P. Tatonetti, Yuval Elovici, Yuval Shahar, George Hripcsak |
AIME | 4 |
| 2015 | Recent Advances in Computational Drug Repositioning
Atul J. Butte, Nigam H. Shah, Nicholas P. Tatonetti, Hua Xu 0001 |
AMIA | 3 |
| 2015 | An Assessment of Family History Information Captured in an Electronic Health Record
Fernanda Polubriaginof, Nicholas P. Tatonetti, David K. Vawdrey |
AMIA | 2 |
| 2015 | Outcomes Prediction via Time Intervals Related PatternsabstractThe increasing availability of multivariate temporal data in many domains, such as biomedical, security and more, provides exceptional opportunities for temporal knowledge discovery, classification and prediction, but also challenges. Temporal variables are often sparse and in many domains, such as in biomedical data, they have huge number of variables. In recent decades in the biomedical domain events, such as conditions, drugs and procedures, are stored as time intervals, which enables to discover Time Intervals Related Patterns (TIRPs) and use for classification or prediction. In this study we present a framework for outcome events prediction, called Maitreya, which includes an algorithm for TIRPs discovery called KarmaLegoD, designed to handle huge number of symbols. Three indexing strategies for pairs of symbolic time intervals are proposed and compared, showing that the use of FullyHashed indexing is only slightly slower but consumes minimal memory. We evaluated Maitreya on eight real datasets for the prediction of clinical procedures as outcome events. The use of TIRPs outperform the use of symbols, especially with horizontal support (number of instances) as TIRPs feature representation. Robert Moskovitch, Colin G. Walsh, Fei Wang 0001, George Hripcsak, Nicholas P. Tatonetti |
ICDM | 5 |
| 2015 | Birth month affects lifetime disease risk: a phenome-wide methodabstractOBJECTIVE: An individual's birth month has a significant impact on the diseases they develop during their lifetime. Previous studies reveal relationships between birth month and several diseases including atherothrombosis, asthma, attention deficit hyperactivity disorder, and myopia, leaving most diseases completely unexplored. This retrospective population study systematically explores the relationship between seasonal affects at birth and lifetime disease risk for 1688 conditions. METHODS: We developed a hypothesis-free method that minimizes publication and disease selection biases by systematically investigating disease-birth month patterns across all conditions. Our dataset includes 1 749 400 individuals with records at New York-Presbyterian/Columbia University Medical Center born between 1900 and 2000 inclusive. We modeled associations between birth month and 1688 diseases using logistic regression. Significance was tested using a chi-squared test with multiplicity correction. RESULTS: We found 55 diseases that were significantly dependent on birth month. Of these 19 were previously reported in the literature (P < .001), 20 were for conditions with close relationships to those reported, and 16 were previously unreported. We found distinct incidence patterns across disease categories. CONCLUSIONS: Lifetime disease risk is affected by birth month. Seasonally dependent early developmental mechanisms may play a role in increasing lifetime risk of disease. Mary Regina Boland, Zach Shahn, David Madigan, George Hripcsak, Nicholas P. Tatonetti |
J. Am. Medical Informatics Assoc. | 5 |
| 2015 | Toward a complete dataset of drug-drug interaction information from publicly available sourcesabstractAlthough potential drug-drug interactions (PDDIs) are a significant source of preventable drug-related harm, there is currently no single complete source of PDDI information. In the current study, all publically available sources of PDDI information that could be identified using a comprehensive and broad search were combined into a single dataset. The combined dataset merged fourteen different sources including 5 clinically-oriented information sources, 4 Natural Language Processing (NLP) Corpora, and 5 Bioinformatics/Pharmacovigilance information sources. As a comprehensive PDDI source, the merged dataset might benefit the pharmacovigilance text mining community by making it possible to compare the representativeness of NLP corpora for PDDI text extraction tasks, and specifying elements that can be useful for future PDDI extraction purposes. An analysis of the overlap between and across the data sources showed that there was little overlap. Even comprehensive PDDI lists such as DrugBank, KEGG, and the NDF-RT had less than 50% overlap with each other. Moreover, all of the comprehensive lists had incomplete coverage of two data sources that focus on PDDIs of interest in most clinical settings. Based on this information, we think that systems that provide access to the comprehensive lists, such as APIs into RxNorm, should be careful to inform users that the lists may be incomplete with respect to PDDIs that drug experts suggest clinicians be aware of. In spite of the low degree of overlap, several dozen cases were identified where PDDI information provided in drug product labeling might be augmented by the merged dataset. Moreover, the combined dataset was also shown to improve the performance of an existing PDDI NLP pipeline and a recently published PDDI pharmacovigilance protocol. Future work will focus on improvement of the methods for mapping between PDDI information sources, identifying methods to improve the use of the merged dataset in PDDI NLP algorithms, integrating high-quality PDDI information from the merged dataset into Wikidata, and making the combined dataset accessible as Semantic Web Linked Data. Serkan Ayvaz, John R. Horn, Oktie Hassanzadeh, Qian Zhu 0003, Johann Stan, Nicholas P. Tatonetti, Santiago Vilar, Mathias Brochhausen, Matthias Samwald, Majid Rastegar-Mojarad, Michel Dumontier, Richard D. Boyce |
J. Biomed. Informatics | 6 |
| 2015 | Connectivity Homology Enables Inter-Species Network Models of Synthetic LethalityabstractSynthetic lethality is a genetic interaction wherein two otherwise nonessential genes cause cellular inviability when knocked out simultaneously. Drugs can mimic genetic knock-out effects; therefore, our understanding of promiscuous drugs, polypharmacology-related adverse drug reactions, and multi-drug therapies, especially cancer combination therapy, may be informed by a deeper understanding of synthetic lethality. However, the colossal experimental burden in humans necessitates in silico methods to guide the identification of synthetic lethal pairs. Here, we present SINaTRA (Species-INdependent TRAnslation), a network-based methodology that discovers genome-wide synthetic lethality in translation between species. SINaTRA uses connectivity homology, defined as biological connectivity patterns that persist across species, to identify synthetic lethal pairs. Importantly, our approach does not rely on genetic homology or structural and functional similarity, and it significantly outperforms models utilizing these data. We validate SINaTRA by predicting synthetic lethality in S. pombe using S. cerevisiae data, then identify over one million putative human synthetic lethal pairs to guide experimental approaches. We highlight the translational applications of our algorithm for drug discovery by identifying clusters of genes significantly enriched for single- and multi-drug cancer therapies. Alexandra Jacunski, Scott J. Dixon, Nicholas P. Tatonetti |
PLoS Comput. Biol. | 3 |
| 2013 | Correspondence: Response to 'Use of an algorithm for identifying hidden drug-drug interactions in adverse event reports' by Gooden et alabstractCritical evaluation of the results of clinical studies is vital to the continued progress of medicine. We appreciate the work performed by Gooden and colleagues1 to evaluate the clinical significance of a drug interaction between paroxetine, a selective serotonin reuptake inhibitor, and pravastatin, a cholesterol-lowering statin, that we published previously.2 Our results demonstrated a 18.5 mg/dl increase in glucose levels in individuals without diabetes, and a 48 mg/dl increase in glucose level for diabetes patients using three electronic medical record systems. In the study, Gooden et al1 did not find a difference in the development of type 2 diabetes using administrative data. We agree that retrospective risk estimates such as ours may be influenced by selection biases, such as confounding by indication. However, in our replication and validation study3 we did not see increased glucose measurements for patients on other combinations of selective serotonin reuptake inhibitors and statins or for the two classes generally—patients who are expected to have the same comorbidities. We were also not able to identify any clinical reason for the existence of clinical confounders for this particular combination of drugs alone. Moreover, we note that prediabetic mice clearly showed a positive biological result and would not be subject to the same possible confounders as the human studies.3 The authors correctly point out that an increase in non-fasting blood glucose measurements may not lead to a clinically significant event, such as type 2 diabetes mellitus (T2DM). It is possible that the increase in random glucose is not sufficiently large result in a patient being newly diagnosed with diabetes. Moreover, our findings were for near-term changes in glucose; it is possible that over the longer term, glucose falls back to normal. This would require further investigation. Finally, patients with T2DM may have the disease for some time before a diagnosis is made. It is possible that the patients enrolled in the study by Gooden et al1 had not been observed long enough to note the development of diabetes if in fact such an observation does exist. To assess the clinical significance of the drug interaction Gooden et al1 evaluated the onset of new T2DM in all patients 18 years or older using claims data. Although administrative data constitute a powerful tool for evaluating disease, accrual of a single billing code for T2DM can falsely label patients as having diabetes (false positives) as well as also falsely excluding others as not having the disease (false negatives). For this reason, Ritchie et al4 and Kho et al5 both used phenotype algorithms for T2DM including laboratory values, medications, and diagnosis billing codes (also see PheKB.org). Using claims data alone may introduce too much noise and undermine the interpretation of the authors' analysis. Gooden et al1 correctly point out that non-fasting glucose values have high variance and are not uniformly collected for all patients. For this reason we performed a paired analysis that required a patient to have glucose laboratory tests run both before and after they began combination treatment with paroxetine and pravastatin.3 We found flat glucose measurements for the single-drug-only groups, which indicate that the variability in glucose laboratory tests is not enough to explain the divergence we see in patients on the combination.3 We fully agree with the authors closing sentiment that there should be careful separation of hypothesis generation (in our case an analysis of the US Food and Drug Administration's adverse event reporting system) and hypothesis testing (in our case replication in three electronic health record systems and validation in a mouse model). It is clear that evaluating the clinical significance of this interaction between these two commonly used drugs will require a deeper understanding of its mechanism, as well as the long-term consequences of exposure. None. Not commissioned; externally peer reviewed. Nicholas P. Tatonetti, Joshua C. Denny, Russ B. Altman |
J. Am. Medical Informatics Assoc. | 1 |
| 2013 | Web-scale pharmacovigilance: listening to signals from the crowdabstractAdverse drug events cause substantial morbidity and mortality and are often discovered after a drug comes to market. We hypothesized that Internet users may provide early clues about adverse drug events via their online information-seeking. We conducted a large-scale study of Web search log data gathered during 2010. We pay particular attention to the specific drug pairing of paroxetine and pravastatin, whose interaction was reported to cause hyperglycemia after the time period of the online logs used in the analysis. We also examine sets of drug pairs known to be associated with hyperglycemia and those not associated with hyperglycemia. We find that anonymized signals on drug interactions can be mined from search logs. Compared to analyses of other sources such as electronic health records (EHR), logs are inexpensive to collect and mine. The results demonstrate that logs of the search activities of populations of computer users can contribute to drug safety surveillance. Ryen W. White, Nicholas P. Tatonetti, Nigam H. Shah, Russ B. Altman, Eric Horvitz |
J. Am. Medical Informatics Assoc. | 2 |
| 2012 | A novel signal detection algorithm for identifying hidden drug-drug interactions in adverse event reportsabstractOBJECTIVE: Adverse drug events (ADEs) are common and account for 770 000 injuries and deaths each year and drug interactions account for as much as 30% of these ADEs. Spontaneous reporting systems routinely collect ADEs from patients on complex combinations of medications and provide an opportunity to discover unexpected drug interactions. Unfortunately, current algorithms for such "signal detection" are limited by underreporting of interactions that are not expected. We present a novel method to identify latent drug interaction signals in the case of underreporting. MATERIALS AND METHODS: We identified eight clinically significant adverse events. We used the FDA's Adverse Event Reporting System to build profiles for these adverse events based on the side effects of drugs known to produce them. We then looked for pairs of drugs that match these single-drug profiles in order to predict potential interactions. We evaluated these interactions in two independent data sets and also through a retrospective analysis of the Stanford Hospital electronic medical records. RESULTS: We identified 171 novel drug interactions (for eight adverse event categories) that are significantly enriched for known drug interactions (p=0.0009) and used the electronic medical record for independently testing drug interaction hypotheses using multivariate statistical models with covariates. CONCLUSION: Our method provides an option for detecting hidden interactions in spontaneous reporting systems by using side effect profiles to infer the presence of unreported adverse events. Nicholas P. Tatonetti, Guy Haskin Fernald, Russ B. Altman |
J. Am. Medical Informatics Assoc. | 1 |
| 2010 | An integrative method for scoring candidate genes from association studies: application to warfarin dosingabstractBACKGROUND: A key challenge in pharmacogenomics is the identification of genes whose variants contribute to drug response phenotypes, which can include severe adverse effects. Pharmacogenomics GWAS attempt to elucidate genotypes predictive of drug response. However, the size of these studies has severely limited their power and potential application. We propose a novel knowledge integration and SNP aggregation approach for identifying genes impacting drug response. Our SNP aggregation method characterizes the degree to which uncommon alleles of a gene are associated with drug response. We first use pre-existing knowledge sources to rank pharmacogenes by their likelihood to affect drug response. We then define a summary score for each gene based on allele frequencies and train linear and logistic regression classifiers to predict drug response phenotypes. RESULTS: We applied our method to a published warfarin GWAS data set comprising 181 individuals. We find that our method can increase the power of the GWAS to identify both VKORC1 and CYP2C9 as warfarin pharmacogenes, where the original analysis had only identified VKORC1. Additionally, we find that our method can be used to discriminate between low-dose (AUROC=0.886) and high-dose (AUROC=0.764) responders. CONCLUSIONS: Our method offers a new route for candidate pharmacogene discovery from pharmacogenomics GWAS, and serves as a foundation for future work in methods for predictive pharmacogenomics. Nicholas P. Tatonetti, Joel Dudley, Hersh Sagreiya, Atul J. Butte, Russ B. Altman |
BMC Bioinform. | 1 |