VLDB 2026 Research / reviewers in the wild / expert
Tianxi Cai
dblp:89/8001
· DBLP profile ↗
48ranked-venue papers
0as first author
27since 2021 · last 2026
0000-0002-5379-2502ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 35 · 18 since 2021Artificial intelligence and machine learning · 13 · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Statistical methods to harmonize electronic health record data across healthcare systems: case study and lessons learnedabstractMOTIVATION: Although common data models for electronic health record (EHR) data can facilitate multi-site data organization and querying, the same medical event may still be coded differently between healthcare systems. In this paper, we present statistical methods to identify and mitigate coding discrepancies using summary-level data, and demonstrate these methods using data from two FDA Sentinel data partners: Kaiser Permanente Washington and Kaiser Permanente Northwest. RESULTS: We first characterize differences in coding patterns, then compute a code mapping matrix to harmonize data between systems. Our findings reveal significant heterogeneity in coded EHR data, even after adopting a common data model with the same coding system, highlighting the importance of data harmonization before downstream analyses. Our study also demonstrates the effectiveness of the data harmonization approaches, which provide a foundational data quality step to promote semantic interoperability, enhance data integration, and improve the integrity of study conclusions. AVAILABILITY AND IMPLEMENTATION: Computation prototypes, including R/Python codes and examples, are included in Section 7, available as supplementary data at Bioinformatics online and will be posted on GitHub upon publication. Yuqi Zhai, Xianshi Yu, Brian L. Hazlehurst, Denis B. Nyongesa, Daniel S. Sapp, Brian D. Williamson, David Carrell, Luesa Healy, Kara L. Cushing-Haugen, Jenna Wong, Shirley V. Wang, James S. Floyd, Kathleen Shattuck, Samuel McGown, Sarah Alam, José J. Hernández-Muñoz, Danijela Stojanovic, Sudha R. Raman, Sharon E. Davis, Tianxi Cai, Jennifer C. Nelson, Patrick J. Heagerty |
Bioinform. | 24 |
| 2025 | CEGA: A Cost-Effective Approach for Graph-Based Model Extraction and AcquisitionabstractGraph Neural Networks (GNNs) have demonstrated remarkable utility across diverse applications, and their growing complexity has made Machine Learning as a Service (MLaaS) a viable platform for scalable deployment. However, this accessibility also exposes GNN to serious security threats, most notably model extraction attacks (MEAs), in which adversaries strategically query a deployed model to construct a high-fidelity replica. In this work, we evaluate the vulnerability of GNNs to MEAs and explore their potential for cost-effective model acquisition in non-adversarial research settings. Importantly, adaptive node querying strategies can also serve a critical role in research, particularly when labeling data is expensive or time-consuming. By selectively sampling informative nodes, researchers can train high-performing GNNs with minimal supervision, which is particularly valuable in domains such as biomedicine, where annotations often require expert input. To address this, we propose a node querying strategy tailored to a highly practical yet underexplored scenario, where bulk queries are prohibited, and only a limited set of initial nodes is available. Our approach iteratively refines the node selection mechanism over multiple learning cycles, leveraging historical feedback to improve extraction efficiency. Extensive experiments on benchmark graph datasets demonstrate our superiority over comparable baselines on accuracy, fidelity, and F1 score under strict query-size constraints. These results highlight both the susceptibility of deployed GNNs to extraction attacks and the promise of ethical, efficient GNN acquisition methods to support low-resource research environments. Our implementation is publicly available at [https://github.com/LabRAI/CEGA](https://github.com/LabRAI/CEGA). Zebin Wang, Menghan Lin, Bolin Shen, Molei Liu, Tianxi Cai, Yushun Dong |
ICML | 6 |
| 2025 | ARCH: Large-scale knowledge graph via aggregated narrative codified health records analysis
Ziming Gan, Doudou Zhou, Everett Neil Rush, Vidul Ayakulangara Panickan, Yuk-Lam Ho, George Ostrouchov, Shuting Shen, Xin Xiong 0006, Kimberly F. Greco, Chuan Hong, Clara-Lea Bonzel, Jun Wen 0001, Lauren Costa, Tianrun A. Cai, Edmon Begoli, Zongqi Xia, John Michael Gaziano, Katherine P. Liao, Kelly Cho, Tianxi Cai |
J. Biomed. Informatics | 21 |
| 2025 | Integrated analysis for electronic health records with structured and sporadic missingness
Jianbin Tan, Chuan Hong, T. Tony Cai, Tianxi Cai, Anru Zhang |
J. Biomed. Informatics | 5 |
| 2025 | DOME: Directional medical embedding vectors from Electronic Health RecordsabstractMOTIVATION: The increasing availability of Electronic Health Record (EHR) systems has created enormous potential for translational research. Recent developments in representation learning techniques have led to effective large-scale representations of EHR concepts along with knowledge graphs that empower downstream EHR studies. However, most existing methods require training with patient-level data, limiting their abilities to expand the training with multi-institutional EHR data. On the other hand, scalable approaches that only require summary-level data do not incorporate temporal dependencies between concepts. METHODS: We introduce a DirectiOnal Medical Embedding (DOME) algorithm to encode temporally directional relationships between medical concepts, using summary-level EHR data. Specifically, DOME first aggregates patient-level EHR data into an asymmetric co-occurrence matrix. Then it computes two Positive Pointwise Mutual Information (PPMI) matrices to correspondingly encode the pairwise prior and posterior dependencies between medical concepts. Following that, a joint matrix factorization is performed on the two PPMI matrices, which results in three vectors for each concept: a semantic embedding and two directional context embeddings. They collectively provide a comprehensive depiction of the temporal relationship between EHR concepts. RESULTS: We highlight the advantages and translational potential of DOME through three sets of validation studies. First, DOME consistently improves existing direction-agnostic embedding vectors for disease risk prediction in several diseases, for example achieving a relative gain of 5.5% in the area under the receiver operating characteristic (AUROC) for lung cancer. Second, DOME excels in directional drug-disease relationship inference by successfully differentiating between drug side effects and indications, correspondingly achieving relative AUROC gain over the state-of-the-art methods by 10.8% and 6.6%. Finally, DOME effectively constructs directional knowledge graphs, which distinguish disease risk factors from comorbidities, thereby revealing disease progression trajectories. The source codes are provided at https://github.com/celehs/Directional-EHR-embedding. Jun Wen 0001, Hao Xue 0005, Everett Neil Rush, Vidul Ayakulangara Panickan, Tianrun A. Cai, Doudou Zhou, Yuk-Lam Ho, Lauren Costa, Edmon Begoli, Chuan Hong, John Michael Gaziano, Kelly Cho, Katherine P. Liao, Tianxi Cai |
J. Biomed. Informatics | 15 |
| 2025 | Efficient and Robust Semi-supervised Estimation of Average Treatment Effect with Partially Annotated Treatment and ResponseabstractA notable challenge of leveraging Electronic Health Records (EHR) for treatment effect assessment is the lack of precise information on important clinical variables, including the treatment received and the response. Both treatment information and response cannot be accurately captured by readily available EHR features in many studies and require labor-intensive manual chart review to precisely annotate, which limits the number of available gold standard labels on these key variables. We considered average treatment effect (ATE) estimation when 1) exact treatment and outcome variables are only observed together in a small labeled subset and 2) noisy surrogates of treatment and outcome, such as relevant prescription and diagnosis codes, along with potential confounders are observed for all subjects. We derived the efficient influence function for ATE and used it to construct a semi-supervised multiple machine learning (SMMAL) estimator. We justified that our SMMAL ATE estimator is semi-parametric efficient with B-spline regression under low-dimensional smooth models. We developed the adaptive sparsity/model doubly robust estimation under high-dimensional logistic propensity score and outcome regression models. Results from simulation studies demonstrated the validity of our SMMAL method and its superiority over supervised and unsupervised benchmarks. We applied SMMAL to the assessment of targeted therapies for metastatic colorectal cancer in comparison to chemotherapy. Jue Hou 0001, Rajarshi Mukherjee, Tianxi Cai |
J. Mach. Learn. Res. | 3 |
| 2024 | A teacher-teacher framework for clinical language representation learningabstractIn recent years, there has been a proliferation of ready-to-use large language models (LLMs) designed for various applications, both general-purpose and domain-specific. Instead of advocating for the development of a new model or continuous pretraining of an existing one, this paper introduces a pragmatic teacher-teacher framework to facilitate mutual learning between two pre-existing models.
By leveraging two teacher models possessing complementary knowledge, we introduce a LIghtweight kNowledge alignmEnt (LINE) module aimed at harmonizing their knowledge within a unified representation space. This framework is particularly valuable in clinical settings, where stringent regulations and privacy considerations dictate the handling of detailed clinical notes. Our trained LINE module excels in capturing critical information from clinical notes, leveraging highly de-identified data. Validation and downstream tasks further demonstrate the effectiveness of the proposed framework. Feiqing Huang, Shenghan Zhang, Sara Morini Sweet, Tianxi Cai |
NeurIPS | 4 |
| 2024 | Centralized Interactive Phenomics Resource: an integrated online phenomics knowledgebase for health data usersabstractOBJECTIVE: Development of clinical phenotypes from electronic health records (EHRs) can be resource intensive. Several phenotype libraries have been created to facilitate reuse of definitions. However, these platforms vary in target audience and utility. We describe the development of the Centralized Interactive Phenomics Resource (CIPHER) knowledgebase, a comprehensive public-facing phenotype library, which aims to facilitate clinical and health services research. MATERIALS AND METHODS: The platform was designed to collect and catalog EHR-based computable phenotype algorithms from any healthcare system, scale metadata management, facilitate phenotype discovery, and allow for integration of tools and user workflows. Phenomics experts were engaged in the development and testing of the site. RESULTS: The knowledgebase stores phenotype metadata using the CIPHER standard, and definitions are accessible through complex searching. Phenotypes are contributed to the knowledgebase via webform, allowing metadata validation. Data visualization tools linking to the knowledgebase enhance user interaction with content and accelerate phenotype development. DISCUSSION: The CIPHER knowledgebase was developed in the largest healthcare system in the United States and piloted with external partners. The design of the CIPHER website supports a variety of front-end tools and features to facilitate phenotype development and reuse. Health data users are encouraged to contribute their algorithms to the knowledgebase for wider dissemination to the research community, and to use the platform as a springboard for phenotyping. CONCLUSION: CIPHER is a public resource for all health data users available at https://phenomics.va.ornl.gov/ which facilitates phenotype reuse, development, and dissemination of phenotyping knowledge. Jacqueline Honerlaw, Yuk-Lam Ho, Francesca Fontin, Michael Murray, Ashley Galloway, David Heise, Keith Connatser, Laura Davies, Jeffrey Gosian, Monika Maripuri, John P. Russo, Rahul Sangar, Vidisha Tanukonda, Edward Zielinski, Maureen Dubreuil, Andrew J. Zimolzak, Vidul Ayakulangara Panickan, Su-Chun Cheng, Stacey B. Whitbourne, David R. Gagnon, Tianxi Cai, Katherine P. Liao, Rachel Badovinac Ramoni, John Michael Gaziano, Sumitra Muralidhar, Kelly Cho |
J. Am. Medical Informatics Assoc. | 21 |
| 2024 | Semi-supervised Double Deep Learning Temporal Risk Prediction (SeDDLeR) with Electronic Health Records
Isabelle-Emmanuella Nogues, Jun Wen 0001, Yihan Zhao, Clara-Lea Bonzel, Victor M. Castro, Yucong Lin, Shike Xu, Jue Hou 0001, Tianxi Cai |
J. Biomed. Informatics | 9 |
| 2023 | Multimodal representation learning for predicting molecule-disease relationsabstractMOTIVATION: Predicting molecule-disease indications and side effects is important for drug development and pharmacovigilance. Comprehensively mining molecule-molecule, molecule-disease and disease-disease semantic dependencies can potentially improve prediction performance. METHODS: We introduce a Multi-Modal REpresentation Mapping Approach to Predicting molecular-disease relations (M2REMAP) by incorporating clinical semantics learned from electronic health records (EHR) of 12.6 million patients. Specifically, M2REMAP first learns a multimodal molecule representation that synthesizes chemical property and clinical semantic information by mapping molecule chemicals via a deep neural network onto the clinical semantic embedding space shared by drugs, diseases and other common clinical concepts. To infer molecule-disease relations, M2REMAP combines multimodal molecule representation and disease semantic embedding to jointly infer indications and side effects. RESULTS: We extensively evaluate M2REMAP on molecule indications, side effects and interactions. Results show that incorporating EHR embeddings improves performance significantly, for example, attaining an improvement over the baseline models by 23.6% in PRC-AUC on indications and 23.9% on side effects. Further, M2REMAP overcomes the limitation of existing methods and effectively predicts drugs for novel diseases and emerging pathogens. AVAILABILITY AND IMPLEMENTATION: The code is available at https://github.com/celehs/M2REMAP, and prediction results are provided at https://shiny.parse-health.org/drugs-diseases-dev/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jun Wen 0001, Xiang Zhang 0012, Everett Neil Rush, Vidul Ayakulangara Panickan, Tianrun A. Cai, Doudou Zhou, Yuk-Lam Ho, Lauren Costa, Edmon Begoli, Chuan Hong, John Michael Gaziano, Kelly Cho, Katherine P. Liao, Marinka Zitnik, Tianxi Cai |
Bioinform. | 17 |
| 2023 | Framework of the Centralized Interactive Phenomics Resource (CIPHER) standard for electronic health data-based phenomics knowledgebaseabstractThe development of phenotypes using electronic health records is a resource-intensive process. Therefore, the cataloging of phenotype algorithm metadata for reuse is critical to accelerate clinical research. The Department of Veterans Affairs (VA) has developed a standard for phenotype metadata collection which is currently used in the VA phenomics knowledgebase library, CIPHER (Centralized Interactive Phenomics Resource), to capture over 5000 phenotypes. The CIPHER standard improves upon existing phenotype library metadata collection by capturing the context of algorithm development, phenotyping method used, and approach to validation. While the standard was iteratively developed with VA phenomics experts, it is applicable to the capture of phenotypes across healthcare systems. We describe the framework of the CIPHER standard for phenotype metadata collection, the rationale for its development, and its current application to the largest healthcare system in the United States. Jacqueline Honerlaw, Yuk-Lam Ho, Francesca Fontin, Jeffrey Gosian, Monika Maripuri, Michael Murray, Rahul Sangar, Ashley Galloway, Andrew J. Zimolzak, Stacey B. Whitbourne, Juan P. Casas, Rachel Badovinac Ramoni, David R. Gagnon, Tianxi Cai, Katherine P. Liao, John Michael Gaziano, Sumitra Muralidhar, Kelly Cho |
J. Am. Medical Informatics Assoc. | 14 |
| 2023 | Semi-supervised calibration of noisy event risk (SCANER) with electronic health records
Chuan Hong, Qianyu Yuan, Kelly Cho, Katherine P. Liao, Michael J. Pencina, David C. Christiani, Tianxi Cai |
J. Biomed. Informatics | 8 |
| 2023 | Multimodal learning on graphs for disease relation extractionabstractDisease knowledge graphs have emerged as a powerful tool for artificial intelligence to connect, organize, and access diverse information about diseases. Relations between disease concepts are often distributed across multiple datasets, including unstructured plain text datasets and incomplete disease knowledge graphs. Extracting disease relations from multimodal data sources is thus crucial for constructing accurate and comprehensive disease knowledge graphs. We introduce REMAP, a multimodal approach for disease relation extraction. The REMAP machine learning approach jointly embeds a partial, incomplete knowledge graph and a medical language dataset into a compact latent vector space, aligning the multimodal embeddings for optimal disease relation extraction. Additionally, REMAP utilizes a decoupled model structure to enable inference in single-modal data, which can be applied under missing modality scenarios. We apply the REMAP approach to a disease knowledge graph with 96,913 relations and a text dataset of 1.24 million sentences. On a dataset annotated by human experts, REMAP improves language-based disease relation extraction by 10.0% (accuracy) and 17.2% (F1-score) by fusing disease knowledge graphs with language information. Furthermore, REMAP leverages text information to recommend new relationships in the knowledge graph, outperforming graph-based methods by 8.4% (accuracy) and 10.4% (F1-score). REMAP is a flexible multimodal approach for extracting disease relations by fusing structured knowledge and language information. This approach provides a powerful model to easily find, access, and evaluate relations between disease concepts. Yucong Lin, Keming Lu, Sheng Yu 0002, Tianxi Cai, Marinka Zitnik |
J. Biomed. Informatics | 4 |
| 2023 | Informative missingness: What can we learn from patterns in missing laboratory data in the electronic health record?
Amelia L. M. Tan, Emily J. Getzen, Meghan Hutch, Zachary H. Strasser, Alba Gutiérrez-Sacristán, Trang T. Le, Arianna Dagliati, Michele Morris, David A. Hanauer, Bertrand Moal, Clara-Lea Bonzel, William Yuan, Lorenzo Chiudinelli, Priyam Das, Harrison G. Zhang, Bruce J. Aronow, Paul Avillach, Gabriel A. Brat, Tianxi Cai, Chuan Hong, William G. La Cava, He Hooi Will Loh, Yuan Luo 0001, Shawn N. Murphy, Kee Yuan Hgiam, Gilbert S. Omenn, Lav P. Patel, Malarkodi J. Samayamuthu, Emily R. Shriver, Zahra Shakeri Hossein Abad, Byorn W. L. Tan, Shyam Visweswaran, Griffin M. Weber, Zongqi Xia, Bertrand Verdy, Qi Long, Danielle L. Mowery, John H. Holmes |
J. Biomed. Informatics | 19 |
| 2023 | Surrogate Assisted Semi-supervised Inference for High Dimensional Risk PredictionabstractRisk modeling with electronic health records (EHR) data is challenging due to no direct observations of the disease outcome and the high-dimensional predictors. In this paper, we develop a surrogate assisted semi-supervised learning approach, leveraging small labeled data with annotated outcomes and extensive unlabeled data of outcome surrogates and high-dimensional predictors. We propose to impute the unobserved outcomes by constructing a sparse imputation model with outcome surrogates and high-dimensional predictors. We further conduct a one-step bias correction to enable interval estimation for the risk prediction. Our inference procedure is valid even if both the imputation and risk prediction models are misspecified. Our novel way of ultilizing unlabelled data enables the high-dimensional statistical inference for the challenging setting with a dense risk prediction model. We present an extensive simulation study to demonstrate the superiority of our approach compared to existing supervised methods. We apply the method to genetic risk prediction of type-2 diabetes mellitus using an EHR biobank cohort. Jue Hou 0001, Zijian Guo 0003, Tianxi Cai |
J. Mach. Learn. Res. | 3 |
| 2023 | Augmented Transfer Regression Learning with Semi-non-parametric Nuisance ModelsabstractWe develop an augmented transfer regression learning (ATReL) approach that introduces an imputation model to augment the importance weighting equation to achieve double robustness for covariate shift correction. More significantly, we propose a novel semi-non-parametric (SNP) construction framework for the two nuisance models. Compared with existing doubly robust approaches relying on fully parametric or fully non-parametric (machine learning) nuisance models, our proposal is more flexible and balanced to address model misspecification and the curse of dimensionality, achieving a better trade-off in terms of model complexity. The SNP construction presents a new technical challenge in controlling the first-order bias caused by the nuisance estimators. To overcome this, we propose a two-step calibrated estimating approach to construct the nuisance models that ensures the effective reduction of potential bias. Under this SNP framework, our ATReL estimator is root-n-consistent when (i) at least one nuisance model is correctly specified and (ii) the nonparametric components are rate-doubly robust. Simulation studies demonstrate that our method is more robust and efficient than existing methods under various configurations. We also examine the utility of our method through a real transfer learning example of the phenotyping algorithm for rheumatoid arthritis across different time windows. Finally, we propose ways to enhance the intrinsic efficiency of our estimator and to incorporate modern machine-learning methods in the proposed SNP framework. Molei Liu, Katherine P. Liao, Tianxi Cai |
J. Mach. Learn. Res. | 4 |
| 2023 | Semi-Supervised Off-Policy Reinforcement Learning and Value Estimation for Dynamic Treatment RegimesabstractReinforcement learning (RL) has shown great promise in estimating dynamic treatment regimes which take into account patient heterogeneity. However, health-outcome information, used as the reward for RL methods, is often not well coded but rather embedded in clinical notes. Extracting precise outcome information is a resource-intensive task, so most of the available well-annotated cohorts are small. To address this issue, we propose a semi-supervised learning (SSL) approach that efficiently leverages a small-sized labeled data set with actual outcomes observed and a large unlabeled data set with outcome surrogates. In particular, we propose a semi-supervised, efficient approach to $Q$-learning and doubly robust off-policy value estimation. Generalizing SSL to dynamic treatment regimes brings interesting challenges: 1) Feature distribution for $Q$-learning is unknown as it includes previous outcomes. 2) The surrogate variables we leverage in the modified SSL framework are predictive of the outcome but not informative of the optimal policy or value function. We provide theoretical results for our $Q$ function and value function estimators to understand the degree of efficiency gained from SSL. Our method is at least as efficient as the supervised approach, and robust to bias from mis-specification of the imputation models. Aaron Sonabend W., Nilanjana Laha, Ashwin N. Ananthakrishnan, Tianxi Cai, Rajarshi Mukherjee |
J. Mach. Learn. Res. | 4 |
| 2023 | Multi-source Learning via Completion of Block-wise Overlapping Noisy MatricesabstractElectronic healthcare records (EHR) provide a rich resource for healthcare research. An important problem for the efficient utilization of the EHR data is the representation of the EHR features, which include the unstructured clinical narratives and the structured codified data. Matrix factorization-based embeddings trained using the summary-level co-occurrence statistics of EHR data have provided a promising solution for feature representation while preserving patients' privacy. However, such methods do not work well with multi-source data when these sources have overlapping but non-identical features. To accommodate multi-sources learning, we propose a novel word embedding generative model. To obtain multi-source embeddings, we design an efficient Block-wise Overlapping Noisy Matrix Integration (BONMI) algorithm to aggregate the multi-source pointwise mutual information matrices optimally with a theoretical guarantee. Our algorithm can also be applied to other multi-source data integration problems with a similar data structure. A by-product of BONMI is the contribution to the field of matrix completion by considering the missing mechanism other than the entry-wise independent missing. We show that the entry-wise missing assumption, despite its prevalence in the works of matrix completion, is not necessary to guarantee recovery. We prove the statistical rate of our estimator, which is comparable to the rate under independent missingness. Simulation studies show that BONMI performs well under a variety of configurations. We further illustrate the utility of BONMI by integrating multi-lingual multi-source medical text and EHR data to perform two tasks: (i) co-training semantic embeddings for medical concepts in both English and Chinese and (ii) the translation between English and Chinese medical concepts. Our method shows an advantage over existing methods. Doudou Zhou, Tianxi Cai |
J. Mach. Learn. Res. | 2 |
| 2022 | Knowledge-Driven Online Multimodal Automated Phenotyping System
Molei Liu, Sara Morini Sweet, Xin Xiong 0006, Chuan Hong, Clara-Lea Bonzel, Vidul Ayakulangara Panickan, Everett Neil Rush, Yuk-Lam Ho, Kelly Cho, John Michael Gaziano, Katherine P. Liao, Tianxi Cai, Tianrun A. Cai |
AMIA | 12 |
| 2022 | Scalable relevance ranking algorithm via semantic similarity assessment improves efficiency of medical chart reviewabstractAccurately assigning phenotype information to individual patients via computational phenotyping using Electronic Health Records (EHRs) has been seen as the first step towards enabling EHRs for precision medicine research. Chart review labels annotated by clinical experts, also known as “gold standard” labels, are essential for the development and validation of computational phenotyping algorithms. However, given the complexity of EHR systems, the process of chart review is both labor intensive and time consuming. We propose a fully automated algorithm, referred to as pGUESS, to rank EHR notes according to their relevance to a given phenotype. By identifying the most relevant notes, pGUESS can greatly improve the efficiency and accuracy of chart reviews. pGUESS uses prior guided semantic similarity to measure the informativeness of a clinical note to a given phenotype. We first select candidate clinical concepts from a pool of comprehensive medical concepts using public knowledge sources and then derive the semantic embedding vector (SEV) for a reference article (SEVref) and each note (SEVnote). The algorithm scores the relevance of a note as the cosine similarity between SEVnote and SEVref. The algorithm was validated against four sets of 200 notes that were manually annotated by clinical experts to assess their informativeness to one of three disease phenotypes. pGUESS algorithm substantially outperforms existing unsupervised approaches for classifying the relevance status with respect to both accuracy and scalability across phenotypes. Averaging over the three phenotypes, the rank correlation between the algorithm ranking and gold standard label was 0.64 for pGUESS, but only 0.47 and 0.35 for the next two best performing algorithms. pGUESS is also much more computationally scalable compared to existing algorithms. pGUESS algorithm can substantially reduce the burden of chart review and holds potential in improving the efficiency and accuracy of human annotation. Tianrun A. Cai, Zeling He, Chuan Hong, Yuk-Lam Ho, Jacqueline Honerlaw, Alon Geva, Vidul Ayakulangara Panickan, Amanda King, David R. Gagnon, John Michael Gaziano, Kelly Cho, Katherine P. Liao, Tianxi Cai |
J. Biomed. Informatics | 14 |
| 2022 | Weakly Semi-supervised phenotyping using Electronic Health recordsabstractOBJECTIVE: Electronic Health Record (EHR) based phenotyping is a crucial yet challenging problem in the biomedical field. Though clinicians typically determine patient-level diagnoses via manual chart review, the sheer volume and heterogeneity of EHR data renders such tasks challenging, time-consuming, and prohibitively expensive, thus leading to a scarcity of clinical annotations in EHRs. Weakly supervised learning algorithms have been successfully applied to various EHR phenotyping problems, due to their ability to leverage information from large quantities of unlabeled samples to better inform predictions based on a far smaller number of patients. However, most weakly supervised methods are subject to the challenge to choose the right cutoff value to generate an optimal classifier. Furthermore, since they only utilize the most informative features (i.e., main ICD and NLP counts) they may fail for episodic phenotypes that cannot be consistently detected via ICD and NLP data. In this paper, we propose a label-efficient, weakly semi-supervised deep learning algorithm for EHR phenotyping (WSS-DL), which overcomes the limitations above. MATERIALS AND METHODS: WSS-DL classifies patient-level disease status through a series of learning stages: 1) generating silver standard labels, 2) deriving enhanced-silver-standard labels by fitting a weakly supervised deep learning model to data with silver standard labels as outcomes and high dimensional EHR features as input, and 3) obtaining the final prediction score and classifier by fitting a supervised learning model to data with a minimal number of gold standard labels as the outcome, and the enhanced-silver-standard labels and a minimal set of most informative EHR features as input. To assess the generalizability of WSS-DL across different phenotypes and medical institutions, we apply WSS-DL to classify a total of 17 diseases, including both acute and chronic conditions, using EHR data from three healthcare systems. Additionally, we determine the minimum quantity of training labels required by WSS-DL to outperform existing supervised and semi-supervised phenotyping methods. RESULTS: The proposed method, in combining the strengths of deep learning and weakly semi-supervised learning, successfully leverages the crucial phenotyping information contained in EHR features from unlabeled samples. Indeed, the deep learning model's ability to handle high-dimensional EHR features allows it to generate strong phenotype status predictions from silver standard labels. These predictions, in turn, provide highly effective features in the final logistic regression stage, leading to high phenotyping accuracy in notably small subsets of labeled data (e.g. n = 40 labeled samples). CONCLUSION: Our method's high performance in EHR datasets with very small numbers of labels indicates its potential value in aiding doctors to diagnose rare diseases as well as conditions susceptible to misdiagnosis. Isabelle-Emmanuella Nogues, Jun Wen 0001, Yucong Lin, Molei Liu, Sara K. Tedeschi, Alon Geva, Tianxi Cai, Chuan Hong |
J. Biomed. Informatics | 7 |
| 2022 | SurvMaximin: Robust federated approach to transporting survival risk prediction models
Harrison G. Zhang, Xin Xiong 0006, Chuan Hong, Griffin M. Weber, Gabriel A. Brat, Clara-Lea Bonzel, Yuan Luo 0001, Rui Duan 0004, Nathan P. Palmer, Meghan Hutch, Alba Gutiérrez-Sacristán, Riccardo Bellazzi, Luca Chiovato, Kelly Cho, Arianna Dagliati, Hossein Estiri, Noelia García-Barrio, Romain Griffier, David A. Hanauer, Yuk-Lam Ho, John H. Holmes, Mark S. Keller, Jeffrey G. Klann, Sehi L'Yi, Sara Lozano-Zahonero, Sarah E. Maidlow, Adeline Makoudjou, Alberto Malovini, Bertrand Moal, Jason H. Moore, Michele Morris, Danielle L. Mowery, Shawn N. Murphy, Antoine Neuraz, Kee Yuan Ngiam, Gilbert S. Omenn, Lav P. Patel, Miguel Pedrera-Jiménez, Andrea Prunotto, Malarkodi J. Samayamuthu, Fernando J. Sanz Vidorreta, Emily Schriver, Petra Schubert, Pablo Serrano-Balazote, Andrew M. South, Amelia L. M. Tan, Byorn W. L. Tan, Valentina Tibollo, Patric Tippmann, Shyam Visweswaran, Zongqi Xia, William Yuan, Daniela Zöller, Isaac S. Kohane, Paul Avillach, Zijian Guo 0003, Tianxi Cai |
J. Biomed. Informatics | 58 |
| 2022 | Multiview Incomplete Knowledge Graph Integration with application to cross-institutional EHR data harmonizationabstractOBJECTIVE: The growing availability of electronic health records (EHR) data opens opportunities for integrative analysis of multi-institutional EHR to produce generalizable knowledge. A key barrier to such integrative analyses is the lack of semantic interoperability across different institutions due to coding differences. We propose a Multiview Incomplete Knowledge Graph Integration (MIKGI) algorithm to integrate information from multiple sources with partially overlapping EHR concept codes to enable translations between healthcare systems. METHODS: The MIKGI algorithm combines knowledge graph information from (i) embeddings trained from the co-occurrence patterns of medical codes within each EHR system and (ii) semantic embeddings of the textual strings of all medical codes obtained from the Self-Aligning Pretrained BERT (SAPBERT) algorithm. Due to the heterogeneity in the coding across healthcare systems, each EHR source provides partial coverage of the available codes. MIKGI synthesizes the incomplete knowledge graphs derived from these multi-source embeddings by minimizing a spherical loss function that combines the pairwise directional similarities of embeddings computed from all available sources. MIKGI outputs harmonized semantic embedding vectors for all EHR codes, which improves the quality of the embeddings and enables direct assessment of both similarity and relatedness between any pair of codes from multiple healthcare systems. RESULTS: With EHR co-occurrence data from Veteran Affairs (VA) healthcare and Mass General Brigham (MGB), MIKGI algorithm produces high quality embeddings for a variety of downstream tasks including detecting known similar or related entity pairs and mapping VA local codes to the relevant EHR codes used at MGB. Based on the cosine similarity of the MIKGI trained embeddings, the AUC was 0.918 for detecting similar entity pairs and 0.809 for detecting related pairs. For cross-institutional medical code mapping, the top 1 and top 5 accuracy were 91.0% and 97.5% when mapping medication codes at VA to RxNorm medication codes at MGB; 59.1% and 75.8% when mapping VA local laboratory codes to LOINC hierarchy. When trained with 500 labels, the lab code mapping attained top 1 and 5 accuracy at 77.7% and 87.9%. MIKGI also attained best performance in selecting VA local lab codes for desired laboratory tests and COVID-19 related features for COVID EHR studies. Compared to existing methods, MIKGI attained the most robust performance with accuracy the highest or near the highest across all tasks. CONCLUSIONS: The proposed MIKGI algorithm can effectively integrate incomplete summary data from biomedical text and EHR data to generate harmonized embeddings for EHR codes for knowledge graph modeling and cross-institutional translation of EHR codes. Doudou Zhou, Ziming Gan, Alina Patwari, Everett Neil Rush, Clara-Lea Bonzel, Vidul Ayakulangara Panickan, Chuan Hong, Yuk-Lam Ho, Tianrun A. Cai, Lauren Costa, Victor M. Castro, Shawn N. Murphy, Gabriel A. Brat, Griffin M. Weber, Paul Avillach, John Michael Gaziano, Kelly Cho, Katherine P. Liao, Tianxi Cai |
J. Biomed. Informatics | 22 |
| 2022 | Prior Adaptive Semi-supervised Learning with Application to EHR PhenotypingabstractElectronic Health Record (EHR) data, a rich source for biomedical research, have been successfully used to gain novel insight into a wide range of diseases. Despite its potential, EHR is currently underutilized for discovery research due to its major limitation in the lack of precise phenotype information. To overcome such difficulties, recent efforts have been devoted to developing supervised algorithms to accurately predict phenotypes based on relatively small training datasets with gold-standard labels extracted via chart review. However, supervised methods typically require a sizable training set to yield generalizable algorithms, especially when the number of candidate features is large. In this paper, we propose a semi-supervised (SS) EHR phenotyping method that borrows information from both a small, labeled dataset (where both the label Y and the feature set X are observed) and a much larger, weakly-labeled dataset in which the feature set X is accompanied only by a surrogate label S that is available to all patients. Under a working prior assumption that S is related to X only through Y and allowing it to hold approximately, we propose a prior adaptive semi-supervised (PASS) estimator that incorporates the prior knowledge by shrinking the estimator towards a direction derived under the prior. We derive asymptotic theory for the proposed estimator and justify its efficiency and robustness to prior information of poor quality. We also demonstrate its superiority over existing estimators under various scenarios via simulation studies and on three real-world EHR phenotyping studies at a large tertiary hospital. Molei Liu, Matey Neykov, Tianxi Cai |
J. Mach. Learn. Res. | 4 |
| 2021 | A high-throughput phenotyping algorithm is portable from adult to pediatric populationsabstractOBJECTIVE: Multimodal automated phenotyping (MAP) is a scalable, high-throughput phenotyping method, developed using electronic health record (EHR) data from an adult population. We tested transportability of MAP to a pediatric population. MATERIALS AND METHODS: Without additional feature engineering or supervised training, we applied MAP to a pediatric population enrolled in a biobank and evaluated performance against physician-reviewed medical records. We also compared performance of MAP at the pediatric institution and the original adult institution where MAP was developed, including for 6 phenotypes validated at both institutions against physician-reviewed medical records. RESULTS: MAP performed equally well in the pediatric setting (average AUC 0.98) as it did at the general adult hospital system (average AUC 0.96). MAP's performance in the pediatric sample was similar across the 6 specific phenotypes also validated against gold-standard labels in the adult biobank. CONCLUSIONS: MAP is highly transportable across diverse populations and has potential for wide-scale use. Alon Geva, Molei Liu, Vidul Ayakulangara Panickan, Paul Avillach, Tianxi Cai, Kenneth D. Mandl |
J. Am. Medical Informatics Assoc. | 5 |
| 2021 | ATLAS: an automated association test using probabilistically linked health records with application to genetic studiesabstractOBJECTIVE: Large amounts of health data are becoming available for biomedical research. Synthesizing information across databases may capture more comprehensive pictures of patient health and enable novel research studies. When no gold standard mappings between patient records are available, researchers may probabilistically link records from separate databases and analyze the linked data. However, previous linked data inference methods are constrained to certain linkage settings and exhibit low power. Here, we present ATLAS, an automated, flexible, and robust association testing algorithm for probabilistically linked data. MATERIALS AND METHODS: Missing variables are imputed at various thresholds using a weighted average method that propagates uncertainty from probabilistic linkage. Next, estimated effect sizes are obtained using a generalized linear model. ATLAS then conducts the threshold combination test by optimally combining P values obtained from data imputed at varying thresholds using Fisher's method and perturbation resampling. RESULTS: In simulations, ATLAS controls for type I error and exhibits high power compared to previous methods. In a real-world genetic association study, meta-analysis of ATLAS-enabled analyses on a linked cohort with analyses using an existing cohort yielded additional significant associations between rheumatoid arthritis genetic risk score and laboratory biomarkers. DISCUSSION: Weighted average imputation weathers false matches and increases contribution of true matches to mitigate linkage error-induced bias. The threshold combination test avoids arbitrarily choosing a threshold to rule a match, thus automating linked data-enabled analyses and preserving power. CONCLUSION: ATLAS promises to enable novel and powerful research studies using linked data to capitalize on all available data sources. Harrison G. Zhang, Boris P. Hejblum, Griffin M. Weber, Nathan P. Palmer, Susanne E. Churchill, Peter Szolovits, Shawn N. Murphy, Katherine P. Liao, Isaac S. Kohane, Tianxi Cai |
J. Am. Medical Informatics Assoc. | 10 |
| 2021 | Integrative High Dimensional Multiple Testing with Heterogeneity under Data Sharing ConstraintsabstractIdentifying informative predictors in a high-dimensional regression model is a critical step for association analysis and predictive modeling. Signal detection in the high dimensional setting often fails due to the limited sample size. One approach to improving power is through meta-analyzing multiple studies which address the same scientific question. However, integrative analysis of high dimensional data from multiple studies is challenging in the presence of between-study heterogeneity. The challenge is even more pronounced with additional data sharing constraints under which only summary data can be shared across different sites. In this paper, we propose a novel data shielding integrative large--scale testing (DSILT) approach to signal detection allowing between-study heterogeneity and not requiring the sharing of individual-level data. Assuming the underlying high dimensional regression models of the data differ across studies yet share similar support, the proposed method incorporates proper integrative estimation and debiasing procedures to construct test statistics for the overall effects of specific covariates. We also develop a multiple testing procedure to identify significant effects while controlling the false discovery rate (FDR) and false discovery proportion (FDP). Theoretical comparisons of the new testing procedure with the ideal individual-level meta-analysis (ILMA) approach and other distributed inference methods are investigated. Simulation studies demonstrate that the proposed testing procedure performs well in both controlling false discovery and attaining power. The new method is applied to a real example detecting interaction effects of the genetic variants for statins and obesity on the risk for type II diabetes. Molei Liu, Yin Xia, Kelly Cho, Tianxi Cai |
J. Mach. Learn. Res. | 4 |
| 2020 | SAMGEP: A Novel Method for Prediction of Phenotype Event Times Using the Electronic Health Record
Yuri Ahuja, Chuan Hong, Zongqi Xia, Tianxi Cai |
AMIA | 4 |
| 2020 | Expert-Supervised Reinforcement Learning for Offline Policy Learning and EvaluationabstractOffline Reinforcement Learning (RL) is a promising approach for learning optimal policies in environments where direct exploration is expensive or unfeasible. However, the adoption of such policies in practice is often challenging, as they are hard to interpret within the application context, and lack measures of uncertainty for the learned policy value and its decisions. To overcome these issues, we propose an Expert-Supervised RL (ESRL) framework which uses uncertainty quantification for offline policy learning. In particular, we have three contributions: 1) the method can learn safe and optimal policies through hypothesis testing, 2) ESRL allows for different levels of risk averse implementations tailored to the application context, and finally, 3) we propose a way to interpret ESRL’s policy at every state through posterior distributions, and use this framework to compute off-policy value function posteriors. We provide theoretical guarantees for our estimators and regret bounds consistent with Posterior Sampling for RL (PSRL). Sample efficiency of ESRL is independent of the chosen risk aversion threshold and quality of the behavior policy. Aaron Sonabend W., Leo A. Celi, Tianxi Cai, Peter Szolovits |
NeurIPS | 4 |
| 2020 | sureLDA: A multidisease automated phenotyping method for the electronic health recordabstractOBJECTIVE: A major bottleneck hindering utilization of electronic health record data for translational research is the lack of precise phenotype labels. Chart review as well as rule-based and supervised phenotyping approaches require laborious expert input, hampering applicability to studies that require many phenotypes to be defined and labeled de novo. Though International Classification of Diseases codes are often used as surrogates for true labels in this setting, these sometimes suffer from poor specificity. We propose a fully automated topic modeling algorithm to simultaneously annotate multiple phenotypes. MATERIALS AND METHODS: Surrogate-guided ensemble latent Dirichlet allocation (sureLDA) is a label-free multidimensional phenotyping method. It first uses the PheNorm algorithm to initialize probabilities based on 2 surrogate features for each target phenotype, and then leverages these probabilities to constrain the LDA topic model to generate phenotype-specific topics. Finally, it combines phenotype-feature counts with surrogates via clustering ensemble to yield final phenotype probabilities. RESULTS: sureLDA achieves reliably high accuracy and precision across a range of simulated and real-world phenotypes. Its performance is robust to phenotype prevalence and relative informativeness of surogate vs nonsurrogate features. It also exhibits powerful feature selection properties. DISCUSSION: sureLDA combines attractive properties of PheNorm and LDA to achieve high accuracy and precision robust to diverse phenotype characteristics. It offers particular improvement for phenotypes insufficiently captured by a few surrogate features. Moreover, sureLDA's feature selection ability enables it to handle high feature dimensions and produce interpretable computational phenotypes. CONCLUSIONS: sureLDA is well suited toward large-scale electronic health record phenotyping for highly multiphenotype applications such as phenome-wide association studies . Yuri Ahuja, Doudou Zhou, Zeling He, Jiehuan Sun, Victor M. Castro, Vivian S. Gainer, Shawn N. Murphy, Chuan Hong, Tianxi Cai |
J. Am. Medical Informatics Assoc. | 9 |
| 2019 | sureLDA: A Novel Multi-Disease Automated Phenotyping Method for the Electronic Health Record
Yuri Ahuja, Doudou Zhou, Zeling He, Jiehuan Sun, Victor M. Castro, Vivian S. Gainer, Shawn N. Murphy, Chuan Hong, Tianxi Cai |
AMIA | 9 |
| 2019 | Feature extraction for phenotyping from semantic and knowledge resources
Wenxin Ning, Stephanie Chan 0002, Andrew L. Beam, Ming Yu 0003, Alon Geva, Katherine P. Liao, Mary Mullen, Kenneth D. Mandl, Isaac S. Kohane, Tianxi Cai, Sheng Yu 0002 |
J. Biomed. Informatics | 10 |
| 2019 | Automated grouping of medical codes via multiview banded spectral clusteringabstractOBJECTIVE: With its increasingly widespread adoption, electronic health records (EHR) have enabled phenotypic information extraction at an unprecedented granularity and scale. However, often a medical concept (e.g. diagnosis, prescription, symptom) is described in various synonyms across different EHR systems, hindering data integration for signal enhancement and complicating dimensionality reduction for knowledge discovery. Despite existing ontologies and hierarchies, tremendous human effort is needed for curation and maintenance - a process that is both unscalable and susceptible to subjective biases. This paper aims to develop a data-driven approach to automate grouping medical terms into clinically relevant concepts by combining multiple up-to-date data sources in an unbiased manner. METHODS: We present a novel data-driven grouping approach - multi-view banded spectral clustering (mvBSC) combining summary data from multiple healthcare systems. The proposed method consists of a banding step that leverages the prior knowledge from the existing coding hierarchy, and a combining step that performs spectral clustering on an optimally weighted matrix. RESULTS: -measure, and were found to consistently exhibit great similarity to the existing manual grouping counterpart. The resulting ICD groupings also enjoy comparable interpretability and are well aligned with the current ICD hierarchy. CONCLUSION: The proposed approach, by systematically leveraging multiple data sources, is able to overcome bias while maximizing consensus to achieve generalizability. It has the advantage of being efficient, scalable, and adaptive to the evolving human knowledge reflected in the data, showing a significant step toward automating medical knowledge integration. Luwan Zhang, Tianrun A. Cai, Yuri Ahuja, Zeling He, Yuk-Lam Ho, Andrew L. Beam, Kelly Cho, Robert J. Carroll, Joshua C. Denny, Isaac S. Kohane, Katherine P. Liao, Tianxi Cai |
J. Biomed. Informatics | 13 |
| 2018 | High-Throughput Multimodal Automated Phenotyping (MAP) Incorporating Natural Language Processing with Application to PheWAS
Katherine P. Liao, Jiehuan Sun, Tianrun A. Cai, Nicholas B. Link, Chuan Hong, Jie Huang 0030, Jennifer E. Huffman, Jessica L. Gronsbell, Lauren Costa, Victor M. Castro, Vivian S. Gainer, Shawn N. Murphy, John Michael Gaziano, Kelly Cho, Peter Szolovits, Isaac S. Kohane, Sheng Yu 0002, Tianxi Cai |
AMIA | 18 |
| 2018 | Enabling phenotypic big data with PheNormabstractObjective: Electronic health record (EHR)-based phenotyping infers whether a patient has a disease based on the information in his or her EHR. A human-annotated training set with gold-standard disease status labels is usually required to build an algorithm for phenotyping based on a set of predictive features. The time intensiveness of annotation and feature curation severely limits the ability to achieve high-throughput phenotyping. While previous studies have successfully automated feature curation, annotation remains a major bottleneck. In this paper, we present PheNorm, a phenotyping algorithm that does not require expert-labeled samples for training. Methods: The most predictive features, such as the number of International Classification of Diseases, Ninth Revision, Clinical Modification (ICD-9-CM) codes or mentions of the target phenotype, are normalized to resemble a normal mixture distribution with high area under the receiver operating curve (AUC) for prediction. The transformed features are then denoised and combined into a score for accurate disease classification. Results: We validated the accuracy of PheNorm with 4 phenotypes: coronary artery disease, rheumatoid arthritis, Crohn's disease, and ulcerative colitis. The AUCs of the PheNorm score reached 0.90, 0.94, 0.95, and 0.94 for the 4 phenotypes, respectively, which were comparable to the accuracy of supervised algorithms trained with sample sizes of 100-300, with no statistically significant difference. Conclusion: The accuracy of the PheNorm algorithms is on par with algorithms trained with annotated samples. PheNorm fully automates the generation of accurate phenotyping algorithms and demonstrates the capacity for EHR-driven annotations to scale to the next level - phenotypic big data. Sheng Yu 0002, Yumeng Ma, Jessica L. Gronsbell, Tianrun A. Cai, Ashwin N. Ananthakrishnan, Vivian S. Gainer, Susanne E. Churchill, Peter Szolovits, Shawn N. Murphy, Isaac S. Kohane, Katherine P. Liao, Tianxi Cai |
J. Am. Medical Informatics Assoc. | 12 |
| 2017 | Improving EHR Chart Review Efficiency via Semantic Similarity Assessment
Tianrun A. Cai, Andrew L. Beam, Stephanie Chan 0002, Jacqueline Honerlaw, David R. Gagnon, Kelly Cho, John Michael Gaziano, Katherine P. Liao, Tianxi Cai |
AMIA | 10 |
| 2017 | Using Electronic Health Records Data for Comparative Effectiveness Research
Ashwin N. Ananthakrishnan, Tianxi Cai |
AMIA | 3 |
| 2017 | High-throughput Phenotyping via Denoised Normal Mixture Transformation
Sheng Yu 0002, Yumeng Ma, Jessica L. Gronsbell, Katherine P. Liao, Tianrun A. Cai, Ashwin N. Ananthakrishnan, Vivian S. Gainer, Susanne E. Churchill, Peter Szolovits, Shawn N. Murphy, Isaac S. Kohane, Tianxi Cai |
AMIA | 12 |
| 2017 | Rare variant association test in family-based sequencing studiesabstractThe objective of this article is to introduce valid and robust methods for the analysis of rare variants for family-based exome chips, whole-exome sequencing or whole-genome sequencing data. Family-based designs provide unique opportunities to detect genetic variants that complement studies of unrelated individuals. Currently, limited methods and software tools have been developed to assist family-based association studies with rare variants, especially for analyzing binary traits. In this article, we address this gap by extending existing burden and kernel-based gene set association tests for population data to related samples, with a particular emphasis on binary phenotypes. The proposed approach blends the strengths of kernel machine methods and generalized estimating equations. Importantly, the efficient generalized kernel score test can be applied as a mega-analysis framework to combine studies with different designs. We illustrate the application of the proposed method using data from an exome sequencing study of autism. Methods discussed in this article are implemented in an R package 'gskat', which is available on CRAN and GitHub. Nathan Morris, Tianxi Cai, Seunggeun Lee, Chaolong Wang, Timothy W. Yu, Christopher A. Walsh, Xihong Lin |
Briefings Bioinform. | 4 |
| 2017 | Surrogate-assisted feature extraction for high-throughput phenotypingabstractOBJECTIVE: Phenotyping algorithms are capable of accurately identifying patients with specific phenotypes from within electronic medical records systems. However, developing phenotyping algorithms in a scalable way remains a challenge due to the extensive human resources required. This paper introduces a high-throughput unsupervised feature selection method, which improves the robustness and scalability of electronic medical record phenotyping without compromising its accuracy. METHODS: The proposed Surrogate-Assisted Feature Extraction (SAFE) method selects candidate features from a pool of comprehensive medical concepts found in publicly available knowledge sources. The target phenotype's International Classification of Diseases, Ninth Revision and natural language processing counts, acting as noisy surrogates to the gold-standard labels, are used to create silver-standard labels. Candidate features highly predictive of the silver-standard labels are selected as the final features. RESULTS: Algorithms were trained to identify patients with coronary artery disease, rheumatoid arthritis, Crohn's disease, and ulcerative colitis using various numbers of labels to compare the performance of features selected by SAFE, a previously published automated feature extraction for phenotyping procedure, and domain experts. The out-of-sample area under the receiver operating characteristic curve and F -score from SAFE algorithms were remarkably higher than those from the other two, especially at small label sizes. CONCLUSION: SAFE advances high-throughput phenotyping methods by automatically selecting a succinct set of informative features for algorithm training, which in turn reduces overfitting and the needed number of gold-standard labels. SAFE also potentially identifies important features missed by automated feature extraction for phenotyping or experts. Sheng Yu 0002, Abhishek Chakrabortty, Katherine P. Liao, Tianrun A. Cai, Ashwin N. Ananthakrishnan, Vivian S. Gainer, Susanne E. Churchill, Peter Szolovits, Shawn N. Murphy, Isaac S. Kohane, Tianxi Cai |
J. Am. Medical Informatics Assoc. | 11 |
| 2016 | On the Characterization of a Class of Fisher-Consistent Loss Functions and its Application to BoostingabstractAccurate classification of categorical outcomes is essential in a wide range of applications. Due to computational issues with minimizing the empirical 0/1 loss, Fisher consistent losses have been proposed as viable proxies. However, even with smooth losses, direct minimization remains a daunting task. To approximate such a minimizer, various boosting algorithms have been suggested. For example, with exponential loss, the AdaBoost algorithm (Freund and Schapire, 1995) is widely used for two- class problems and has been extended to the multi-class setting (Zhu et al., 2009). Alternative loss functions, such as the logistic and the hinge losses, and their corresponding boosting algorithms have also been proposed (Zou et al., 2008; Wang, 2012). In this paper we demonstrate that a broad class of losses, including non-convex functions, achieve Fisher consistency, and in addition can be used for explicit estimation of the conditional class probabilities. Furthermore, we provide a generic boosting algorithm that is not loss-specific. Extensive simulation results suggest that the proposed boosting algorithms could outperform existing methods with properly chosen losses and bags of weak learners. Matey Neykov, Jun S. Liu, Tianxi Cai |
J. Mach. Learn. Res. | 3 |
| 2016 | L1-Regularized Least Squares for Support Recovery of High Dimensional Single Index Models with Gaussian DesignsabstractIt is known that for a certain class of single index models (SIMs) $Y = f(X_{p \times 1}^\top\beta_0, \varepsilon)$, support recovery is impossible when $X \sim \mathcal{N}(0, I_{p \times p})$ and a model complexity adjusted sample size is below a critical threshold. Recently, optimal algorithms based on Sliced Inverse Regression (SIR) were suggested. These algorithms work provably under the assumption that the design $X$ comes from an i.i.d. Gaussian distribution. In the present paper we analyze algorithms based on covariance screening and least squares with $L_1$ penalization (i.e. LASSO) and demonstrate that they can also enjoy optimal (up to a scalar) rescaled sample size in terms of support recovery, albeit under slightly different assumptions on $f$ and $\varepsilon$ compared to the SIR based algorithms. Furthermore, we show more generally, that LASSO succeeds in recovering the signed support of $\beta_0$ if $X \sim \mathcal{N}(0, \Sigma)$, and the covariance $\Sigma$ satisfies the irrepresentable condition. Our work extends existing results on the support recovery of LASSO for the linear model, to a more general class of SIMs. Matey Neykov, Jun S. Liu, Tianxi Cai |
J. Mach. Learn. Res. | 3 |
| 2015 | Demonstrating the Advantages of Applying Data Mining Techniques on Time-Dependent Electronic Medical Records
Uri Kartoun, Vishesh Kumar, Su-Chun Cheng, Sheng Yu 0002, Katherine P. Liao, Elizabeth W. Karlson, Ashwin N. Ananthakrishnan, Zongqi Xia, Vivian S. Gainer, Andrew Cagan, Guergana K. Savova, Pei J. Chen, Shawn N. Murphy, Susanne E. Churchill, Isaac S. Kohane, Peter Szolovits, Tianxi Cai, Stanley Y. Shaw |
AMIA | 17 |
| 2015 | Computable Phenotypes enabled by the i2b2 Validation Platform
Shawn N. Murphy, Vivian S. Gainer, Victor M. Castro, Alyssa P. Goodson, Lori C. Phillips, Sheng Yu 0002, Tianxi Cai |
AMIA | 7 |
| 2015 | Toward high-throughput phenotyping: unbiased automated feature extraction and selection from knowledge sourcesabstractOBJECTIVE: Analysis of narrative (text) data from electronic health records (EHRs) can improve population-scale phenotyping for clinical and genetic research. Currently, selection of text features for phenotyping algorithms is slow and laborious, requiring extensive and iterative involvement by domain experts. This paper introduces a method to develop phenotyping algorithms in an unbiased manner by automatically extracting and selecting informative features, which can be comparable to expert-curated ones in classification accuracy. MATERIALS AND METHODS: Comprehensive medical concepts were collected from publicly available knowledge sources in an automated, unbiased fashion. Natural language processing (NLP) revealed the occurrence patterns of these concepts in EHR narrative notes, which enabled selection of informative features for phenotype classification. When combined with additional codified features, a penalized logistic regression model was trained to classify the target phenotype. RESULTS: The authors applied our method to develop algorithms to identify patients with rheumatoid arthritis and coronary artery disease cases among those with rheumatoid arthritis from a large multi-institutional EHR. The area under the receiver operating characteristic curves (AUC) for classifying RA and CAD using models trained with automated features were 0.951 and 0.929, respectively, compared to the AUCs of 0.938 and 0.929 by models trained with expert-curated features. DISCUSSION: Models trained with NLP text features selected through an unbiased, automated procedure achieved comparable or slightly higher accuracy than those trained with expert-curated features. The majority of the selected model features were interpretable. CONCLUSION: The proposed automated feature extraction method, generating highly accurate phenotyping algorithms with improved efficiency, is a significant step toward high-throughput phenotyping. Sheng Yu 0002, Katherine P. Liao, Stanley Y. Shaw, Vivian S. Gainer, Susanne E. Churchill, Peter Szolovits, Shawn N. Murphy, Isaac S. Kohane, Tianxi Cai |
J. Am. Medical Informatics Assoc. | 9 |
| 2014 | Classification of CT pulmonary angiography reports by presence, chronicity, and location of pulmonary embolism with natural language processing
Sheng Yu 0002, Kanako K. Kumamaru, Elizabeth George, Ruth M. Dunne, Arash Bedayat, Matey Neykov, Andetta R. Hunsaker, Karin E. Dill, Tianxi Cai, Frank J. Rybicki |
J. Biomed. Informatics | 9 |
| 2012 | Portability of an algorithm to identify rheumatoid arthritis in electronic health recordsabstractOBJECTIVES: Electronic health records (EHR) can allow for the generation of large cohorts of individuals with given diseases for clinical and genomic research. A rate-limiting step is the development of electronic phenotype selection algorithms to find such cohorts. This study evaluated the portability of a published phenotype algorithm to identify rheumatoid arthritis (RA) patients from EHR records at three institutions with different EHR systems. MATERIALS AND METHODS: Physicians reviewed charts from three institutions to identify patients with RA. Each institution compiled attributes from various sources in the EHR, including codified data and clinical narratives, which were searched using one of two natural language processing (NLP) systems. The performance of the published model was compared with locally retrained models. RESULTS: Applying the previously published model from Partners Healthcare to datasets from Northwestern and Vanderbilt Universities, the area under the receiver operating characteristic curve was found to be 92% for Northwestern and 95% for Vanderbilt, compared with 97% at Partners. Retraining the model improved the average sensitivity at a specificity of 97% to 72% from the original 65%. Both the original logistic regression models and locally retrained models were superior to simple billing code count thresholds. DISCUSSION: These results show that a previously published algorithm for RA is portable to two external hospitals using different EHR systems, different NLP systems, and different target NLP vocabularies. Retraining the algorithm primarily increased the sensitivity at each site. CONCLUSION: Electronic phenotype algorithms allow rapid identification of case populations in multiple sites with little retraining. Robert J. Carroll, William K. Thompson, Anne E. Eyler, Arthur M. Mandelin, Tianxi Cai, Raquel M. Zink, Jennifer A. Pacheco, Chad S. Boomershine, Thomas A. Lasko, Hua Xu 0001, Elizabeth W. Karlson, Raúl G. Pérez, Vivian S. Gainer, Shawn N. Murphy, Eric M. Ruderman, Richard M. Pope, Robert M. Plenge, Abel N. Kho, Katherine P. Liao, Joshua C. Denny |
J. Am. Medical Informatics Assoc. | 5 |
| 2008 | Estimating the Confidence Interval for Prediction Errors of Support Vector Machine Classifiers
Bo Jiang 0005, Xuegong Zhang, Tianxi Cai |
J. Mach. Learn. Res. | 3 |