VLDB 2026 Research / reviewers in the wild / expert
Yong Chen 0016
dblp:67/6351-16
· DBLP profile ↗
58ranked-venue papers
1as first author
37since 2021 · last 2026
0000-0003-0835-0788ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 46 · 1 first-author · 31 since 2021Artificial intelligence and machine learning · 9 · 5 since 2021Databases, data management, data science and information retrieval · 5 · 1 since 2021Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | dGAMLSS: an exact, distributed algorithm to fit Generalized Additive Models for Location, Scale, and Shape for privacy-preserving population reference chartsabstractMOTIVATION: There is growing interest in estimating population reference ranges across age and sex to better identify atypical clinically-relevant measurements throughout the lifespan. For this task, the World Health Organization recommends using Generalized Additive Models for Location, Scale, and Shape (GAMLSS), which can model non-linear growth trajectories under complex distributions that address the heterogeneity in human populations.Fitting GAMLSS models requires large, generalizable sample sizes, especially for accurate estimation of extreme quantiles, but obtaining such multi-site data can be challenging due to privacy concerns and practical considerations. In settings where patient data cannot be shared, privacy-preserving distributed algorithms for federated learning can be used, but no such algorithm exists for GAMLSS. RESULTS: We propose distributed GAMLSS (dGAMLSS), a distributed algorithm that can fit GAMLSS models across multiple sites without sharing patient-level data. This includes specific considerations for the fitting of smooth functions at varying levels of communication efficiency. We demonstrate the effectiveness of dGAMLSS in constructing population reference charts across clinical, genomics, and neuroimaging settings and show that dGAMLSS is able to reproduce pooled reference charts and inference down to numerical differences. AVAILABILITY AND IMPLEMENTATION: An R package providing examples of the dGAMLSS algorithm, as well as functions for sharing and aggregating site-specific parameters, is available at https://github.com/hufengling/dGAMLSS. Fengling Hu, Jiayi Tong, Margaret Gardner, Lifespan Brain Chart Consortium, Andrew A. Chen, Richard A. I. Bethlehem, Jakob Seidlitz, Hongzhe Li, Aaron Alexander-Bloch, Yong Chen 0016, Russell T. Shinohara |
Bioinform. | 10 |
| 2026 | A lossless one-shot distributed algorithm for addressing heterogeneity in multi-site generalized linear modelsabstractOBJECTIVE: We propose Heterogeneity-aware Collaborative One-shot Lossless Algorithm for Generalized Linear Model (COLA-GLM-H), a novel one-shot lossless distributed algorithm that enables the integration of heterogeneous multi-institutional data while relying solely on instituion-level summary information rather than patient-level data. MATERIALS AND METHODS: Generalized Linear Models (GLMs) are widely used in medical research for analyzing diverse outcome types. In multi-institution settings, we demonstrated that the global likelihood can be reconstructed using only institution-level summary statistics, enabling lossless estimation without accessing individual records. We validated COLA-GLM-H in two real-world studies: (1) an emulated U.S. pediatric centralized network (719,383 patients) evaluating long-term cardiovascular risks following COVID-19, and (2) an internationally decentralized network of 120,429 hospitalized patients from seven databases across three countries assessing risk factors for COVID-19 mortality. RESULTS: In the centralized network, COLA-GLM-H produced estimates identical to those from pooled analyses. In the decentralized setting, the algorithm effectively integrated heterogeneous data across multiple clinical institutions using a single communication round. CONCLUSIONS: COLA-GLM-H provides a lossless, communication-efficient, and computation-efficient solution for multi-institutional research using only institution-level summary data. It accounts for between-institution heterogeneity and supports all outcome types within the exponential family, enabling secure, scalable, and accurate analysis in collaborative clinical research. Bingyu Zhang, Jenna Reps, Jiayi Tong, Dazheng Zhang, Juan Manuel Ramírez-Anguita, Jiang Bian 0001, Milou T. Brand, Thomas Falconer, Miguel A. Mayer, Ross D. Williams, Yong Chen 0016 |
J. Am. Medical Informatics Assoc. | 14 |
| 2025 | MAP: Low-compute Model Merging with Amortized Pareto Fronts via Quadratic ApproximationabstractModel merging has emerged as an effective approach to combining multiple single-task models into a multitask model. This process typically involves computing a weighted average of the model parameters without additional training. Existing model-merging methods focus on improving average task accuracy. However, interference and conflicts between the objectives of different tasks can lead to trade-offs during the merging process. In real-world applications, a set of solutions with various trade-offs can be more informative, helping practitioners make decisions based on diverse preferences. In this paper, we introduce a novel and low-compute algorithm, Model Merging with Amortized Pareto Front (MAP). MAP efficiently identifies a Pareto set of scaling coefficients for merging multiple models, reflecting the trade-offs involved. It amortizes the substantial computational cost of evaluations needed to estimate the Pareto front by using quadratic approximation surrogate models derived from a preselected set of scaling coefficients. Experimental results on vision and natural language processing tasks demonstrate that MAP can accurately identify the Pareto front, providing practitioners with flexible solutions to balance competing task objectives. We also introduce Bayesian MAP for scenarios with a relatively low number of tasks and Nested MAP for situations with a high number of tasks, further reducing the computational cost of evaluation. Zhiqi Bu, Suyuchen Wang, Jie Fu 0001, Yonghui Wu 0001, Jiang Bian 0002, Yong Chen 0016, Yoshua Bengio |
ICLR | 9 |
| 2025 | Objective study validity diagnostics: a framework requiring pre-specified, empirical verification to increase trust in the reliability of real-world evidenceabstractOBJECTIVE: Propose a framework to empirically evaluate and report validity of findings from observational studies using pre-specified objective diagnostics, increasing trust in real-world evidence (RWE). MATERIALS AND METHODS: The framework employs objective diagnostic measures to assess the appropriateness of study designs, analytic assumptions, and threats to validity in generating reliable evidence addressing causal questions. Diagnostic evaluations should be interpreted before the unblinding of study results or, alternatively, only unblind results from analyses that pass pre-specified thresholds. We provide a conceptual overview of objective diagnostic measures and demonstrate their impact on the validity of RWE from a large-scale comparative new-user study of various antihypertensive medications. We evaluated expected absolute systematic error (EASE) before and after applying diagnostic thresholds, using a large set of negative control outcomes. RESULTS: Applying objective diagnostics reduces bias and improves evidence reliability in observational studies. Among 11 716 analyses (EASE = 0.38), 13.9% met pre-specified diagnostic thresholds which reduced EASE to zero. Objective diagnostics provide a comprehensive and empirical set of tests that increase confidence when passed and raise doubts when failed. DISCUSSION: The increasing use of real-world data presents a scientific opportunity; however, the complexity of the evidence generation process poses challenges for understanding study validity and trusting RWE. Deploying objective diagnostics is crucial to reducing bias and improving reliability in RWE generation. Under ideal conditions, multiple study designs pass diagnostics and generate consistent results, deepening understanding of causal relationships. Open-source, standardized programs can facilitate implementation of diagnostic analyses. CONCLUSION: Objective diagnostics are a valuable addition to the RWE generation process. Mitchell Conover, Patrick B. Ryan, Yong Chen 0016, Marc A. Suchard, George Hripcsak, Martijn J. Schuemie |
J. Am. Medical Informatics Assoc. | 3 |
| 2025 | Communication-efficient federated learning of temporal effects on opioid use disorder with data from distributed research networksabstractOBJECTIVE: To develop a distributed algorithm to fit multi-center Cox regression models with time-varying coefficients to facilitate privacy-preserving data integration across multiple health systems. MATERIALS AND METHODS: The Cox model with time-varying coefficients relaxes the proportional hazards assumption of the usual Cox model and is particularly useful to model time-to-event outcomes. We proposed a One-shot Distributed Algorithm to fit multi-center Cox regression models with Time varying coefficients (ODACT). This algorithm constructed a surrogate likelihood function to approximate the Cox partial likelihood function, using patient-level data from a lead site and aggregated data from other sites. The performance of ODACT was demonstrated by simulation and a real-world study of opioid use disorder (OUD) using decentralized data from a large clinical research network across 5 sites with 69 163 subjects. RESULTS: The ODACT method precisely estimated the time-varying effects over time. In the simulation study, ODACT always achieved estimation close to that of the pooled analysis, while the meta-estimator showed considerable amount of bias. In the OUD study, the bias of the estimated hazard ratios by ODACT are smaller than those of the meta-estimator for all 7 risk factors at almost all of the time points from 0 to 2.5 years. The greatest bias of the meta-estimator was for the effects of age ≥65 years, and smoking. CONCLUSION: ODACT is a privacy-preserving and communication-efficient method for analyzing multi-center time-to-event data which allows the covariates' effects to be time-varying. ODACT provides estimates close to the pooled estimator and substantially outperforms the meta-analysis estimator. DISCUSSION: The proposed ODACT is a privacy-preserving distributed algorithm for fitting Cox models with time-varying coefficients. The limitations of ODACT include that privacy-preserving via aggregate data does rely on relatively large number of data at each individual site, and rigorous quantification of the risk of privacy leaks requires further investigation. C. Jason Liang, Chongliang Luo, Henry R. Kranzler, Jiang Bian 0001, Yong Chen 0016 |
J. Am. Medical Informatics Assoc. | 5 |
| 2025 | Incorporating preprints in systematic reviews: a preliminary study of a novel method for rapid evidence synthesisabstractOBJECTIVES: By October 1, 2024, over 450,000 COVID-19 manuscripts were published, with 10% posted as unreviewed preprints. While they accelerate knowledge sharing, their inconsistent quality complicates systematic studies. MATERIALS AND METHODS: We propose a 2-stage method to include preprints in meta-analyses. In Stage A, preprints are integrated through restriction or imputation and weighted by a confidence score reflecting their publication likelihood. In Stage B, we assess and adjust for potential publication or reporting biases. RESULTS: This preliminary study employed a 2-stage procedure validated with 2 COVID-19 treatment case studies. For hydroxychloroquine, the relative risk (RR) was 1.06 [95% CI: 0.62, 1.80], suggesting no mortality benefit over placebo. For corticosteroids, the RR was 0.88 [95% CI: 0.62, 1.27], which, while not statistically significant, aligns with evidence supporting a mortality benefit. DISCUSSION: Our research aims to bridge a significant methodological gap by providing a solution for timely evidence synthesis, particularly in the face of the overwhelming number of publications surrounding COVID-19. CONCLUSION: This preliminary study presents a method to efficiently synthesize COVID-19 research, including non-peer-reviewed preprints, to support clinical and policy decisions amidst the information surge. Jiayi Tong, Yifei Sun 0007, Rebecca A. Hubbard, M. Elle Saine, Hua Xu 0001, Xu Zuo, Chunhua Weng, Christopher H. Schmid, Stephen E. Kimmel, Craig A. Umscheid, Adam Cuker, Yong Chen 0016 |
J. Am. Medical Informatics Assoc. | 13 |
| 2025 | A communication-efficient federated learning algorithm to assess racial disparities in post-transplantation survival timeabstractOBJECTIVE: Patients of different race have different outcomes following renal transplantation. Patients of different race also undergo renal transplantation at different hospitals. We used a novel decentralized multisite approach to quantitatively assess the effect of site of care on racial disparities between non-Hispanic Black (NHB) and non-Hispanic White (NHW) patients in post-transplantation survival times. MATERIALS AND METHODS: In this study, we develop a communication-efficient federated learning algorithm to assess site-of-care associated racial disparities based on decentralized time-to-event data, called Communication-Efficient Distributed Analysis for Racial Disparity in Time-to-event Data (CEDAR-t2e). The algorithm includes 2 modules. Module I is to estimate the site-specific proportional hazards model for time-to-event outcomes in a distributed manner, in which the Poissonization is used to simplify the estimation procedure. Based on the estimated results from Module I, Module II calculates how long the kidney failure time of NHB patients would be extended had they been admitted to transplant centers in the same distribution as NHW patients were admitted. RESULTS: With application to United States Renal Data System data covering 39 043 patients across 73 transplant centers, we found no evidence suggesting the presence of site-of-care associated racial disparities in post-transplantation survival times. In particular, restricting to one year after transplantation, the counterfactual graft failure time would have been extended by only 0.61 days on average if NHB had the same admission distribution to transplant centers as NHW patients. DISCUSSION: The proposed approach offers a quantitative measure to evaluate site-of-care associated racial disparities. CONCLUSION: Our approach has the potential to be extended to investigate site-of-care related disparities in other time-to-event outcomes, thus promoting health equity and improving patient health in various fields. Dazheng Zhang, Jiayi Tong, Xing He 0003, Liang Li 0026, Lichao Sun 0001, Ashutosh M. Shukla, Jiang Bian 0001, David A. Asch, Yong Chen 0016 |
J. Am. Medical Informatics Assoc. | 10 |
| 2025 | Enabling inclusive systematic reviews: incorporating preprint articles with large language model-driven evaluationsabstractOBJECTIVES: Systematic reviews in comparative effectiveness research require timely evidence synthesis. With the rapid advancement of medical research, preprint articles play an increasingly important role in accelerating knowledge dissemination. However, as preprint articles are not peer-reviewed before publication, their quality varies significantly, posing challenges for evidence inclusion in systematic reviews. MATERIALS AND METHODS: We developed AutoConfidenceScore (automated confidence score assessment), an advanced framework for predicting preprint publication, which reduces reliance on manual curation and expands the range of predictors, including three key advancements: (1) automated data extraction using natural language processing techniques, (2) semantic embeddings of titles and abstracts, and (3) large language model (LLM)-driven evaluation scores. Additionally, we employed two prediction models: a random forest classifier for binary outcome and a survival cure model that predicts both binary outcome and publication risk over time. RESULTS: The random forest classifier achieved an area under the receiver operating characteristic curve (AUROC) of 0.747 using all features. The survival cure model achieved an AUROC of 0.731 for binary outcome prediction and a concordance index of 0.667 for time-to-publication risk. DISCUSSION: Our study advances the framework for preprint publication prediction through automated data extraction and multiple feature integration. By combining semantic embeddings with LLM-driven evaluations, AutoConfidenceScore significantly enhances predictive performance while reducing manual annotation burden. CONCLUSION: AutoConfidenceScore has the potential to facilitate incorporation of preprint articles during the appraisal phase of systematic reviews, supporting researchers in more effective utilization of preprint resources. Rui Yang 0016, Jiayi Tong, Nan Liu 0003, Christopher J. Lindsell, Michael J. Pencina, Yong Chen 0016, Chuan Hong |
J. Am. Medical Informatics Assoc. | 10 |
| 2025 | Leveraging undecided cases in chart-reviewed phenotypes to enhance EHR-based association studies
Xinyao Jian, Dazheng Zhang, Zehao Yu 0001, Hua Xu 0001, Jiang Bian 0001, Yonghui Wu 0001, Jiayi Tong, Yong Chen 0016 |
J. Biomed. Informatics | 8 |
| 2025 | Evaluating the Bias, type I error and statistical power of the prior Knowledge-Guided integrated likelihood estimation (PIE) for bias reduction in EHR based association studiesabstract• Question: How does PIE perform in various types of real-world scenarios, in terms of estimation and hypothesis testing? • Findings: Under non-differential misclassification, PIE had a smaller bias in estimated associations compared to the naïve method, but it had similar type I error and power. • The bias reduction of PIE was superior when the prior distribution of sensitivity and specificity of the phenotyping algorithm is more accurate (i.e., close to the true operating characteristics of the phenotyping algorithm). The impact of prior is relatively small when the outcome has low prevalence and is larger when the outcome is common. • PIE can effectively reduce the bias due to phenotyping error under a wide spectrum of real-world settings. However, its main advantage is in the reduction of bias in estimation but not in hypothesis testing. Binary outcomes in electronic health records (EHR) derived using automated phenotype algorithms may suffer from phenotyping error, resulting in bias in association estimation. Huang et al. [1] proposed the Prior Knowledge-Guided Integrated Likelihood Estimation (PIE) method to mitigate the estimation bias, however, their investigation focused on point estimation without statistical inference, and the evaluation of PIE therein using simulation was a proof-of-concept with only a limited scope of scenarios. This study aims to comprehensively assess PIE’s performance including (1) how well PIE performs under a wide spectrum of operating characteristics of phenotyping algorithms under real-world scenarios (e. g., low prevalence, low sensitivity, high specificity); (2) beyond point estimation, how much variation of the PIE estimator was introduced by the prior distribution; and (3) from a hypothesis testing point of view, if PIE improves type I error and statistical power relative to the naïve method (i.e., ignoring the phenotyping error). Synthetic data and use-case analysis were utilized to evaluate PIE. The synthetic data were generated under diverse outcome prevalence, phenotyping algorithm sensitivity, and association effect sizes. Simulation studies compared PIE under different prior distributions with the naïve method, assessing bias, variance, type I error, and power. Use-case analysis compared the performance of PIE and the naïve method in estimating the association of multiple predictors with COVID-19 infection. PIE exhibited reduced bias compared to the naïve method across varied simulation settings, with comparable type I error and power. As the effect size became larger, the bias reduced by PIE was larger. PIE has superior performance when prior distributions aligned closely with true phenotyping algorithm characteristics. Impact of prior quality was minor for low-prevalence outcomes but large for common outcomes. In use-case analysis, PIE maintains a relatively accurate estimation across different scenarios, particularly outperforming the naïve approach under large effect sizes. PIE effectively mitigates estimation bias in a wide spectrum of real-world settings, particularly with accurate prior information. Its main benefit lies in bias reduction rather than hypothesis testing. The impact of the prior is small for low-prevalence outcomes. Naimin Jing, Jiayi Tong, James Weaver, Patrick B. Ryan, Hua Xu 0001, Yong Chen 0016 |
J. Biomed. Informatics | 7 |
| 2025 | DisC2o-HD: Distributed causal inference with covariates shift for analyzing real-world high-dimensional dataabstractHigh-dimensional healthcare data, such as electronic health records (EHR) data and claims data, present two primary challenges due to the large number of variables and the need to consolidate data from multiple clinical sites. The third key challenge is the potential existence of heterogeneity in terms of covariate shift. In this paper, we propose a distributed learning algorithm accounting for covariate shift to estimate the average treatment effect (ATE) for high-dimensional data, named DisC2o-HD. Leveraging the surrogate likelihood method, our method calibrates the estimates of the propensity score and outcome models to approximately attain the desired covariate balancing property, while accounting for the covariate shift across multiple clinical sites. We show that our distributed covariate balancing propensity score estimator can approximate the pooled estimator, which is obtained by pooling the data from multiple sites together. The proposed estimator remains consistent if either the propensity score model or the outcome regression model is correctly specified. The semiparametric efficiency bound is achieved when both the propensity score and the outcome models are correctly specified. We conduct simulation studies to demonstrate the performance of the proposed algorithm; additionally, we conduct an empirical study to present the readiness of implementation and validity. Jiayi Tong, George Hripcsak, Yang Ning, Yong Chen 0016 |
J. Mach. Learn. Res. | 5 |
| 2024 | VGG-ST: A VGG-Swin Transformer-Based Model for ROP Disease DiagnosisabstractThe detection of Referral-Warranted Retinopathy of Prematurity (RW-ROP) is crucial for preventing severe visual impairment in premature infants. Recent studies have demonstrated that deep learning models, particularly CNNs, are effective in classifying ROP. However, the comprehensive extraction and integration of ROP-relevant features from retinal images for accurate classification remains challenging. In this paper, we propose a hybrid model based on Very Deep Convolutional Networks (VGG) and Swin Transformer (VGG-ST) for identifying RW-ROP using retinal images. The VGG-ST model first employs the VGG19 architecture to extract detailed local image features and the Swin Transformer V2 architecture to capture comprehensive global contextual information. It then introduces a novel feature enhancement module that combines these local and global features through an adaptive integration strategy, optimizing the feature representation for more accurate RW-ROP detection. Finally, the integrated features are processed through a classification module to predict the probability of RW-ROP, distinguishing between normal and RW-ROP cases. We also present customized data preprocessing techniques to address class imbalance and retinal image blur issues inherent in the e-ROP dataset. Experimental results demonstrate that the VGG-ST model offers improved classification sensitivity compared to existing methods, making it a promising tool for automated ROP screening in clinical settings. The source code is available at https://github.com/hawk-sudo/VGG-ST. Xinwei Luo, Songlin Zhao, Yong Chen 0016, Gui-Shuang Ying, Lifang He 0001 |
BIBM | 3 |
| 2024 | A Flexible Generative Model for Heterogeneous Tabular EHR with Missing ModalityabstractRealistic synthetic electronic health records (EHRs) can be leveraged to acceler- ate methodological developments for research purposes while mitigating privacy concerns associated with data sharing. However, the training of Generative Ad- versarial Networks remains challenging, often resulting in issues like mode col- lapse. While diffusion models have demonstrated progress in generating qual- ity synthetic samples for tabular EHRs given ample denoising steps, their perfor- mance wanes when confronted with missing modalities in heterogeneous tabular EHRs data. For example, some EHRs contain solely static measurements, and some contain only contain temporal measurements, or a blend of both data types. To bridge this gap, we introduce FLEXGEN-EHR– a versatile diffusion model tai- lored for heterogeneous tabular EHRs, equipped with the capability of handling missing modalities in an integrative learning framework. We define an optimal transport module to align and accentuate the common feature space of hetero- geneity of EHRs. We empirically show that our model consistently outperforms existing state-of-the-art synthetic EHR generation methods both in fidelity by up to 3.10% and utility by up to 7.16%. Additionally, we show that our method can be successfully used in privacy-sensitive settings, where the original patient-level data cannot be shared. William Hao, Yuanzhe Xi, Yong Chen 0016, Bradley A. Malin, Joyce C. Ho |
ICLR | 4 |
| 2024 | Position: TrustLLM: Trustworthiness in Large Language ModelsabstractLarge language models (LLMs) have gained considerable attention for their excellent natural language processing capabilities. Nonetheless, these LLMs present many challenges, particularly in the realm of trustworthiness. This paper introduces TrustLLM, a comprehensive study of trustworthiness in LLMs, including principles for different dimensions of trustworthiness, established benchmark, evaluation, and analysis of trustworthiness for mainstream LLMs, and discussion of open challenges and future directions. Specifically, we first propose a set of principles for trustworthy LLMs that span eight different dimensions. Based on these principles, we further establish a benchmark across six dimensions including truthfulness, safety, fairness, robustness, privacy, and machine ethics. We then present a study evaluating 16 mainstream LLMs in TrustLLM, consisting of over 30 datasets. Our findings firstly show that in general trustworthiness and capability (i.e., functional effectiveness) are positively related. Secondly, our observations reveal that proprietary LLMs generally outperform most open-source counterparts in terms of trustworthiness, raising concerns about the potential risks of widely accessible open-source LLMs. However, a few open-source LLMs come very close to proprietary ones, suggesting that open-source models can achieve high levels of trustworthiness without additional mechanisms like moderator, offering valuable insights for developers in this field. Thirdly, it is important to note that some LLMs may be overly calibrated towards exhibiting trustworthiness, to the extent that they compromise their utility by mistakenly treating benign prompts as harmful and consequently not responding. Besides these observations, we’ve uncovered key insights into the multifaceted trustworthiness in LLMs. We emphasize the importance of ensuring transparency not only in the models themselves but also in the technologies that underpin trustworthiness. We advocate that the establishment of an AI alliance between industry, academia, the open-source community to foster collaboration is imperative to advance the trustworthiness of LLMs. Yue Huang 0001, Lichao Sun 0001, Haoran Wang 0005, Siyuan Wu 0001, Qihui Zhang, Chujie Gao, Wenhan Lyu, Yixuan Zhang 0001, Xiner Li, Hanchi Sun, Zhengliang Liu, Yixin Liu 0002, Yijue Wang, Bertie Vidgen, Bhavya Kailkhura, Caiming Xiong, Chaowei Xiao, Chunyuan Li, Eric P. Xing, Furong Huang, Heng Ji 0001, Hongyi Wang 0001, Huan Zhang 0001, Huaxiu Yao, Manolis Kellis, Marinka Zitnik, Meng Jiang 0001, Mohit Bansal, James Zou 0001, Jian Pei 0001, Jianfeng Gao 0001, Jiawei Han 0001, Jieyu Zhao 0001, Jiliang Tang, Jindong Wang 0001, Joaquin Vanschoren, John C. Mitchell, Kai Shu, Kaidi Xu, Kai-Wei Chang 0001, Lifang He 0001, Lifu Huang, Michael Backes 0001, Neil Zhenqiang Gong, Philip S. Yu, Quanquan Gu, Ran Xu 0001, Rex Ying, Shuiwang Ji, Suman Jana, Tianlong Chen 0001, Tianming Liu 0001, Tianyi Zhou 0001, William Yang Wang, Xiang Li 0001, Xiangliang Zhang 0001, Xiao Wang 0012, Xing Xie 0001, Xuyu Wang, Yan Liu 0002, Yanfang Ye 0001, Yinzhi Cao, Yong Chen 0016, Yue Zhao 0016 |
ICML | 70 |
| 2024 | Confidence score: a data-driven measure for inclusive systematic reviews considering unpublished preprintsabstractOBJECTIVES: COVID-19, since its emergence in December 2019, has globally impacted research. Over 360 000 COVID-19-related manuscripts have been published on PubMed and preprint servers like medRxiv and bioRxiv, with preprints comprising about 15% of all manuscripts. Yet, the role and impact of preprints on COVID-19 research and evidence synthesis remain uncertain. MATERIALS AND METHODS: We propose a novel data-driven method for assigning weights to individual preprints in systematic reviews and meta-analyses. This weight termed the "confidence score" is obtained using the survival cure model, also known as the survival mixture model, which takes into account the time elapsed between posting and publication of a preprint, as well as metadata such as the number of first 2-week citations, sample size, and study type. RESULTS: Using 146 preprints on COVID-19 therapeutics posted from the beginning of the pandemic through April 30, 2021, we validated the confidence scores, showing an area under the curve of 0.95 (95% CI, 0.92-0.98). Through a use case on the effectiveness of hydroxychloroquine, we demonstrated how these scores can be incorporated practically into meta-analyses to properly weigh preprints. DISCUSSION: It is important to note that our method does not aim to replace existing measures of study quality but rather serves as a supplementary measure that overcomes some limitations of current approaches. CONCLUSION: Our proposed confidence score has the potential to improve systematic reviews of evidence related to COVID-19 and other clinical conditions by providing a data-driven approach to including unpublished manuscripts. Jiayi Tong, Chongliang Luo, Yifei Sun 0007, Rui Duan 0004, M. Elle Saine, Yifan Peng 0002, Anchita Batra, Anni Pan, Olivia Wang, Ruowang Li, Arielle Marks-Anglin, Xu Zuo, Yulun Liu 0004, Jiang Bian 0001, Stephen E. Kimmel, Keith Hamilton, Adam Cuker, Rebecca A. Hubbard, Hua Xu 0001, Yong Chen 0016 |
J. Am. Medical Informatics Assoc. | 23 |
| 2024 | Evaluating site-of-care-related racial disparities in kidney graft failure using a novel federated learning frameworkabstractOBJECTIVES: Racial disparities in kidney transplant access and posttransplant outcomes exist between non-Hispanic Black (NHB) and non-Hispanic White (NHW) patients in the United States, with the site of care being a key contributor. Using multi-site data to examine the effect of site of care on racial disparities, the key challenge is the dilemma in sharing patient-level data due to regulations for protecting patients' privacy. MATERIALS AND METHODS: We developed a federated learning framework, named dGEM-disparity (decentralized algorithm for Generalized linear mixed Effect Model for disparity quantification). Consisting of 2 modules, dGEM-disparity first provides accurately estimated common effects and calibrated hospital-specific effects by requiring only aggregated data from each center and then adopts a counterfactual modeling approach to assess whether the graft failure rates differ if NHB patients had been admitted at transplant centers in the same distribution as NHW patients were admitted. RESULTS: Utilizing United States Renal Data System data from 39 043 adult patients across 73 transplant centers over 10 years, we found that if NHB patients had followed the distribution of NHW patients in admissions, there would be 38 fewer deaths or graft failures per 10 000 NHB patients (95% CI, 35-40) within 1 year of receiving a kidney transplant on average. DISCUSSION: The proposed framework facilitates efficient collaborations in clinical research networks. Additionally, the framework, by using counterfactual modeling to calculate the event rate, allows us to investigate contributions to racial disparities that may occur at the level of site of care. CONCLUSIONS: Our framework is broadly applicable to other decentralized datasets and disparities research related to differential access to care. Ultimately, our proposed framework will advance equity in human health by identifying and addressing hospital-level racial disparities. Jiayi Tong, Yishan Shen, Alice Xu, Xing He 0003, Chongliang Luo, Mackenzie J. Edmondson, Dazheng Zhang, Chao Yan 0004, Ruowang Li, Lianne Siegel, Lichao Sun 0001, Elizabeth Shenkman, Sally C. Morton, Bradley A. Malin, Jiang Bian 0001, David A. Asch, Yong Chen 0016 |
J. Am. Medical Informatics Assoc. | 18 |
| 2024 | Learning competing risks across multiple hospitals: one-shot distributed algorithmsabstractOBJECTIVES: To characterize the complex interplay between multiple clinical conditions in a time-to-event analysis framework using data from multiple hospitals, we developed two novel one-shot distributed algorithms for competing risk models (ODACoR). By applying our algorithms to the EHR data from eight national children's hospitals, we quantified the impacts of a wide range of risk factors on the risk of post-acute sequelae of SARS-COV-2 (PASC) among children and adolescents. MATERIALS AND METHODS: Our ODACoR algorithms are effectively executed due to their devised simplicity and communication efficiency. We evaluated our algorithms via extensive simulation studies as applications to quantification of the impacts of risk factors for PASC among children and adolescents using data from eight children's hospitals including the Children's Hospital of Philadelphia, Cincinnati Children's Hospital Medical Center, Children's Hospital of Colorado covering over 6.5 million pediatric patients. The accuracy of the estimation was assessed by comparing the results from our ODACoR algorithms with the estimators derived from the meta-analysis and the pooled data. RESULTS: The meta-analysis estimator showed a high relative bias (∼40%) when the clinical condition is relatively rare (∼0.5%), whereas ODACoR algorithms exhibited a substantially lower relative bias (∼0.2%). The estimated effects from our ODACoR algorithms were identical on par with the estimates from the pooled data, suggesting the high reliability of our federated learning algorithms. In contrast, the meta-analysis estimate failed to identify risk factors such as age, gender, chronic conditions history, and obesity, compared to the pooled data. DISCUSSION: Our proposed ODACoR algorithms are communication-efficient, highly accurate, and suitable to characterize the complex interplay between multiple clinical conditions. CONCLUSION: Our study demonstrates that our ODACoR algorithms are communication-efficient and can be widely applicable for analyzing multiple clinical conditions in a time-to-event analysis framework. Dazheng Zhang, Jiayi Tong, Naimin Jing, Chongliang Luo, Dimitri A. Christakis, Diana Güthe, Mady Hornig, Kelly J. Kelleher, Keith E. Morse, Colin M. Rogerson, Jasmin Divers, Raymond J. Carroll, Christopher B. Forrest, Yong Chen 0016 |
J. Am. Medical Informatics Assoc. | 16 |
| 2024 | Balancing the efforts of chart review and gains in PRS prediction accuracy: An empirical study
Yuqing Lei, Adam Christian Naj, Hua Xu 0001, Ruowang Li, Yong Chen 0016 |
J. Biomed. Informatics | 5 |
| 2024 | Leveraging error-prone algorithm-derived phenotypes: Enhancing association studies for risk factors in EHR data
Jiayi Tong, Jessica Chubak, Thomas Lumley, Rebecca A. Hubbard, Hua Xu 0001, Yong Chen 0016 |
J. Biomed. Informatics | 7 |
| 2024 | Leveraging generative AI for clinical evidence synthesis needs to ensure trustworthiness
Qiao Jin 0001, Denis Jered McInerney, Yong Chen 0016, Fei Wang 0001, Curtis L. Cole, Qian Yang 0004, Yanshan Wang, Bradley A. Malin, Mor Peleg, Byron C. Wallace, Zhiyong Lu, Chunhua Weng, Yifan Peng 0002 |
J. Biomed. Informatics | 4 |
| 2024 | One-shot distributed algorithms for addressing heterogeneity in competing risks data across clinical sites
Dazheng Zhang, Jiayi Tong, Ronen Stein, Naimin Jing, Mary Regina Boland, Chongliang Luo, Robert N. Baldassano, Raymond J. Carroll, Christopher B. Forrest, Yong Chen 0016 |
J. Biomed. Informatics | 12 |
| 2023 | Towards precise PICO extraction from abstracts of randomized controlled trials using a section-specific learning approachabstractMOTIVATION: Automated extraction of participants, intervention, comparison/control, and outcome (PICO) from the randomized controlled trial (RCT) abstracts is important for evidence synthesis. Previous studies have demonstrated the feasibility of applying natural language processing (NLP) for PICO extraction. However, the performance is not optimal due to the complexity of PICO information in RCT abstracts and the challenges involved in their annotation. RESULTS: We propose a two-step NLP pipeline to extract PICO elements from RCT abstracts: (i) sentence classification using a prompt-based learning model and (ii) PICO extraction using a named entity recognition (NER) model. First, the sentences in abstracts were categorized into four sections namely background, methods, results, and conclusions. Next, the NER model was applied to extract the PICO elements from the sentences within the title and methods sections that include >96% of PICO information. We evaluated our proposed NLP pipeline on three datasets, the EBM-NLPmoddataset, a randomly selected and reannotated dataset of 500 RCT abstracts from the EBM-NLP corpus, a dataset of 150 COVID-19 RCT abstracts, and a dataset of 150 Alzheimer's disease (AD) RCT abstracts. The end-to-end evaluation reveals that our proposed approach achieved an overall micro F1 score of 0.833 on the EBM-NLPmod dataset, 0.928 on the COVID-19 dataset, and 0.899 on the AD dataset when measured at the token-level and an overall micro F1 score of 0.712 on EBM-NLPmod dataset, 0.850 on the COVID-19 dataset, and 0.805 on the AD dataset when measured at the entity-level. AVAILABILITY: Our codes and datasets are publicly available at https://github.com/BIDS-Xu-Lab/section_specific_annotation_of_PICO. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Vipina Kuttichi Keloth, Kalpana Raja, Yong Chen 0016, Hua Xu 0001 |
Bioinform. | 4 |
| 2023 | Missing data matter: an empirical evaluation of the impacts of missing EHR data in comparative effectiveness researchabstractOBJECTIVES: The impacts of missing data in comparative effectiveness research (CER) using electronic health records (EHRs) may vary depending on the type and pattern of missing data. In this study, we aimed to quantify these impacts and compare the performance of different imputation methods. MATERIALS AND METHODS: We conducted an empirical (simulation) study to quantify the bias and power loss in estimating treatment effects in CER using EHR data. We considered various missing scenarios and used the propensity scores to control for confounding. We compared the performance of the multiple imputation and spline smoothing methods to handle missing data. RESULTS: When missing data depended on the stochastic progression of disease and medical practice patterns, the spline smoothing method produced results that were close to those obtained when there were no missing data. Compared to multiple imputation, the spline smoothing generally performed similarly or better, with smaller estimation bias and less power loss. The multiple imputation can still reduce study bias and power loss in some restrictive scenarios, eg, when missing data did not depend on the stochastic process of disease progression. DISCUSSION AND CONCLUSION: Missing data in EHRs could lead to biased estimates of treatment effects and false negative findings in CER even after missing data were imputed. It is important to leverage the temporal information of disease trajectory to impute missing values when using EHRs as a data resource for CER and to consider the missing rate and the effect size when choosing an imputation method. Yizhao Zhou, Jiasheng Shi, Ronen Stein, Robert N. Baldassano, Christopher B. Forrest, Yong Chen 0016, Jing Huang 0021 |
J. Am. Medical Informatics Assoc. | 7 |
| 2023 | FedScore: A privacy-preserving framework for federated scoring system development
Siqi Li 0004, Yilin Ning, Marcus Eng Hock Ong, Bibhas Chakraborty, Chuan Hong, Feng Xie 0004, Mingxuan Liu 0005, Daniel M. Buckland, Yong Chen 0016, Nan Liu 0003 |
J. Biomed. Informatics | 10 |
| 2023 | Padé approximant meets federated learning: A nearly lossless, one-shot algorithm for evidence synthesis in distributed research networks with rare outcomes
Martijn J. Schuemie, Marc A. Suchard, Patrick B. Ryan, George Hripcsak, Charles A. Rohde, Yong Chen 0016 |
J. Biomed. Informatics | 7 |
| 2023 | Scalable high-dimensional Bayesian varying coefficient models with unknown within-subject covarianceabstractNonparametric varying coefficient (NVC) models are useful for modeling time-varying effects on responses that are measured repeatedly for the same subjects. When the number of covariates is moderate or large, it is desirable to perform variable selection from the varying coefficient functions. However, existing methods for variable selection in NVC models either fail to account for within-subject correlations or require the practitioner to specify a parametric form for the correlation structure. In this paper, we introduce the nonparametric varying coefficient spike-and-slab lasso (NVC-SSL) for Bayesian high dimensional NVC models. Through the introduction of functional random effects, our method allows for flexible modeling of within-subject correlations without needing to specify a parametric covariance function. We further propose several scalable optimization and Markov chain Monte Carlo (MCMC) algorithms. For variable selection, we propose an Expectation Conditional Maximization (ECM) algorithm to rapidly obtain maximum a posteriori (MAP) estimates. Our ECM algorithm scales linearly in the total number of observations $N$ and the number of covariates $p$. For uncertainty quantification, we introduce an approximate MCMC algorithm that also scales linearly in both $N$ and $p$. We demonstrate the scalability, variable selection performance, and inferential capabilities of our method through simulations and a real data application. These algorithms are implemented in the publicly available R package NVCSSL on the Comprehensive R Archive Network. Ray Bai, Mary Regina Boland, Yong Chen 0016 |
J. Mach. Learn. Res. | 3 |
| 2022 | Tensor-Based Multi-Modal Multi-Target Regression for Alzheimer's Disease PredictionabstractThe assessment of Alzheimer’s Disease (AD) progression via the analysis of physical changes within the brain has attracted great interest from the fields of healthcare, computational medicine, and machine learning alike. Recent studies have demonstrated that using both multi-modal data and multiple AD assessment scores in a predictive model can better reflect pathological characteristics and enhance prediction performance. However, using such high-dimensional structure information to model inter-correlation between multiple targets remains a challenging task. In this paper, we propose a Tensor-based Multi-modal Multi-Target Regression (TMMTR) method for AD detection and prediction, which enables simultaneously modeling multilinear structure information as well as intrinsic inter-target correlations in a general learning framework. We also investigate the tensor-structured sparsity that supports the interpretability of our prediction. Experiments conducted on the ADNI dataset validate the superior performance of our method when compared to other state-of-the-art methods. Benjamin Zalatan, Yong Chen 0016, Li Shen 0001, Lifang He 0001 |
BIBM | 3 |
| 2022 | Mining on Alzheimer's diseases related knowledge graph to identity potential AD-related semantic triples for drug repurposingabstractBACKGROUND: To date, there are no effective treatments for most neurodegenerative diseases. Knowledge graphs can provide comprehensive and semantic representation for heterogeneous data, and have been successfully leveraged in many biomedical applications including drug repurposing. Our objective is to construct a knowledge graph from literature to study the relations between Alzheimer's disease (AD) and chemicals, drugs and dietary supplements in order to identify opportunities to prevent or delay neurodegenerative progression. We collected biomedical annotations and extracted their relations using SemRep via SemMedDB. We used both a BERT-based classifier and rule-based methods during data preprocessing to exclude noise while preserving most AD-related semantic triples. The 1,672,110 filtered triples were used to train with knowledge graph completion algorithms (i.e., TransE, DistMult, and ComplEx) to predict candidates that might be helpful for AD treatment or prevention. RESULTS: Among three knowledge graph completion models, TransE outperformed the other two (MR = 10.53, Hits@1 = 0.28). We leveraged the time-slicing technique to further evaluate the prediction results. We found supporting evidence for most highly ranked candidates predicted by our model which indicates that our approach can inform reliable new knowledge. CONCLUSION: This paper shows that our graph mining model can predict reliable new relationships between AD and other entities (i.e., dietary supplements, chemicals, and drugs). The knowledge graph constructed can facilitate data-driven knowledge discoveries and the generation of novel hypotheses. Yi Nian, Xinyue Hu 0002, Rui Zhang 0028, Jingna Feng, Jingcheng Du, Fang Li 0011, Larry Bu, Yuji Zhang 0001, Yong Chen 0016, Cui Tao |
BMC Bioinform. | 9 |
| 2022 | SAT: a Surrogate-Assisted Two-wave case boosting sampling method, with application to EHR-based association studiesabstractOBJECTIVES: Electronic health records (EHRs) enable investigation of the association between phenotypes and risk factors. However, studies solely relying on potentially error-prone EHR-derived phenotypes (ie, surrogates) are subject to bias. Analyses of low prevalence phenotypes may also suffer from poor efficiency. Existing methods typically focus on one of these issues but seldom address both. This study aims to simultaneously address both issues by developing new sampling methods to select an optimal subsample to collect gold standard phenotypes for improving the accuracy of association estimation. MATERIALS AND METHODS: We develop a surrogate-assisted two-wave (SAT) sampling method, where a surrogate-guided sampling (SGS) procedure and a modified optimal subsampling procedure motivated from A-optimality criterion (OSMAC) are employed sequentially, to select a subsample for outcome validation through manual chart review subject to budget constraints. A model is then fitted based on the subsample with the true phenotypes. Simulation studies and an application to an EHR dataset of breast cancer survivors are conducted to demonstrate the effectiveness of SAT. RESULTS: We found that the subsample selected with the proposed method contains informative observations that effectively reduce the mean squared error of the resultant estimator of the association. CONCLUSIONS: The proposed approach can handle the problem brought by the rarity of cases and misclassification of the surrogate in phenotype-absent EHR-based association studies. With a well-behaved surrogate, SAT successfully boosts the case prevalence in the subsample and improves the efficiency of estimation. Jessica Chubak, Rebecca A. Hubbard, Yong Chen 0016 |
J. Am. Medical Informatics Assoc. | 4 |
| 2022 | dPQL: a lossless distributed algorithm for generalized linear mixed model with application to privacy-preserving hospital profilingabstractOBJECTIVE: To develop a lossless distributed algorithm for generalized linear mixed model (GLMM) with application to privacy-preserving hospital profiling. MATERIALS AND METHODS: The GLMM is often fitted to implement hospital profiling, using clinical or administrative claims data. Due to individual patient data (IPD) privacy regulations and the computational complexity of GLMM, a distributed algorithm for hospital profiling is needed. We develop a novel distributed penalized quasi-likelihood (dPQL) algorithm to fit GLMM when only aggregated data, rather than IPD, can be shared across hospitals. We also show that the standardized mortality rates, which are often reported as the results of hospital profiling, can also be calculated distributively without sharing IPD. We demonstrate the applicability of the proposed dPQL algorithm by ranking 929 hospitals for coronavirus disease 2019 (COVID-19) mortality or referral to hospice that have been previously studied. RESULTS: The proposed dPQL algorithm is mathematically proven to be lossless, that is, it obtains identical results as if IPD were pooled from all hospitals. In the example of hospital profiling regarding COVID-19 mortality, the dPQL algorithm reached convergence with only 5 iterations, and the estimation of fixed effects, random effects, and mortality rates were identical to that of the PQL from pooled data. CONCLUSION: The dPQL algorithm is lossless, privacy-preserving and fast-converging for fitting GLMM. It provides an extremely suitable and convenient distributed approach for hospital profiling. Chongliang Luo, Md. Nazmul Islam, Natalie E. Sheils, John Buresh, Martijn J. Schuemie, Jalpa A. Doshi, Rachel M. Werner, David A. Asch, Yong Chen 0016 |
J. Am. Medical Informatics Assoc. | 9 |
| 2022 | Distributed Quasi-Poisson regression algorithm for modeling multi-site count outcomes in distributed data networks
Mackenzie J. Edmondson, Chongliang Luo, Md. Nazmul Islam, Natalie E. Sheils, John Buresh, Zhaoyi Chen, Jiang Bian 0001, Yong Chen 0016 |
J. Biomed. Informatics | 8 |
| 2022 | Federated Multi-view Learning for Private Medical Data Integration and AnalysisabstractAlong with the rapid expansion of information technology and digitalization of health data, there is an increasing concern on maintaining data privacy while garnering the benefits in the medical field. Two critical challenges are identified: First, medical data is naturally distributed across multiple local sites, making it difficult to collectively train machine learning models without data leakage. Second, in medical applications, data are often collected from different sources and views, resulting in heterogeneity and complexity that requires reconciliation. In this article, we present a generic Federated Multi-view Learning (FedMV) framework for multi-view data leakage prevention. Specifically, we apply this framework to two types of problems based on local data availability: Vertical Federated Multi-view Learning (V-FedMV) and Horizontal Federated Multi-view Learning (H-FedMV). We experimented with real-world keyboard data collected from BiAffect study. Our results demonstrated that the proposed approach can make full use of multi-view data in a privacy-preserving way, and both V-FedMV and H-FedMV perform better than their single-view and pairwise counterparts. Besides, the framework can be easily adapted to deal with multi-view sequential data. We have developed a sequential model (S-FedMV) that takes sequence of multi-view data as input and demonstrated it experimentally. To the best of our knowledge, this framework is the first to consider both vertical and horizontal diversification in the multi-view setting, as well as their sequential federated learning. Sicong Che, Zhaoming Kong, Hao Peng 0001, Lichao Sun 0001, Alex D. Leow, Yong Chen 0016, Lifang He 0001 |
ACM Trans. Intell. Syst. Technol. | 6 |
| 2021 | Validation of Real-World Data-based Endpoint Measures of Cancer Treatment Outcomes
Qian Li 0034, Hansi Zhang, Zhaoyi Chen, Yi Guo 0005, Thomas J. George, Yong Chen 0016, Fei Wang 0001, Jiang Bian 0001 |
AMIA | 6 |
| 2021 | How do we share data in COVID-19 research? A systematic review of COVID-19 datasets in PubMed Central ArticlesabstractOBJECTIVE: This study aims at reviewing novel coronavirus disease (COVID-19) datasets extracted from PubMed Central articles, thus providing quantitative analysis to answer questions related to dataset contents, accessibility and citations. METHODS: We downloaded COVID-19-related full-text articles published until 31 May 2020 from PubMed Central. Dataset URL links mentioned in full-text articles were extracted, and each dataset was manually reviewed to provide information on 10 variables: (1) type of the dataset, (2) geographic region where the data were collected, (3) whether the dataset was immediately downloadable, (4) format of the dataset files, (5) where the dataset was hosted, (6) whether the dataset was updated regularly, (7) the type of license used, (8) whether the metadata were explicitly provided, (9) whether there was a PubMed Central paper describing the dataset and (10) the number of times the dataset was cited by PubMed Central articles. Descriptive statistics about these seven variables were reported for all extracted datasets. RESULTS: We found that 28.5% of 12 324 COVID-19 full-text articles in PubMed Central provided at least one dataset link. In total, 128 unique dataset links were mentioned in 12 324 COVID-19 full text articles in PubMed Central. Further analysis showed that epidemiological datasets accounted for the largest portion (53.9%) in the dataset collection, and most datasets (84.4%) were available for immediate download. GitHub was the most popular repository for hosting COVID-19 datasets. CSV, XLSX and JSON were the most popular data formats. Additionally, citation patterns of COVID-19 datasets varied depending on specific datasets. CONCLUSION: PubMed Central articles are an important source of COVID-19 datasets, but there is significant heterogeneity in the way these datasets are mentioned, shared, updated and cited. Xu Zuo, Yong Chen 0016, Lucila Ohno-Machado, Hua Xu 0001 |
Briefings Bioinform. | 2 |
| 2021 | Extracting postmarketing adverse events from safety reports in the vaccine adverse event reporting system (VAERS) using deep learningabstractOBJECTIVE: Automated analysis of vaccine postmarketing surveillance narrative reports is important to understand the progression of rare but severe vaccine adverse events (AEs). This study implemented and evaluated state-of-the-art deep learning algorithms for named entity recognition to extract nervous system disorder-related events from vaccine safety reports. MATERIALS AND METHODS: We collected Guillain-Barré syndrome (GBS) related influenza vaccine safety reports from the Vaccine Adverse Event Reporting System (VAERS) from 1990 to 2016. VAERS reports were selected and manually annotated with major entities related to nervous system disorders, including, investigation, nervous_AE, other_AE, procedure, social_circumstance, and temporal_expression. A variety of conventional machine learning and deep learning algorithms were then evaluated for the extraction of the above entities. We further pretrained domain-specific BERT (Bidirectional Encoder Representations from Transformers) using VAERS reports (VAERS BERT) and compared its performance with existing models. RESULTS AND CONCLUSIONS: Ninety-one VAERS reports were annotated, resulting in 2512 entities. The corpus was made publicly available to promote community efforts on vaccine AEs identification. Deep learning-based methods (eg, bi-long short-term memory and BERT models) outperformed conventional machine learning-based methods (ie, conditional random fields with extensive features). The BioBERT large model achieved the highest exact match F-1 scores on nervous_AE, procedure, social_circumstance, and temporal_expression; while VAERS BERT large models achieved the highest exact match F-1 scores on investigation and other_AE. An ensemble of these 2 models achieved the highest exact match microaveraged F-1 score at 0.6802 and the second highest lenient match microaveraged F-1 score at 0.8078 among peer models. Jingcheng Du, Yang Xiang 0003, Madhuri Sankaranarayanapillai, Yuqi Si, Huy Anh Pham, Hua Xu 0001, Yong Chen 0016, Cui Tao |
J. Am. Medical Informatics Assoc. | 9 |
| 2021 | A cost-effective chart review sampling design to account for phenotyping error in electronic health records (EHR) dataabstractOBJECTIVES: Electronic health records (EHR) are commonly used for the identification of novel risk factors for disease, often referred to as an association study. A major challenge to EHR-based association studies is phenotyping error in EHR-derived outcomes. A manual chart review of phenotypes is necessary for unbiased evaluation of risk factor associations. However, this process is time-consuming and expensive. The objective of this paper is to develop an outcome-dependent sampling approach for designing manual chart review, where EHR-derived phenotypes can be used to guide the selection of charts to be reviewed in order to maximize statistical efficiency in the subsequent estimation of risk factor associations. MATERIALS AND METHODS: After applying outcome-dependent sampling, an augmented estimator can be constructed by optimally combining the chart-reviewed phenotypes from the selected patients with the error-prone EHR-derived phenotype. We conducted simulation studies to evaluate the proposed method and applied our method to data on colon cancer recurrence in a cohort of patients treated for a primary colon cancer in the Kaiser Permanente Washington (KPW) healthcare system. RESULTS: Simulations verify the coverage probability of the proposed method and show that, when disease prevalence is less than 30%, the proposed method has smaller variance than an existing method where the validation set for chart review is uniformly sampled. In addition, from design perspective, the proposed method is able to achieve the same statistical power with 50% fewer charts to be validated than the uniform sampling method, thus, leading to a substantial efficiency gain in chart review. These findings were also confirmed by the application of the competing methods to the KPW colon cancer data. DISCUSSION: Our simulation studies and analysis of data from KPW demonstrate that, compared to an existing uniform sampling method, the proposed outcome-dependent method can lead to a more efficient chart review sampling design and unbiased association estimates with higher statistical efficiency. CONCLUSION: The proposed method not only optimally combines phenotypes from chart review with EHR-derived phenotypes but also suggests an efficient design for conducting chart review, with the goal of improving the efficiency of estimated risk factor associations using EHR data. Ziyan Yin, Jiayi Tong, Yong Chen 0016, Rebecca A. Hubbard, Cheng Yong Tang |
J. Am. Medical Informatics Assoc. | 3 |
| 2021 | Studying pediatric health outcomes with electronic health records using Bayesian clustering and trajectory analysisabstractUse of routinely collected data from electronic health records (EHR) can expedite longitudinal studies that investigate childhood exposures and rare pediatric health outcomes. For instance, characteristics of the body mass index (BMI) trajectory early in life may be associated with subsequent development of type 2 diabetes. Past studies investigating these relationships have used longitudinal cohort data collected over the course of many years to investigate the connection between BMI trajectory and subsequent development of diabetes. In contrast, EHR data from routine clinical care can provide longitudinal information on early-life BMI trajectories as well as subsequent health outcomes without requiring any additional data collection. In this study, we introduce a Bayesian joint phenotyping and BMI trajectory model to address data quality challenges in an EHR-based study of early-life BMI and type 2 diabetes in adolescence. We compared this joint modeling approach to traditional approaches using a computable phenotype for type 2 diabetes or separately estimated BMI trajectories and type 2 diabetes phenotypes. In a sample of 49,062 children derived from the PEDSnet consortium of pediatric healthcare systems, a median 8 (interquartile range [IQR] 5-13) BMI measurements were available to characterize the early-life BMI trajectory. The joint modeling and computable phenotype approaches found that age at adiposity rebound between 5 and 9 years was associated with higher odds of type 2 diabetes in adolescence compared to age at adiposity rebound between 2 and 5 years (joint model odds ratio [OR] = 1.77; computable phenotype OR = 1.88) and that BMI in excess of 140% of the 95th percentile for age and sex at age 9 years was associated with higher odds of type 2 diabetes in adolescence relative to children with BMI from 100 to 120% of the 95th percentile (joint model OR = 6.22; computable phenotype OR = 13.25). Estimates from the separate phenotyping and trajectory model were substantially attenuated towards the null. These results demonstrate that EHR data coupled with modern methodologic approaches can improve efficiency and timeliness of studies of childhood exposures and rare health outcomes. Rebecca A. Hubbard, Robert Siegel, Yong Chen 0016, Ihuoma Eneli |
J. Biomed. Informatics | 4 |
| 2020 | Leverage Real-World Longitudinal Data in Large Clinical Research Networks for Alzheimer's Disease and Related Dementia (ADRD)
Rui Duan 0004, Zhaoyi Chen, Jiayi Tong, Chongliang Luo, Tianchen Lyu, Cui Tao, Demetrius Maraganore, Jiang Bian 0001, Yong Chen 0016 |
AMIA | 9 |
| 2020 | Identifying Clinical Risk Factors for Opioid Use Disorder using a Distributed Algorithm to Combine Real-World Data from a Large Clinical Data Research Network
Jiayi Tong, Zhaoyi Chen, Rui Duan 0004, Wei-Hsuan Lo-Ciganic, Tianchen Lyu, Cui Tao, Peter A. Merkel, Henry R. Kranzler, Jiang Bian 0001, Yong Chen 0016 |
AMIA | 10 |
| 2020 | How Computational Experiments Can Improve Our Understanding of the Genetic Architecture of Common Human DiseasesabstractSusceptibility to common human diseases such as cancer is influenced by many genetic and environmental factors that work together in a complex manner. The state of the art is to perform a genome-wide association study (GWAS) that measures millions of single-nucleotide polymorphisms (SNPs) throughout the genome followed by a one-SNP-at-a-time statistical analysis to detect univariate associations. This approach has identified thousands of genetic risk factors for hundreds of diseases. However, the genetic risk factors detected have very small effect sizes and collectively explain very little of the overall heritability of the disease. Nonetheless, it is assumed that the genetic component of risk is due to many independent risk factors that contribute additively. The fact that many genetic risk factors with small effects can be detected is taken as evidence to support this notion. It is our working hypothesis that the genetic architecture of common diseases is partly driven by non-additive interactions. To test this hypothesis, we developed a heuristic simulation-based method for conducting experiments about the complexity of genetic architecture. We show that a genetic architecture driven by complex interactions is highly consistent with the magnitude and distribution of univariate effects seen in real data. We compare our results with measures of univariate and interaction effects from two large-scale GWASs of sporadic breast cancer and find evidence to support our hypothesis that is consistent with the results of our computational experiment. Jason H. Moore, Randal S. Olson, Peter Schmitt, Yong Chen 0016, Elisabetta Manduchi |
Artif. Life | 4 |
| 2020 | Learning from electronic health records across multiple sites: A communication-efficient and privacy-preserving distributed algorithmabstractOBJECTIVES: We propose a one-shot, privacy-preserving distributed algorithm to perform logistic regression (ODAL) across multiple clinical sites. MATERIALS AND METHODS: ODAL effectively utilizes the information from the local site (where the patient-level data are accessible) and incorporates the first-order (ODAL1) and second-order (ODAL2) gradients of the likelihood function from other sites to construct an estimator without requiring iterative communication across sites or transferring patient-level data. We evaluated ODAL via extensive simulation studies and an application to a dataset from the University of Pennsylvania Health System. The estimation accuracy was evaluated by comparing it with the estimator based on the combined individual participant data or pooled data (ie, gold standard). RESULTS: Our simulation studies revealed that the relative estimation bias of ODAL1 compared with the pooled estimates was <3%, and the ratio of standard errors was <1.25 for all scenarios. ODAL2 achieved higher accuracy (with relative bias <0.1% and ratio of standard errors <1.05). In real data analysis, we investigated the associations of 100 medications with fetal loss during pregnancy. We found that ODAL1 provided estimates with relative bias <10% for 85% of medications, and ODAL2 has relative bias <10% for 99% of medications. For communication cost, ODAL1 requires transferring p numbers from each site to the local site and ODAL2 requires transferring (p×p+p) numbers from each site to the local site, where p is the number of parameters in the regression model. CONCLUSIONS: This study demonstrates that ODAL is privacy-preserving and communication-efficient with small bias and high statistical efficiency. Rui Duan 0004, Mary Regina Boland, Howard H. Chang, Hua Xu 0001, Haitao Chu, Christopher H. Schmid, Christopher B. Forrest, John H. Holmes, Martijn J. Schuemie, Jesse A. Berlin, Jason H. Moore, Yong Chen 0016 |
J. Am. Medical Informatics Assoc. | 14 |
| 2020 | Learning from local to global: An efficient distributed algorithm for modeling time-to-event dataabstractOBJECTIVE: We developed and evaluated a privacy-preserving One-shot Distributed Algorithm to fit a multicenter Cox proportional hazards model (ODAC) without sharing patient-level information across sites. MATERIALS AND METHODS: Using patient-level data from a single site combined with only aggregated information from other sites, we constructed a surrogate likelihood function, approximating the Cox partial likelihood function obtained using patient-level data from all sites. By maximizing the surrogate likelihood function, each site obtained a local estimate of the model parameter, and the ODAC estimator was constructed as a weighted average of all the local estimates. We evaluated the performance of ODAC with (1) a simulation study and (2) a real-world use case study using 4 datasets from the Observational Health Data Sciences and Informatics network. RESULTS: On the one hand, our simulation study showed that ODAC provided estimates nearly the same as the estimator obtained by analyzing, in a single dataset, the combined patient-level data from all sites (ie, the pooled estimator). The relative bias was <0.1% across all scenarios. The accuracy of ODAC remained high across different sample sizes and event rates. On the other hand, the meta-analysis estimator, which was obtained by the inverse variance weighted average of the site-specific estimates, had substantial bias when the event rate is <5%, with the relative bias reaching 20% when the event rate is 1%. In the Observational Health Data Sciences and Informatics network application, the ODAC estimates have a relative bias <5% for 15 out of 16 log hazard ratios, whereas the meta-analysis estimates had substantially higher bias than ODAC. CONCLUSIONS: ODAC is a privacy-preserving and noniterative method for implementing time-to-event analyses across multiple sites. It provides estimates on par with the pooled estimator and substantially outperforms the meta-analysis estimator when the event is uncommon, making it extremely suitable for studying rare events and diseases in a distributed manner. Rui Duan 0004, Chongliang Luo, Martijn J. Schuemie, Jiayi Tong, C. Jason Liang, Howard H. Chang, Mary Regina Boland, Jiang Bian 0001, Hua Xu 0001, John H. Holmes, Christopher B. Forrest, Sally C. Morton, Jesse A. Berlin, Jason H. Moore, Kevin B. Mahoney, Yong Chen 0016 |
J. Am. Medical Informatics Assoc. | 16 |
| 2020 | SCOR: A secure international informatics infrastructure to investigate COVID-19abstractGlobal pandemics call for large and diverse healthcare data to study various risk factors, treatment options, and disease progression patterns. Despite the enormous efforts of many large data consortium initiatives, scientific community still lacks a secure and privacy-preserving infrastructure to support auditable data sharing and facilitate automated and legally compliant federated analysis on an international scale. Existing health informatics systems do not incorporate the latest progress in modern security and federated machine learning algorithms, which are poised to offer solutions. An international group of passionate researchers came together with a joint mission to solve the problem with our finest models and tools. The SCOR Consortium has developed a ready-to-deploy secure infrastructure using world-class privacy and security technologies to reconcile the privacy/utility conflicts. We hope our effort will make a change and accelerate research in future pandemics with broad and diverse samples on an international scale. Jean Louis Raisaro, Juan Ramón Troncoso-Pastoriza, Raphaelle Beau-Lejdstrom, Riccardo Bellazzi, Robert Murphy, Elmer V. Bernstam, Henry Wang, Mauro Bucalo, Yong Chen 0016, Assaf Gottlieb, Arif Ozgun Harmanci, Miran Kim, Yejin Kim 0001, Jeffrey G. Klann, Catherine Klersy, Bradley A. Malin, Marie Méan, Fabian Prasser, Luigia Scudeller, Ali Torkamani, Julien Vaucher, Mamta Puppala, Stephen T. C. Wong, Milana Frenkel-Morgenstern, Hua Xu 0001, Baba Maiyaki Musa, Abdulrazaq G. Habib, Trevor Cohen, Adam B. Wilcox, Hamisu M. Salihu, Heidi Sofia, Xiaoqian Jiang, Jean-Pierre Hubaux |
J. Am. Medical Informatics Assoc. | 10 |
| 2020 | An augmented estimation procedure for EHR-based association studies accounting for differential misclassificationabstractOBJECTIVES: The ability to identify novel risk factors for health outcomes is a key strength of electronic health record (EHR)-based research. However, the validity of such studies is limited by error in EHR-derived phenotypes. The objective of this study was to develop a novel procedure for reducing bias in estimated associations between risk factors and phenotypes in EHR data. MATERIALS AND METHODS: The proposed method combines the strengths of a gold-standard phenotype obtained through manual chart review for a small validation set of patients and an automatically-derived phenotype that is available for all patients but is potentially error-prone (hereafter referred to as the algorithm-derived phenotype). An augmented estimator of associations is obtained by optimally combining these 2 phenotypes. We conducted simulation studies to evaluate the performance of the augmented estimator and conducted an analysis of risk factors for second breast cancer events using data on a cohort from Kaiser Permanente Washington. RESULTS: The proposed method was shown to reduce bias relative to an estimator using only the algorithm-derived phenotype and reduce variance compared to an estimator using only the validation data. DISCUSSION: Our simulation studies and real data application demonstrate that, compared to the estimator using validation data only, the augmented estimator has lower variance (ie, higher statistical efficiency). Compared to the estimator using error-prone EHR-derived phenotypes, the augmented estimator has smaller bias. CONCLUSIONS: The proposed estimator can effectively combine an error-prone phenotype with gold-standard data from a limited chart review in order to improve analyses of risk factors using EHR data. Jiayi Tong, Jing Huang 0021, Jessica Chubak, Jason H. Moore, Rebecca A. Hubbard, Yong Chen 0016 |
J. Am. Medical Informatics Assoc. | 7 |
| 2019 | Integration of genetic and clinical information to improve imputation of data missing from electronic health recordsabstractOBJECTIVE: Clinical data of patients' measurements and treatment history stored in electronic health record (EHR) systems are starting to be mined for better treatment options and disease associations. A primary challenge associated with utilizing EHR data is the considerable amount of missing data. Failure to address this issue can introduce significant bias in EHR-based research. Currently, imputation methods rely on correlations among the structured phenotype variables in the EHR. However, genetic studies have shown that many EHR-based phenotypes have a heritable component, suggesting that measured genetic variants might be useful for imputing missing data. In this article, we developed a computational model that incorporates patients' genetic information to perform EHR data imputation. MATERIALS AND METHODS: We used the individual single nucleotide polymorphism's association with phenotype variables in the EHR as input to construct a genetic risk score that quantifies the genetic contribution to the phenotype. Multiple approaches to constructing the genetic risk score were evaluated for optimal performance. The genetic score, along with phenotype correlation, is then used as a predictor to impute the missing values. RESULTS: To demonstrate the method performance, we applied our model to impute missing cardiovascular related measurements including low-density lipoprotein, heart failure, and aortic aneurysm disease in the electronic Medical Records and Genomics data. The integration method improved imputation's area-under-the-curve for binary phenotypes and decreased root-mean-square error for continuous phenotypes. CONCLUSION: Compared with standard imputation approaches, incorporating genetic information offers a novel approach that can utilize more of the EHR data for better performance in missing data imputation. Ruowang Li, Yong Chen 0016, Jason H. Moore |
J. Am. Medical Informatics Assoc. | 2 |
| 2019 | A regression framework to uncover pleiotropy in large-scale electronic health record dataabstractOBJECTIVE: Pleiotropy, where 1 genetic locus affects multiple phenotypes, can offer significant insights in understanding the complex genotype-phenotype relationship. Although individual genotype-phenotype associations have been thoroughly explored, seemingly unrelated phenotypes can be connected genetically through common pleiotropic loci or genes. However, current analyses of pleiotropy have been challenged by both methodologic limitations and a lack of available suitable data sources. MATERIALS AND METHODS: In this study, we propose to utilize a new regression framework, reduced rank regression, to simultaneously analyze multiple phenotypes and genotypes to detect pleiotropic effects. We used a large-scale biobank linked electronic health record data from the Penn Medicine BioBank to select 5 cardiovascular diseases (hypertension, cardiac dysrhythmias, ischemic heart disease, congestive heart failure, and heart valve disorders) and 5 mental disorders (mood disorders; anxiety, phobic and dissociative disorders; alcohol-related disorders; neurological disorders; and delirium dementia) to validate our framework. RESULTS: Compared with existing methods, reduced rank regression showed a higher power to distinguish known associated single-nucleotide polymorphisms from random single-nucleotide polymorphisms. In addition, genome-wide gene-based investigation of pleiotropy showed that reduced rank regression was able to identify candidate genetic variants with novel pleiotropic effects compared to existing methods. CONCLUSION: The proposed regression framework offers a new approach to account for the phenotype and genotype correlations when identifying pleiotropic effects. By jointly modeling multiple phenotypes and genotypes together, the method has the potential to distinguish confounding from causal genotype and phenotype associations. Ruowang Li, Rui Duan 0004, Daniel J. Rader, Scott M. Damrauer, Jason H. Moore, Yong Chen 0016 |
J. Am. Medical Informatics Assoc. | 6 |
| 2018 | Identification of Rare Adverse Events with Year-varying Reporting Rates for FLU4 Vaccine in VAERS
Jiayi Tong, Jing Huang 0021, Jingcheng Du, Cui Tao, Yong Chen 0016 |
AMIA | 6 |
| 2018 | A Two-Stage Working Model Strategy for Network Analysis Under Hierarchical Exponential Random Graph ModelsabstractSocial network data are complex and dependent data. At the macro-level, social networks often exhibit clustering in the sense that social networks consist of communities; and at the micro-level, social networks often exhibit complex network features such as transitivity within communities. Modeling real-world social networks requires modeling both the macro- and micro-level, but many existing models focus on one of them while neglecting the other. In recent work, [28] introduced a class of Exponential Random Graph Models (ERGMs) capturing community structure as well as microlevel features within communities. While attractive, existing approaches to estimating ERGMs with community structure are not scalable. We propose here a scalable two-stage strategy to estimate an important class of ERGMs with community structure, which induces transitivity within communities. At the first stage, we use an approximate model, called working model, to estimate the community structure. At the second stage, we use ERGMs with geometrically weighted dyadwise and edgewise shared partner terms to capture refined forms of transitivity within communities. We use simulations to demonstrate the performance of the two-stage strategy in terms of the estimated community structure. In addition, we show that the estimated ERGMs with geometrically weighted dyadwise and edgewise shared partner terms within communities outperform the working model in terms of goodness-of-fit. Last, but not least, we present an application to high-resolution human contact network data. Ming Cao 0005, Yong Chen 0016, Kayo Fujimoto, Michael Schweinberger |
ASONAM | 2 |
| 2018 | Comparing adverse effects of Hepatitis C drugs using FAERS data
Jing Huang 0021, Xinyuan Zhang 0003, Jiayi Tong, Jingcheng Du, Rui Duan 0004, Liu Yang 0026, Jason H. Moore, Yong Chen 0016, Cui Tao |
BIBM | 8 |
| 2018 | A Self-Organizing Tensor Architecture for Multi-view ClusteringabstractIn many real-world applications, data are often unlabeled and comprised of different representations/views which often provide information complementary to each other. Although several multi-view clustering methods have been proposed, most of them routinely assume one weight for one view of features, and thus inter-view correlations are only considered at the view-level. These approaches, however, fail to explore the explicit correlations between features across multiple views. In this paper, we introduce a tensor-based approach to incorporate the higher-order interactions among multiple views as a tensor structure. Specifically, we propose a multi-linear multi-view clustering (MMC) method that can efficiently explore the full-order structural information among all views and reveal the underlying subspace structure embedded within the tensor. Extensive experiments on realworld datasets demonstrate that our proposed MMC algorithm clearly outperforms other related state-of-the-art methods. Lifang He 0001, Chun-Ta Lu, Yong Chen 0016, Jiawei Zhang 0001, LinLin Shen, Philip S. Yu, Fei Wang 0001 |
ICDM | 3 |
| 2018 | PIE: A prior knowledge guided integrated likelihood estimation method for bias reduction in association studies using electronic health records dataabstractOBJECTIVES: This study proposes a novel Prior knowledge guided Integrated likelihood Estimation (PIE) method to correct bias in estimations of associations due to misclassification of electronic health record (EHR)-derived binary phenotypes, and evaluates the performance of the proposed method by comparing it to 2 methods in common practice. METHODS: We conducted simulation studies and data analysis of real EHR-derived data on diabetes from Kaiser Permanente Washington to compare the estimation bias of associations using the proposed method, the method ignoring phenotyping errors, the maximum likelihood method with misspecified sensitivity and specificity, and the maximum likelihood method with correctly specified sensitivity and specificity (gold standard). The proposed method effectively leverages available information on phenotyping accuracy to construct a prior distribution for sensitivity and specificity, and incorporates this prior information through the integrated likelihood for bias reduction. RESULTS: Our simulation studies and real data application demonstrated that the proposed method effectively reduces the estimation bias compared to the 2 current methods. It performed almost as well as the gold standard method when the prior had highest density around true sensitivity and specificity. The analysis of EHR data from Kaiser Permanente Washington showed that the estimated associations from PIE were very close to the estimates from the gold standard method and reduced bias by 60%-100% compared to the 2 commonly used methods in current practice for EHR data. CONCLUSIONS: This study demonstrates that the proposed method can effectively reduce estimation bias caused by imperfect phenotyping in EHR-derived data by incorporating prior information through integrated likelihood. Jing Huang 0021, Rui Duan 0004, Rebecca A. Hubbard, Yonghui Wu 0001, Jason H. Moore, Hua Xu 0001, Yong Chen 0016 |
J. Am. Medical Informatics Assoc. | 7 |
| 2017 | pETM: a penalized Exponential Tilt Model for analysis of correlated high-dimensional DNA methylation dataabstractMOTIVATION: DNA methylation plays an important role in many biological processes and cancer progression. Recent studies have found that there are also differences in methylation variations in different groups other than differences in methylation means. Several methods have been developed that consider both mean and variance signals in order to improve statistical power of detecting differentially methylated loci. Moreover, as methylation levels of neighboring CpG sites are known to be strongly correlated, methods that incorporate correlations have also been developed. We previously developed a network-based penalized logistic regression for correlated methylation data, but only focusing on mean signals. We have also developed a generalized exponential tilt model that captures both mean and variance signals but only examining one CpG site at a time. RESULTS: In this article, we proposed a penalized Exponential Tilt Model (pETM) using network-based regularization that captures both mean and variance signals in DNA methylation data and takes into account the correlations among nearby CpG sites. By combining the strength of the two models we previously developed, we demonstrated the superior power and better performance of the pETM method through simulations and the applications to the 450K DNA methylation array data of the four breast invasive carcinoma cancer subtypes from The Cancer Genome Atlas (TCGA) project. The developed pETM method identifies many cancer-related methylation loci that were missed by our previously developed method that considers correlations among nearby methylation loci but not variance signals. AVAILABILITY AND IMPLEMENTATION: The R package 'pETM' is publicly available through CRAN: http://cran.r-project.org . CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hokeun Sun, Yong Chen 0016 |
Bioinform. | 3 |
| 2017 | SeqCNV: a novel method for identification of copy number variations in targeted next-generation sequencing dataabstractBACKGROUND: Targeted next-generation sequencing (NGS) has been widely used as a cost-effective way to identify the genetic basis of human disorders. Copy number variations (CNVs) contribute significantly to human genomic variability, some of which can lead to disease. However, effective detection of CNVs from targeted capture sequencing data remains challenging. RESULTS: Here we present SeqCNV, a novel CNV calling method designed to use capture NGS data. SeqCNV extracts the read depth information and utilizes the maximum penalized likelihood estimation (MPLE) model to identify the copy number ratio and CNV boundary. We applied SeqCNV to both bacterial artificial clone (BAC) and human patient NGS data to identify CNVs. These CNVs were validated by array comparative genomic hybridization (aCGH). CONCLUSIONS: SeqCNV is able to robustly identify CNVs of different size using capture NGS data. Compared with other CNV-calling methods, SeqCNV shows a significant improvement in both sensitivity and specificity. Yong Chen 0016, Ming Cao 0005, Violet Gelowani, Mingchu Xu, Smriti A. Agrawal, Yumei Li 0007, Stephen P. Daiger, Richard A. Gibbs, Fei Wang 0017, Rui Chen 0013 |
BMC Bioinform. | 1 |
| 2016 | An Empirical Study for Impacts of Measurement Errors on EHR based Association Studies
Rui Duan 0004, Ming Cao 0005, Yonghui Wu 0001, Jing Huang 0021, Joshua C. Denny, Hua Xu 0001, Yong Chen 0016 |
AMIA | 7 |
| 2013 | SIBER: systematic identification of bimodally expressed genes using RNAseq dataabstractMOTIVATION: Identification of bimodally expressed genes is an important task, as genes with bimodal expression play important roles in cell differentiation, signalling and disease progression. Several useful algorithms have been developed to identify bimodal genes from microarray data. Currently, no method can deal with data from next-generation sequencing, which is emerging as a replacement technology for microarrays. RESULTS: We present SIBER (systematic identification of bimodally expressed genes using RNAseq data) for effectively identifying bimodally expressed genes from next-generation RNAseq data. We evaluate several candidate methods for modelling RNAseq count data and compare their performance in identifying bimodal genes through both simulation and real data analysis. We show that the lognormal mixture model performs best in terms of power and robustness under various scenarios. We also compare our method with alternative approaches, including profile analysis using clustering and kurtosis (PACK) and cancer outlier profile analysis (COPA). Our method is robust, powerful, invariant to shifting and scaling, has no blind spots and has a sample-size-free interpretation. AVAILABILITY: The R package SIBER is available at the website http://bioinformatics.mdanderson.org/main/OOMPA:Overview. Pan Tong, Yong Chen 0016, Kevin R. Coombes |
Bioinform. | 2 |
| 2012 | Ensemble Clustering for Internet Security ApplicationsabstractDue to their damage to Internet security, malware and phishing website detection has been the Internet security topics that are of great interests. Compared with malware attacks, phishing website fraud is a relatively new Internet crime. However, they share some common properties: 1) both malware samples and phishing websites are created at a rate of thousands per day driven by economic benefits; and 2) phishing websites represented by the term frequencies of the webpage content share similar characteristics with malware samples represented by the instruction frequencies of the program. Over the past few years, many clustering techniques have been employed for automatic malware and phishing website detection. In these techniques, the detection process is generally divided into two steps: 1) feature extraction, where representative features are extracted to capture the characteristics of the file samples or the websites; and 2) categorization, where intelligent techniques are used to automatically group the file samples or websites into different classes based on computational analysis of the feature representations. However, few have been applied in real industry products. In this paper, we develop an automatic categorization system to automatically group phishing websites or malware samples using a cluster ensemble by aggregating the clustering solutions that are generated by different base clustering algorithms. We propose a principled cluster ensemble framework to combine individual clustering solutions that are based on the consensus partition, which can not only be applied for malware categorization, but also for phishing website clustering. In addition, the domain knowledge in the form of sample-level/website-level constraints can be naturally incorporated into the ensemble framework. The case studies on large and real daily phishing websites and malware collection from the Kingsoft Internet Security Laboratory demonstrate the effectiveness and efficiency of our proposed method. Weiwei Zhuang, Yanfang Ye 0001, Yong Chen 0016, Tao Li 0001 |
IEEE Trans. Syst. Man Cybern. Part C | 3 |
| 2010 | Automatic malware categorization using cluster ensembleabstractIn this paper, resting on the analysis of instruction frequency and function-based instruction sequences, we develop an Automatic Malware Categorization System (AMCS) for automatically grouping malware samples into families that share some common characteristics using a cluster ensemble by aggregating the clustering solutions generated by different base clustering algorithms. We propose a principled cluster ensemble framework for combining individual clustering solutions based on the consensus partition. The domain knowledge in the form of sample-level constraints can be naturally incorporated in the ensemble framework. In addition, to account for the characteristics of feature representations, we propose a hybrid hierarchical clustering algorithm which combines the merits of hierarchical clustering and k-medoids algorithms and a weighted subspace K-medoids algorithm to generate base clusterings. The categorization results of our AMCS system can be used to generate signatures for malware families that are useful for malware detection. The case studies on large and real daily malware collection from Kingsoft Anti-Virus Lab demonstrate the effectiveness and efficiency of our AMCS system. Yanfang Ye 0001, Tao Li 0001, Yong Chen 0016, Qingshan Jiang |
KDD | 3 |
| 2010 | Hierarchical associative classifier (HAC) for malware detection from the large and imbalanced gray list
Yanfang Ye 0001, Tao Li 0001, Qingshan Jiang, Yong Chen 0016 |
J. Intell. Inf. Syst. | 5 |