EDBT 2026 Demo / reviewers in the wild / expert
Pengyue Zhang
dblp:136/9293
· DBLP profile ↗
14ranked-venue papers
3as first author
7since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Integrating explainable artificial intelligence with multiomics systems biology and electronic health record data mining for personalized drug repurposing in Alzheimer's diseaseabstractAlzheimer's disease (AD) is characterized by region- and patient-specific molecular heterogeneity, which hinders therapeutic design. In this study, we introduce PRISM-ML (PRecision-medicine using Interpretable Systems and Multiomics with Machine Learning), an open-source integrated analysis pipeline that combines interpretable machine learning with systems biology and electronic health records data mining to elucidate the molecular diversity of AD and predict promising drug repurposing opportunities. First, we integrated and harmonized transcriptomic (bulk RNA-seq) and genomic (genome-wide association study) data from 2105 brain samples, each with matched data from the same individual (1363 AD patients, 742 controls; 9 tissues), sourced from three independent studies. Random forest classifiers with SHapley Additive exPlanations identified patient-specific biomarkers; unsupervised clustering resolved 36 molecularly distinct subtissues (defined as clusters of samples within a brain tissue that share a specific expression pattern); and gene-gene coexpression networks prioritized 262 high-centrality bottleneck genes as putative regulators of dysregulated pathways. Next, knowledge graph-based drug repurposing predicted six Food and Drug Administration (FDA)-approved drugs that simultaneously target multiple bottleneck genes and multiple AD-relevant pathways. Notably, in a large US de-identified insurance-claims database (n = 364 733), exposure to promethazine, one of the candidate drugs, was associated with a 57%-62% lower incidence of AD versus an active antihistamine comparator (adjusted hazard ratio 0.38; inverse-probability weighted 0.43; both P < .001), providing real-world support for its repurposing potential. In summary, PRISM-ML, as an explainable multiomics analysis pipeline, is readily transferable to other complex diseases, advancing precision medicine. Mohammadsadeq Mottaqi, Pengyue Zhang |
Briefings Bioinform. | 2 |
| 2025 | A trajectory-informed model for detecting drug-drug-host interaction from real-world dataabstractOBJECTIVE: Adverse drug event (ADE) is a significant challenge to public health. Since data mining methods have been developed to identify signals of drug-drug interaction-induced (DDI-induced) or drug-host interaction-induced (DHI-induced) ADE from real-world data, we aim to develop a new method to detect adverse drug-drug interaction with a special awareness on patient characteristics. METHODS: We developed a trajectory-informed model (TIM) to identify signals of adverse DDI with a special awareness on patient characteristics (i.e., drug-drug-host interaction [DDHI]). We also proposed a study design based on an optimal selection of within-subject and between-subjects controls for detecting ADEs from real-world data. We analyzed a large-scale US administrative claims data and conducted a simulation study. RESULTS: In administrative claims data analysis, we developed optimally matched case-control datasets for potential ADEs including acute kidney injury and gastrointestinal bleeding. We identified that an optimal selection of controls had a higher AUC compared to traditional designs for ADE detection (AUCs: 0.79-0.80 vs. 0.56-0.76). We observed that TIM detected more signals than reference methods (odds ratios: 1.13-3.18, P < 0.01), and found that 36 % of all signals generated by TIM were DDHI signals. In a simulation study, we demonstrated that TIM had an empirical false discovery rate (FDR) less than the desired value of 0.05, as well as > 1.4-fold higher probabilities of detection of DDHI signals than reference methods. CONCLUSIONS: TIM had a high probability to identify signals of adverse DDI and DDHI in a high-throughput ADE mining while controlling false positive rate. A significant portion of drug-drug combinations were associated with an increased risk of ADEs only in specific patient subpopulations. Optimal selection of within-subject and between-subjects controls could improve the performance of ADE data mining. Anna Sun, Hongmei Nan, Yuedi Yang, Michael Eadon, Jing Su 0003, Pengyue Zhang |
J. Biomed. Informatics | 8 |
| 2024 | Health disparities in the risk of severe acidosis: real-world evidence from the All of Us cohortabstractOBJECTIVE: To assess the health disparities across social determinants of health (SDoH) domains for the risk of severe acidosis independent of demographical and clinical factors. MATERIALS AND METHODS: A retrospective case-control study (n = 13 310, 1:4 matching) is performed using electronic health records (EHRs), SDoH surveys, and genomics data from the All of Us participants. The propensity score matching controls confounding effects due to EHR data availability. Conditional logistic regressions are used to estimate odds ratios describing associations between SDoHs and the risk of acidosis events, adjusted for demographic features, and clinical conditions. RESULTS: Those with employer-provided insurance and those with Medicaid plans show dramatically different risks [adjusted odds ratio (AOR): 0.761 vs 1.41]. Low-income groups demonstrate higher risk (household income less than $25k, AOR: 1.3-1.57) than high-income groups ($100-$200k, AOR: 0.597-0.867). Other high-risk factors include impaired mobility (AOR: 1.32), unemployment (AOR: 1.32), renters (AOR: 1.41), other non-house-owners (AOR: 1.7), and house instability (AOR: 1.25). Education was negatively associated with acidosis risk. DISCUSSION: Our work provides real-world evidence of the comprehensive health disparities due to socioeconomic and behavioral contributors in a cohort enriched in minority groups or underrepresented populations. CONCLUSIONS: SDoHs are strongly associated with systematic health disparities in the risk of severe metabolic acidosis. Types of health insurance, household income levels, housing status and stability, employment status, educational level, and mobility disability play significant roles after being adjusted for demographic features and clinical conditions. Comprehensive solutions are needed to improve equity in healthcare and reduce the risk of severe acidosis. Allison E. Gatz, Chenxi Xiong, Shihui Jiang, Chi Mai Nguyen, Qianqian Song 0002, Xiaochun Li 0003, Pengyue Zhang, Michael Eadon, Jing Su 0003 |
J. Am. Medical Informatics Assoc. | 8 |
| 2024 | TrajVis: a visual clinical decision support system to translate artificial intelligence trajectory models in the precision management of chronic kidney diseaseabstractOBJECTIVE: Our objective is to develop and validate TrajVis, an interactive tool that assists clinicians in using artificial intelligence (AI) models to leverage patients' longitudinal electronic medical records (EMRs) for personalized precision management of chronic disease progression. MATERIALS AND METHODS: We first perform requirement analysis with clinicians and data scientists to determine the visual analytics tasks of the TrajVis system as well as its design and functionalities. A graph AI model for chronic kidney disease (CKD) trajectory inference named DisEase PrOgression Trajectory (DEPOT) is used for system development and demonstration. TrajVis is implemented as a full-stack web application with synthetic EMR data derived from the Atrium Health Wake Forest Baptist Translational Data Warehouse and the Indiana Network for Patient Care research database. A case study with a nephrologist and a user experience survey of clinicians and data scientists are conducted to evaluate the TrajVis system. RESULTS: The TrajVis clinical information system is composed of 4 panels: the Patient View for demographic and clinical information, the Trajectory View to visualize the DEPOT-derived CKD trajectories in latent space, the Clinical Indicator View to elucidate longitudinal patterns of clinical features and interpret DEPOT predictions, and the Analysis View to demonstrate personal CKD progression trajectories. System evaluations suggest that TrajVis supports clinicians in summarizing clinical data, identifying individualized risk predictors, and visualizing patient disease progression trajectories, overcoming the barriers of AI implementation in healthcare. DISCUSSION: The TrajVis system provides a novel visualization solution which is complimentary to other risk estimators such as the Kidney Failure Risk Equations. CONCLUSION: TrajVis bridges the gap between the fast-growing AI/ML modeling and the clinical use of such models for personalized and precision management of chronic diseases. Zuotian Li, Xiang Liu 0016, Ziyang Tang, Nanxin Jin, Pengyue Zhang, Michael Eadon, Qianqian Song 0002, Victor Y. Chen, Jing Su 0003 |
J. Am. Medical Informatics Assoc. | 5 |
| 2021 | PASC CKD: revealing the progression trajectories of sustained COVID-19-related renal injury using real-world evidence
Jing Su 0003, Pengyue Zhang, Zuoyi Zhang, Michael Eadon, Xiaochun Li 0003, Stanley Taylor, Travis Johnson, Zhaorui Liu, Ziyang Tang, Baijian Yang 0001, Qianqian Song 0002, Kun Huang 0001 |
AMIA | 2 |
| 2021 | Improved Adverse Drug Event Prediction Through Information Component Guided Pharmacological Network Model (IC-PNM)abstractImproving adverse drug event (ADE) prediction is highly critical in pharmacovigilance research. We propose a novel information component guided pharmacological network model (IC-PNM) to predict drug-ADE signals. This new method combines the pharmacological network model and information component, a Bayes statistics method. We use 33,947 drug-ADE pairs from the FDA Adverse Event Reporting System (FAERS) 2010 data as the training data, and the new 21,065 drug-ADE pairs from FAERS 2011-2015 as the validations samples. The IC-PNM data analysis suggests that both large and small sample size drug-ADE pairs are needed in training the predictive model for its prediction performance to reach an area under the receiver operating characteristic curve [Formula: see text]. On the other hand, the IC-PNM prediction performance improved to [Formula: see text] if we removed the small sample size drug-ADE pairs from the prediction model during validation. Xiangmin Ji, Lei Wang 0169, Liyan Hua, Pengyue Zhang, Aditi Shendre, Weixing Feng, Jin Li 0012, Lang Li 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2021 | A Fast and Furious Bayesian Network and Its Application of Identifying Colon Cancer to Liver Metastasis Gene Regulatory NetworksabstractBayesian networks is a powerful method for identifying causal relationships among variables. However, as the network size increases, the time complexity of searching the optimal structure grows exponentially. We proposed a novel search algorithm - Fast and Furious Bayesian Network (FFBN). Compared to the existing greedy search algorithm, FFBN uses significantly fewer model configuration rules to determine the causal direction of edges when constructing the Bayesian network, which leads to greatly improved computational speed. We benchmarked the performance of FFBN by reconstructing gene regulatory networks (GRNs) from two DREAM5 challenge datasets: a synthetic dataset and a larger yeast transcriptome dataset. In both datasets, FFBN shows a much faster speed than the existing greedy search algorithm, while maintaining equally good or better performance in recall and precision. We then constructed three whole transcriptome GRNs for primary liver cancer (PL), primary colon cancer (PC) and colon to liver metastasis (CLM) expression data, which the existing greedy search algorithms failed. Three GRNs contain 12,099 common genes. Unprecedentedly, our newly developed FFBN algorithm is able to build up GRNs at a scale larger than 10,000 genes. Using FFBN, we discovered that CLM has its unique cancer molecular mechanisms and shares a certain degree of similarity with both PL and PC. Enze Liu 0002, Jin Li 0024, Garrett Kinnebrew, Pengyue Zhang, Yan Zhang 0032, Lijun Cheng, Lang Li 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2020 | Network-based prediction of drug-target interactions using an arbitrary-order proximity embedded deep forestabstractMOTIVATION: Systematic identification of molecular targets among known drugs plays an essential role in drug repurposing and understanding of their unexpected side effects. Computational approaches for prediction of drug-target interactions (DTIs) are highly desired in comparison to traditional experimental assays. Furthermore, recent advances of multiomics technologies and systems biology approaches have generated large-scale heterogeneous, biological networks, which offer unexpected opportunities for network-based identification of new molecular targets among known drugs. RESULTS: In this study, we present a network-based computational framework, termed AOPEDF, an arbitrary-order proximity embedded deep forest approach, for prediction of DTIs. AOPEDF learns a low-dimensional vector representation of features that preserve arbitrary-order proximity from a highly integrated, heterogeneous biological network connecting drugs, targets (proteins) and diseases. In total, we construct a heterogeneous network by uniquely integrating 15 networks covering chemical, genomic, phenotypic and network profiles among drugs, proteins/targets and diseases. Then, we build a cascade deep forest classifier to infer new DTIs. Via systematic performance evaluation, AOPEDF achieves high accuracy in identifying molecular targets among known drugs on two external validation sets collected from DrugCentral [area under the receiver operating characteristic curve (AUROC) = 0.868] and ChEMBL (AUROC = 0.768) databases, outperforming several state-of-the-art methods. In a case study, we showcase that multiple molecular targets predicted by AOPEDF are associated with mechanism-of-action of substance abuse disorder for several marketed drugs (such as aripiprazole, risperidone and haloperidol). AVAILABILITY AND IMPLEMENTATION: Source code and data can be downloaded from https://github.com/ChengF-Lab/AOPEDF. Xiangxiang Zeng, Siyi Zhu, Yuan Hou, Pengyue Zhang, Lang Li 0001, L. Frank Huang, Stephen J. Lewis, Ruth Nussinov, Feixiong Cheng |
Bioinform. | 4 |
| 2019 | Mining Directional Drug Interaction Effects on Myopathy Using the FAERS DatabaseabstractMining high-order drug-drug interaction (DDI) induced adverse drug effects from electronic health record databases is an emerging area, and very few studies have explored the relationships between high-order drug combinations. We investigate a novel pharmacovigilance problem for mining directional DDI effects on myopathy using the FDA Adverse Event Reporting System (FAERS) database. Our paper provides information on the risk of myopathy associated with adding new drugs on the already prescribed medication, and visualizes the identified directional DDI patterns as user-friendly graphical representation. We utilize the Apriori algorithm to extract frequent drug combinations from the FAERS database. We use odds ratio to estimate the risk of myopathy associated with directional DDI. We create a tree-structured graph to visualize the findings for easy interpretation. Our method confirmed myopathy association with previously reported HMG-CoA reductase inhibitors like rosuvastatin, fluvastatin, simvastatin, and atorvastatin. New, previously unidentified but mechanistically plausible associations with myopathy were also observed, such as the DDI between pamidronate and levofloxacin. Additional top findings are gadolinium-based imaging agents, which however are often used in myopathy diagnosis. Other DDIs with no obvious mechanism are also reported, such as that of sulfamethoxazole with trimethoprim and potassium chloride. This study shows the feasibility to estimate high-order directional DDIs in a fast and accurate manner. The results of the analysis could become a useful tool in the specialists' hands through an easy-to-understand graphic visualization. Danai Chasioti, Xiaohui Yao, Pengyue Zhang, Samuel Lerner, Sara K. Quinney, Xia Ning, Lang Li 0001, Li Shen 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2018 | Multi-channel Generative Adversarial Network for Parallel Magnetic Resonance Image Reconstruction in K-space
Pengyue Zhang, Fusheng Wang 0001, Wei Xu 0020 |
MICCAI (1) | 1 |
| 2018 | Deep Reinforcement Learning for Vessel Centerline Tracing in Multi-modality 3D Volumes
Pengyue Zhang, Fusheng Wang 0001, Yefeng Zheng 0001 |
MICCAI (4) | 1 |
| 2016 | Sparse discriminative multi-manifold embedding for one-sample face identification
Pengyue Zhang, Xinge You, Weihua Ou, C. L. Philip Chen, Yiu-Ming Cheung |
Pattern Recognit. | 1 |
| 2014 | Robust face recognition via occlusion dictionary learning
Weihua Ou, Xinge You, Dacheng Tao, Pengyue Zhang, Yuan Yan Tang |
Pattern Recognit. | 4 |
| 2013 | Learning a Sparse Representation for Robust Face Recognition
Weihua Ou, Xinge You, Pengyue Zhang, Xiubao Jiang, Duanquan Xu |
ICONIP (3) | 3 |