Lang Li 0001

dblp:l/LangLi · DBLP profile ↗
← Back
24ranked-venue papers
1as first author
11since 2021 · last 2026
0000-0002-0746-1809ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 23 · 1 first-author · 10 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Derivation and validation of an algorithm for maternal-child linkage in electronic health records
abstract
INTRODUCTION: We created a probabilistic maternal-child electronic health record (EHR) linkage algorithm to promote clinical research in maternal-child health. METHODS: We used EHR data from 1994 to 2024 to create an XGBoost model to predict maternal-child linkages. The model used standard EHR elements as predictor variables, including first name, last name, birthdate, address, phone number, email, and an EHR-embedded maternal-child indicator as the deterministic outcome. RESULTS: From 82 million unique records, 6.2 billion potential pairs met blocking criteria. Of the potential pairs, 33 364 674 contained the deterministic indicator and were used as cases, and an equal number of controls were randomly sampled. The final model obtained an accuracy of 92%, a precision of 98%, a recall of 87%, and an F1-score of 92%. CONCLUSION: We derived and validated a probabilistic maternal-child linkage algorithm using routinely collected EHR data elements that could benefit future observational research in maternal-child health.
Colin M. Rogerson, Christopher W. Bartlett, John P. Price, Lang Li 0001, Eneida A. Mendonça, Shaun J. Grannis
J. Am. Medical Informatics Assoc.4
2025 A multi-layer encoder prediction model for individual sample specific gene combination effect (MLEC-iGeneCombo)
abstract
Using data from gene combination double knockout (CDKO) experiments, top ranked synthetic lethal (SL) gene pairs were highly inconsistent among different SL scores. This leads to a significant concern that SL prediction models highly depend on SL scores. In this paper, we introduce a new gene combination effect (GCE) measurement, log-fold change of dual-gRNA expression before and after CRISPR-cas9 lentivirus transfection. We show it is a direct and highly consistent measurement of GCE in all CDKO experiments. We therefore develop a multi-layer encoder model for individual sample specific GCE prediction, MLEC-iGeneCombo. Under a deep learning framework, MLEC-iGeneCombo is a systems biology model that contains sample specific multi-omics encoder, network encoder and cell-line encoder. For the first time, MLEC-iGeneCombo predicts GCE for a new cell. Using data from 18 CDKO experiments, MLEC-iGeneCombo achieves an average GCE prediction performance, 71.9%. All three encoders significantly improve the model's prediction performance (p[Formula: see text]), and their combined use yields the best GCE prediction performance. Our source code is available at https://github.com/karenyun/MLEC-iGeneCombo.
Kunjie Fan, Birkan Gokbag, Nuo Sun, Lijun Cheng, Lang Li 0001
PLoS Comput. Biol.7
2025 SubgroupTE: Advancing Treatment Effect Estimation with Subgroup Identification
abstract
Precise estimation of treatment effects is crucial for accurately evaluating the intervention. While deep learning models have exhibited promising performance in learning counterfactual representations for treatment effect estimation (TEE), a major limitation in most of these models is that they often overlook the diversity of treatment effects across potential subgroups that have varying treatment effects and characteristics, treating the entire population as a homogeneous group. This limitation restricts the ability to precisely estimate treatment effects and provide targeted treatment recommendations. In this paper, we propose a novel treatment effect estimation model, named SubgroupTE, which incorporates subgroup identification in TEE. SubgroupTE identifies heterogeneous subgroups with different responses and more precisely estimates treatment effects by considering subgroup-specific treatment effects in the estimation process. In addition, we introduce an expectation-maximization (EM)-based training process that iteratively optimizes estimation and subgrouping networks to improve both estimation and subgroup identification. Comprehensive experiments on the synthetic and semi-synthetic datasets demonstrate the outstanding performance of SubgroupTE compared to the existing works for treatment effect estimation and subgrouping models. Additionally, a real-world study demonstrates the capabilities of SubgroupTE in enhancing targeted treatment recommendations for patients with opioid use disorder (OUD) by incorporating subgroup identification with treatment effect estimation.
Seungyeon Lee 0002, Ruoqi Liu, Wenyu Song, Lang Li 0001, Ping Zhang 0016
ACM Trans. Intell. Syst. Technol.4
2024 Discovering clinical drug-drug interactions with known pharmacokinetics mechanisms using spontaneous reporting systems and electronic health records
abstract
OBJECTIVE: Although the mechanisms behind pharmacokinetic (PK) drug-drug interactions (DDIs) are well-documented, bridging the gap between this knowledge and clinical evidence of DDIs, especially for serious adverse drug reactions (SADRs), remains challenging. While leveraging the FDA Adverse Event Reporting System (FAERS) database along with disproportionality analysis tends to detect a vast number of DDI signals, this abundance complicates further investigation, such as validation through clinical trials. Our study proposed a framework to efficiently prioritize these signals and assessed their reliability using multi-source Electronic Health Records (EHR) to identify top candidates for further investigation. METHODS: We analyzed FAERS data spanning from January 2004 to March 2023, employing four established disproportionality methods: Proportional Reporting Ratio (PRR), Reporting Odds Ratio (ROR), Multi-item Gamma Poisson Shrinker (MGPS), and Bayesian Confidence Propagating Neural Network (BCPNN). Building upon these models, we developed four ranking models to prioritize DDI-SADR signals and cross-referenced signals with DrugBank. To validate the top-ranked signals, we employed longitudinal EHRs from Vanderbilt University Medical Center and the All of Us research program. The performance of each model was assessed by counting how many of the top-ranked signals were confirmed by EHRs and calculating the average ranking of these confirmed signals. RESULTS: Out of 189 DDI-SADR signals identified by all four disproportionality methods, only two were documented in the DrugBank database. By prioritizing the top 20 signals as determined by each of the four disproportionality methods and our four ranking models, 58 unique DDI-SADR signals were selected for EHR validations. Of these, five signals were confirmed. The ranking model, which integrated the MGPS and BCPNN, demonstrated superior performance by assigning the highest priority to those five EHR-confirmed signals. CONCLUSION: The fusion of disproportionality analysis with ranking models, validated through multi-source EHRs, presents a groundbreaking approach to pharmacovigilance. Our study's confirmation of five significant DDI-SADRs, previously unrecorded in the DrugBank database, highlights the essential role of advanced data analysis techniques in identifying ADRs.
Eugene Jeong, Yu Su 0001, Lang Li 0001, You Chen 0001
J. Biomed. Informatics3
2022 CEDA: integrating gene expression data with CRISPR-pooled screen data identifies essential genes with higher expression
abstract
MOTIVATION: Clustered regularly interspaced short palindromic repeats (CRISPR)-based genetic perturbation screen is a powerful tool to probe gene function. However, experimental noises, especially for the lowly expressed genes, need to be accounted for to maintain proper control of false positive rate. METHODS: We develop a statistical method, named CRISPR screen with Expression Data Analysis (CEDA), to integrate gene expression profiles and CRISPR screen data for identifying essential genes. CEDA stratifies genes based on expression level and adopts a three-component mixture model for the log-fold change of single-guide RNAs (sgRNAs). Empirical Bayesian prior and expectation-maximization algorithm are used for parameter estimation and false discovery rate inference. RESULTS: Taking advantage of gene expression data, CEDA identifies essential genes with higher expression. Compared to existing methods, CEDA shows comparable reliability but higher sensitivity in detecting essential genes with moderate sgRNA fold change. Therefore, using the same CRISPR data, CEDA generates an additional hit gene list. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Lianbo Yu, Haoran Li 0015, Kevin R. Coombes, Kin Fai Au, Lijun Cheng, Lang Li 0001
Bioinform.8
2022 DGCyTOF: Deep learning with graphic cluster visualization to predict cell types of single cell mass cytometry data
abstract
Single-cell mass cytometry, also known as cytometry by time of flight (CyTOF) is a powerful high-throughput technology that allows analysis of up to 50 protein markers per cell for the quantification and classification of single cells. Traditional manual gating utilized to identify new cell populations has been inadequate, inefficient, unreliable, and difficult to use, and no algorithms to identify both calibration and new cell populations has been well established. A deep learning with graphic cluster (DGCyTOF) visualization is developed as a new integrated embedding visualization approach in identifying canonical and new cell types. The DGCyTOF combines deep-learning classification and hierarchical stable-clustering methods to sequentially build a tri-layer construct for known cell types and the identification of new cell types. First, deep classification learning is constructed to distinguish calibration cell populations from all cells by softmax classification assignment under a probability threshold, and graph embedding clustering is then used to identify new cell populations sequentially. In the middle of two-layer, cell labels are automatically adjusted between new and unknown cell populations via a feedback loop using an iteration calibration system to reduce the rate of error in the identification of cell types, and a 3-dimensional (3D) visualization platform is finally developed to display the cell clusters with all cell-population types annotated. Utilizing two benchmark CyTOF databases comprising up to 43 million cells, we compared accuracy and speed in the identification of cell types among DGCyTOF, DeepCyTOF, and other technologies including dimension reduction with clustering, including Principal Component Analysis (PCA), Factor Analysis (FA), Independent Component Analysis (ICA), Isometric Feature Mapping (Isomap), t-distributed Stochastic Neighbor Embedding (t-SNE), and Uniform Manifold Approximation and Projection (UMAP) with k-means clustering and Gaussian mixture clustering. We observed the DGCyTOF represents a robust complete learning system with high accuracy, speed and visualization by eight measurement criteria. The DGCyTOF displayed F-scores of 0.9921 for CyTOF1 and 0.9992 for CyTOF2 datasets, whereas those scores were only 0.507 and 0.529 for the t-SNE+k-means; 0.565 and 0.59, for UMAP+ k-means. Comparison of DGCyTOF with t-SNE and UMAP visualization in accuracy demonstrated its approximately 35% superiority in predicting cell types. In addition, observation of cell-population distribution was more intuitive in the 3D visualization in DGCyTOF than t-SNE and UMAP visualization. The DGCyTOF model can automatically assign known labels to single cells with high accuracy using deep-learning classification assembling with traditional graph-clustering and dimension-reduction strategies. Guided by a calibration system, the model seeks optimal accuracy balance among calibration cell populations and unknown cell types, yielding a complete and robust learning system that is highly accurate in the identification of cell populations compared to results using other methods in the analysis of single-cell CyTOF data. Application of the DGCyTOF method to identify cell populations could be extended to the analysis of single-cell RNASeq data and other omics data.
Lijun Cheng, Pratik Karkhanis, Birkan Gokbag, Yueze Liu, Lang Li 0001
PLoS Comput. Biol.5
2022 DSCN: Double-target selection guided by CRISPR screening and network
abstract
Cancer is a complex disease with usually multiple disease mechanisms. Target combination is a better strategy than a single target in developing cancer therapies. However, target combinations are generally more difficult to be predicted. Current CRISPR-cas9 technology enables genome-wide screening for potential targets, but only a handful of genes have been screend as target combinations. Thus, an effective computational approach for selecting candidate target combinations is highly desirable. Selected target combinations also need to be translational between cell lines and cancer patients. We have therefore developed DSCN (double-target selection guided by CRISPR screening and network), a method that matches expression levels in patients and gene essentialities in cell lines through spectral-clustered protein-protein interaction (PPI) network. In DSCN, a sub-sampling approach is developed to model first-target knockdown and its impact on the PPI network, and it also facilitates the selection of a second target. Our analysis first demonstrated a high correlation of the DSCN sub-sampling-based gene knockdown model and its predicted differential gene expressions using observed gene expression in 22 pancreatic cell lines before and after MAP2K1 and MAP2K2 inhibition (R2 = 0.75). In DSCN algorithm, various scoring schemes were evaluated. The 'diffusion-path' method showed the most significant statistical power of differentialting known synthetic lethal (SL) versus non-SL gene pairs (P = 0.001) in pancreatic cancer. The superior performance of DSCN over existing network-based algorithms, such as OptiCon and VIPER, in the selection of target combinations is attributable to its ability to calculate combinations for any gene pairs, whereas other approaches focus on the combinations among optimized regulators in the network. DSCN's computational speed is also at least ten times fast than that of other methods. Finally, in applying DSCN to predict target combinations and drug combinations for individual samples (DSCNi), DSCNi showed high correlation between target combinations predicted and real synergistic combinations (P = 1e-5) in pancreatic cell lines. In summary, DSCN is a highly effective computational method for the selection of target combinations.
Enze Liu 0002, Lei Wang 0169, Yang Huo, Huanmei Wu, Lang Li 0001, Lijun Cheng
PLoS Comput. Biol.6
2021 Discovering Drug-Drug Interactions in COVID-19 Patients
Eugene Jeong, Anna K. Person, Joanna L. Stollings, Lang Li 0001, You Chen 0001
AMIA4
2021 Artificial intelligence and machine learning methods in predicting anti-cancer drug combination effects
abstract
Drug combinations have exhibited promising therapeutic effects in treating cancer patients with less toxicity and adverse side effects. However, it is infeasible to experimentally screen the enormous search space of all possible drug combinations. Therefore, developing computational models to efficiently and accurately identify potential anti-cancer synergistic drug combinations has attracted a lot of attention from the scientific community. Hypothesis-driven explicit mathematical methods or network pharmacology models have been popular in the last decade and have been comprehensively reviewed in previous surveys. With the surge of artificial intelligence and greater availability of large-scale datasets, machine learning especially deep learning methods are gaining popularity in the field of computational models for anti-cancer drug synergy prediction. Machine learning-based methods can be derived without strong assumptions about underlying mechanisms and have achieved state-of-the-art prediction performances, promoting much greater growth of the field. Here, we present a structured overview of available large-scale databases and machine learning especially deep learning methods in computational predictive models for anti-cancer drug synergy prediction. We provide a unified framework for machine learning models and detail existing model architectures as well as their contributions and limitations, shedding light into the future design of computational models. Besides, unbiased experiments are conducted to provide in-depth comparisons between reviewed papers in terms of their prediction performance.
Kunjie Fan, Lijun Cheng, Lang Li 0001
Briefings Bioinform.3
2021 Improved Adverse Drug Event Prediction Through Information Component Guided Pharmacological Network Model (IC-PNM)
abstract
Improving adverse drug event (ADE) prediction is highly critical in pharmacovigilance research. We propose a novel information component guided pharmacological network model (IC-PNM) to predict drug-ADE signals. This new method combines the pharmacological network model and information component, a Bayes statistics method. We use 33,947 drug-ADE pairs from the FDA Adverse Event Reporting System (FAERS) 2010 data as the training data, and the new 21,065 drug-ADE pairs from FAERS 2011-2015 as the validations samples. The IC-PNM data analysis suggests that both large and small sample size drug-ADE pairs are needed in training the predictive model for its prediction performance to reach an area under the receiver operating characteristic curve [Formula: see text]. On the other hand, the IC-PNM prediction performance improved to [Formula: see text] if we removed the small sample size drug-ADE pairs from the prediction model during validation.
Xiangmin Ji, Lei Wang 0169, Liyan Hua, Pengyue Zhang, Aditi Shendre, Weixing Feng, Jin Li 0012, Lang Li 0001
IEEE ACM Trans. Comput. Biol. Bioinform.9
2021 A Fast and Furious Bayesian Network and Its Application of Identifying Colon Cancer to Liver Metastasis Gene Regulatory Networks
abstract
Bayesian networks is a powerful method for identifying causal relationships among variables. However, as the network size increases, the time complexity of searching the optimal structure grows exponentially. We proposed a novel search algorithm - Fast and Furious Bayesian Network (FFBN). Compared to the existing greedy search algorithm, FFBN uses significantly fewer model configuration rules to determine the causal direction of edges when constructing the Bayesian network, which leads to greatly improved computational speed. We benchmarked the performance of FFBN by reconstructing gene regulatory networks (GRNs) from two DREAM5 challenge datasets: a synthetic dataset and a larger yeast transcriptome dataset. In both datasets, FFBN shows a much faster speed than the existing greedy search algorithm, while maintaining equally good or better performance in recall and precision. We then constructed three whole transcriptome GRNs for primary liver cancer (PL), primary colon cancer (PC) and colon to liver metastasis (CLM) expression data, which the existing greedy search algorithms failed. Three GRNs contain 12,099 common genes. Unprecedentedly, our newly developed FFBN algorithm is able to build up GRNs at a scale larger than 10,000 genes. Using FFBN, we discovered that CLM has its unique cancer molecular mechanisms and shares a certain degree of similarity with both PL and PC.
Enze Liu 0002, Jin Li 0024, Garrett Kinnebrew, Pengyue Zhang, Yan Zhang 0032, Lijun Cheng, Lang Li 0001
IEEE ACM Trans. Comput. Biol. Bioinform.7
2020 Network-based prediction of drug-target interactions using an arbitrary-order proximity embedded deep forest
abstract
MOTIVATION: Systematic identification of molecular targets among known drugs plays an essential role in drug repurposing and understanding of their unexpected side effects. Computational approaches for prediction of drug-target interactions (DTIs) are highly desired in comparison to traditional experimental assays. Furthermore, recent advances of multiomics technologies and systems biology approaches have generated large-scale heterogeneous, biological networks, which offer unexpected opportunities for network-based identification of new molecular targets among known drugs. RESULTS: In this study, we present a network-based computational framework, termed AOPEDF, an arbitrary-order proximity embedded deep forest approach, for prediction of DTIs. AOPEDF learns a low-dimensional vector representation of features that preserve arbitrary-order proximity from a highly integrated, heterogeneous biological network connecting drugs, targets (proteins) and diseases. In total, we construct a heterogeneous network by uniquely integrating 15 networks covering chemical, genomic, phenotypic and network profiles among drugs, proteins/targets and diseases. Then, we build a cascade deep forest classifier to infer new DTIs. Via systematic performance evaluation, AOPEDF achieves high accuracy in identifying molecular targets among known drugs on two external validation sets collected from DrugCentral [area under the receiver operating characteristic curve (AUROC) = 0.868] and ChEMBL (AUROC = 0.768) databases, outperforming several state-of-the-art methods. In a case study, we showcase that multiple molecular targets predicted by AOPEDF are associated with mechanism-of-action of substance abuse disorder for several marketed drugs (such as aripiprazole, risperidone and haloperidol). AVAILABILITY AND IMPLEMENTATION: Source code and data can be downloaded from https://github.com/ChengF-Lab/AOPEDF.
Xiangxiang Zeng, Siyi Zhu, Yuan Hou, Pengyue Zhang, Lang Li 0001, L. Frank Huang, Stephen J. Lewis, Ruth Nussinov, Feixiong Cheng
Bioinform.5
2019 Mining Directional Drug Interaction Effects on Myopathy Using the FAERS Database
abstract
Mining high-order drug-drug interaction (DDI) induced adverse drug effects from electronic health record databases is an emerging area, and very few studies have explored the relationships between high-order drug combinations. We investigate a novel pharmacovigilance problem for mining directional DDI effects on myopathy using the FDA Adverse Event Reporting System (FAERS) database. Our paper provides information on the risk of myopathy associated with adding new drugs on the already prescribed medication, and visualizes the identified directional DDI patterns as user-friendly graphical representation. We utilize the Apriori algorithm to extract frequent drug combinations from the FAERS database. We use odds ratio to estimate the risk of myopathy associated with directional DDI. We create a tree-structured graph to visualize the findings for easy interpretation. Our method confirmed myopathy association with previously reported HMG-CoA reductase inhibitors like rosuvastatin, fluvastatin, simvastatin, and atorvastatin. New, previously unidentified but mechanistically plausible associations with myopathy were also observed, such as the DDI between pamidronate and levofloxacin. Additional top findings are gadolinium-based imaging agents, which however are often used in myopathy diagnosis. Other DDIs with no obvious mechanism are also reported, such as that of sulfamethoxazole with trimethoprim and potassium chloride. This study shows the feasibility to estimate high-order directional DDIs in a fast and accurate manner. The results of the analysis could become a useful tool in the specialists' hands through an easy-to-understand graphic visualization.
Danai Chasioti, Xiaohui Yao, Pengyue Zhang, Samuel Lerner, Sara K. Quinney, Xia Ning, Lang Li 0001, Li Shen 0001
IEEE J. Biomed. Health Informatics7
2017 Identifying articles relevant to drug-drug interaction: Addressing class imbalance
abstract
Interactions between drugs (also known as drug-drug interactions or DDIs), which may cause adverse affects, are of much concern; predicting, anticipating and avoiding them is key for improving patient safety and treatment outcome. Knowledge of DDIs is important for physicians to avoid adverse effects when prescribing two drugs simultaneously. DDIs are often published in the biomedical literature; however, gathering information about DDIs is time consuming given the shear volume of publications. Automatic text classification can speed up access to documents related to DDIs. However, the biomedical literature contains a relatively small number of publications relevant to DDIs, compared to the vast amount of irrelevant publications. This imbalance can lead to incorrect classification. While methods addressing class imbalance have been introduced to correctly identify items in the minority (relevant) class to improve recall, they often misclassify items in the majority (irrelevant) class, which leads to low precision. To reduce the number of irrelevant documents misclassified as relevant (false positive), we develop a two-stage cascade classifier. In each step, we separate publication abstracts that are DDI-relevant from those that are either drug-irrelevant or drug-relevant but DDI-irrelevant. We compare our classifier with other popular learning methods that aim to handle imbalance, applying the methods to a well-curated corpus consisting of DDI-relevant and DDI-irrelevant PubMed abstracts. Our method achieves higher precision and F1 measure than other methods while maintaining similar recall.
Moumita Bhattacharya, Heng-Yi Wu, Pengyuan Li 0001, Lang Li 0001, Hagit Shatkay
BIBM5
2017 Using machine learning algorithms to identify genes essential for cell survival
abstract
BACKGROUND: With the explosion of data comes a proportional opportunity to identify novel knowledge with the potential for application in targeted therapies. In spite of this huge amounts of data, the solutions to treating complex disease is elusive. One reason being that these diseases are driven by a network of genes that need to be targeted in order to understand and treat them effectively. Part of the solution lies in mining and integrating information from various disciplines. Here we propose a machine learning method to mining through publicly available literature on RNA interference with the goal of identifying genes essential for cell survival. RESULTS: A total of 32,164 RNA interference abstracts were identified from 10.5 million pubmed abstracts (2001 - 2015). These abstracts spanned over 1467 cancer cell lines and 4373 genes representing a total of 25,891 cell gene associations. Among the 1467 cell lines 88% of them had at least 1 or up to 25 genes studied in a given cell line. Among the 4373 genes 96% of them were studied in at least 1 or up to 25 different cell lines. CONCLUSIONS: Identifying genes that are crucial for cell survival can be a critical piece of information especially in treating complex diseases, such as cancer. The efficacy of a therapeutic intervention is multifactorial in nature and in many cases the source of therapeutic disruption could be from an unsuspected source. Machine learning algorithms helps to narrow down the search and provides information about essential genes in different cancer types. It also provides the building blocks to generate a network of interconnected genes and processes. The information thus gained can be used to generate hypothesis which can be experimentally validated to improve our understanding of what triggers and maintains the growth of cancerous cells.
Santosh Philips, Heng-Yi Wu, Lang Li 0001
BMC Bioinform.3
2017 Uncovering direct and indirect molecular determinants of chromatin loops using a computational integrative approach
abstract
Chromosomal organization in 3D plays a central role in regulating cell-type specific transcriptional and DNA replication timing programs. Yet it remains unclear to what extent the resulting long-range contacts depend on specific molecular drivers. Here we propose a model that comprehensively assesses the influence on contacts of DNA-binding proteins, cis-regulatory elements and DNA consensus motifs. Using real data, we validate a large number of predictions for long-range contacts involving known architectural proteins and DNA motifs. Our model outperforms existing approaches including enrichment test, random forests and correlation, and it uncovers numerous novel long-range contacts in Drosophila and human. The model uncovers the orientation-dependent specificity for long-range contacts between CTCF motifs in Drosophila, highlighting its conserved property in 3D organization of metazoan genomes. Our model further unravels long-range contacts depending on co-factors recruited to DNA indirectly, as illustrated by the influence of cohesin in stabilizing long-range contacts between CTCF sites. It also reveals asymmetric contacts such as enhancer-promoter contacts that highlight opposite influences of the transcription factors EBF1, EGR1 or MEF2C depending on RNA Polymerase II pausing.
Raphaël Mourad, Lang Li 0001, Olivier Cuvier
PLoS Comput. Biol.2
2016 A bioinformatics approach for precision medicine off-label drug drug selection among triple negative breast cancer patients
abstract
BACKGROUND: Cancer has been extensively characterized on the basis of genomics. The integration of genetic information about cancers with data on how the cancers respond to target based therapy to help to optimum cancer treatment. OBJECTIVE: The increasing usage of sequencing technology in cancer research and clinical practice has enormously advanced our understanding of cancer mechanisms. The cancer precision medicine is becoming a reality. Although off-label drug usage is a common practice in treating cancer, it suffers from the lack of knowledge base for proper cancer drug selections. This eminent need has become even more apparent considering the upcoming genomics data. METHODS: In this paper, a personalized medicine knowledge base is constructed by integrating various cancer drugs, drug-target database, and knowledge sources for the proper cancer drugs and their target selections. Based on the knowledge base, a bioinformatics approach for cancer drugs selection in precision medicine is developed. It integrates personal molecular profile data, including copy number variation, mutation, and gene expression. RESULTS: By analyzing the 85 triple negative breast cancer (TNBC) patient data in the Cancer Genome Altar, we have shown that 71.7% of the TNBC patients have FDA approved drug targets, and 51.7% of the patients have more than one drug target. Sixty-five drug targets are identified as TNBC treatment targets and 85 candidate drugs are recommended. Many existing TNBC candidate targets, such as Poly (ADP-Ribose) Polymerase 1 (PARP1), Cell division protein kinase 6 (CDK6), epidermal growth factor receptor, etc., were identified. On the other hand, we found some additional targets that are not yet fully investigated in the TNBC, such as Gamma-Glutamyl Hydrolase (GGH), Thymidylate Synthetase (TYMS), Protein Tyrosine Kinase 6 (PTK6), Topoisomerase (DNA) I, Mitochondrial (TOP1MT), Smoothened, Frizzled Class Receptor (SMO), etc. Our additional analysis of target and drug selection strategy is also fully supported by the drug screening data on TNBC cell lines in the Cancer Cell Line Encyclopedia. CONCLUSIONS: The proposed bioinformatics approach lays a foundation for cancer precision medicine. It supplies much needed knowledge base for the off-label cancer drug usage in clinics.
Lijun Cheng, Bryan P. Schneider, Lang Li 0001
J. Am. Medical Informatics Assoc.3
2013 An integrated pharmacokinetics ontology and corpus for text mining
abstract
BACKGROUND: Drug pharmacokinetics parameters, drug interaction parameters, and pharmacogenetics data have been unevenly collected in different databases and published extensively in the literature. Without appropriate pharmacokinetics ontology and a well annotated pharmacokinetics corpus, it will be difficult to develop text mining tools for pharmacokinetics data collection from the literature and pharmacokinetics data integration from multiple databases. DESCRIPTION: A comprehensive pharmacokinetics ontology was constructed. It can annotate all aspects of in vitro pharmacokinetics experiments and in vivo pharmacokinetics studies. It covers all drug metabolism and transportation enzymes. Using our pharmacokinetics ontology, a PK-corpus was constructed to present four classes of pharmacokinetics abstracts: in vivo pharmacokinetics studies, in vivo pharmacogenetic studies, in vivo drug interaction studies, and in vitro drug interaction studies. A novel hierarchical three level annotation scheme was proposed and implemented to tag key terms, drug interaction sentences, and drug interaction pairs. The utility of the pharmacokinetics ontology was demonstrated by annotating three pharmacokinetics studies; and the utility of the PK-corpus was demonstrated by a drug interaction extraction text mining analysis. CONCLUSIONS: The pharmacokinetics ontology annotates both in vitro pharmacokinetics experiments and in vivo pharmacokinetics studies. The PK-corpus is a highly valuable resource for the text mining of pharmacokinetics parameters and drug interactions.
Heng-Yi Wu, Shreyas D. Karnik, Abhinita Subhadarshini, Santosh Philips, Chienwei Chiang, Malaz Boustani, Luis M. Rocha, Sara K. Quinney, David A. Flockhart, Lang Li 0001
BMC Bioinform.13
2012 Literature Based Drug Interaction Prediction with Clinical Assessment Using Electronic Medical Records: Novel Myopathy Associated Drug Interactions
abstract
Drug-drug interactions (DDIs) are a common cause of adverse drug events. In this paper, we combined a literature discovery approach with analysis of a large electronic medical record database method to predict and evaluate novel DDIs. We predicted an initial set of 13197 potential DDIs based on substrates and inhibitors of cytochrome P450 (CYP) metabolism enzymes identified from published in vitro pharmacology experiments. Using a clinical repository of over 800,000 patients, we narrowed this theoretical set of DDIs to 3670 drug pairs actually taken by patients. Finally, we sought to identify novel combinations that synergistically increased the risk of myopathy. Five pairs were identified with their p-values less than 1E-06: loratadine and simvastatin (relative risk or RR = 1.69); loratadine and alprazolam (RR = 1.86); loratadine and duloxetine (RR = 1.94); loratadine and ropinirole (RR = 3.21); and promethazine and tegaserod (RR = 3.00). When taken together, each drug pair showed a significantly increased risk of myopathy when compared to the expected additive myopathy risk from taking either of the drugs alone. Based on additional literature data on in vitro drug metabolism and inhibition potency, loratadine and simvastatin and tegaserod and promethazine were predicted to have a strong DDI through the CYP3A4 and CYP2D6 enzymes, respectively. This new translational biomedical informatics approach supports not only detection of new clinically significant DDI signals, but also evaluation of their potential molecular mechanisms.
Jon D. Duke, Abhinita Subhadarshini, Shreyas D. Karnik, Stephen D. Hall, J. Thomas Callaghan, J. Marc Overhage, David A. Flockhart, R. Matthew Strother, Sara K. Quinney, Lang Li 0001
PLoS Comput. Biol.14
2009 Enabling Data Analysis on High-Throughput Data in Large Data Depository Using Web-Based Analysis Platform - A Case Study on Integrating QUEST with GenePattern in Epigenetics Research
abstract
Enabling data analysis in large data depositories for high throughput experimental data such as gene microarrays and ChIP-seq is challenging. In this paper, we discuss three methods for integrating QUEST, a data depository for epigenetic experiments, with a web-based data analysis platform GenePattern. These methods are universal and can serve as an exemplary implementation resolving the dilemma facing many similar database systems in integrating data analysis tools.
Terry Camerlengo, Hatice Gulcin Ozer, Pearlly Yan, Jeffrey D. Parvin, Tim Hui-Ming Huang, Kun Huang 0001, Mingxiang Teng, Lang Li 0001, Francisco Perez, Tahsin M. Kurç
BIBM8
2009 Literature mining on pharmacokinetics numerical data: A feasibility study
Sara K. Quinney, Stephen D. Hall, Luis M. Rocha, Lang Li 0001
J. Biomed. Informatics7
2008 Enriched transcription factor binding sites in hypermethylated gene promoters in drug resistant cancer cells
abstract
MOTIVATION: In the human genome, 'CpG islands', CG-rich regions located in or near gene promoters, are normally unmethylated. However, in cancer cells, CpG islands frequently gain methylation, resulting in silencing of growth-limiting tumor suppressor genes. To our knowledge, the potential relationship between CpG island hypermethylation, transcription factor (TF) binding in local promoter regions and transcriptional control has not been previously explored in a genome-wide context. RESULTS: In this study, we utilized bioinformatics tools and TF binding site(TFBs) databases to globally analyze sequences methylated in a laboratory model for the development of drug-resistant cancer. Our results demonstrated that four TFBS were enriched in hypermethylated sequences. More interestingly, overrepresentation of these TFBS was observed in hyper-/hypo-methylated sequences where signi.cant changes in methylation levels were observed in drug-resistant cancer cells. In summary, we believe that these.ndings offer a means to further explore the relationship between DNA methylation and gene expression in drug resistance and tumorigenesis.
Meng Li 0013, Hyun-il Henry Paik, Curtis Balch, Yoosung Kim, Lang Li 0001, Tim Hui-Ming Huang, Kenneth P. Nephew, Sun Kim
Bioinform.5
2008 A hierarchical statistical model to assess the confidence of peptides and proteins inferred from tandem mass spectrometry
abstract
MOTIVATION: Statistical evaluation of the confidence of peptide and protein identifications made by tandem mass spectrometry is a critical component for appropriately interpreting the experimental data and conducting downstream analysis. Although many approaches have been developed to assign confidence measure from different perspectives, a unified statistical framework that integrates the uncertainty of peptides and proteins is still missing. RESULTS: We developed a hierarchical statistical model (HSM) that jointly models the uncertainty of the identified peptides and proteins and can be applied to any scoring system. With data sets of a standard mixture and the yeast proteome, we demonstrate that the HSM offers a reliable or at least conservative false discovery rate (FDR) estimate for peptide and protein identifications. The probability measure of HSM also offers a powerful discriminating score for peptide identification. AVAILABILITY: The algorithm is available upon request from the authors.
Changyu Shen, Ganesh Shankar, Xiang Zhang 0003, Lang Li 0001
Bioinform.5
2006 A mixture model-based discriminate analysis for identifying ordered transcription factor binding site pairs in gene promoters directly regulated by estrogen receptor-alpha
abstract
MOTIVATION: To detect and select patterns of transcription factor binding sites (TFBSs) which distinguish genes directly regulated by estrogen receptor-alpha (ERalpha), we developed an innovative mixture model-based discriminate analysis for identifying ordered TFBS pairs. RESULTS: Biologically, our proposed new algorithm clearly suggests that TFBSs are not randomly distributed within ERalpha target promoters (P-value < 0.001). The up-regulated targets significantly (P-value < 0.01) possess TFBS pairs, (DBP, MYC), (DBP, MYC/MAX heterodimer), (DBP, USF2) and (DBP, MYOGENIN); and down-regulated ERalpha target genes significantly (P-value < 0.01) possess TFBS pairs, such as (DBP, c-ETS1-68), (DBP, USF2) and (DBP, MYOGENIN). Statistically, our proposed mixture model-based discriminate analysis can simultaneously perform TFBS pattern recognition, TFBS pattern selection, and target class prediction; such integrative power cannot be achieved by current methods. AVAILABILITY: The software is available on request from the authors. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Lang Li 0001, Alfred S. L. Cheng, Victor X. Jin, Henry H. Paik, Meiyun Fan, Xiaoman Shawn Li, Jason Robarge, Curtis Balch, Ramana V. Davuluri, Sun Kim, Tim Hui-Ming Huang, Kenneth P. Nephew
Bioinform.1