VLDB 2026 Research / reviewers in the wild / expert
Feixiong Cheng
dblp:115/9180
· DBLP profile ↗
25ranked-venue papers
6as first author
7since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 21 · 6 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
6 papers |
Bioinformatics and computational biology · 72% Computational social science and digital humanities · 20% Medical and health informatics · 8% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Emerging computing paradigms · 100% | |
| Artificial intelligence
1 paper |
Probabilistic and Bayesian machine learning · 100% |
Topics — the 15 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › drug discovery
drug repurposing |
1.0 | 2 | 2025 | A Deep Subgrouping Framework for Precision Drug Repurposing via Emulating Clinical Trials on Real-world Patient Data · KDD (1) 2025 Network-based prediction of drug-target interactions using an arbitrary-order proximity embedded deep forest · Bioinform. 2020 |
Bioinformatics and computational biology
protein structure prediction |
0.9 | 1 | 2025 | QDockBank: A dataset for Ligand Docking on Protein Fragments Predicted on Utility-Level Quantum Computers · SC 2025 |
Computational social science and digital humanities › causal inference
treatment effect estimation |
0.9 | 1 | 2025 | A Deep Subgrouping Framework for Precision Drug Repurposing via Emulating Clinical Trials on Real-world Patient Data · KDD (1) 2025 |
Emerging computing paradigms
quantum computing |
0.9 | 1 | 2025 | QDockBank: A dataset for Ligand Docking on Protein Fragments Predicted on Utility-Level Quantum Computers · SC 2025 |
Emerging computing paradigms › quantum computing
quantum simulation |
0.9 | 1 | 2025 | QDockBank: A dataset for Ligand Docking on Protein Fragments Predicted on Utility-Level Quantum Computers · SC 2025 |
Bioinformatics and computational biology
cancer genomics |
0.5 | 1 | 2021 | A network-based deep learning methodology for stratification of tumor mutations · Bioinform. 2021 |
Bioinformatics and computational biology › drug discovery
drug-target interaction prediction |
0.4 | 1 | 2020 | Network-based prediction of drug-target interactions using an arbitrary-order proximity embedded deep forest · Bioinform. 2020 |
Bioinformatics and computational biology › drug discovery
drug repositioning |
0.4 | 1 | 2019 | deepDR: a network-based deep learning approach to in silico drug repositioning · Bioinform. 2019 |
Bioinformatics and computational biology
consensus clustering |
0.3 | 1 | 2017 | Entropy-based consensus clustering for patient stratification · Bioinform. 2017 |
Medical and health informatics › precision medicine
patient stratification |
0.3 | 1 | 2017 | Entropy-based consensus clustering for patient stratification · Bioinform. 2017 |
Machine learning › Probabilistic and Bayesian machine learning
subgroup analysis |
0.3 | 1 | 2025 | A Deep Subgrouping Framework for Precision Drug Repurposing via Emulating Clinical Trials on Real-world Patient Data · KDD (1) 2025 |
Bioinformatics and computational biology › cancer genomics
somatic mutation analysis |
0.1 | 1 | 2021 | A network-based deep learning methodology for stratification of tumor mutations · Bioinform. 2021 |
Bioinformatics and computational biology › drug discovery › drug repositioning
drug-disease association prediction |
0.1 | 1 | 2019 | deepDR: a network-based deep learning approach to in silico drug repositioning · Bioinform. 2019 |
Medical and health informatics › clinical data analysis › phenotyping
disease subtyping |
0.1 | 1 | 2017 | Entropy-based consensus clustering for patient stratification · Bioinform. 2017 |
Medical and health informatics
precision medicine |
0.1 | 1 | 2017 | Entropy-based consensus clustering for patient stratification · Bioinform. 2017 |
Methods — techniques the papers use, named apart from their topics
treatment effect estimation · 1.7subgroup analysis · 1.7quantum computing · 1.7clinical trial emulation · 1.7network embedding · 0.9unsupervised clustering · 0.5LightGBM · 0.5deep forest · 0.4arbitrary-order proximity · 0.4multimodal deep autoencoder · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Deep Subgrouping Framework for Precision Drug Repurposing via Emulating Clinical Trials on Real-world Patient Dataabstractdrug discovery. Most existing drug repurposing studies using real-world patient data often treat the entire population as homogeneous, ignoring the heterogeneity of treatment responses across patient subgroups. This approach may overlook promising drugs that benefit specific subgroups but lack notable treatment effects across the entire population, potentially limiting the number of repurposable candidates identified. To address this, we introduce STEDR, a novel drug repurposing framework that integrates subgroup analysis with treatment effect estimation. Our approach first identifies repurposing candidates by emulating multiple clinical trials on real-world patient data and then characterizes patient subgroups by learning subgroup-specific treatment effects. We deploy STEDR to Alzheimer's Disease (AD), a condition with few approved drugs and known heterogeneity in treatment responses. We emulate trials for over one thousand medications on a large-scale real-world database covering over 8 million patients, identifying 14 drug candidates with beneficial effects to AD in characterized subgroups. Experiments demonstrate STEDR's superior capability in identifying repurposing candidates compared to existing approaches. Additionally, our method can characterize clinically relevant patient subgroups associated with important AD-related risk factors, paving the way for precision drug repurposing. Seungyeon Lee 0002, Ruoqi Liu, Feixiong Cheng, Ping Zhang 0016 |
KDD (1) | 3 |
| 2025 | QDockBank: A dataset for Ligand Docking on Protein Fragments Predicted on Utility-Level Quantum ComputersabstractProtein structure prediction is a core challenge in computational biology, particularly for fragments within ligand-binding regions, where accurate modeling is still difficult. Quantum computing offers a novel first-principles modeling paradigm, but its application is currently limited by hardware constraints, high computational cost, and the lack of a standardized benchmarking dataset. In this work, we present QDockBank—the first large-scale protein fragment structure dataset generated entirely using utility-level quantum computers, specifically designed for protein–ligand docking tasks. QDockBank comprises 55 protein fragments extracted from ligand-binding pockets. The dataset was generated through tens of hours of execution on superconducting quantum processors, making it the first quantum-based protein structure dataset with a total computational cost exceeding one million USD. Experimental evaluations demonstrate that structures predicted by QDockBank outperform those predicted by AlphaFold2 and AlphaFold3 in terms of both RMSD and docking affinity scores. QDockBank serves as a new benchmark for evaluating quantum-based protein structure prediction. Yuxin Yang 0001, Cheng-Chang Lu, Weiwen Jiang, Feixiong Cheng, Bo Fang 0002, Qiang Guan |
SC | 5 |
| 2024 | A Deep Multimodal Representation Learning Framework for Accurate Molecular Properties PredictionabstractDrug discovery is a challenging process, requiring the optimization of compounds to become safe and effective. Predicting molecular properties is an indispensable step in the drug discovery pipeline. Traditionally, this process is costly, involving multiple rounds of experiments, rendering it impractical for every candidate compound. Deep learning techniques have emerged as a promising approach to drug discovery to reduce the cost during the process. However, prevalent research in deep learning models focused on predicting molecular properties has primarily fixated on single-modal models, neglecting the potential benefits of combining different data modalities. To overcome this limitation, we introduce MRL-Mol: a deep Multimodal Representation Learning framework for accurate Molecular properties prediction. MRL-Mol harnesses three data modalities: sequence, graph, and image, augmenting the depth of comprehension. Leveraging a large-scale unlabeled dataset ( 1M unique molecules), we pretrain MRL-Mol to extract inter- and intra-modal information. Our study demonstrates the superior performance of MRL-Mol in predicting molecular properties across six benchmark datasets. Notably, MRL-Mol outperforms other state-of-the-art molecular properties prediction models. These findings suggest that by combining information from multiple data modalities, MRL-Mol can comprehend molecules better than single-modal deep learning models and identify molecular properties with better accuracy. Yuxin Yang 0001, Pegah Ahadian, Abby Jerger, Jeremy Zucker, Feixiong Cheng, Qiang Guan |
ACM Great Lakes Symposium on VLSI | 7 |
| 2022 | Comprehensively modeling heterogeneous symptom progression for Parkinson's disease subtyping
Chang Su 0002, Jielin Xu, Matthew Brendel, Jacqueline R. M. A. Maasch, Zilong Bai, Yingying Zhu 0003, Claire Henchcliffe, Feixiong Cheng, Fei Wang 0001 |
AMIA | 9 |
| 2022 | Single-cell network biology characterizes cell type gene regulation for drug repurposing and phenotype prediction in Alzheimer's diseaseabstractDysregulation of gene expression in Alzheimer's disease (AD) remains elusive, especially at the cell type level. Gene regulatory network, a key molecular mechanism linking transcription factors (TFs) and regulatory elements to govern gene expression, can change across cell types in the human brain and thus serve as a model for studying gene dysregulation in AD. However, AD-induced regulatory changes across brain cell types remains uncharted. To address this, we integrated single-cell multi-omics datasets to predict the gene regulatory networks of four major cell types, excitatory and inhibitory neurons, microglia and oligodendrocytes, in control and AD brains. Importantly, we analyzed and compared the structural and topological features of networks across cell types and examined changes in AD. Our analysis shows that hub TFs are largely common across cell types and AD-related changes are relatively more prominent in some cell types (e.g., microglia). The regulatory logics of enriched network motifs (e.g., feed-forward loops) further uncover cell type-specific TF-TF cooperativities in gene regulation. The cell type networks are also highly modular and several network modules with cell-type-specific expression changes in AD pathology are enriched with AD-risk genes. The further disease-module-drug association analysis suggests cell-type candidate drugs and their potential target genes. Finally, our network-based machine learning analysis systematically prioritized cell type risk genes likely involved in AD. Our strategy is validated using an independent dataset which showed that top ranked genes can predict clinical phenotypes (e.g., cognitive impairment) of AD with reasonable accuracy. Overall, this single-cell network biology analysis provides a comprehensive map linking genes, regulatory networks, cell types and drug targets and reveals cell-type gene dysregulation in AD. Jielin Xu, Saniya Khullar, Sayali Alatkar, Feixiong Cheng, Daifeng Wang |
PLoS Comput. Biol. | 7 |
| 2021 | Artificial Intelligence for Drug DiscoveryabstractDrug discovery is a long and costly process, taking on average 10 years and 2.5 billion dollars to develop a new drug. Artificial intelligence has the potential to significantly accelerate the process of drug discovery by analyzing a large amount of data generated in the biomedical domain such as bioassays, chemical experiments, and biomedical literature. Recently, there is a growing interesting in developing AI techniques for drug discovery in many different communities including machine learning, data mining, and biomedical community. In this tutorial, we will provide a detailed introduction to key problems in drug discovery such as molecular property prediction, de novo molecular design and molecular optimization, retrosynthesis reaction and prediction, and drug repurposing and combination, and also key technique advancements with artificial intelligence for these problems. This tutorial can be served as introduction materials for both computer scientist interested in drug discovery as well as drug discovery practitioners for learning the latest AI techniques along this direction. Jian Tang 0005, Fei Wang 0001, Feixiong Cheng |
KDD | 3 |
| 2021 | A network-based deep learning methodology for stratification of tumor mutationsabstractMOTIVATION: Tumor stratification has a wide range of biomedical and clinical applications, including diagnosis, prognosis and personalized treatment. However, cancer is always driven by the combination of mutated genes, which are highly heterogeneous across patients. Accurately subdividing the tumors into subtypes is challenging. RESULTS: We developed a network-embedding based stratification (NES) methodology to identify clinically relevant patient subtypes from large-scale patients' somatic mutation profiles. The central hypothesis of NES is that two tumors would be classified into the same subtypes if their somatic mutated genes located in the similar network regions of the human interactome. We encoded the genes on the human protein-protein interactome with a network embedding approach and constructed the patients' vectors by integrating the somatic mutation profiles of 7344 tumor exomes across 15 cancer types. We firstly adopted the lightGBM classification algorithm to train the patients' vectors. The AUC value is around 0.89 in the prediction of the patient's cancer type and around 0.78 in the prediction of the tumor stage within a specific cancer type. The high classification accuracy suggests that network embedding-based patients' features are reliable for dividing the patients. We conclude that we can cluster patients with a specific cancer type into several subtypes by using an unsupervised clustering algorithm to learn the patients' vectors. Among the 15 cancer types, the new patient clusters (subtypes) identified by the NES are significantly correlated with patient survival across 12 cancer types. In summary, this study offers a powerful network-based deep learning methodology for personalized cancer medicine. AVAILABILITY AND IMPLEMENTATION: Source code and data can be downloaded from https://github.com/ChengF-Lab/NES. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Chuang Liu 0001, Zi-Ke Zhang, Ruth Nussinov, Feixiong Cheng |
Bioinform. | 5 |
| 2020 | Network-based prediction of drug-target interactions using an arbitrary-order proximity embedded deep forestabstractMOTIVATION: Systematic identification of molecular targets among known drugs plays an essential role in drug repurposing and understanding of their unexpected side effects. Computational approaches for prediction of drug-target interactions (DTIs) are highly desired in comparison to traditional experimental assays. Furthermore, recent advances of multiomics technologies and systems biology approaches have generated large-scale heterogeneous, biological networks, which offer unexpected opportunities for network-based identification of new molecular targets among known drugs. RESULTS: In this study, we present a network-based computational framework, termed AOPEDF, an arbitrary-order proximity embedded deep forest approach, for prediction of DTIs. AOPEDF learns a low-dimensional vector representation of features that preserve arbitrary-order proximity from a highly integrated, heterogeneous biological network connecting drugs, targets (proteins) and diseases. In total, we construct a heterogeneous network by uniquely integrating 15 networks covering chemical, genomic, phenotypic and network profiles among drugs, proteins/targets and diseases. Then, we build a cascade deep forest classifier to infer new DTIs. Via systematic performance evaluation, AOPEDF achieves high accuracy in identifying molecular targets among known drugs on two external validation sets collected from DrugCentral [area under the receiver operating characteristic curve (AUROC) = 0.868] and ChEMBL (AUROC = 0.768) databases, outperforming several state-of-the-art methods. In a case study, we showcase that multiple molecular targets predicted by AOPEDF are associated with mechanism-of-action of substance abuse disorder for several marketed drugs (such as aripiprazole, risperidone and haloperidol). AVAILABILITY AND IMPLEMENTATION: Source code and data can be downloaded from https://github.com/ChengF-Lab/AOPEDF. Xiangxiang Zeng, Siyi Zhu, Yuan Hou, Pengyue Zhang, Lang Li 0001, L. Frank Huang, Stephen J. Lewis, Ruth Nussinov, Feixiong Cheng |
Bioinform. | 10 |
| 2020 | iDrug: Integration of drug repositioning and drug-target prediction via cross-network embeddingabstractComputational drug repositioning and drug-target prediction have become essential tasks in the early stage of drug discovery. In previous studies, these two tasks have often been considered separately. However, the entities studied in these two tasks (i.e., drugs, targets, and diseases) are inherently related. On one hand, drugs interact with targets in cells to modulate target activities, which in turn alter biological pathways to promote healthy functions and to treat diseases. On the other hand, both drug repositioning and drug-target prediction involve the same drug feature space, which naturally connects these two problems and the two domains (diseases and targets). By using the wisdom of the crowds, it is possible to transfer knowledge from one of the domains to the other. The existence of relationships among drug-target-disease motivates us to jointly consider drug repositioning and drug-target prediction in drug discovery. In this paper, we present a novel approach called iDrug, which seamlessly integrates drug repositioning and drug-target prediction into one coherent model via cross-network embedding. In particular, we provide a principled way to transfer knowledge from these two domains and to enhance prediction performance for both tasks. Using real-world datasets, we demonstrate that iDrug achieves superior performance on both learning tasks compared to several state-of-the-art approaches. Our code and datasets are available at: https://github.com/Case-esaC/iDrug. Huiyuan Chen, Feixiong Cheng, Jing Li 0002 |
PLoS Comput. Biol. | 2 |
| 2020 | Individualized genetic network analysis reveals new therapeutic vulnerabilities in 6, 700 cancer genomesabstractTumor-specific genomic alterations allow systematic identification of genetic interactions that promote tumorigenesis and tumor vulnerabilities, offering novel strategies for development of targeted therapies for individual patients. We develop an Individualized Network-based Co-Mutation (INCM) methodology by inspecting over 2.5 million nonsynonymous somatic mutations derived from 6,789 tumor exomes across 14 cancer types from The Cancer Genome Atlas. Our INCM analysis reveals a higher genetic interaction burden on the significantly mutated genes, experimentally validated cancer genes, chromosome regulatory factors, and DNA damage repair genes, as compared to human pan-cancer essential genes identified by CRISPR-Cas9 screenings on 324 cancer cell lines. We find that genes involved in the cancer type-specific genetic subnetworks identified by INCM are significantly enriched in established cancer pathways, and the INCM-inferred putative genetic interactions are correlated with patient survival. By analyzing drug pharmacogenomics profiles from the Genomics of Drug Sensitivity in Cancer database, we show that the network-predicted putative genetic interactions (e.g., BRCA2-TP53) are significantly correlated with sensitivity/resistance of multiple therapeutic agents. We experimentally validated that afatinib has the strongest cytotoxic activity on BT474 (IC50 = 55.5 nM, BRCA2 and TP53 co-mutant) compared to MCF7 (IC50 = 7.7 μM, both BRCA2 and TP53 wild type) and MDA-MB-231 (IC50 = 7.9 μM, BRCA2 wild type but TP53 mutant). Finally, drug-target network analysis reveals several potential druggable genetic interactions by targeting tumor vulnerabilities. This study offers a powerful network-based methodology for identification of candidate therapeutic pathways that target tumor vulnerabilities and prioritization of potential pharmacogenomics biomarkers for development of personalized cancer medicine. Chuang Liu 0001, Junfei Zhao, Weiqiang Lu, Yao Dai, Jennifer Hockings, Yadi Zhou, Ruth Nussinov, Charis Eng, Feixiong Cheng |
PLoS Comput. Biol. | 9 |
| 2019 | deepDR: a network-based deep learning approach to in silico drug repositioningabstractMOTIVATION: Traditional drug discovery and development are often time-consuming and high risk. Repurposing/repositioning of approved drugs offers a relatively low-cost and high-efficiency approach toward rapid development of efficacious treatments. The emergence of large-scale, heterogeneous biological networks has offered unprecedented opportunities for developing in silico drug repositioning approaches. However, capturing highly non-linear, heterogeneous network structures by most existing approaches for drug repositioning has been challenging. RESULTS: In this study, we developed a network-based deep-learning approach, termed deepDR, for in silico drug repurposing by integrating 10 networks: one drug-disease, one drug-side-effect, one drug-target and seven drug-drug networks. Specifically, deepDR learns high-level features of drugs from the heterogeneous networks by a multi-modal deep autoencoder. Then the learned low-dimensional representation of drugs together with clinically reported drug-disease pairs are encoded and decoded collectively via a variational autoencoder to infer candidates for approved drugs for which they were not originally approved. We found that deepDR revealed high performance [the area under receiver operating characteristic curve (AUROC) = 0.908], outperforming conventional network-based or machine learning-based approaches. Importantly, deepDR-predicted drug-disease associations were validated by the ClinicalTrials.gov database (AUROC = 0.826) and we showcased several novel deepDR-predicted approved drugs for Alzheimer's disease (e.g. risperidone and aripiprazole) and Parkinson's disease (e.g. methylphenidate and pergolide). AVAILABILITY AND IMPLEMENTATION: Source code and data can be downloaded from https://github.com/ChengF-Lab/deepDR. SUPPLEMENTARY INFORMATION: Supplementary data are available online at Bioinformatics. Xiangxiang Zeng, Siyi Zhu, Xiangrong Liu, Yadi Zhou, Ruth Nussinov, Feixiong Cheng |
Bioinform. | 6 |
| 2019 | Review: Precision medicine and driver mutations: Computational methods, functional assays and conformational principles for interpreting cancer driversabstractAt the root of the so-called precision medicine or precision oncology, which is our focus here, is the hypothesis that cancer treatment would be considerably better if therapies were guided by a tumor's genomic alterations. This hypothesis has sparked major initiatives focusing on whole-genome and/or exome sequencing, creation of large databases, and developing tools for their statistical analyses-all aspiring to identify actionable alterations, and thus molecular targets, in a patient. At the center of the massive amount of collected sequence data is their interpretations that largely rest on statistical analysis and phenotypic observations. Statistics is vital, because it guides identification of cancer-driving alterations. However, statistics of mutations do not identify a change in protein conformation; therefore, it may not define sufficiently accurate actionable mutations, neglecting those that are rare. Among the many thematic overviews of precision oncology, this review innovates by further comprehensively including precision pharmacology, and within this framework, articulating its protein structural landscape and consequences to cellular signaling pathways. It provides the underlying physicochemical basis, thereby also opening the door to a broader community. Ruth Nussinov, Hyunbum Jang, Chung-Jung Tsai, Feixiong Cheng |
PLoS Comput. Biol. | 4 |
| 2019 | Correction: Review: Precision medicine and driver mutations: Computational methods, functional assays and conformational principles for interpreting cancer driversabstract[This corrects the article DOI: 10.1371/journal.pcbi.1006658.]. Ruth Nussinov, Hyunbum Jang, Chung-Jung Tsai, Feixiong Cheng |
PLoS Comput. Biol. | 4 |
| 2019 | A component overlapping attribute clustering (COAC) algorithm for single-cell RNA sequencing data analysis and potential pathobiological implicationsabstractRecent advances in next-generation sequencing and computational technologies have enabled routine analysis of large-scale single-cell ribonucleic acid sequencing (scRNA-seq) data. However, scRNA-seq technologies have suffered from several technical challenges, including low mean expression levels in most genes and higher frequencies of missing data than bulk population sequencing technologies. Identifying functional gene sets and their regulatory networks that link specific cell types to human diseases and therapeutics from scRNA-seq profiles are daunting tasks. In this study, we developed a Component Overlapping Attribute Clustering (COAC) algorithm to perform the localized (cell subpopulation) gene co-expression network analysis from large-scale scRNA-seq profiles. Gene subnetworks that represent specific gene co-expression patterns are inferred from the components of a decomposed matrix of scRNA-seq profiles. We showed that single-cell gene subnetworks identified by COAC from multiple time points within cell phases can be used for cell type identification with high accuracy (83%). In addition, COAC-inferred subnetworks from melanoma patients' scRNA-seq profiles are highly correlated with survival rate from The Cancer Genome Atlas (TCGA). Moreover, the localized gene subnetworks identified by COAC from individual patients' scRNA-seq data can be used as pharmacogenomics biomarkers to predict drug responses (The area under the receiver operating characteristic curves ranges from 0.728 to 0.783) in cancer cell lines from the Genomics of Drug Sensitivity in Cancer (GDSC) database. In summary, COAC offers a powerful tool to identify potential network-based diagnostic and pharmacogenomics biomarkers from large-scale scRNA-seq profiles. COAC is freely available at https://github.com/ChengF-Lab/COAC. He Peng, Xiangxiang Zeng, Yadi Zhou, Ruth Nussinov, Feixiong Cheng |
PLoS Comput. Biol. | 6 |
| 2018 | In silico polypharmacology of natural productsabstractNatural products with polypharmacological profiles have demonstrated promise as novel therapeutics for various complex diseases, including cancer. Currently, many gaps exist in our knowledge of which compounds interact with which targets, and experimentally testing all possible interactions is infeasible. Recent advances and developments of systems pharmacology and computational (in silico) approaches provide powerful tools for exploring the polypharmacological profiles of natural products. In this review, we introduce recent progresses and advances of computational tools and systems pharmacology approaches for identifying drug targets of natural products by focusing on the development of targeted cancer therapy. We survey the polypharmacological and systems immunology profiles of five representative natural products that are being considered as cancer therapies. We summarize various chemoinformatics, bioinformatics and systems biology resources for reconstructing drug-target networks of natural products. We then review currently available computational approaches and tools for prediction of drug-target interactions by focusing on five domains: target-based, ligand-based, chemogenomics-based, network-based and omics-based systems biology approaches. In addition, we describe a practical example of the application of systems pharmacology approaches by integrating the polypharmacology of natural products and large-scale cancer genomics data for the development of precision oncology under the systems biology framework. Finally, we highlight the promise of cancer immunotherapies and combination therapies that target tumor ecosystems (e.g. clones or 'selfish' sub-clones) via exploiting the immunological and inflammatory 'side' effects of natural products in the cancer post-genomics era. Jiansong Fang, Qi Wang 0170, Feixiong Cheng |
Briefings Bioinform. | 5 |
| 2017 | Individualized network-based drug repositioning infrastructure for precision oncology in the panomics eraabstractAdvances in next-generation sequencing technologies have generated the data supporting a large volume of somatic alterations in several national and international cancer genome projects, such as The Cancer Genome Atlas and the International Cancer Genome Consortium. These cancer genomics data have facilitated the revolution of a novel oncology drug discovery paradigm from candidate target or gene studies toward targeting clinically relevant driver mutations or molecular features for precision cancer therapy. This focuses on identifying the most appropriately targeted therapy to an individual patient harboring a particularly genetic profile or molecular feature. However, traditional experimental approaches that are used to develop new chemical entities for targeting the clinically relevant driver mutations are costly and high-risk. Drug repositioning, also known as drug repurposing, re-tasking or re-profiling, has been demonstrated as a promising strategy for drug discovery and development. Recently, computational techniques and methods have been proposed for oncology drug repositioning and identifying pharmacogenomics biomarkers, but overall progress remains to be seen. In this review, we focus on introducing new developments and advances of the individualized network-based drug repositioning approaches by targeting the clinically relevant driver events or molecular features derived from cancer panomics data for the development of precision oncology drug therapies (e.g. one-person trials) to fully realize the promise of precision medicine. We discuss several potential challenges (e.g. tumor heterogeneity and cancer subclones) for precision oncology. Finally, we highlight several new directions for the precision oncology drug discovery via biotherapies (e.g. gene therapy and immunotherapy) that target the 'undruggable' cancer genome in the functional genomics era. Feixiong Cheng, Huixiao Hong, Sheng-Yong Yang, Yuquan Wei |
Briefings Bioinform. | 1 |
| 2017 | SDTNBI: an integrated network and chemoinformatics tool for systematic prediction of drug-target interactions and drug repositioningabstractComputational prediction of drug-target interactions (DTIs) and drug repositioning provides a low-cost and high-efficiency approach for drug discovery and development. The traditional social network-derived methods based on the naïve DTI topology information cannot predict potential targets for new chemical entities or failed drugs in clinical trials. There are currently millions of commercially available molecules with biologically relevant representations in chemical databases. It is urgent to develop novel computational approaches to predict targets for new chemical entities and failed drugs on a large scale. In this study, we developed a useful tool, namely substructure-drug-target network-based inference (SDTNBI), to prioritize potential targets for old drugs, failed drugs and new chemical entities. SDTNBI incorporates network and chemoinformatics to bridge the gap between new chemical entities and known DTI network. High performance was yielded in 10-fold and leave-one-out cross validations using four benchmark data sets, covering G protein-coupled receptors, kinases, ion channels and nuclear receptors. Furthermore, the highest areas under the receiver operating characteristic curve were 0.797 and 0.863 for two external validation sets, respectively. Finally, we identified thousands of new potential DTIs via implementing SDTNBI on a global network. As a proof-of-principle, we showcased the use of SDTNBI to identify novel anticancer indications for nonsteroidal anti-inflammatory drugs by inhibiting AKR1C3, CA9 or CA12. In summary, SDTNBI is a powerful network-based approach that predicts potential targets for new chemical entities on a large scale and will provide a new tool for DTI prediction and drug repositioning. The program and predicted DTIs are available on request. Zengrui Wu, Feixiong Cheng, Weihua Li 0005, Guixia Liu, Yun Tang 0001 |
Briefings Bioinform. | 2 |
| 2017 | Entropy-based consensus clustering for patient stratificationabstractMOTIVATION: Patient stratification or disease subtyping is crucial for precision medicine and personalized treatment of complex diseases. The increasing availability of high-throughput molecular data provides a great opportunity for patient stratification. Many clustering methods have been employed to tackle this problem in a purely data-driven manner. Yet, existing methods leveraging high-throughput molecular data often suffers from various limitations, e.g. noise, data heterogeneity, high dimensionality or poor interpretability. RESULTS: Here we introduced an Entropy-based Consensus Clustering (ECC) method that overcomes those limitations all together. Our ECC method employs an entropy-based utility function to fuse many basic partitions to a consensus one that agrees with the basic ones as much as possible. Maximizing the utility function in ECC has a much more meaningful interpretation than any other consensus clustering methods. Moreover, we exactly map the complex utility maximization problem to the classic K -means clustering problem, which can then be efficiently solved with linear time and space complexity. Our ECC method can also naturally integrate multiple molecular data types measured from the same set of subjects, and easily handle missing values without any imputation. We applied ECC to 110 synthetic and 48 real datasets, including 35 cancer gene expression benchmark datasets and 13 cancer types with four molecular data types from The Cancer Genome Atlas. We found that ECC shows superior performance against existing clustering methods. Our results clearly demonstrate the power of ECC in clinically relevant patient stratification. AVAILABILITY AND IMPLEMENTATION: The Matlab package is available at http://scholar.harvard.edu/yyl/ecc . CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hongfu Liu 0001, Hongsheng Fang, Feixiong Cheng, Yun Fu 0001, Yang-Yu Liu |
Bioinform. | 4 |
| 2016 | Advances in computational approaches for prioritizing driver mutations and significantly mutated genes in cancer genomesabstractCancer is often driven by the accumulation of genetic alterations, including single nucleotide variants, small insertions or deletions, gene fusions, copy-number variations, and large chromosomal rearrangements. Recent advances in next-generation sequencing technologies have helped investigators generate massive amounts of cancer genomic data and catalog somatic mutations in both common and rare cancer types. So far, the somatic mutation landscapes and signatures of >10 major cancer types have been reported; however, pinpointing driver mutations and cancer genes from millions of available cancer somatic mutations remains a monumental challenge. To tackle this important task, many methods and computational tools have been developed during the past several years and, thus, a review of its advances is urgently needed. Here, we first summarize the main features of these methods and tools for whole-exome, whole-genome and whole-transcriptome sequencing data. Then, we discuss major challenges like tumor intra-heterogeneity, tumor sample saturation and functionality of synonymous mutations in cancer, all of which may result in false-positive discoveries. Finally, we highlight new directions in studying regulatory roles of noncoding somatic mutations and quantitatively measuring circulating tumor DNA in cancer. This review may help investigators find an appropriate tool for detecting potential driver or actionable mutations in rapidly emerging precision cancer medicine. Feixiong Cheng, Junfei Zhao, Zhongming Zhao |
Briefings Bioinform. | 1 |
| 2016 | Systematic dissection of dysregulated transcription factor-miRNA feed-forward loops across tumor typesabstractTranscription factor and microRNA (miRNA) can mutually regulate each other and jointly regulate their shared target genes to form feed-forward loops (FFLs). While there are many studies of dysregulated FFLs in a specific cancer, a systematic investigation of dysregulated FFLs across multiple tumor types (pan-cancer FFLs) has not been performed yet. In this study, using The Cancer Genome Atlas data, we identified 26 pan-cancer FFLs, which were dysregulated in at least five tumor types. These pan-cancer FFLs could communicate with each other and form functionally consistent subnetworks, such as epithelial to mesenchymal transition-related subnetwork. Many proteins and miRNAs in each subnetwork belong to the same protein and miRNA family, respectively. Importantly, cancer-associated genes and drug targets were enriched in these pan-cancer FFLs, in which the genes and miRNAs also tended to be hubs and bottlenecks. Finally, we identified potential anticancer indications for existing drugs with novel mechanism of action. Collectively, this study highlights the potential of pan-cancer FFLs as a novel paradigm in elucidating pathogenesis of cancer and developing anticancer drugs. Wei Jiang 0023, Ramkrishna Mitra, Chen-Ching Lin, Quan Wang 0004, Feixiong Cheng, Zhongming Zhao |
Briefings Bioinform. | 5 |
| 2016 | A network-based drug repositioning infrastructure for precision cancer medicine through targeting significantly mutated genes in the human cancer genomesabstractOBJECTIVE: Development of computational approaches and tools to effectively integrate multidomain data is urgently needed for the development of newly targeted cancer therapeutics. METHODS: We proposed an integrative network-based infrastructure to identify new druggable targets and anticancer indications for existing drugs through targeting significantly mutated genes (SMGs) discovered in the human cancer genomes. The underlying assumption is that a drug would have a high potential for anticancer indication if its up-/down-regulated genes from the Connectivity Map tended to be SMGs or their neighbors in the human protein interaction network. RESULTS: We assembled and curated 693 SMGs in 29 cancer types and found 121 proteins currently targeted by known anticancer or noncancer (repurposed) drugs. We found that the approved or experimental cancer drugs could potentially target these SMGs in 33.3% of the mutated cancer samples, and this number increased to 68.0% by drug repositioning through surveying exome-sequencing data in approximately 5000 normal-tumor pairs from The Cancer Genome Atlas. Furthermore, we identified 284 potential new indications connecting 28 cancer types and 48 existing drugs (adjusted P < .05), with a 66.7% success rate validated by literature data. Several existing drugs (e.g., niclosamide, valproic acid, captopril, and resveratrol) were predicted to have potential indications for multiple cancer types. Finally, we used integrative analysis to showcase a potential mechanism-of-action for resveratrol in breast and lung cancer treatment whereby it targets several SMGs (ARNTL, ASPM, CTTN, EIF4G1, FOXP1, and STIP1). CONCLUSIONS: In summary, we demonstrated that our integrative network-based infrastructure is a promising strategy to identify potential druggable targets and uncover new indications for existing drugs to speed up molecularly targeted cancer therapeutics. Feixiong Cheng, Junfei Zhao, Michaela Fooksa, Zhongming Zhao |
J. Am. Medical Informatics Assoc. | 1 |
| 2016 | Systems Biology-Based Investigation of Cellular Antiviral Drug Targets Identified by Gene-Trap Insertional MutagenesisabstractViruses require host cellular factors for successful replication. A comprehensive systems-level investigation of the virus-host interactome is critical for understanding the roles of host factors with the end goal of discovering new druggable antiviral targets. Gene-trap insertional mutagenesis is a high-throughput forward genetics approach to randomly disrupt (trap) host genes and discover host genes that are essential for viral replication, but not for host cell survival. In this study, we used libraries of randomly mutagenized cells to discover cellular genes that are essential for the replication of 10 distinct cytotoxic mammalian viruses, 1 gram-negative bacterium, and 5 toxins. We herein reported 712 candidate cellular genes, characterizing distinct topological network and evolutionary signatures, and occupying central hubs in the human interactome. Cell cycle phase-specific network analysis showed that host cell cycle programs played critical roles during viral replication (e.g. MYC and TAF4 regulating G0/1 phase). Moreover, the viral perturbation of host cellular networks reflected disease etiology in that host genes (e.g. CTCF, RHOA, and CDKN1B) identified were frequently essential and significantly associated with Mendelian and orphan diseases, or somatic mutations in cancer. Computational drug repositioning framework via incorporating drug-gene signatures from the Connectivity Map into the virus-host interactome identified 110 putative druggable antiviral targets and prioritized several existing drugs (e.g. ajmaline) that may be potential for antiviral indication (e.g. anti-Ebola). In summary, this work provides a powerful methodology with a tight integration of gene-trap insertional mutagenesis testing and systems biology to identify new antiviral targets and drugs for the development of broadly acting and targeted clinical antiviral therapeutics. Feixiong Cheng, James L. Murray, Junfei Zhao, Jinsong Sheng, Zhongming Zhao, Donald H. Rubin |
PLoS Comput. Biol. | 1 |
| 2015 | SGDriver: a novel structural genomics-based approach to prioritize cancer related and potentially druggable somatic mutationsabstractBackground A huge volume of somatic mutations have been generated through large cancer genome sequencing projects such as The Cancer Genome Atlas (TCGA) and the International Cancer Genome Consortium (ICGC). However, understanding the functional consequences of somatic mutations in cancer and translating the results into clinical use remains a major challenge in cancer genomic studies. Thanks to the rapid development of structural genomic technologies, such as X-ray and NMR, large amounts of protein structure data have been generated during the past decade, which enables us to map somatic mutations to protein functional features (i.e., protein-ligand binding sites) and investigate their potential impacts[1,2]. Junfei Zhao, Feixiong Cheng, Zhongming Zhao |
BMC Bioinform. | 2 |
| 2015 | A Gene Gravity Model for the Evolution of Cancer Genomes: A Study of 3, 000 Cancer Genomes across 9 Cancer TypesabstractCancer development and progression result from somatic evolution by an accumulation of genomic alterations. The effects of those alterations on the fitness of somatic cells lead to evolutionary adaptations such as increased cell proliferation, angiogenesis, and altered anticancer drug responses. However, there are few general mathematical models to quantitatively examine how perturbations of a single gene shape subsequent evolution of the cancer genome. In this study, we proposed the gene gravity model to study the evolution of cancer genomes by incorporating the genome-wide transcription and somatic mutation profiles of ~3,000 tumors across 9 cancer types from The Cancer Genome Atlas into a broad gene network. We found that somatic mutations of a cancer driver gene may drive cancer genome evolution by inducing mutations in other genes. This functional consequence is often generated by the combined effect of genetic and epigenetic (e.g., chromatin regulation) alterations. By quantifying cancer genome evolution using the gene gravity model, we identified six putative cancer genes (AHNAK, COL11A1, DDX3X, FAT4, STAG2, and SYNE1). The tumor genomes harboring the nonsynonymous somatic mutations in these genes had a higher mutation density at the genome level compared to the wild-type groups. Furthermore, we provided statistical evidence that hypermutation of cancer driver genes on inactive X chromosomes is a general feature in female cancer genomes. In summary, this study sheds light on the functional consequences and evolutionary characteristics of somatic mutations during tumorigenesis by propelling adaptive cancer genome evolution, which would provide new perspectives for cancer research and therapeutics. Feixiong Cheng, Chen-Ching Lin, Junfei Zhao, Peilin Jia, Wen-Hsiung Li, Zhongming Zhao |
PLoS Comput. Biol. | 1 |
| 2012 | Prediction of Drug-Target Interactions and Drug Repositioning via Network-Based InferenceabstractDrug-target interaction (DTI) is the basis of drug discovery and design. It is time consuming and costly to determine DTI experimentally. Hence, it is necessary to develop computational methods for the prediction of potential DTI. Based on complex network theory, three supervised inference methods were developed here to predict DTI and used for drug repositioning, namely drug-based similarity inference (DBSI), target-based similarity inference (TBSI) and network-based inference (NBI). Among them, NBI performed best on four benchmark data sets. Then a drug-target network was created with NBI based on 12,483 FDA-approved and experimental drug-target binary links, and some new DTIs were further predicted. In vitro assays confirmed that five old drugs, namely montelukast, diclofenac, simvastatin, ketoconazole, and itraconazole, showed polypharmacological features on estrogen receptors or dipeptidyl peptidase-IV with half maximal inhibitory or effective concentration ranged from 0.2 to 10 µM. Moreover, simvastatin and ketoconazole showed potent antiproliferative activities on human MDA-MB-231 breast cancer cell line in MTT assays. The results indicated that these methods could be powerful tools in prediction of DTIs and drug repositioning. Feixiong Cheng, Weiqiang Lu, Weihua Li 0005, Guixia Liu, Wei-Xing Zhou, Yun Tang 0001 |
PLoS Comput. Biol. | 1 |