EDBT 2026 Demo / reviewers in the wild / expert
Jing Tang 0002
dblp:83/663-2
· DBLP profile ↗
27ranked-venue papers
2as first author
16since 2021 · last 2026
0000-0001-7480-7710ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 25 · 2 first-author · 16 since 2021Artificial intelligence and machine learning · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A causal inference framework for identifying essential genes to enhance drug synergy predictionabstractMOTIVATION: Identifying synergistic drug combinations holds promise for more effective treatment strategies. Recent deep learning methods such as Transformers and Graph Neural Networks have shown improved predictive performance, but most of them integrate drug and cell line representations without explicitly modelling the causal effects of genes in mediating drug responses. RESULTS: We introduce CADS (Causal Adjustment for Drug Synergy), a deep learning framework that explicitly models the gene-drug causal relationships to improve both prediction accuracy and biological interpretability. CADS integrates multi-omics data with a learnable gene-selection mechanism that performs causal backdoor adjustment, enabling both drug synergy prediction and causal gene discovery. Across multiple benchmark datasets, CADS consistently achieves superior performance compared with state-of-the-art drug synergy prediction models. In addition, downstream analyses on case studies demonstrate that the inferred gene causal scores can recover clinically validated cancer-related genes involved in drug combinations. These results demonstrate that explicitly modelling causal genetic effects can enhance the reliability and interpretability of drug synergy prediction. AVAILABILITY AND IMPLEMENTATION: The source code of CADS can be found at https://github.com/HuaiwuZhang/causalDC. Huaiwu Zhang, Xinliang Sun, Jianxin Wang 0001, Min Li 0007, Jing Tang 0002 |
Bioinform. | 5 |
| 2025 | CADS: Causal Inference for Dissecting Essential Genes to Predict Drug Synergy
Huaiwu Zhang, Jing Tang 0002 |
ISBRA (1) | 2 |
| 2025 | TransScore: A Graph Model for Pose Scoring and Affinity Prediction Based on Transformer Convolution NetworkabstractPredicting the interaction of protein and compound is an important task in drug discovery. Molecular docking has been a fundamental and vital computer-aid tool for digging potential interaction of the protein-compound pair. With the recent great success of artificial intelligence (AI), the scoring function, as a fundamental part of molecular docking, has been achieving much better performance by incorporating AI-based models. However, the AI-based models usually focus on a single prediction task (e.g., affinity prediction), which is limited by their lack of extensibility. Moreover, the performance of AI-based models usually declines in cold start scenarios, thus compromising the robustness. To this end, we propose a novel deep learning-based graph model based on the transformer convolution network for pose scoring and affinity prediction. TransScore captures the intrinsic characteristics of protein-compound poses by employing the self-attention mechanism, which achieves superior performances in both cold and warm scenarios for the pose-scoring task. The outstanding performance is also shown in imbalanced datasets, which demonstrates the robustness of TransScore. In addition, the gated residual algorithm in TransScore enhances the model to adapt to diverse related tasks. In particular, in the affinity prediction task, we have observed consistent improvements in warm/cold start scenarios. Moreover, it is noticeable that TransScore excels in both accuracy and precision, accurately predicting affinities and their relative ordering. We also conducted an analysis on carbonic anhydrase II, which bears out that TransScore can elaborate the interaction mechanism of the protein-ligand pair, suggesting the potential application of TransScore in drug discovery. Chuqi Lei, Wenkang Wang, Wei Fan 0010, Zhangli Lu, Jing Tang 0002, Min Li 0007 |
IEEE J. Biomed. Health Informatics | 5 |
| 2024 | Herb-CMap: a multimodal fusion framework for deciphering the mechanisms of action in traditional Chinese medicine using Suhuang antitussive capsule as a case studyabstractHerbal medicines, particularly traditional Chinese medicines (TCMs), are a rich source of natural products with significant therapeutic potential. However, understanding their mechanisms of action is challenging due to the complexity of their multi-ingredient compositions. We introduced Herb-CMap, a multimodal fusion framework leveraging protein-protein interactions and herb-perturbed gene expression signatures. Utilizing a network-based heat diffusion algorithm, Herb-CMap creates a connectivity map linking herb perturbations to their therapeutic targets, thereby facilitating the prioritization of active ingredients. As a case study, we applied Herb-CMap to Suhuang antitussive capsule (Suhuang), a TCM formula used for treating cough variant asthma (CVA). Using in vivo rat models, our analysis established the transcriptomic signatures of Suhuang and identified its key compounds, such as quercetin and luteolin, and their target genes, including IL17A, PIK3CB, PIK3CD, AKT1, and TNF. These drug-target interactions inhibit the IL-17 signaling pathway and deactivate PI3K, AKT, and NF-κB, effectively reducing lung inflammation and alleviating CVA. The study demonstrates the efficacy of Herb-CMap in elucidating the molecular mechanisms of herbal medicines, offering valuable insights for advancing drug discovery in TCM. Yinyin Wang, Yihang Sui, Qimeng Tian, Yun Tang 0001, Yongyu Ou, Jing Tang 0002, Ninghua Tan |
Briefings Bioinform. | 8 |
| 2024 | Drug repositioning with adaptive graph convolutional networksabstractMOTIVATION: Drug repositioning is an effective strategy to identify new indications for existing drugs, providing the quickest possible transition from bench to bedside. With the rapid development of deep learning, graph convolutional networks (GCNs) have been widely adopted for drug repositioning tasks. However, prior GCNs based methods exist limitations in deeply integrating node features and topological structures, which may hinder the capability of GCNs. RESULTS: In this study, we propose an adaptive GCNs approach, termed AdaDR, for drug repositioning by deeply integrating node features and topological structures. Distinct from conventional graph convolution networks, AdaDR models interactive information between them with adaptive graph convolution operation, which enhances the expression of model. Concretely, AdaDR simultaneously extracts embeddings from node features and topological structures and then uses the attention mechanism to learn adaptive importance weights of the embeddings. Experimental results show that AdaDR achieves better performance than multiple baselines for drug repositioning. Moreover, in the case study, exploratory analyses are offered for finding novel drug-disease associations. AVAILABILITY AND IMPLEMENTATION: The soure code of AdaDR is available at: https://github.com/xinliangSun/AdaDR. Xinliang Sun, Xiao Jia 0020, Zhangli Lu, Jing Tang 0002, Min Li 0007 |
Bioinform. | 4 |
| 2024 | SAFER: sub-hypergraph attention-based neural network for predicting effective responses to dose combinationsabstractBACKGROUND: The potential benefits of drug combination synergy in cancer medicine are significant, yet the risks must be carefully managed due to the possibility of increased toxicity. Although artificial intelligence applications have demonstrated notable success in predicting drug combination synergy, several key challenges persist: (1) Existing models often predict average synergy values across a restricted range of testing dosages, neglecting crucial dose amounts and the mechanisms of action of the drugs involved. (2) Many graph-based models rely on static protein-protein interactions, failing to adapt to dynamic and higher-order relationships. These limitations constrain the applicability of current methods. RESULTS: We introduce SAFER, a Sub-hypergraph Attention-based graph model, addressing these issues by incorporating complex relationships among biological knowledge networks and considering dosing effects on subject-specific networks. SAFER outperformed previous models on the benchmark and the independent test set. The analysis of subgraph attention weight for the lung cancer cell line highlighted JAK-STAT signaling pathway, PRDM12, ZNF781, and CDC5L that have been implicated in lung fibrosis. CONCLUSIONS: SAFER presents an interpretable framework designed to identify drug-responsive signals. Tailored for comprehending dose effects on subject-specific molecular contexts, our model uniquely captures dose-level drug combination responses. This capability unlocks previously inaccessible avenues of investigation compared to earlier models. Furthermore, the SAFER framework can be leveraged by future inquiries to investigate molecular networks that uniquely characterize individual patients and can be applied to prioritize personalized effective treatment based on safe dose combinations. Yi-Ching Tang, Rongbin Li, Jing Tang 0002, W. Jim Zheng, Xiaoqian Jiang |
BMC Bioinform. | 3 |
| 2023 | GraphscoreDTA: optimized graph neural network for protein-ligand binding affinity predictionabstractMOTIVATION: Computational approaches for identifying the protein-ligand binding affinity can greatly facilitate drug discovery and development. At present, many deep learning-based models are proposed to predict the protein-ligand binding affinity and achieve significant performance improvement. However, protein-ligand binding affinity prediction still has fundamental challenges. One challenge is that the mutual information between proteins and ligands is hard to capture. Another challenge is how to find and highlight the important atoms of the ligands and residues of the proteins. RESULTS: To solve these limitations, we develop a novel graph neural network strategy with the Vina distance optimization terms (GraphscoreDTA) for predicting protein-ligand binding affinity, which takes the combination of graph neural network, bitransport information mechanism and physics-based distance terms into account for the first time. Unlike other methods, GraphscoreDTA can not only effectively capture the protein-ligand pairs' mutual information but also highlight the important atoms of the ligands and residues of the proteins. The results show that GraphscoreDTA significantly outperforms existing methods on multiple test sets. Furthermore, the tests of drug-target selectivity on the cyclin-dependent kinase and the homologous protein families demonstrate that GraphscoreDTA is a reliable tool for protein-ligand binding affinity prediction. AVAILABILITY AND IMPLEMENTATION: The resource codes are available at https://github.com/CSUBioGroup/GraphscoreDTA. Renyi Zhou, Jing Tang 0002, Min Li 0007 |
Bioinform. | 3 |
| 2023 | The Impact of Computational Drug Discovery on SocietyabstractGreetings and welcome to the fifth issue of IEEE Transactions on Computational Social Systems (TCSS) for 2023. This edition presents a collection of 55 diverse regular articles that illuminate various facets of the interaction between computer technology and society. Jianxin Wang 0001, Min Li 0007, Edwin Wang, Jing Tang 0002, Bin Hu 0001 |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2022 | Minimal information for chemosensitivity assays (MICHA): a next-generation pipeline to enable the FAIRification of drug screening experimentsabstractChemosensitivity assays are commonly used for preclinical drug discovery and clinical trial optimization. However, data from independent assays are often discordant, largely attributed to uncharacterized variation in the experimental materials and protocols. We report here the launching of Minimal Information for Chemosensitivity Assays (MICHA), accessed via https://micha-protocol.org. Distinguished from existing efforts that are often lacking support from data integration tools, MICHA can automatically extract publicly available information to facilitate the assay annotation including: 1) compounds, 2) samples, 3) reagents and 4) data processing methods. For example, MICHA provides an integrative web server and database to obtain compound annotation including chemical structures, targets and disease indications. In addition, the annotation of cell line samples, assay protocols and literature references can be greatly eased by retrieving manually curated catalogues. Once the annotation is complete, MICHA can export a report that conforms to the FAIR principle (Findable, Accessible, Interoperable and Reusable) of drug screening studies. To consolidate the utility of MICHA, we provide FAIRified protocols from five major cancer drug screening studies as well as six recently conducted COVID-19 studies. With the MICHA web server and database, we envisage a wider adoption of a community-driven effort to improve the open access of drug sensitivity assays. ZiaurRehman Tanoli, Jehad Aldahdooh, Muhammad Farhan Alam, Yinyin Wang, Umair Seemab, Maddalena Fratelli, Petr Pavlis, Marián Hajdúch, Florence Bietrix, Philip Gribbon, Andrea Zaliani, Matthew D. Hall, Kyle R. Brimacombe, Evgeny Kulesskiy, Saarela Jani, Krister Wennerberg, Markus Vähä-Koskela, Jing Tang 0002 |
Briefings Bioinform. | 19 |
| 2022 | The ENDS of assumptions: an online tool for the epistemic non-parametric drug-response scoringabstractMOTIVATION: The drug sensitivity analysis is often elucidated from drug dose-response curves. These curves capture the degree of cell viability (or inhibition) over a range of induced drugs, often with parametric assumptions that are rarely validated. RESULTS: We present a class of non-parametric models for the curve fitting and scoring of drug dose-responses. To allow a more objective representation of the drug sensitivity, these epistemic models devoid of any parametric assumptions attached to the linear fit, allow the parallel indexing such as half-maximal inhibitory concentration and area under curve. Specifically, three non-parametric models including spline (npS), monotonic and Bayesian and the parametric logistic are implemented. Other indices including maximum effective dose and drug-response span gradient pertinent to the npS are also provided to facilitate the interpretation of the fit. The collection of these models is implemented in an online app, standing as useful resource for drug dose-response curve fitting and analysis. AVAILABILITY AND IMPLEMENTATION: The ENDS is freely available online at https://irscope.shinyapps.io/ENDS/ and source codes can be obtained from https://github.com/AmiryousefiLab/ENDS. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ali Amiryousefi, Bernardo Williams, Mohieddin Jafari, Jing Tang 0002 |
Bioinform. | 4 |
| 2022 | Using BERT to identify drug-target interactions from whole PubMedabstractBACKGROUND: Drug-target interactions (DTIs) are critical for drug repurposing and elucidation of drug mechanisms, and are manually curated by large databases, such as ChEMBL, BindingDB, DrugBank and DrugTargetCommons. However, the number of curated articles likely constitutes only a fraction of all the articles that contain experimentally determined DTIs. Finding such articles and extracting the experimental information is a challenging task, and there is a pressing need for systematic approaches to assist the curation of DTIs. To this end, we applied Bidirectional Encoder Representations from Transformers (BERT) to identify such articles. Because DTI data intimately depends on the type of assays used to generate it, we also aimed to incorporate functions to predict the assay format. RESULTS: Our novel method identified 0.6 million articles (along with drug and protein information) which are not previously included in public DTI databases. Using 10-fold cross-validation, we obtained ~ 99% accuracy for identifying articles containing quantitative drug-target profiles. The F1 micro for the prediction of assay format is 88%, which leaves room for improvement in future studies. CONCLUSION: The BERT model in this study is robust and the proposed pipeline can be used to identify previously overlooked articles containing quantitative DTIs. Overall, our method provides a significant advancement in machine-assisted DTI extraction and curation. We expect it to be a useful addition to drug mechanism discovery and repurposing. Jehad Aldahdooh, Markus Vähä-Koskela, Jing Tang 0002, ZiaurRehman Tanoli |
BMC Bioinform. | 3 |
| 2021 | Network-guided identification of cancer-selective combinatorial therapies in ovarian cancerabstractEach patient's cancer consists of multiple cell subpopulations that are inherently heterogeneous and may develop differing phenotypes such as drug sensitivity or resistance. A personalized treatment regimen should therefore target multiple oncoproteins in the cancer cell populations that are driving the treatment resistance or disease progression in a given patient to provide maximal therapeutic effect, while avoiding severe co-inhibition of non-malignant cells that would lead to toxic side effects. To address the intra- and inter-tumoral heterogeneity when designing combinatorial treatment regimens for cancer patients, we have implemented a machine learning-based platform to guide identification of safe and effective combinatorial treatments that selectively inhibit cancer-related dysfunctions or resistance mechanisms in individual patients. In this case study, we show how the platform enables prediction of cancer-selective drug combinations for patients with high-grade serous ovarian cancer using single-cell imaging cytometry drug response assay, combined with genome-wide transcriptomic and genetic profiles. The platform makes use of drug-target interaction networks to prioritize those combinations that warrant further preclinical testing in scarce patient-derived primary cells. During the case study in ovarian cancer patients, we investigated (i) the relative performance of various ensemble learning algorithms for drug response prediction, (ii) the use of matched single-cell RNA-sequencing data to deconvolute cell population-specific transcriptome profiles from bulk RNA-seq data, (iii) and whether multi-patient or patient-specific predictive models lead to better predictive accuracy. The general platform and the comparison results are expected to become useful for future studies that use similar predictive approaches also in other cancer types. Liye He, Daria Bulanova, Jaana Oikkonen, Antti Häkkinen, Kaiyang Zhang, Erdogan Pekcan Erkan, Olli Carpén, Titta Joutsiniemi, Sakari Hietanen, Johanna Hynninen, Kaisa Huhtinen, Sampsa Hautaniemi, Anna Vähärautio, Jing Tang 0002, Krister Wennerberg, Tero Aittokallio |
Briefings Bioinform. | 16 |
| 2021 | Exploration of databases and methods supporting drug repurposing: a comprehensive surveyabstractDrug development involves a deep understanding of the mechanisms of action and possible side effects of each drug, and sometimes results in the identification of new and unexpected uses for drugs, termed as drug repurposing. Both in case of serendipitous observations and systematic mechanistic explorations, confirmation of new indications for a drug requires hypothesis building around relevant drug-related data, such as molecular targets involved, and patient and cellular responses. These datasets are available in public repositories, but apart from sifting through the sheer amount of data imposing computational bottleneck, a major challenge is the difficulty in selecting which databases to use from an increasingly large number of available databases. The database selection is made harder by the lack of an overview of the types of data offered in each database. In order to alleviate these problems and to guide the end user through the drug repurposing efforts, we provide here a survey of 102 of the most promising and drug-relevant databases reported to date. We summarize the target coverage and types of data available in each database and provide several examples of how multi-database exploration can facilitate drug repurposing. ZiaurRehman Tanoli, Umair Seemab, Andreas Scherer, Krister Wennerberg, Jing Tang 0002, Markus Vähä-Koskela |
Briefings Bioinform. | 5 |
| 2021 | Network-based modeling of herb combinations in traditional Chinese medicineabstractTraditional Chinese medicine (TCM) has been practiced for thousands of years for treating human diseases. In comparison to modern medicine, one of the advantages of TCM is the principle of herb compatibility, known as TCM formulae. A TCM formula usually consists of multiple herbs to achieve the maximum treatment effects, where their interactions are believed to elicit the therapeutic effects. Despite being a fundamental component of TCM, the rationale of combining specific herb combinations remains unclear. In this study, we proposed a network-based method to quantify the interactions in herb pairs. We constructed a protein-protein interaction network for a given herb pair by retrieving the associated ingredients and protein targets, and determined multiple network-based distances including the closest, shortest, center, kernel, and separation, both at the ingredient and at the target levels. We found that the frequently used herb pairs tend to have shorter distances compared to random herb pairs, suggesting that a therapeutic herb pair is more likely to affect neighboring proteins in the human interactome. Furthermore, we found that the center distance determined at the ingredient level improves the discrimination of top-frequent herb pairs from random herb pairs, suggesting the rationale of considering the topologically important ingredients for inferring the mechanisms of action of TCM. Taken together, we have provided a network pharmacology framework to quantify the degree of herb interactions, which shall help explore the space of herb combinations more effectively to identify the synergistic compound interactions based on network topology. Yinyin Wang, Hongbin Yang 0002, Linxiao Chen, Mohieddin Jafari, Jing Tang 0002 |
Briefings Bioinform. | 5 |
| 2021 | Comparative analysis of molecular fingerprints in prediction of drug combination effectsabstractApplication of machine and deep learning methods in drug discovery and cancer research has gained a considerable amount of attention in the past years. As the field grows, it becomes crucial to systematically evaluate the performance of novel computational solutions in relation to established techniques. To this end, we compare rule-based and data-driven molecular representations in prediction of drug combination sensitivity and drug synergy scores using standardized results of 14 high-throughput screening studies, comprising 64 200 unique combinations of 4153 molecules tested in 112 cancer cell lines. We evaluate the clustering performance of molecular representations and quantify their similarity by adapting the Centered Kernel Alignment metric. Our work demonstrates that to identify an optimal molecular representation type, it is necessary to supplement quantitative benchmark results with qualitative considerations, such as model interpretability and robustness, which may vary between and throughout preclinical drug development projects. Bulat Zagidullin, Yuanfang Guan, Esa Pitkänen, Jing Tang 0002 |
Briefings Bioinform. | 5 |
| 2021 | Anticancer drug synergy prediction in understudied tissues using transfer learningabstractOBJECTIVE: Drug combination screening has advantages in identifying cancer treatment options with higher efficacy without degradation in terms of safety. A key challenge is that the accumulated number of observations in in-vitro drug responses varies greatly among different cancer types, where some tissues are more understudied than the others. Thus, we aim to develop a drug synergy prediction model for understudied tissues as a way of overcoming data scarcity problems. MATERIALS AND METHODS: We collected a comprehensive set of genetic, molecular, phenotypic features for cancer cell lines. We developed a drug synergy prediction model based on multitask deep neural networks to integrate multimodal input and multiple output. We also utilized transfer learning from data-rich tissues to data-poor tissues. RESULTS: We showed improved accuracy in predicting synergy in both data-rich tissues and understudied tissues. In data-rich tissue, the prediction model accuracy was 0.9577 AUROC for binarized classification task and 174.3 mean squared error for regression task. We observed that an adequate transfer learning strategy significantly increases accuracy in the understudied tissues. CONCLUSIONS: Our synergy prediction model can be used to rank synergistic drug combinations in understudied tissues and thus help to prioritize future in-vitro experiments. Code is available at https://github.com/yejinjkim/synergy-transfer. Yejin Kim 0001, Jing Tang 0002, W. Jim Zheng, Xiaoqian Jiang |
J. Am. Medical Informatics Assoc. | 3 |
| 2020 | SynergyFinder: a web application for analyzing drug combination dose-response matrix dataabstractBioinformatics (2017), 33, 2017, 2413–2415, doi: 10.1093/bioinformatics/btx162 The following funding source was inadvertently omitted from the above article: European Research Council (ERC) starting grant DrugComb (Informatics approaches for the rationale selection of personaliszed cancer drug combinations) [No. 716063] This has now been added. Aleksandr Ianevski, Liye He, Tero Aittokallio, Jing Tang 0002 |
Bioinform. | 4 |
| 2019 | Drug combination sensitivity scoring facilitates the discovery of synergistic and efficacious drug combinations in cancerabstractHigh-throughput drug screening has facilitated the discovery of drug combinations in cancer. Many existing studies adopted a full matrix design, aiming for the characterization of drug pair effects for cancer cells. However, the full matrix design may be suboptimal as it requires a drug pair to be combined at multiple concentrations in a full factorial manner. Furthermore, many of the computational tools assess only the synergy but not the sensitivity of drug combinations, which might lead to false positive discoveries. We proposed a novel cross design to enable a more cost-effective and simultaneous testing of drug combination sensitivity and synergy. We developed a drug combination sensitivity score (CSS) to determine the sensitivity of a drug pair, and showed that the CSS is highly reproducible between the replicates and thus supported its usage as a robust metric. We further showed that CSS can be predicted using machine learning approaches which determined the top pharmaco-features to cluster cancer cell lines based on their drug combination sensitivity profiles. To assess the degree of drug interactions using the cross design, we developed an S synergy score based on the difference between the drug combination and the single drug dose-response curves. We showed that the S score is able to detect true synergistic and antagonistic drug combinations at an accuracy level comparable to that using the full matrix design. Taken together, we showed that the cross design coupled with the CSS sensitivity and S synergy scoring methods may provide a robust and accurate characterization of both drug combination sensitivity and synergy levels, with minimal experimental materials required. Our experimental-computational approach could be utilized as an efficient pipeline for improving the discovery rate in high-throughput drug combination screening, particularly for primary patient samples which are difficult to obtain. Alina Malyutina, Muntasir Mamun Majumder, Alberto Pessia, Caroline Heckman, Jing Tang 0002 |
PLoS Comput. Biol. | 6 |
| 2019 | Predicting Meridian in Chinese traditional medicine using machine learning approachesabstractPlant-derived nature products, known as herb formulas, have been commonly used in Traditional Chinese Medicine (TCM) for disease prevention and treatment. The herbs have been traditionally classified into different categories according to the TCM Organ systems known as Meridians. Despite the increasing knowledge on the active components of the herbs, the rationale of Meridian classification remains poorly understood. In this study, we took a machine learning approach to explore the classification of Meridian. We determined the molecule features for 646 herbs and their active components including structure-based fingerprints and ADME properties (absorption, distribution, metabolism and excretion), and found that the Meridian can be predicted by machine learning approaches with a top accuracy of 0.83. We also identified the top compound features that were important for the Meridian prediction. To the best of our knowledge, this is the first time that molecular properties of the herb compounds are associated with the TCM Meridians. Taken together, the machine learning approach may provide novel insights for the understanding of molecular evidence of Meridians in TCM. Yinyin Wang, Mohieddin Jafari, Yun Tang 0001, Jing Tang 0002 |
PLoS Comput. Biol. | 4 |
| 2017 | SynergyFinder: a web application for analyzing drug combination dose-response matrix dataabstractSUMMARY: Rational design of drug combinations has become a promising strategy to tackle the drug sensitivity and resistance problem in cancer treatment. To systematically evaluate the pre-clinical significance of pairwise drug combinations, functional screening assays that probe combination effects in a dose-response matrix assay are commonly used. To facilitate the analysis of such drug combination experiments, we implemented a web application that uses key functions of R-package SynergyFinder, and provides not only the flexibility of using multiple synergy scoring models, but also a user-friendly interface for visualizing the drug combination landscapes in an interactive manner. AVAILABILITY AND IMPLEMENTATION: The SynergyFinder web application is freely accessible at https://synergyfinder.fimm.fi ; The R-package and its source-code are freely available at http://bioconductor.org/packages/release/bioc/html/synergyfinder.html . CONTACT: [email protected]. Aleksandr Ianevski, Liye He, Tero Aittokallio, Jing Tang 0002 |
Bioinform. | 4 |
| 2015 | Toward more realistic drug-target interaction predictionsabstractA number of supervised machine learning models have recently been introduced for the prediction of drug-target interactions based on chemical structure and genomic sequence information. Although these models could offer improved means for many network pharmacology applications, such as repositioning of drugs for new therapeutic uses, the prediction models are often being constructed and evaluated under overly simplified settings that do not reflect the real-life problem in practical applications. Using quantitative drug-target bioactivity assays for kinase inhibitors, as well as a popular benchmarking data set of binary drug-target interactions for enzyme, ion channel, nuclear receptor and G protein-coupled receptor targets, we illustrate here the effects of four factors that may lead to dramatic differences in the prediction results: (i) problem formulation (standard binary classification or more realistic regression formulation), (ii) evaluation data set (drug and target families in the application use case), (iii) evaluation procedure (simple or nested cross-validation) and (iv) experimental setting (whether training and test sets share common drugs and targets, only drugs or targets or neither). Each of these factors should be taken into consideration to avoid reporting overoptimistic drug-target interaction prediction results. We also suggest guidelines on how to make the supervised drug-target interaction prediction studies more realistic in terms of such model formulations and evaluation setups that better address the inherent complexity of the prediction task in the practical applications, as well as novel benchmarking data sets that capture the continuous nature of the drug-target interactions for kinase inhibitors. Tapio Pahikkala, Antti Airola, Sami Pietilä, Sushil Kumar Shakyawar, Agnieszka Szwajda, Jing Tang 0002, Tero Aittokallio |
Briefings Bioinform. | 6 |
| 2015 | TIMMA-R: an R package for predicting synergistic multi-targeted drug combinations in cancer cell lines or patient-derived samplesabstractUNLABELLED: Network pharmacology-based prediction of multi-targeted drug combinations is becoming a promising strategy to improve anticancer efficacy and safety. We developed a logic-based network algorithm, called Target Inhibition Interaction using Maximization and Minimization Averaging (TIMMA), which predicts the effects of drug combinations based on their binary drug-target interactions and single-drug sensitivity profiles in a given cancer sample. Here, we report the R implementation of the algorithm (TIMMA-R), which is much faster than the original MATLAB code. The major extensions include modeling of multiclass drug-target profiles and network visualization. We also show that the TIMMA-R predictions are robust to the intrinsic noise in the experimental data, thus making it a promising high-throughput tool to prioritize drug combinations in various cancer types for follow-up experimentation or clinical applications. AVAILABILITY AND IMPLEMENTATION: TIMMA-R source code is freely available at http://cran.r-project.org/web/packages/timma/. Liye He, Krister Wennerberg, Tero Aittokallio, Jing Tang 0002 |
Bioinform. | 4 |
| 2015 | A Bayesian Predictive Model for Clustering Data of Mixed Discrete and Continuous TypeabstractAdvantages of model-based clustering methods over heuristic alternatives have been widely demonstrated in the literature. Most model-based clustering algorithms assume that the data are either discrete or continuous, possibly allowing both types to be present in separate features. In this paper, we introduce a model-based approach for clustering feature vectors of mixed type, allowing each feature to simultaneously take on both categorical and real values. Such data may be encountered, for instance, in chemical and biological analyses, in the analysis of survey data, as well as in image analysis. Our model is formulated within a Bayesian predictive framework, where clustering solutions correspond to random partitions of the data. Using conjugate analysis, the posterior probability for each possible partition can be determined analytically, enabling the utilization of efficient computational search strategies for finding the posterior optimal partition. The derived model is illustrated using several synthetic and real datasets. Paul Blomstedt, Jing Tang 0002, Christian Granlund, Jukka Corander |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2013 | Target Inhibition Networks: Predicting Selective Combinations of Druggable Targets to Block Cancer Survival PathwaysabstractA recent trend in drug development is to identify drug combinations or multi-target agents that effectively modify multiple nodes of disease-associated networks. Such polypharmacological effects may reduce the risk of emerging drug resistance by means of attacking the disease networks through synergistic and synthetic lethal interactions. However, due to the exponentially increasing number of potential drug and target combinations, systematic approaches are needed for prioritizing the most potent multi-target alternatives on a global network level. We took a functional systems pharmacology approach toward the identification of selective target combinations for specific cancer cells by combining large-scale screening data on drug treatment efficacies and drug-target binding affinities. Our model-based prediction approach, named TIMMA, takes advantage of the polypharmacological effects of drugs and infers combinatorial drug efficacies through system-level target inhibition networks. Case studies in MCF-7 and MDA-MB-231 breast cancer and BxPC-3 pancreatic cancer cells demonstrated how the target inhibition modeling allows systematic exploration of functional interactions between drugs and their targets to maximally inhibit multiple survival pathways in a given cancer type. The TIMMA prediction results were experimentally validated by means of systematic siRNA-mediated silencing of the selected targets and their pairwise combinations, showing increased ability to identify not only such druggable kinase targets that are essential for cancer survival either individually or in combination, but also synergistic interactions indicative of non-additive drug efficacies. These system-level analyses were enabled by a novel model construction method utilizing maximization and minimization rules, as well as a model selection algorithm based on sequential forward floating search. Compared with an existing computational solution, TIMMA showed both enhanced prediction accuracies in cross validation as well as significant reduction in computation times. Such cost-effective computational-experimental design strategies have the potential to greatly speed-up the drug testing efforts by prioritizing those interventions and interactions warranting further study in individual cancer cases. Jing Tang 0002, Leena Karhinen, Agnieszka Szwajda, Bhagwan Yadav, Krister Wennerberg, Tero Aittokallio |
PLoS Comput. Biol. | 1 |
| 2009 | Bayesian Clustering of Fuzzy Feature Vectors Using a Quasi-Likelihood ApproachabstractBayesian model-based classifiers, both unsupervised and supervised, have been studied extensively and their value and versatility have been demonstrated on a wide spectrum of applications within science and engineering. A majority of the classifiers are built on the assumption of intrinsic discreteness of the considered data features or on the discretization of them prior to the modeling. On the other hand, Gaussian mixture classifiers have also been utilized to a large extent for continuous features in the Bayesian framework. Often the primary reason for discretization in the classification context is the simplification of the analytical and numerical properties of the models. However, the discretization can be problematic due to its \textit{ad hoc} nature and the decreased statistical power to detect the correct classes in the resulting procedure. We introduce an unsupervised classification approach for fuzzy feature vectors that utilizes a discrete model structure while preserving the continuous characteristics of data. This is achieved by replacing the ordinary likelihood by a binomial quasi-likelihood to yield an analytical expression for the posterior probability of a given clustering solution. The resulting model can be justified from an information-theoretic perspective. Our method is shown to yield highly accurate clusterings for challenging synthetic and empirical data sets. Pekka Marttinen, Jing Tang 0002, Bernard De Baets, Peter Dawyndt, Jukka Corander |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2009 | Identifying Currents in the Gene Pool for Bacterial Populations Using an Integrative ApproachabstractThe evolution of bacterial populations has recently become considerably better understood due to large-scale sequencing of population samples. It has become clear that DNA sequences from a multitude of genes, as well as a broad sample coverage of a target population, are needed to obtain a relatively unbiased view of its genetic structure and the patterns of ancestry connected to the strains. However, the traditional statistical methods for evolutionary inference, such as phylogenetic analysis, are associated with several difficulties under such an extensive sampling scenario, in particular when a considerable amount of recombination is anticipated to have taken place. To meet the needs of large-scale analyses of population structure for bacteria, we introduce here several statistical tools for the detection and representation of recombination between populations. Also, we introduce a model-based description of the shape of a population in sequence space, in terms of its molecular variability and affinity towards other populations. Extensive real data from the genus Neisseria are utilized to demonstrate the potential of an approach where these population genetic tools are combined with an phylogenetic analysis. The statistical tools introduced here are freely available in BAPS 5.2 software, which can be downloaded from http://web.abo.fi/fak/mnf/mate/jc/software/baps.html. Jing Tang 0002, William P. Hanage, Christophe Fraser, Jukka Corander |
PLoS Comput. Biol. | 1 |
| 2008 | Enhanced Bayesian modelling in BAPS software for learning genetic structures of populationsabstractBACKGROUND: During the most recent decade many Bayesian statistical models and software for answering questions related to the genetic structure underlying population samples have appeared in the scientific literature. Most of these methods utilize molecular markers for the inferences, while some are also capable of handling DNA sequence data. In a number of earlier works, we have introduced an array of statistical methods for population genetic inference that are implemented in the software BAPS. However, the complexity of biological problems related to genetic structure analysis keeps increasing such that in many cases the current methods may provide either inappropriate or insufficient solutions. RESULTS: We discuss the necessity of enhancing the statistical approaches to face the challenges posed by the ever-increasing amounts of molecular data generated by scientists over a wide range of research areas and introduce an array of new statistical tools implemented in the most recent version of BAPS. With these methods it is possible, e.g., to fit genetic mixture models using user-specified numbers of clusters and to estimate levels of admixture under a genetic linkage model. Also, alleles representing a different ancestry compared to the average observed genomic positions can be tracked for the sampled individuals, and a priori specified hypotheses about genetic population structure can be directly compared using Bayes' theorem. In general, we have improved further the computational characteristics of the algorithms behind the methods implemented in BAPS facilitating the analyses of large and complex datasets. In particular, analysis of a single dataset can now be spread over multiple computers using a script interface to the software. CONCLUSION: The Bayesian modelling methods introduced in this article represent an array of enhanced tools for learning the genetic structure of populations. Their implementations in the BAPS software are designed to meet the increasing need for analyzing large-scale population genetics data. The software is freely downloadable for Windows, Linux and Mac OS X systems at http://web.abo.fi/fak/mnf//mate/jc/software/baps.html. Jukka Corander, Pekka Marttinen, Jukka Sirén, Jing Tang 0002 |
BMC Bioinform. | 4 |