EDBT 2026 Demo / reviewers in the wild / expert
Tero Aittokallio
dblp:41/6124
· DBLP profile ↗
46ranked-venue papers
3as first author
11since 2021 · last 2025
0000-0002-0886-9769ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 42 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 3Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Scaling up drug combination surface predictionabstractDrug combinations are required to treat advanced cancers and other complex diseases. Compared with monotherapy, combination treatments can enhance efficacy and reduce toxicity by lowering the doses of single drugs-and there especially synergistic combinations are of interest. Since drug combination screening experiments are costly and time-consuming, reliable machine learning models are needed for prioritizing potential combinations for further studies. Most of the current machine learning models are based on scalar-valued approaches, which predict individual response values or synergy scores for drug combinations. We take a functional output prediction approach, in which full, continuous dose-response combination surfaces are predicted for each drug combination on the cell lines. We investigate the predictive power of the recently proposed comboKR method, which is based on a powerful input-output kernel regression technique and functional modeling of the response surface. In this work, we develop a scaled-up formulation of the comboKR, which also implements improved modeling choices: we (1) incorporate new modeling choices for the output drug combination response surfaces to the comboKR framework, and (2) propose a projected gradient descent method to solve the challenging pre-image problem that is traditionally solved with simple candidate set approaches. We provide thorough experimental analysis of comboKR 2.0 with three real-word datasets within various challenging experimental settings, including cases where drugs or cell lines have not been encountered in the training data. Our comparison with synergy score prediction methods further highlights the relevance of dose-response prediction approaches, instead of relying on simple scoring methods. Riikka Huusari, Tianduanyi Wang, Sándor Szedmák, Diogo Dias, Tero Aittokallio, Juho Rousu |
Briefings Bioinform. | 5 |
| 2024 | RepurposeDrugs: an interactive web-portal and predictive platform for repurposing mono- and combination therapiesabstractRepurposeDrugs (https://repurposedrugs.org/) is a comprehensive web-portal that combines a unique drug indication database with a machine learning (ML) predictor to discover new drug-indication associations for approved as well as investigational mono and combination therapies. The platform provides detailed information on treatment status, disease indications and clinical trials across 25 indication categories, including neoplasms and cardiovascular conditions. The current version comprises 4314 compounds (approved, terminated or investigational) and 161 drug combinations linked to 1756 indications/conditions, totaling 28 148 drug-disease pairs. By leveraging data on both approved and failed indications, RepurposeDrugs provides ML-based predictions for the approval potential of new drug-disease indications, both for mono- and combinatorial therapies, demonstrating high predictive accuracy in cross-validation. The validity of the ML predictor is validated through a number of real-world case studies, demonstrating its predictive power to accurately identify repurposing candidates with a high likelihood of future approval. To our knowledge, RepurposeDrugs web-portal is the first integrative database and ML-based predictor for interactive exploration and prediction of both single-drug and combination approval likelihood across indications. Given its broad coverage of indication areas and therapeutic options, we expect it accelerates many future drug repurposing projects. Aleksandr Ianevski, Aleksandr Kushnir, Kristen Nader, Mitro Miihkinen, Henri Xhaard, Tero Aittokallio, ZiaurRehman Tanoli |
Briefings Bioinform. | 6 |
| 2024 | ScType enables fast and accurate cell type identification from spatial transcriptomics dataabstractSUMMARY: The limited resolution of spatial transcriptomics (ST) assays in the past has led to the development of cell type annotation methods that separate the convolved signal based on available external atlas data. In light of the rapidly increasing resolution of the ST assay technologies, we made available and investigated the performance of a deconvolution-free marker-based cell annotation method called scType. In contrast to existing methods, the spatial application of scType does not require computationally strenuous deconvolution, nor large single-cell reference atlases. We show that scType enables ultra-fast and accurate identification of abundant cell types from ST data, especially when a large enough panel of genes is detected. Examples of such assays are Visium and Slide-seq, which currently offer the best trade-off between high resolution and number of genes detected by the assay for cell type annotation. AVAILABILITY AND IMPLEMENTATION: scType source R and python codes for spatial data are openly available in GitHub (https://github.com/kris-nader/sp-type or https://github.com/kris-nader/sc-type-py). Step-by-step tutorials for R and python spatial data analysis can be found in https://github.com/kris-nader/sp-type and https://github.com/kris-nader/sc-type-py/blob/main/spatial_tutorial.md, respectively. Kristen Nader, Misra Tasci, Aleksandr Ianevski, Andrew Erickson, Emmy W. Verschuren, Tero Aittokallio, Mitro Miihkinen |
Bioinform. | 6 |
| 2024 | Attention-based approach to predict drug-target interactions across seven target superfamiliesabstractMOTIVATION: Drug-target interactions (DTIs) hold a pivotal role in drug repurposing and elucidation of drug mechanisms of action. While single-targeted drugs have demonstrated clinical success, they often exhibit limited efficacy against complex diseases, such as cancers, whose development and treatment is dependent on several biological processes. Therefore, a comprehensive understanding of primary, secondary and even inactive targets becomes essential in the quest for effective and safe treatments for cancer and other indications. The human proteome offers over a thousand druggable targets, yet most FDA-approved drugs bind to only a small fraction of these targets. RESULTS: This study introduces an attention-based method (called as MMAtt-DTA) to predict drug-target bioactivities across human proteins within seven superfamilies. We meticulously examined nine different descriptor sets to identify optimal signature descriptors for predicting novel DTIs. Our testing results demonstrated Spearman correlations exceeding 0.72 (P < 0.001) for six out of seven superfamilies. The proposed method outperformed fourteen state-of-the-art machine learning, deep learning and graph-based methods and maintained relatively high performance for most target superfamilies when tested with independent bioactivity data sources. We computationally validated 185 676 drug-target pairs from ChEMBL-V33 that were not available during model training, achieving a reasonable performance with Spearman correlation >0.57 (P < 0.001) for most superfamilies. This underscores the robustness of the proposed method for predicting novel DTIs. Finally, we applied our method to predict missing bioactivities among 3492 approved molecules in ChEMBL-V33, offering a valuable tool for advancing drug mechanism discovery and repurposing existing drugs for new indications. AVAILABILITY AND IMPLEMENTATION: https://github.com/AronSchulman/MMAtt-DTA. Aron Schulman, Juho Rousu, Tero Aittokallio, ZiaurRehman Tanoli |
Bioinform. | 3 |
| 2024 | Tutorial on survival modeling with applications to omics dataabstractMOTIVATION: Identification of genomic, molecular and clinical markers prognostic of patient survival is important for developing personalized disease prevention, diagnostic and treatment approaches. Modern omics technologies have made it possible to investigate the prognostic impact of markers at multiple molecular levels, including genomics, epigenomics, transcriptomics, proteomics and metabolomics, and how these potential risk factors complement clinical characterization of patient outcomes for survival prognosis. However, the massive sizes of the omics datasets, along with their correlation structures, pose challenges for studying relationships between the molecular information and patients' survival outcomes. RESULTS: We present a general workflow for survival analysis that is applicable to high-dimensional omics data as inputs when identifying survival-associated features and validating survival models. In particular, we focus on the commonly used Cox-type penalized regressions and hierarchical Bayesian models for feature selection in survival analysis, which are especially useful for high-dimensional data, but the framework is applicable more generally. AVAILABILITY AND IMPLEMENTATION: A step-by-step R tutorial using The Cancer Genome Atlas survival and omics data for the execution and evaluation of survival models has been made available at https://ocbe-uio.github.io/survomics. John Zobolas, Manuela Zucknick, Tero Aittokallio |
Bioinform. | 4 |
| 2023 | OSCAR: Optimal subset cardinality regression using the L0-pseudonorm with applications to prognostic modelling of prostate cancerabstractIn many real-world applications, such as those based on electronic health records, prognostic prediction of patient survival is based on heterogeneous sets of clinical laboratory measurements. To address the trade-off between the predictive accuracy of a prognostic model and the costs related to its clinical implementation, we propose an optimized L0-pseudonorm approach to learn sparse solutions in multivariable regression. The model sparsity is maintained by restricting the number of nonzero coefficients in the model with a cardinality constraint, which makes the optimization problem NP-hard. In addition, we generalize the cardinality constraint for grouped feature selection, which makes it possible to identify key sets of predictors that may be measured together in a kit in clinical practice. We demonstrate the operation of our cardinality constraint-based feature subset selection method, named OSCAR, in the context of prognostic prediction of prostate cancer patients, where it enables one to determine the key explanatory predictors at different levels of model sparsity. We further explore how the model sparsity affects the model accuracy and implementation cost. Lastly, we demonstrate generalization of the presented methodology to high-dimensional transcriptomics data. Anni S. Halkola, Kaisa Joki, Tuomas Mirtti, Marko M. Mäkelä, Tero Aittokallio, Teemu D. Laajala |
PLoS Comput. Biol. | 5 |
| 2022 | Discovery of host-directed modulators of virus infection by probing the SARS-CoV-2-host protein-protein interaction networkabstractThe ongoing coronavirus disease 2019 (COVID-19) pandemic has highlighted the need to better understand virus-host interactions. We developed a network-based method that expands the severe acute respiratory syndrome coronavirus-2 (SARS-CoV-2)-host protein interaction network and identifies host targets that modulate viral infection. To disrupt the SARS-CoV-2 interactome, we systematically probed for potent compounds that selectively target the identified host proteins with high expression in cells relevant to COVID-19. We experimentally tested seven chemical inhibitors of the identified host proteins for modulation of SARS-CoV-2 infection in human cells that express ACE2 and TMPRSS2. Inhibition of the epigenetic regulators bromodomain-containing protein 4 (BRD4) and histone deacetylase 2 (HDAC2), along with ubiquitin-specific peptidase (USP10), enhanced SARS-CoV-2 infection. Such proviral effect was observed upon treatment with compounds JQ1, vorinostat, romidepsin and spautin-1, when measured by cytopathic effect and validated by viral RNA assays, suggesting that the host proteins HDAC2, BRD4 and USP10 have antiviral functions. We observed marked differences in antiviral effects across cell lines, which may have consequences for identification of selective modulators of viral infection or potential antiviral therapeutics. While network-based approaches enable systematic identification of host targets and selective compounds that may modulate the SARS-CoV-2 interactome, further developments are warranted to increase their accuracy and cell-context specificity. Vandana Ravindran, Jessica Wagoner, Paschalis Athanasiadis, Andreas B. Den Hartigh, Julia M. Sidorova, Aleksandr Ianevski, Susan L. Fink, Arnoldo Frigessi, Judith White, Stephen J. Polyak, Tero Aittokallio |
Briefings Bioinform. | 11 |
| 2022 | Evaluation of statistical approaches for association testing in noisy drug screening dataabstractBACKGROUND: Identifying associations among biological variables is a major challenge in modern quantitative biological research, particularly given the systemic and statistical noise endemic to biological systems. Drug sensitivity data has proven to be a particularly challenging field for identifying associations to inform patient treatment. RESULTS: To address this, we introduce two semi-parametric variations on the commonly used concordance index: the robust concordance index and the kernelized concordance index (rCI, kCI), which incorporate measurements about the noise distribution from the data. We demonstrate that common statistical tests applied to the concordance index and its variations fail to control for false positives, and introduce efficient implementations to compute p-values using adaptive permutation testing. We then evaluate the statistical power of these coefficients under simulation and compare with Pearson and Spearman correlation coefficients. Finally, we evaluate the various statistics in matching drugs across pharmacogenomic datasets. CONCLUSIONS: We observe that the rCI and kCI are better powered than the concordance index in simulation and show some improvement on real data. Surprisingly, we observe that the Pearson correlation was the most robust to measurement noise among the different metrics. Petr Smirnov, Ian Smith, Zhaleh Safikhani, Wail Ba-alawi, Farnoosh Khodakarami, Eva Lin, Yihong Yu, Scott Martin, Janosch Ortmann, Tero Aittokallio, Marc Hafner, Benjamin Haibe-Kains |
BMC Bioinform. | 10 |
| 2021 | Network-guided identification of cancer-selective combinatorial therapies in ovarian cancerabstractEach patient's cancer consists of multiple cell subpopulations that are inherently heterogeneous and may develop differing phenotypes such as drug sensitivity or resistance. A personalized treatment regimen should therefore target multiple oncoproteins in the cancer cell populations that are driving the treatment resistance or disease progression in a given patient to provide maximal therapeutic effect, while avoiding severe co-inhibition of non-malignant cells that would lead to toxic side effects. To address the intra- and inter-tumoral heterogeneity when designing combinatorial treatment regimens for cancer patients, we have implemented a machine learning-based platform to guide identification of safe and effective combinatorial treatments that selectively inhibit cancer-related dysfunctions or resistance mechanisms in individual patients. In this case study, we show how the platform enables prediction of cancer-selective drug combinations for patients with high-grade serous ovarian cancer using single-cell imaging cytometry drug response assay, combined with genome-wide transcriptomic and genetic profiles. The platform makes use of drug-target interaction networks to prioritize those combinations that warrant further preclinical testing in scarce patient-derived primary cells. During the case study in ovarian cancer patients, we investigated (i) the relative performance of various ensemble learning algorithms for drug response prediction, (ii) the use of matched single-cell RNA-sequencing data to deconvolute cell population-specific transcriptome profiles from bulk RNA-seq data, (iii) and whether multi-patient or patient-specific predictive models lead to better predictive accuracy. The general platform and the comparison results are expected to become useful for future studies that use similar predictive approaches also in other cancer types. Liye He, Daria Bulanova, Jaana Oikkonen, Antti Häkkinen, Kaiyang Zhang, Erdogan Pekcan Erkan, Olli Carpén, Titta Joutsiniemi, Sakari Hietanen, Johanna Hynninen, Kaisa Huhtinen, Sampsa Hautaniemi, Anna Vähärautio, Jing Tang 0002, Krister Wennerberg, Tero Aittokallio |
Briefings Bioinform. | 18 |
| 2021 | Modeling drug combination effects via latent tensor reconstructionabstractMOTIVATION: Combination therapies have emerged as a powerful treatment modality to overcome drug resistance and improve treatment efficacy. However, the number of possible drug combinations increases very rapidly with the number of individual drugs in consideration, which makes the comprehensive experimental screening infeasible in practice. Machine-learning models offer time- and cost-efficient means to aid this process by prioritizing the most effective drug combinations for further pre-clinical and clinical validation. However, the complexity of the underlying interaction patterns across multiple drug doses and in different cellular contexts poses challenges to the predictive modeling of drug combination effects. RESULTS: We introduce comboLTR, highly time-efficient method for learning complex, non-linear target functions for describing the responses of therapeutic agent combinations in various doses and cancer cell-contexts. The method is based on a polynomial regression via powerful latent tensor reconstruction. It uses a combination of recommender system-style features indexing the data tensor of response values in different contexts, and chemical and multi-omics features as inputs. We demonstrate that comboLTR outperforms state-of-the-art methods in terms of predictive performance and running time, and produces highly accurate results even in the challenging and practical inference scenario where full dose-response matrices are predicted for completely new drug combinations with no available combination and monotherapy response measurements in any training cell line. AVAILABILITY AND IMPLEMENTATION: comboLTR code is available at https://github.com/aalto-ics-kepaco/ComboLTR. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Tianduanyi Wang, Sándor Szedmák, Haishan Wang 0001, Tero Aittokallio, Tapio Pahikkala, Anna Cichonska, Juho Rousu |
Bioinform. | 4 |
| 2021 | Characterizing the Quality of Insight by Interactions: A Case StudyabstractUnderstanding the quality of insight has become increasingly important with the trend of allowing users to post comments during visual exploration, yet approaches for qualifying insight are rare. This article presents a case study to investigate the possibility of characterizing the quality of insight via the interactions performed. To do this, we devised the interaction of a visualization tool-MediSyn-for insight generation. MediSyn supports five types of interactions: selecting, connecting, elaborating, exploring, and sharing. We evaluated MediSyn with 14 participants by allowing them to freely explore the data and generate insights. We then extracted seven interaction patterns from their interaction logs and correlated the patterns to four aspects of insight quality. The results show the possibility of qualifying insights via interactions. Among other findings, exploration actions can lead to unexpected insights; the drill-down pattern tends to increase the domain values of insights. A qualitative analysis shows that using domain knowledge to guide exploration can positively affect the domain value of derived insights. We discuss the study's implications, lessons learned, and future research opportunities. Chen He 0003, Luana Micallef, Liye He, Gopal Peddinti, Tero Aittokallio, Giulio Jacucci |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2020 | Interactive visual analysis of drug-target interaction networks using Drug Target Profiler, with applications to precision medicine and drug repurposingabstractKnowledge of the full target space of drugs (or drug-like compounds) provides important insights into the potential therapeutic use of the agents to modulate or avoid their various on- and off-targets in drug discovery and precision medicine. However, there is a lack of consolidated databases and associated data exploration tools that allow for systematic profiling of drug target-binding potencies of both approved and investigational agents using a network-centric approach. We recently initiated a community-driven platform, Drug Target Commons (DTC), which is an open-data crowdsourcing platform designed to improve the management, reproducibility and extended use of compound-target bioactivity data for drug discovery and repurposing, as well as target identification applications. In this work, we demonstrate an integrated use of the rich bioactivity data from DTC and related drug databases using Drug Target Profiler (DTP), an open-source software and web tool for interactive exploration of drug-target interaction networks. DTP was designed for network-centric modeling of mode-of-action of multi-targeting anticancer compounds, especially for precision oncology applications. DTP enables users to construct an interaction network based on integrated bioactivity data across selected chemical compounds and their protein targets, further customizable using various visualization and filtering options, as well as cross-links to several drug and protein databases to provide comprehensive information of the network nodes and interactions. We demonstrate here the operation of the DTP tool and its unique features by several use cases related to both drug discovery and drug repurposing applications, using examples of anticancer drugs with shared target profiles. DTP is freely accessible at http://drugtargetprofiler.fimm.fi/. ZiaurRehman Tanoli, Zaid Alam, Aleksandr Ianevski, Krister Wennerberg, Markus Vähä-Koskela, Tero Aittokallio |
Briefings Bioinform. | 6 |
| 2020 | SynergyFinder: a web application for analyzing drug combination dose-response matrix dataabstractBioinformatics (2017), 33, 2017, 2413–2415, doi: 10.1093/bioinformatics/btx162 The following funding source was inadvertently omitted from the above article: European Research Council (ERC) starting grant DrugComb (Informatics approaches for the rationale selection of personaliszed cancer drug combinations) [No. 716063] This has now been added. Aleksandr Ianevski, Liye He, Tero Aittokallio, Jing Tang 0002 |
Bioinform. | 3 |
| 2020 | Breeze: an integrated quality control and data analysis application for high-throughput drug screeningabstractSUMMARY: High-throughput screening (HTS) enables systematic testing of thousands of chemical compounds for potential use as investigational and therapeutic agents. HTS experiments are often conducted in multi-well plates that inherently bear technical and experimental sources of error. Thus, HTS data processing requires the use of robust quality control procedures before analysis and interpretation. Here, we have implemented an open-source analysis application, Breeze, an integrated quality control and data analysis application for HTS data. Furthermore, Breeze enables a reliable way to identify individual drug sensitivity and resistance patterns in cell lines or patient-derived samples for functional precision medicine applications. The Breeze application provides a complete solution for data quality assessment, dose-response curve fitting and quantification of the drug responses along with interactive visualization of the results. AVAILABILITY AND IMPLEMENTATION: The Breeze application with video tutorial and technical documentation is accessible at https://breeze.fimm.fi; the R source code is publicly available at https://github.com/potdarswapnil/Breeze under GNU General Public License v3.0. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Potdar Swapnil, Aleksandr Ianevski, John Mpindi, Dmitrii Bychkov, Clément Fiere, Philipp Ianevski, Bhagwan Yadav, Krister Wennerberg, Tero Aittokallio, Olli-P. Kallioniemi, Saarela Jani, Päivi Östling |
Bioinform. | 9 |
| 2020 | SynToxProfiler: An interactive analysis of drug combination synergy, toxicity and efficacyabstractDrug combinations are becoming a standard treatment of many complex diseases due to their capability to overcome resistance to monotherapy. In the current preclinical drug combination screening, the top combinations for further study are often selected based on synergy alone, without considering the combination efficacy and toxicity effects, even though these are critical determinants for the clinical success of a therapy. To promote the prioritization of drug combinations based on integrated analysis of synergy, efficacy and toxicity profiles, we implemented a web-based open-source tool, SynToxProfiler (Synergy-Toxicity-Profiler). When applied to 20 anti-cancer drug combinations tested both in healthy control and T-cell prolymphocytic leukemia (T-PLL) patient cells, as well as to 77 anti-viral drug pairs tested in Huh7 liver cell line with and without Ebola virus infection, SynToxProfiler prioritized as top hits those synergistic drug pairs that showed higher selective efficacy (difference between efficacy and toxicity), which offers an improved likelihood for clinical success. Aleksandr Ianevski, Sanna Timonen, Alexander Kononov, Tero Aittokallio, Anil K. Giri |
PLoS Comput. Biol. | 4 |
| 2020 | Multiobjective optimization identifies cancer-selective combination therapiesabstractCombinatorial therapies are required to treat patients with advanced cancers that have become resistant to monotherapies through rewiring of redundant pathways. Due to a massive number of potential drug combinations, there is a need for systematic approaches to identify safe and effective combinations for each patient, using cost-effective methods. Here, we developed an exact multiobjective optimization method for identifying pairwise or higher-order combinations that show maximal cancer-selectivity. The prioritization of patient-specific combinations is based on Pareto-optimization in the search space spanned by the therapeutic and nonselective effects of combinations. We demonstrate the performance of the method in the context of BRAF-V600E melanoma treatment, where the optimal solutions predicted a number of co-inhibition partners for vemurafenib, a selective BRAF-V600E inhibitor, approved for advanced melanoma. We experimentally validated many of the predictions in BRAF-V600E melanoma cell line, and the results suggest that one can improve selective inhibition of BRAF-V600E melanoma cells by combinatorial targeting of MAPK/ERK and other compensatory pathways using pairwise and third-order drug combinations. Our mechanism-agnostic optimization method is widely applicable to various cancer types, and it takes as input only measurements of a subset of pairwise drug combinations, without requiring target information or genomic profiles. Such data-driven approaches may become useful for functional precision oncology applications that go beyond the cancer genetic dependency paradigm to optimize cancer-selective combinatorial treatments. Otto Pulkkinen, Prson Gautam, Ville Mustonen, Tero Aittokallio |
PLoS Comput. Biol. | 4 |
| 2018 | Global proteomics profiling improves drug sensitivity prediction: results from a multi-omics, pan-cancer modeling approachabstractMotivation: Proteomics profiling is increasingly being used for molecular stratification of cancer patients and cell-line panels. However, systematic assessment of the predictive power of large-scale proteomic technologies across various drug classes and cancer types is currently lacking. To that end, we carried out the first pan-cancer, multi-omics comparative analysis of the relative performance of two proteomic technologies, targeted reverse phase protein array (RPPA) and global mass spectrometry (MS), in terms of their accuracy for predicting the sensitivity of cancer cells to both cytotoxic chemotherapeutics and molecularly targeted anticancer compounds. Results: Our results in two cell-line panels demonstrate how MS profiling improves drug response predictions beyond that of the RPPA or the other omics profiles when used alone. However, frequent missing MS data values complicate its use in predictive modeling and required additional filtering, such as focusing on completely measured or known oncoproteins, to obtain maximal predictive performance. Rather strikingly, the two proteomics profiles provided complementary predictive signal both for the cytotoxic and targeted compounds. Further, information about the cellular-abundance of primary target proteins was found critical for predicting the response of targeted compounds, although the non-target features also contributed significantly to the predictive power. The clinical relevance of the selected protein markers was confirmed in cancer patient data. These results provide novel insights into the relative performance and optimal use of the widely applied proteomic technologies, MS and RPPA, which should prove useful in translational applications, such as defining the best combination of omics technologies and marker panels for understanding and predicting drug sensitivities in cancer patients. Availability and implementation: Processed datasets, R as well as Matlab implementations of the methods are available at https://github.com/mehr-een/bemkl-rbps. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Mehreen Ali, Suleiman A. Khan, Krister Wennerberg, Tero Aittokallio |
Bioinform. | 4 |
| 2018 | Learning with multiple pairwise kernels for drug bioactivity predictionabstractMotivation: Many inference problems in bioinformatics, including drug bioactivity prediction, can be formulated as pairwise learning problems, in which one is interested in making predictions for pairs of objects, e.g. drugs and their targets. Kernel-based approaches have emerged as powerful tools for solving problems of that kind, and especially multiple kernel learning (MKL) offers promising benefits as it enables integrating various types of complex biomedical information sources in the form of kernels, along with learning their importance for the prediction task. However, the immense size of pairwise kernel spaces remains a major bottleneck, making the existing MKL algorithms computationally infeasible even for small number of input pairs. Results: We introduce pairwiseMKL, the first method for time- and memory-efficient learning with multiple pairwise kernels. pairwiseMKL first determines the mixture weights of the input pairwise kernels, and then learns the pairwise prediction function. Both steps are performed efficiently without explicit computation of the massive pairwise matrices, therefore making the method applicable to solving large pairwise learning problems. We demonstrate the performance of pairwiseMKL in two related tasks of quantitative drug bioactivity prediction using up to 167 995 bioactivity measurements and 3120 pairwise kernels: (i) prediction of anticancer efficacy of drug compounds across a large panel of cancer cell lines; and (ii) prediction of target profiles of anticancer compounds across their kinome-wide target spaces. We show that pairwiseMKL provides accurate predictions using sparse solutions in terms of selected kernels, and therefore it automatically identifies also data sources relevant for the prediction problem. Availability and implementation: Code is available at https://github.com/aalto-ics-kepaco. Supplementary information: Supplementary data are available at Bioinformatics online. Anna Cichonska, Tapio Pahikkala, Sándor Szedmák, Heli Julkunen, Antti Airola, Markus Heinonen, Tero Aittokallio, Juho Rousu |
Bioinform. | 7 |
| 2018 | ePCR: an R-package for survival and time-to-event prediction in advanced prostate cancer, applied to real-world patient cohortsabstractMotivation: Prognostic models are widely used in clinical decision-making, such as risk stratification and tailoring treatment strategies, with the aim to improve patient outcomes while reducing overall healthcare costs. While prognostic models have been adopted into clinical use, benchmarking their performance has been difficult due to lack of open clinical datasets. The recent DREAM 9.5 Prostate Cancer Challenge carried out an extensive benchmarking of prognostic models for metastatic Castration-Resistant Prostate Cancer (mCRPC), based on multiple cohorts of open clinical trial data. Results: We make available an open-source implementation of the top-performing model, ePCR, along with an extended toolbox for its further re-use and development, and demonstrate how to best apply the implemented model to real-world data cohorts of advanced prostate cancer patients. Availability and implementation: The open-source R-package ePCR and its reference documentation are available at the Central R Archive Network (CRAN): https://CRAN.R-project.org/package=ePCR. R-vignette provides step-by-step examples for the ePCR usage. Supplementary information: Supplementary data are available at Bioinformatics online. Teemu D. Laajala, Mika Murtojärvi, Arho Virkki, Tero Aittokallio |
Bioinform. | 4 |
| 2017 | Systematic identification of feature combinations for predicting drug response with Bayesian multi-view multi-task linear regressionabstractMOTIVATION: A prime challenge in precision cancer medicine is to identify genomic and molecular features that are predictive of drug treatment responses in cancer cells. Although there are several computational models for accurate drug response prediction, these often lack the ability to infer which feature combinations are the most predictive, particularly for high-dimensional molecular datasets. As increasing amounts of diverse genome-wide data sources are becoming available, there is a need to build new computational models that can effectively combine these data sources and identify maximally predictive feature combinations. RESULTS: We present a novel approach that leverages on systematic integration of data sources to identify response predictive features of multiple drugs. To solve the modeling task we implement a Bayesian linear regression method. To further improve the usefulness of the proposed model, we exploit the known human cancer kinome for identifying biologically relevant feature combinations. In case studies with a synthetic dataset and two publicly available cancer cell line datasets, we demonstrate the improved accuracy of our method compared to the widely used approaches in drug response analysis. As key examples, our model identifies meaningful combinations of features for the well known EGFR, ALK, PLK and PDGFR inhibitors. AVAILABILITY AND IMPLEMENTATION: The source code of the method is available at https://github.com/suleimank/mvlr . CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Muhammad Ammad-ud-din, Suleiman A. Khan, Krister Wennerberg, Tero Aittokallio |
Bioinform. | 4 |
| 2017 | SynergyFinder: a web application for analyzing drug combination dose-response matrix dataabstractSUMMARY: Rational design of drug combinations has become a promising strategy to tackle the drug sensitivity and resistance problem in cancer treatment. To systematically evaluate the pre-clinical significance of pairwise drug combinations, functional screening assays that probe combination effects in a dose-response matrix assay are commonly used. To facilitate the analysis of such drug combination experiments, we implemented a web application that uses key functions of R-package SynergyFinder, and provides not only the flexibility of using multiple synergy scoring models, but also a user-friendly interface for visualizing the drug combination landscapes in an interactive manner. AVAILABILITY AND IMPLEMENTATION: The SynergyFinder web application is freely accessible at https://synergyfinder.fimm.fi ; The R-package and its source-code are freely available at http://bioconductor.org/packages/release/bioc/html/synergyfinder.html . CONTACT: [email protected]. Aleksandr Ianevski, Liye He, Tero Aittokallio, Jing Tang 0002 |
Bioinform. | 3 |
| 2017 | MediSyn: uncertainty-aware visualization of multiple biomedical datasets to support drug treatment selectionabstractBACKGROUND: Dispersed biomedical databases limit user exploration to generate structured knowledge. Linked Data unifies data structures and makes the dispersed data easy to search across resources, but it lacks supporting human cognition to achieve insights. In addition, potential errors in the data are difficult to detect in their free formats. Devising a visualization that synthesizes multiple sources in such a way that links between data sources are transparent, and uncertainties, such as data conflicts, are salient is challenging. RESULTS: To investigate the requirements and challenges of uncertainty-aware visualizations of linked data, we developed MediSyn, a system that synthesizes medical datasets to support drug treatment selection. It uses a matrix-based layout to visually link drugs, targets (e.g., mutations), and tumor types. Data uncertainties are salient in MediSyn; for example, (i) missing data are exposed in the matrix view of drug-target relations; (ii) inconsistencies between datasets are shown via overlaid layers; and (iii) data credibility is conveyed through links to data provenance. CONCLUSIONS: Through the synthesis of two manually curated datasets, cancer treatment biomarkers and drug-target bioactivities, a use case shows how MediSyn effectively supports the discovery of drug-repurposing opportunities. A study with six domain experts indicated that MediSyn benefited the drug selection and data inconsistency discovery. Though linked publication sources supported user exploration for further information, the causes of inconsistencies were not easy to find. Additionally, MediSyn could embrace more patient data to increase its informativeness. We derive design implications from the findings. Chen He 0003, Luana Micallef, ZiaurRehman Tanoli, Samuel Kaski, Tero Aittokallio, Giulio Jacucci |
BMC Bioinform. | 5 |
| 2017 | Computational-experimental approach to drug-target interaction mapping: A case study on kinase inhibitorsabstractDue to relatively high costs and labor required for experimental profiling of the full target space of chemical compounds, various machine learning models have been proposed as cost-effective means to advance this process in terms of predicting the most potent compound-target interactions for subsequent verification. However, most of the model predictions lack direct experimental validation in the laboratory, making their practical benefits for drug discovery or repurposing applications largely unknown. Here, we therefore introduce and carefully test a systematic computational-experimental framework for the prediction and pre-clinical verification of drug-target interactions using a well-established kernel-based regression algorithm as the prediction model. To evaluate its performance, we first predicted unmeasured binding affinities in a large-scale kinase inhibitor profiling study, and then experimentally tested 100 compound-kinase pairs. The relatively high correlation of 0.77 (p < 0.0001) between the predicted and measured bioactivities supports the potential of the model for filling the experimental gaps in existing compound-target interaction maps. Further, we subjected the model to a more challenging task of predicting target interactions for such a new candidate drug compound that lacks prior binding profile information. As a specific case study, we used tivozanib, an investigational VEGF receptor inhibitor with currently unknown off-target profile. Among 7 kinases with high predicted affinity, we experimentally validated 4 new off-targets of tivozanib, namely the Src-family kinases FRK and FYN A, the non-receptor tyrosine kinase ABL1, and the serine/threonine kinase SLK. Our sub-sequent experimental validation protocol effectively avoids any possible information leakage between the training and validation data, and therefore enables rigorous model validation for practical applications. These results demonstrate that the kernel-based modeling approach offers practical benefits for probing novel insights into the mode of action of investigational compounds, and for the identification of new target selectivities for drug repurposing applications. Anna Cichonska, Balaguru Ravikumar, Elina Parri, Sanna Timonen, Tapio Pahikkala, Antti Airola, Krister Wennerberg, Juho Rousu, Tero Aittokallio |
PLoS Comput. Biol. | 9 |
| 2016 | Drug response prediction by inferring pathway-response associations with kernelized Bayesian matrix factorizationabstractMOTIVATION: A key goal of computational personalized medicine is to systematically utilize genomic and other molecular features of samples to predict drug responses for a previously unseen sample. Such predictions are valuable for developing hypotheses for selecting therapies tailored for individual patients. This is especially valuable in oncology, where molecular and genetic heterogeneity of the cells has a major impact on the response. However, the prediction task is extremely challenging, raising the need for methods that can effectively model and predict drug responses. RESULTS: In this study, we propose a novel formulation of multi-task matrix factorization that allows selective data integration for predicting drug responses. To solve the modeling task, we extend the state-of-the-art kernelized Bayesian matrix factorization (KBMF) method with component-wise multiple kernel learning. In addition, our approach exploits the known pathway information in a novel and biologically meaningful fashion to learn the drug response associations. Our method quantitatively outperforms the state of the art on predicting drug responses in two publicly available cancer datasets as well as on a synthetic dataset. In addition, we validated our model predictions with lab experiments using an in-house cancer cell line panel. We finally show the practical applicability of the proposed method by utilizing prior knowledge to infer pathway-drug response associations, opening up the opportunity for elucidating drug action mechanisms. We demonstrate that pathway-response associations can be learned by the proposed model for the well-known EGFR and MEK inhibitors. AVAILABILITY AND IMPLEMENTATION: The source code implementing the method is available at http://research.cs.aalto.fi/pml/software/cwkbmf/ CONTACTS: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Muhammad Ammad-ud-din, Suleiman A. Khan, Disha Malani, Astrid Murumägi, Olli-P. Kallioniemi, Tero Aittokallio, Samuel Kaski |
Bioinform. | 6 |
| 2015 | Toward more realistic drug-target interaction predictionsabstractA number of supervised machine learning models have recently been introduced for the prediction of drug-target interactions based on chemical structure and genomic sequence information. Although these models could offer improved means for many network pharmacology applications, such as repositioning of drugs for new therapeutic uses, the prediction models are often being constructed and evaluated under overly simplified settings that do not reflect the real-life problem in practical applications. Using quantitative drug-target bioactivity assays for kinase inhibitors, as well as a popular benchmarking data set of binary drug-target interactions for enzyme, ion channel, nuclear receptor and G protein-coupled receptor targets, we illustrate here the effects of four factors that may lead to dramatic differences in the prediction results: (i) problem formulation (standard binary classification or more realistic regression formulation), (ii) evaluation data set (drug and target families in the application use case), (iii) evaluation procedure (simple or nested cross-validation) and (iv) experimental setting (whether training and test sets share common drugs and targets, only drugs or targets or neither). Each of these factors should be taken into consideration to avoid reporting overoptimistic drug-target interaction prediction results. We also suggest guidelines on how to make the supervised drug-target interaction prediction studies more realistic in terms of such model formulations and evaluation setups that better address the inherent complexity of the prediction task in the practical applications, as well as novel benchmarking data sets that capture the continuous nature of the drug-target interactions for kinase inhibitors. Tapio Pahikkala, Antti Airola, Sami Pietilä, Sushil Kumar Shakyawar, Agnieszka Szwajda, Jing Tang 0002, Tero Aittokallio |
Briefings Bioinform. | 7 |
| 2015 | TIMMA-R: an R package for predicting synergistic multi-targeted drug combinations in cancer cell lines or patient-derived samplesabstractUNLABELLED: Network pharmacology-based prediction of multi-targeted drug combinations is becoming a promising strategy to improve anticancer efficacy and safety. We developed a logic-based network algorithm, called Target Inhibition Interaction using Maximization and Minimization Averaging (TIMMA), which predicts the effects of drug combinations based on their binary drug-target interactions and single-drug sensitivity profiles in a given cancer sample. Here, we report the R implementation of the algorithm (TIMMA-R), which is much faster than the original MATLAB code. The major extensions include modeling of multiclass drug-target profiles and network visualization. We also show that the TIMMA-R predictions are robust to the intrinsic noise in the experimental data, thus making it a promising high-throughput tool to prioritize drug combinations in various cancer types for follow-up experimentation or clinical applications. AVAILABILITY AND IMPLEMENTATION: TIMMA-R source code is freely available at http://cran.r-project.org/web/packages/timma/. Liye He, Krister Wennerberg, Tero Aittokallio, Jing Tang 0002 |
Bioinform. | 3 |
| 2015 | Impact of normalization methods on high-throughput screening data with high hit rates and drug testing with dose-response dataabstractMOTIVATION: Most data analysis tools for high-throughput screening (HTS) seek to uncover interesting hits for further analysis. They typically assume a low hit rate per plate. Hit rates can be dramatically higher in secondary screening, RNAi screening and in drug sensitivity testing using biologically active drugs. In particular, drug sensitivity testing on primary cells is often based on dose-response experiments, which pose a more stringent requirement for data quality and for intra- and inter-plate variation. Here, we compared common plate normalization and noise-reduction methods, including the B-score and the Loess a local polynomial fit method under high hit-rate scenarios of drug sensitivity testing. We generated simulated 384-well plate HTS datasets, each with 71 plates having a range of 20 (5%) to 160 (42%) hits per plate, with controls placed either at the edge of the plates or in a scattered configuration. RESULTS: We identified 20% (77/384) as the critical hit-rate after which the normalizations started to perform poorly. Results from real drug testing experiments supported this estimation. In particular, the B-score resulted in incorrect normalization of high hit-rate plates, leading to poor data quality, which could be attributed to its dependency on the median polish algorithm. We conclude that a combination of a scattered layout of controls per plate and normalization using a polynomial least squares fit method, such as Loess helps to reduce column, row and edge effects in HTS experiments with high hit-rates and is optimal for generating accurate dose-response curves. CONTACT: [email protected]. AVAILABILITY AND IMPLEMENTATION: Supplementary information: R code and Supplementary data are available at Bioinformatics online. John Mpindi, Potdar Swapnil, Dmitrii Bychkov, Saarela Jani, Khalid Saeed 0002, Krister Wennerberg, Tero Aittokallio, Päivi Östling, Olli-P. Kallioniemi |
Bioinform. | 7 |
| 2014 | A Two-Step Learning Approach for Solving Full and Almost Full Cold Start Problems in Dyadic Prediction
Tapio Pahikkala, Michiel Stock, Antti Airola, Tero Aittokallio, Bernard De Baets, Willem Waegeman |
ECML/PKDD (2) | 4 |
| 2013 | Target Inhibition Networks: Predicting Selective Combinations of Druggable Targets to Block Cancer Survival PathwaysabstractA recent trend in drug development is to identify drug combinations or multi-target agents that effectively modify multiple nodes of disease-associated networks. Such polypharmacological effects may reduce the risk of emerging drug resistance by means of attacking the disease networks through synergistic and synthetic lethal interactions. However, due to the exponentially increasing number of potential drug and target combinations, systematic approaches are needed for prioritizing the most potent multi-target alternatives on a global network level. We took a functional systems pharmacology approach toward the identification of selective target combinations for specific cancer cells by combining large-scale screening data on drug treatment efficacies and drug-target binding affinities. Our model-based prediction approach, named TIMMA, takes advantage of the polypharmacological effects of drugs and infers combinatorial drug efficacies through system-level target inhibition networks. Case studies in MCF-7 and MDA-MB-231 breast cancer and BxPC-3 pancreatic cancer cells demonstrated how the target inhibition modeling allows systematic exploration of functional interactions between drugs and their targets to maximally inhibit multiple survival pathways in a given cancer type. The TIMMA prediction results were experimentally validated by means of systematic siRNA-mediated silencing of the selected targets and their pairwise combinations, showing increased ability to identify not only such druggable kinase targets that are essential for cancer survival either individually or in combination, but also synergistic interactions indicative of non-additive drug efficacies. These system-level analyses were enabled by a novel model construction method utilizing maximization and minimization rules, as well as a model selection algorithm based on sequential forward floating search. Compared with an existing computational solution, TIMMA showed both enhanced prediction accuracies in cross validation as well as significant reduction in computation times. Such cost-effective computational-experimental design strategies have the potential to greatly speed-up the drug testing efforts by prioritizing those interventions and interactions warranting further study in individual cancer cases. Jing Tang 0002, Leena Karhinen, Agnieszka Szwajda, Bhagwan Yadav, Krister Wennerberg, Tero Aittokallio |
PLoS Comput. Biol. | 7 |
| 2011 | Probabilistic Analysis of Probe Reliability in Differential Gene Expression Studies with Short Oligonucleotide ArraysabstractProbe defects are a major source of noise in gene expression studies. While existing approaches detect noisy probes based on external information such as genomic alignments, we introduce and validate a targeted probabilistic method for analyzing probe reliability directly from expression data and independently of the noise source. This provides insights into the various sources of probe-level noise and gives tools to guide probe design. Leo Lahti, Laura Elo, Tero Aittokallio, Samuel Kaski |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2010 | Dealing with missing values in large-scale studies: microarray data imputation and beyondabstractHigh-throughput biotechnologies, such as gene expression microarrays or mass-spectrometry-based proteomic assays, suffer from frequent missing values due to various experimental reasons. Since the missing data points can hinder downstream analyses, there exists a wide variety of ways in which to deal with missing values in large-scale data sets. Nowadays, it has become routine to estimate (or impute) the missing values prior to the actual data analysis. After nearly a decade since the publication of the first missing value imputation methods for gene expression microarray data, new imputation approaches are still being developed at an increasing rate. However, what is lagging behind is a systematic and objective evaluation of the strengths and weaknesses of the different approaches when faced with different types of data sets and experimental questions. In this review, the present strategies for missing value imputation and the measures for evaluating their performance are described. The imputation methods are first reviewed in the context of gene expression microarray data, since most of the methods have been developed for estimating gene expression levels; then, we turn to other large-scale data sets that also suffer from the problems posed by missing values, together with pointers to possible imputation approaches in these settings. Along with a description of the basic principles behind the different imputation approaches, the review tries to provide practical guidance for the users of high-throughput technologies on how to choose the imputation tool for their data and questions, and some additional research directions for the developers of imputation methodologies. Tero Aittokallio |
Briefings Bioinform. | 1 |
| 2009 | Optimized detection of differential expression in global profiling experiments: case studies in clinical transcriptomic and quantitative proteomic datasetsabstractIdentification of reliable molecular markers that show differential expression between distinct groups of samples has remained a fundamental research problem in many large-scale profiling studies, such as those based on DNA microarray or mass-spectrometry technologies. Despite the availability of a wide spectrum of statistical procedures, the users of the high-throughput platforms are still facing the crucial challenge of deciding which test statistic is best adapted to the intrinsic properties of their own datasets. To meet this challenge, we recently introduced an adaptive procedure, named ROTS (Reproducibility-Optimized Test Statistic), which learns an optimal statistic directly from the given data, and whose relative benefits have previously been shown in comparison with state-of-the-art procedures for detecting differential expression. Using gene expression microarray and mass-spectrometry (MS)-based protein expression datasets as case studies, we illustrate here the practical usage and advantages of ROTS toward detecting reliable marker lists in clinical transcriptomic and proteomic studies. In a public leukemia microarray dataset, the procedure could improve the sensitivity of the gene marker lists detected with high specificity. When applied to a recent LC-MS dataset, involving plasma samples from severe burn patients, the procedure could identify several peptide markers that remained undetected in the conventional analysis, thus demonstrating the effectiveness of ROTS also for global quantitative proteomic studies. To promote its widespread usage, we have made freely available efficient implementations of ROTS, which are easily accessible either as a stand-alone R-package or as integrated in the open-source data analysis software Chipster. Laura Elo, Jukka Hiissa, Jarno Tuimala, Aleksi Kallio, Eija Korpelainen, Tero Aittokallio |
Briefings Bioinform. | 6 |
| 2009 | Genoscape: a Cytoscape plug-in to automate the retrieval and integration of gene expression data and molecular networksabstractAbstract Summary: Genoscape is an open-source Cytoscape plug-in that visually integrates gene expression data sets from GenoScript, a transcriptomic database, and KEGG pathways into Cytoscape networks. The generated visualisation highlights gene expression changes and their statistical significance. The plug-in also allows one to browse GenoScript or import transcriptomic data from other sources through tab-separated text files. Genoscape has been successfully used by researchers to investigate the results of gene expression profiling experiments. Availability: Genoscape is an open-source software freely available from the Genoscape webpage (http://www.pasteur.fr/recherche/unites/Gim/genoscape/). Installation instructions and tutorial can also be found at this URL. Contact: [email protected]; [email protected] Supplementary information: Supplementary data are available at Bioinformatics online. Mathieu Clément-Ziza, Christophe Malabat, Ivan Moszer, Tero Aittokallio, Catherine Letondal, Sandrine Moreira |
Bioinform. | 5 |
| 2008 | Overnight features of transcutaneous carbon dioxide measurement as predictors of metabolic status
Arho Virkki, Olli Polo, Tarja Saaresranta, Anne Laapotti-Salo, Mats Gyllenberg, Tero Aittokallio |
Artif. Intell. Medicine | 6 |
| 2008 | Model-based prediction of sequence alignment qualityabstractMOTIVATION: Multiple sequence alignment (MSA) is an essential prerequisite for many sequence analysis methods and valuable tool itself for describing relationships between protein sequences. Since the success of the sequence analysis is highly dependent on the reliability of alignments, measures for assessing the quality of alignments are highly requisite. RESULTS: We present a statistical model-based alignment quality score. Unlike other quality scores, it does not require several parallel alignments for the same set of sequences or additional structural information. Our quality score is based on measuring the conservation level of reference alignments in Homstrad. Reference sequences were realigned with the Mafft, Muscle and Probcons alignment programs, and a sum-of-pairs (SP) score was used to measure the quality of the realignments. Statistical modelling of the SP score as a function of conservation level and other alignment characteristics makes it possible to predict the SP score for any global MSA. The predicted SP scores are highly correlated with the correct SP scores, when tested on the Homstrad and SABmark databases. The results are comparable to that of multiple overlap score (MOS) and better than those of normalized mean distance (NorMD) and normalized iRMSD (NiRMSD) alignment quality criteria. Furthermore, the predicted SP score is able to detect alignments with badly aligned or unrelated sequences. AVAILABILITY: The method is freely available at http://www.mtt.fi/AlignmentQuality/. Virpi Ahola, Tero Aittokallio, Mauno Vihinen, Esa Uusipaikka |
Bioinform. | 2 |
| 2008 | Missing value imputation improves clustering and interpretation of gene expression microarray dataabstractBACKGROUND: Missing values frequently pose problems in gene expression microarray experiments as they can hinder downstream analysis of the datasets. While several missing value imputation approaches are available to the microarray users and new ones are constantly being developed, there is no general consensus on how to choose between the different methods since their performance seems to vary drastically depending on the dataset being used. RESULTS: We show that this discrepancy can mostly be attributed to the way in which imputation methods have traditionally been developed and evaluated. By comparing a number of advanced imputation methods on recent microarray datasets, we show that even when there are marked differences in the measurement-level imputation accuracies across the datasets, these differences become negligible when the methods are evaluated in terms of how well they can reproduce the original gene clusters or their biological interpretations. Regardless of the evaluation approach, however, imputation always gave better results than ignoring missing data points or replacing them with zeros or average values, emphasizing the continued importance of using more advanced imputation methods. CONCLUSION: The results demonstrate that, while missing values are still severely complicating microarray data analysis, their impact on the discovery of biologically meaningful gene groups can - up to a certain degree - be reduced by using readily available and relatively fast imputation methods, such as the Bayesian Principal Components Algorithm (BPCA). Johannes Tuikkala, Laura Elo, Olli Nevalainen, Tero Aittokallio |
BMC Bioinform. | 4 |
| 2008 | Reproducibility-Optimized Test Statistic for Ranking Genes in Microarray StudiesabstractA principal goal of microarray studies is to identify the genes showing differential expression under distinct conditions. In such studies, the selection of an optimal test statistic is a crucial challenge, which depends on the type and amount of data under analysis. While previous studies on simulated or spike-in datasets do not provide practical guidance on how to choose the best method for a given real dataset, we introduce an enhanced reproducibility-optimization procedure, which enables the selection of a suitable gene- anking statistic directly from the data. In comparison with existing ranking methods, the reproducibilityoptimized statistic shows good performance consistently under various simulated conditions and on Affymetrix spike-in dataset. Further, the feasibility of the novel statistic is confirmed in a practical research setting using data from an in-house cDNA microarray study of asthma-related gene expression changes. These results suggest that the procedure facilitates the selection of an appropriate test statistic for a given dataset without relying on a priori assumptions, which may bias the findings and their interpretation. Moreover, the general reproducibilityoptimization procedure is not limited to detecting differential expression only but could be extended to a wide range of other applications as well. Laura Elo, Sanna Filen, Riitta Lahesmaa, Tero Aittokallio |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2007 | Systematic construction of gene coexpression networks with applications to human T helper cell differentiation processabstractMOTIVATION: Coexpression networks have recently emerged as a novel holistic approach to microarray data analysis and interpretation. Choosing an appropriate cutoff threshold, above which a gene-gene interaction is considered as relevant, is a critical task in most network-centric applications, especially when two or more networks are being compared. RESULTS: We demonstrate that the performance of traditional approaches, which are based on a pre-defined cutoff or significance level, can vary drastically depending on the type of data and application. Therefore, we introduce a systematic procedure for estimating a cutoff threshold of coexpression networks directly from their topological properties. Both synthetic and real datasets show clear benefits of our data-driven approach under various practical circumstances. In particular, the procedure provides a robust estimate of individual degree distributions, even from multiple microarray studies performed with different array platforms or experimental designs, which can be used to discriminate the corresponding phenotypes. Application to human T helper cell differentiation process provides useful insights into the components and interactions controlling this process, many of which would have remained unidentified on the basis of expression change alone. Moreover, several human-mouse orthologs showed conserved topological changes in both systems, suggesting their potential importance in the differentiation process. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Laura Elo, Henna Järvenpää, Matej Oresic, Riitta Lahesmaa, Tero Aittokallio |
Bioinform. | 5 |
| 2007 | GOlorize: a Cytoscape plug-in for network visualization with Gene Ontology-based layout and coloringabstractUNLABELLED: We have implemented a graph layout algorithm that exposes Gene Ontology (GO) class structure on the network nodes. It can be used in conjunction with BiNGO plug-in to Cytoscape, which finds the GO categories over-represented in a given network. Our plug-in, named GOlorize, first highlights the class members with category-specific color-coding and then constructs an enhanced visualization of the network using a class-directed layout algorithm. AVAILABILITY: http://www.cytoscape.org/plugins2.php. SUPPLEMENTARY INFORMATION: Installation instructions and tutorial at http://www.cytoscape.org/plugins/GOlorize/GOlorizeUserGuide.pdf. Olivier Garcia, Cosmin Saveanu, Melissa S. Cline, Micheline Fromont-Racine, Alain Jacquier, Benno Schwikowski, Tero Aittokallio |
Bioinform. | 7 |
| 2006 | Graph-based methods for analysing networks in cell biologyabstractAvailability of large-scale experimental data for cell biology is enabling computational methods to systematically model the behaviour of cellular networks. This review surveys the recent advances in the field of graph-driven methods for analysing complex cellular networks. The methods are outlined on three levels of increasing complexity, ranging from methods that can characterize global or local structural properties of networks to methods that can detect groups of interconnected nodes, called motifs or clusters, potentially involved in common elementary biological functions. We also briefly summarize recent approaches to data integration and network inference through graph-based formalisms. Finally, we highlight some challenges in the field and offer our personal view of the key future trends and developments in graph-based analysis of large-scale datasets. Tero Aittokallio, Benno Schwikowski |
Briefings Bioinform. | 1 |
| 2006 | Quality classification of tandem mass spectrometry dataabstractUNLABELLED: Peptide identification by tandem mass spectrometry is an important tool in proteomic research. Powerful identification programs exist, such as SEQUEST, ProICAT and Mascot, which can relate experimental spectra to the theoretical ones derived from protein databases, thus removing much of the manual input needed in the identification process. However, the time-consuming validation of the peptide identifications is still the bottleneck of many proteomic studies. One way to further streamline this process is to remove those spectra that are unlikely to provide a confident or valid peptide identification, and in this way to reduce the labour from the validation phase. RESULTS: We propose a prefiltering scheme for evaluating the quality of spectra before the database search. The spectra are classified into two classes: spectra which contain valuable information for peptide identification and spectra that are not derived from peptides or contain insufficient information for interpretation. The different spectral features developed for the classification are tested on a real-life material originating from human lymphoblast samples and on a standard mixture of 9 proteins, both labelled with the ICAT-reagent. The results show that the prefiltering scheme efficiently separates the two spectra classes. Jussi Salmi, Robert Moulder, Jan-Jonas Filén, Olli Nevalainen, Tuula A. Nyman, Riitta Lahesmaa, Tero Aittokallio |
Bioinform. | 7 |
| 2006 | Improving missing value estimation in microarray data with gene ontologyabstractMOTIVATION: Gene expression microarray experiments produce datasets with frequent missing expression values. Accurate estimation of missing values is an important prerequisite for efficient data analysis as many statistical and machine learning techniques either require a complete dataset or their results are significantly dependent on the quality of such estimates. A limitation of the existing estimation methods for microarray data is that they use no external information but the estimation is based solely on the expression data. We hypothesized that utilizing a priori information on functional similarities available from public databases facilitates the missing value estimation. RESULTS: We investigated whether semantic similarity originating from gene ontology (GO) annotations could improve the selection of relevant genes for missing value estimation. The relative contribution of each information source was automatically estimated from the data using an adaptive weight selection procedure. Our experimental results in yeast cDNA microarray datasets indicated that by considering GO information in the k-nearest neighbor algorithm we can enhance its performance considerably, especially when the number of experimental conditions is small and the percentage of missing values is high. The increase of performance was less evident with a more sophisticated estimation method. We conclude that even a small proportion of annotated genes can provide improvements in data quality significant for the eventual interpretation of the microarray experiments. AVAILABILITY: Java and Matlab codes are available on request from the authors. SUPPLEMENTARY MATERIAL: Available online at http://users.utu.fi/jotatu/GOImpute.html. Johannes Tuikkala, Laura Elo, Olli Nevalainen, Tero Aittokallio |
Bioinform. | 4 |
| 2006 | A statistical score for assessing the quality of multiple sequence alignmentsabstractBACKGROUND: Multiple sequence alignment is the foundation of many important applications in bioinformatics that aim at detecting functionally important regions, predicting protein structures, building phylogenetic trees etc. Although the automatic construction of a multiple sequence alignment for a set of remotely related sequences cause a very challenging and error-prone task, many downstream analyses still rely heavily on the accuracy of the alignments. RESULTS: To address the need for an objective evaluation framework, we introduce a statistical score that assesses the quality of a given multiple sequence alignment. The quality assessment is based on counting the number of significantly conserved positions in the alignment using importance sampling method in conjunction with statistical profile analysis framework. We first evaluate a novel objective function used in the alignment quality score for measuring the positional conservation. The results for the Src homology 2 (SH2) domain, Ras-like proteins, peptidase M13, subtilase and beta-lactamase families demonstrate that the score can distinguish sequence patterns with different degrees of conservation. Secondly, we evaluate the quality of the alignments produced by several widely used multiple sequence alignment programs using a novel alignment quality score and a commonly used sum of pairs method. According to these results, the Mafft strategy L-INS-i outperforms the other methods, although the difference between the Probcons, TCoffee and Muscle is mostly insignificant. The novel alignment quality score provides similar results than the sum of pairs method. CONCLUSION: The results indicate that the proposed statistical score is useful in assessing the quality of multiple sequence alignments. Virpi Ahola, Tero Aittokallio, Mauno Vihinen, Esa Uusipaikka |
BMC Bioinform. | 2 |
| 2003 | Efficient estimation of emission probabilities in profile hidden Markov modelsabstractMOTIVATION: Profile hidden Markov models provide a sensitive method for performing sequence database search and aligning multiple sequences. One of the drawbacks of the hidden Markov model is that the conserved amino acids are not emphasized, but signal and noise are treated equally. For this reason, the number of estimated emission parameters is often enormous. Focusing the analysis on conserved residues only should increase the accuracy of sequence database search. RESULTS: We address this issue with a new method for efficient emission probability (EEP) estimation, in which amino acids are divided into effective and ineffective residues at each conserved alignment position. A practical study with 20 protein families demonstrated that the EEP method is capable of detecting family members from other proteins with sensitivity of 98% and specificity of 99% on the average, even if the number of free emission parameters was decreased to 15% of the original. In the database search for TIM barrel sequences, EEP recognizes the family members nearly as accurately as HMMER or Blast, but the number of false positive sequences was significantly less than that obtained with the other methods. AVAILABILITY: The algorithms written in C language are available on request from the authors. Virpi Ahola, Tero Aittokallio, Esa Uusipaikka, Mauno Vihinen |
Bioinform. | 2 |
| 2003 | Feature learning with a genetic algorithm for fluorescence fingerprinting of plant species
Marius C. Codrea, Tero Aittokallio, Mika Keränen, Esa Tyystjärvi, Olli Nevalainen |
Pattern Recognit. Lett. | 2 |
| 1999 | Classification of Nasal Inspiratory Flow Shapes by Attributed Finite Automata
Tero Aittokallio, Olli Nevalainen, U. Pursiheimo, Tarja Saaresranta, Olli Polo |
Comput. Biomed. Res. | 1 |