EDBT 2026 Demo / reviewers in the wild / expert
Mehmet Gönen
dblp:37/6563
· DBLP profile ↗
41ranked-venue papers
24as first author
4since 2021 · last 2023
0000-0002-2483-075XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 19 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 first-authorDatabases, data management, data science and information retrieval · 2 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
12 papers |
Bioinformatics and computational biology · 98% Computational science and engineering · 2% | |
| Artificial intelligence
7 papers |
Kernel, tree and ensemble methods · 53% Representation and self-supervised learning · 27% Transfer learning and domain adaptation · 20% | |
| Databases, data mining, and information retrieval
2 papers |
Data mining · 61% Data integration and cleaning · 31% Recommender systems · 8% |
Topics — the 28 heaviest of 32, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
cancer genomics |
1.6 | 5 | 2020 | A multitask multiple kernel learning formulation for discriminating early- and late-stage cancers · Bioinform. 2020 An efficient framework to identify key miRNA-mRNA regulatory modules in cancer · Bioinform. 2020 A Multitask Multiple Kernel Learning Algorithm for Survival Analysis with Application to Cancer Biology · ICML 2019 |
Bioinformatics and computational biology
survival analysis |
0.8 | 2 | 2019 | Path2Surv: Pathway/gene set-based survival analysis using multiple kernel learning · Bioinform. 2019 A Multitask Multiple Kernel Learning Algorithm for Survival Analysis with Application to Cancer Biology · ICML 2019 |
Bioinformatics and computational biology › genomics
genomic data analysis |
0.6 | 1 | 2022 | Fast and interpretable genomic data analysis using multiple approximate kernel learning · Bioinform. 2022 |
Machine learning › Kernel, tree and ensemble methods › kernel methods › kernel learning
multiple kernel learning |
0.5 | 5 | 2014 | Bayesian Efficient Multiple Kernel Learning · ICML 2012 Multiple Kernel Learning Algorithms · J. Mach. Learn. Res. 2011 Localized multiple kernel learning · ICML 2008 |
Machine learning › Kernel, tree and ensemble methods
kernel methods |
0.4 | 4 | 2013 | Bayesian Efficient Multiple Kernel Learning · ICML 2012 Multiple Kernel Learning Algorithms · J. Mach. Learn. Res. 2011 Localized multiple kernel learning · ICML 2008 |
Bioinformatics and computational biology › genomics
machine learning for genomics |
0.4 | 1 | 2019 | Path2Surv: Pathway/gene set-based survival analysis using multiple kernel learning · Bioinform. 2019 |
Bioinformatics and computational biology
drug discovery |
0.3 | 2 | 2014 | Drug susceptibility prediction against a panel of drugs using kernelized Bayesian multitask learning · Bioinform. 2014 Predicting drug-target interactions from chemical and genomic kernels using Bayesian matrix factorization · Bioinform. 2012 |
Bioinformatics and computational biology › drug discovery
drug response prediction |
0.3 | 1 | 2017 | Modeling gene-wise dependencies improves the identification of drug response biomarkers in cancer studies · Bioinform. 2017 |
Bioinformatics and computational biology
multi-omics data integration |
0.3 | 1 | 2017 | Modeling gene-wise dependencies improves the identification of drug response biomarkers in cancer studies · Bioinform. 2017 |
Machine learning › Representation and self-supervised learning › matrix factorization
bayesian matrix factorization |
0.2 | 1 | 2014 | Kernelized Bayesian Matrix Factorization · IEEE Trans. Pattern Anal. Mach. Intell. 2014 |
Machine learning › Transfer learning and domain adaptation
domain adaptation |
0.2 | 1 | 2014 | Kernelized Bayesian Transfer Learning · AAAI 2014 |
Machine learning › Transfer learning and domain adaptation › domain adaptation
supervised domain adaptation |
0.2 | 1 | 2014 | Kernelized Bayesian Transfer Learning · AAAI 2014 |
Bioinformatics and computational biology
cancer biology |
0.2 | 1 | 2014 | Localized Data Fusion for Kernel k-Means Clustering with Application to Cancer Biology · NIPS 2014 |
Bioinformatics and computational biology › drug discovery › drug response prediction
drug sensitivity prediction |
0.2 | 1 | 2014 | Drug susceptibility prediction against a panel of drugs using kernelized Bayesian multitask learning · Bioinform. 2014 |
Bioinformatics and computational biology › genomics
pharmacogenomics |
0.2 | 1 | 2014 | Drug susceptibility prediction against a panel of drugs using kernelized Bayesian multitask learning · Bioinform. 2014 |
Data mining
clustering |
0.2 | 1 | 2014 | Localized Data Fusion for Kernel k-Means Clustering with Application to Cancer Biology · NIPS 2014 |
Data integration and cleaning
data fusion |
0.2 | 1 | 2014 | Localized Data Fusion for Kernel k-Means Clustering with Application to Cancer Biology · NIPS 2014 |
Data mining › clustering › kernel clustering
kernel k-means |
0.2 | 1 | 2014 | Localized Data Fusion for Kernel k-Means Clustering with Application to Cancer Biology · NIPS 2014 |
Bioinformatics and computational biology
biomarker discovery |
0.2 | 1 | 2022 | Fast and interpretable genomic data analysis using multiple approximate kernel learning · Bioinform. 2022 |
Machine learning › Representation and self-supervised learning › representation learning
dimensionality reduction |
0.2 | 1 | 2013 | Supervised Multiple Kernel Embedding for Learning Predictive Subspaces · IEEE Trans. Knowl. Data Eng. 2013 |
Machine learning › Kernel, tree and ensemble methods
kernel embedding |
0.2 | 1 | 2013 | Supervised Multiple Kernel Embedding for Learning Predictive Subspaces · IEEE Trans. Knowl. Data Eng. 2013 |
Machine learning › Representation and self-supervised learning
matrix factorization |
0.2 | 1 | 2013 | Kernelized Bayesian Matrix Factorization · ICML (3) 2013 |
Bioinformatics and computational biology › molecular informatics › cheminformatics
chemogenomics |
0.1 | 1 | 2012 | Predicting drug-target interactions from chemical and genomic kernels using Bayesian matrix factorization · Bioinform. 2012 |
Bioinformatics and computational biology › drug discovery
drug-target interaction prediction |
0.1 | 1 | 2012 | Predicting drug-target interactions from chemical and genomic kernels using Bayesian matrix factorization · Bioinform. 2012 |
Computational science and engineering
multi-task learning |
0.1 | 1 | 2020 | A multitask multiple kernel learning formulation for discriminating early- and late-stage cancers · Bioinform. 2020 |
Bioinformatics and computational biology › functional genomics › functional enrichment analysis
gene set enrichment analysis |
0.1 | 1 | 2018 | Discriminating early- and late-stage cancers using multiple kernel learning on gene sets · Bioinform. 2018 |
Bioinformatics and computational biology › systems bioinformatics
pathway analysis |
0.1 | 1 | 2018 | Discriminating early- and late-stage cancers using multiple kernel learning on gene sets · Bioinform. 2018 |
Recommender systems
cold-start recommendation |
0.0 | 1 | 2013 | Kernelized Bayesian Matrix Factorization · ICML (3) 2013 |
Methods — techniques the papers use, named apart from their topics
multiple kernel learning · 2.6variational approximation · 0.7kernel approximation · 0.6group lasso · 0.6regularized factor regression · 0.4linear programming · 0.4cutting-plane algorithm · 0.4co-clustering · 0.4survival random forest · 0.4multi-task learning · 0.4Survival-SVM · 0.4localized fusion · 0.2kernel k-means · 0.2kernel dimensionality reduction · 0.2full-bayesian treatment · 0.2bayesian modeling · 0.2joint optimization · 0.2bayesian inference · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | MOKPE: drug-target interaction prediction via manifold optimization based kernel preserving embeddingabstractBACKGROUND: In many applications of bioinformatics, data stem from distinct heterogeneous sources. One of the well-known examples is the identification of drug-target interactions (DTIs), which is of significant importance in drug discovery. In this paper, we propose a novel framework, manifold optimization based kernel preserving embedding (MOKPE), to efficiently solve the problem of modeling heterogeneous data. Our model projects heterogeneous drug and target data into a unified embedding space by preserving drug-target interactions and drug-drug, target-target similarities simultaneously. RESULTS: We performed ten replications of ten-fold cross validation on four different drug-target interaction network data sets for predicting DTIs for previously unseen drugs. The classification evaluation metrics showed better or comparable performance compared to previous similarity-based state-of-the-art methods. We also evaluated MOKPE on predicting unknown DTIs of a given network. Our implementation of the proposed algorithm in R together with the scripts that replicate the reported experiments is publicly available at https://github.com/ocbinatli/mokpe . Oguz C. Binatli, Mehmet Gönen |
BMC Bioinform. | 2 |
| 2022 | Fast and interpretable genomic data analysis using multiple approximate kernel learningabstractMOTIVATION: Dataset sizes in computational biology have been increased drastically with the help of improved data collection tools and increasing size of patient cohorts. Previous kernel-based machine learning algorithms proposed for increased interpretability started to fail with large sample sizes, owing to their lack of scalability. To overcome this problem, we proposed a fast and efficient multiple kernel learning (MKL) algorithm to be particularly used with large-scale data that integrates kernel approximation and group Lasso formulations into a conjoint model. Our method extracts significant and meaningful information from the genomic data while conjointly learning a model for out-of-sample prediction. It is scalable with increasing sample size by approximating instead of calculating distinct kernel matrices. RESULTS: To test our computational framework, namely, Multiple Approximate Kernel Learning (MAKL), we demonstrated our experiments on three cancer datasets and showed that MAKL is capable to outperform the baseline algorithm while using only a small fraction of the input features. We also reported selection frequencies of approximated kernel matrices associated with feature subsets (i.e. gene sets/pathways), which helps to see their relevance for the given classification task. Our fast and interpretable MKL algorithm producing sparse solutions is promising for computational biology applications considering its scalability and highly correlated structure of genomic datasets, and it can be used to discover new biomarkers and new therapeutic guidelines. AVAILABILITY AND IMPLEMENTATION: MAKL is available at https://github.com/begumbektas/makl together with the scripts that replicate the reported experiments. MAKL is also available as an R package at https://cran.r-project.org/web/packages/MAKL. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ayyüce Begüm Bektas, Çigdem Ak, Mehmet Gönen |
Bioinform. | 3 |
| 2022 | Efficient Multitask Multiple Kernel Learning With Application to Cancer ResearchabstractMultitask multiple kernel learning (MKL) algorithms combine the capabilities of incorporating different data sources into the prediction model and using the data from one task to improve the accuracy on others. However, these methods do not necessarily produce interpretable results. Restricting the solutions to the set of interpretable solutions increases the computational burden of the learning problem significantly, leading to computationally prohibitive run times for some important biomedical applications. That is why we propose a multitask MKL formulation with a clustering of tasks and develop a highly time-efficient solution approach for it. Our solution method is based on the Benders decomposition and treating the clustering problem as finding a given number of tree structures in a graph; hence, it is called the forest formulation. We use our method to discriminate early-stage and late-stage cancers using genomic data and gene sets and compare our algorithm against two other algorithms. The two other algorithms are based on different approaches for linearization of the problem while all algorithms make use of the cutting-plane method. Our results indicate that as the number of tasks and/or the number of desired clusters increase, the forest formulation becomes increasingly favorable in terms of computational performance. Arezou Rahimi, Mehmet Gönen |
IEEE Trans. Cybern. | 2 |
| 2021 | PrognosiT: Pathway/gene set-based tumour volume prediction using multiple kernel learningabstractBACKGROUND: Identification of molecular mechanisms that determine tumour progression in cancer patients is a prerequisite for developing new disease treatment guidelines. Even though the predictive performance of current machine learning models is promising, extracting significant and meaningful knowledge from the data simultaneously during the learning process is a difficult task considering the high-dimensional and highly correlated nature of genomic datasets. Thus, there is a need for models that not only predict tumour volume from gene expression data of patients but also use prior information coming from pathway/gene sets during the learning process, to distinguish molecular mechanisms which play crucial role in tumour progression and therefore, disease prognosis. RESULTS: In this study, instead of initially choosing several pathways/gene sets from an available set and training a model on this previously chosen subset of genomic features, we built a novel machine learning algorithm, PrognosiT, that accomplishes both tasks together. We tested our algorithm on thyroid carcinoma patients using gene expression profiles and cancer-specific pathways/gene sets. Predictive performance of our novel multiple kernel learning algorithm (PrognosiT) was comparable or even better than random forest (RF) and support vector regression (SVR). It is also notable that, to predict tumour volume, PrognosiT used gene expression features less than one-tenth of what RF and SVR algorithms used. CONCLUSIONS: PrognosiT was able to obtain comparable or even better predictive performance than SVR and RF. Moreover, we demonstrated that during the learning process, our algorithm managed to extract relevant and meaningful pathway/gene sets information related to the studied cancer type, which provides insights about its progression and aggressiveness. We also compared gene expressions of the selected genes by our algorithm in tumour and normal tissues, and we then discussed up- and down-regulated genes selected by our algorithm while learning, which could be beneficial for determining new biomarkers. Ayyüce Begüm Bektas, Mehmet Gönen |
BMC Bioinform. | 2 |
| 2020 | Improving Fraud Detection and Concept Drift Adaptation in Credit Card Transactions Using Incremental Gradient Boosting TreesabstractDue to the increase in the use of credit cards in electronic shopping, card payments for online commerce have rapidly become a popular trend, which also led to the growth in the number of retailers. Because of these various online shopping options, a more frequent variation in the spending behaviors of the customers and purchasing trends of online markets, known as concept drift problem, can be observed over time which also causes an increase in the need for novel fraudulent strategies. This drifting problem may significantly hinder the effective performance of state-of-the-art fraud detection approaches in real credit card transaction data, which also has the imbalanced class distribution problem. In this study, a card-based incremental Gradient Boosting Tree (GBT) is investigated to detect credit card frauds and to adapt in real-time to drifts occurred in online transactions. The card-based incremental learning is achieved in which the transactions of the fraudulent credit cards reported in each day are incrementally learned by the GBT model. Therefore, the card-based incremental GBT model is compared with the regular GBT model, and retraining of a new transaction set formed by combining the previous set and the transactions of the cards reported as fraudulent. The experiments have been carried out on the 4-month real transaction data from December 2019 to March 2020 in which the concept drift problem occurred in December, dramatically affecting the performance of the GBT model. In these experiments, the improvements in the fraud detection performance have been realized in all months, and also the effectiveness of the card-based increment has been verified by comparing it with the transaction-based incremental learning that may cause catastrophic forgetting problem. Baris Bayram, Bilge Koroglu, Mehmet Gönen |
ICMLA | 3 |
| 2020 | An efficient framework to identify key miRNA-mRNA regulatory modules in cancerabstractMOTIVATION: Micro-RNAs (miRNAs) are known as the important components of RNA silencing and post-transcriptional gene regulation, and they interact with messenger RNAs (mRNAs) either by degradation or by translational repression. miRNA alterations have a significant impact on the formation and progression of human cancers. Accordingly, it is important to establish computational methods with high predictive performance to identify cancer-specific miRNA-mRNA regulatory modules. RESULTS: We presented a two-step framework to model miRNA-mRNA relationships and identify cancer-specific modules between miRNAs and mRNAs from their matched expression profiles of more than 9000 primary tumors. We first estimated the regulatory matrix between miRNA and mRNA expression profiles by solving multiple linear programming problems. We then formulated a unified regularized factor regression (RFR) model that simultaneously estimates the effective number of modules (i.e. latent factors) and extracts modules by decomposing regulatory matrix into two low-rank matrices. Our RFR model groups correlated miRNAs together and correlated mRNAs together, and also controls sparsity levels of both matrices. These attributes lead to interpretable results with high predictive performance. We applied our method on a very comprehensive data collection by including 32 TCGA cancer types. To find the biological relevance of our approach, we performed functional gene set enrichment and survival analyses. A large portion of the identified modules are significantly enriched in Hallmark, PID and KEGG pathways/gene sets. To validate the identified modules, we also performed literature validation as well as validation using experimentally supported miRTarBase database. AVAILABILITY AND IMPLEMENTATION: Our implementation of proposed two-step RFR algorithm in R is available at https://github.com/MiladMokhtaridoost/2sRFR together with the scripts that replicate the reported experiments. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Milad Mokhtaridoost, Mehmet Gönen |
Bioinform. | 2 |
| 2020 | A multitask multiple kernel learning formulation for discriminating early- and late-stage cancersabstractMOTIVATION: Genomic information is increasingly being used in diagnosis, prognosis and treatment of cancer. The severity of the disease is usually measured by the tumor stage. Therefore, identifying pathways playing an important role in progression of the disease stage is of great interest. Given that there are similarities in the underlying mechanisms of different cancers, in addition to the considerable correlation in the genomic data, there is a need for machine learning methods that can take these aspects of genomic data into account. Furthermore, using machine learning for studying multiple cancer cohorts together with a collection of molecular pathways creates an opportunity for knowledge extraction. RESULTS: We studied the problem of discriminating early- and late-stage tumors of several cancers using genomic information while enforcing interpretability on the solutions. To this end, we developed a multitask multiple kernel learning (MTMKL) method with a co-clustering step based on a cutting-plane algorithm to identify the relationships between the input tasks and kernels. We tested our algorithm on 15 cancer cohorts and observed that, in most cases, MTMKL outperforms other algorithms (including random forests, support vector machine and single-task multiple kernel learning) in terms of predictive power. Using the aggregate results from multiple replications, we also derived similarity matrices between cancer cohorts, which are, in many cases, in agreement with available relationships reported in the relevant literature. AVAILABILITY AND IMPLEMENTATION: Our implementations of support vector machine and multiple kernel learning algorithms in R are available at https://github.com/arezourahimi/mtgsbc together with the scripts that replicate the reported experiments. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Arezou Rahimi, Mehmet Gönen |
Bioinform. | 2 |
| 2019 | A Multitask Multiple Kernel Learning Algorithm for Survival Analysis with Application to Cancer BiologyabstractPredictive performance of machine learning algorithms on related problems can be improved using multitask learning approaches. Rather than performing survival analysis on each data set to predict survival times of cancer patients, we developed a novel multitask approach based on multiple kernel learning (MKL). Our multitask MKL algorithm both works on multiple cancer data sets and integrates cancer-related pathways/gene sets into survival analysis. We tested our algorithm, which is named as Path2MSurv, on the Cancer Genome Atlas data sets analyzing gene expression profiles of 7,655 patients from 20 cancer types together with cancer-specific pathway/gene set collections. Path2MSurv obtained better or comparable predictive performance when benchmarked against random survival forest, survival support vector machine, and single-task variant of our algorithm. Path2MSurv has the ability to identify key pathways/gene sets in predicting survival times of patients from different cancer types. Onur Dereli, Ceyda Oguz, Mehmet Gönen |
ICML | 3 |
| 2019 | Path2Surv: Pathway/gene set-based survival analysis using multiple kernel learningabstractMOTIVATION: Survival analysis methods that integrate pathways/gene sets into their learning model could identify molecular mechanisms that determine survival characteristics of patients. Rather than first picking the predictive pathways/gene sets from a given collection and then training a predictive model on the subset of genomic features mapped to these selected pathways/gene sets, we developed a novel machine learning algorithm (Path2Surv) that conjointly performs these two steps using multiple kernel learning. RESULTS: We extensively tested our Path2Surv algorithm on 7655 patients from 20 cancer types using cancer-specific pathway/gene set collections and gene expression profiles of these patients. Path2Surv statistically significantly outperformed survival random forest (RF) on 12 out of 20 datasets and obtained comparable predictive performance against survival support vector machine (SVM) using significantly fewer gene expression features (i.e. less than 10% of what survival RF and survival SVM used). AVAILABILITY AND IMPLEMENTATION: Our implementations of survival SVM and Path2Surv algorithms in R are available at https://github.com/mehmetgonen/path2surv together with the scripts that replicate the reported experiments. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Onur Dereli, Ceyda Oguz, Mehmet Gönen |
Bioinform. | 3 |
| 2018 | Structured Gaussian Processes with Twin Multiple Kernel LearningabstractVanilla Gaussian processes (GPs) have prohibitive computational needs for very large data sets. To overcome this difficulty, special structures in the covariance matrix, if exist, should be exploited using decomposition methods such as the Kronecker product. In this paper, we integrated the Kronecker decomposition approach into a multiple kernel learning (MKL) framework for GP regression. We first formulated a regression algorithm with the Kronecker decomposition of structured kernels for spatiotemporal modeling to learn the contribution of spatial and temporal features as well as learning a model for out-of-sample prediction. We then evaluated the performance of our proposed computational framework, namely, structured GPs with twin MKL, on two different real data sets to show its efficiency and effectiveness. MKL helped us extract relative importance of input features by assigning weights to kernels calculated on different subsets of temporal and spatial features. Çigdem Ak, Önder Ergönül, Mehmet Gönen |
ACML | 3 |
| 2018 | Discriminating early- and late-stage cancers using multiple kernel learning on gene setsabstractMotivation: Identifying molecular mechanisms that drive cancers from early to late stages is highly important to develop new preventive and therapeutic strategies. Standard machine learning algorithms could be used to discriminate early- and late-stage cancers from each other using their genomic characterizations. Even though these algorithms would get satisfactory predictive performance, their knowledge extraction capability would be quite restricted due to highly correlated nature of genomic data. That is why we need algorithms that can also extract relevant information about these biological mechanisms using our prior knowledge about pathways/gene sets. Results: In this study, we addressed the problem of separating early- and late-stage cancers from each other using their gene expression profiles. We proposed to use a multiple kernel learning (MKL) formulation that makes use of pathways/gene sets (i) to obtain satisfactory/improved predictive performance and (ii) to identify biological mechanisms that might have an effect in cancer progression. We extensively compared our proposed MKL on gene sets algorithm against two standard machine learning algorithms, namely, random forests and support vector machines, on 20 diseases from the Cancer Genome Atlas cohorts for two different sets of experiments. Our method obtained statistically significantly better or comparable predictive performance on most of the datasets using significantly fewer gene expression features. We also showed that our algorithm was able to extract meaningful and disease-specific information that gives clues about the progression mechanism. Availability and implementation: Our implementations of support vector machine and multiple kernel learning algorithms in R are available at https://github.com/mehmetgonen/gsbc together with the scripts that replicate the reported experiments. Arezou Rahimi, Mehmet Gönen |
Bioinform. | 2 |
| 2017 | Modeling gene-wise dependencies improves the identification of drug response biomarkers in cancer studiesabstractMotivation: In recent years, vast advances in biomedical technologies and comprehensive sequencing have revealed the genomic landscape of common forms of human cancer in unprecedented detail. The broad heterogeneity of the disease calls for rapid development of personalized therapies. Translating the readily available genomic data into useful knowledge that can be applied in the clinic remains a challenge. Computational methods are needed to aid these efforts by robustly analyzing genome-scale data from distinct experimental platforms for prioritization of targets and treatments. Results: We propose a novel, biologically motivated, Bayesian multitask approach, which explicitly models gene-centric dependencies across multiple and distinct genomic platforms. We introduce a gene-wise prior and present a fully Bayesian formulation of a group factor analysis model. In supervised prediction applications, our multitask approach leverages similarities in response profiles of groups of drugs that are more likely to be related to true biological signal, which leads to more robust performance and improved generalization ability. We evaluate the performance of our method on molecularly characterized collections of cell lines profiled against two compound panels, namely the Cancer Cell Line Encyclopedia and the Cancer Therapeutics Response Portal. We demonstrate that accounting for the gene-centric dependencies enables leveraging information from multi-omic input data and improves prediction and feature selection performance. We further demonstrate the applicability of our method in an unsupervised dimensionality reduction application by inferring genes essential to tumorigenesis in the pancreatic ductal adenocarcinoma and lung adenocarcinoma patient cohorts from The Cancer Genome Atlas. Availability and Implementation: : The code for this work is available at https://github.com/olganikolova/gbgfa. Contact: : [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Olga Nikolova, Russell Moser, Christopher Kemp, Mehmet Gönen, Adam A. Margolin |
Bioinform. | 4 |
| 2016 | AUC Maximization in Bayesian Hierarchical ModelsabstractThe area under the curve (AUC) measures such as the area under the receiver operating characteristics curve (AUROC) and the area under the precision-recall curve (AUPR) are known to be more appropriate than the error rate, especially, for imbalanced data sets. There are several algorithms to optimize AUC measures instead of minimizing the error rate. However, this idea has not been fully exploited in Bayesian hierarchical models owing to the difficulties in inference. Here, we formulate a general Bayesian inference framework, called Bayesian AUC Maximization (BAM), to integrate AUC maximization into Bayesian hierarchical models by borrowing the pairwise and listwise ranking ideas from the information retrieval literature. To showcase our BAM framework, we develop two Bayesian linear classifier variants for two ranking approaches and derive their variational inference procedures. We perform validation experiments on four biomedical data sets to demonstrate the better predictive performance of our framework over its error-minimizing counterpart in terms of average AUROC and AUPR values. Mehmet Gönen |
ECAI | 1 |
| 2016 | Integrating gene set analysis and nonlinear predictive modeling of disease phenotypes using a Bayesian multitask formulationabstractBACKGROUND: Identifying molecular signatures of disease phenotypes is studied using two mainstream approaches: (i) Predictive modeling methods such as linear classification and regression algorithms are used to find signatures predictive of phenotypes from genomic data, which may not be robust due to limited sample size or highly correlated nature of genomic data. (ii) Gene set analysis methods are used to find gene sets on which phenotypes are linearly dependent by bringing prior biological knowledge into the analysis, which may not capture more complex nonlinear dependencies. Thus, formulating an integrated model of gene set analysis and nonlinear predictive modeling is of great practical importance. RESULTS: In this study, we propose a Bayesian binary classification framework to integrate gene set analysis and nonlinear predictive modeling. We then generalize this formulation to multitask learning setting to model multiple related datasets conjointly. Our main novelty is the probabilistic nonlinear formulation that enables us to robustly capture nonlinear dependencies between genomic data and phenotype even with small sample sizes. We demonstrate the performance of our algorithms using repeated random subsampling validation experiments on two cancer and two tuberculosis datasets by predicting important disease phenotypes from genome-wide gene expression data. CONCLUSIONS: We are able to obtain comparable or even better predictive performance than a baseline Bayesian nonlinear algorithm and to identify sparse sets of relevant genes and gene sets on all datasets. We also show that our multitask learning formulation enables us to further improve the generalization performance and to better understand biological processes behind disease phenotypes. Mehmet Gönen |
BMC Bioinform. | 1 |
| 2015 | Understanding emotional impact of images using Bayesian multiple kernel learning
He Zhang 0009, Mehmet Gönen, Zhirong Yang, Erkki Oja |
Neurocomputing | 2 |
| 2014 | Kernelized Bayesian Transfer LearningabstractTransfer learning considers related but distinct tasks defined on heterogenous domains and tries to transfer knowledge between these tasks to improve generalization performance. It is particularly useful when we do not have sufficient amount of labeled training data in some tasks, which may be very costly, laborious, or even infeasible to obtain. Instead, learning the tasks jointly enables us to effectively increase the amount of labeled training data. In this paper, we formulate a kernelized Bayesian transfer learning framework that is a principled combination of kernel-based dimensionality reduction models with task-specific projection matrices to find a shared subspace and a coupled classification model for all of the tasks in this subspace. Our two main contributions are: (i) two novel probabilistic models for binary and multiclass classification, and (ii) very efficient variational approximation procedures for these models. We illustrate the generalization performance of our algorithms on two different applications. In computer vision experiments, our method outperforms the state-of-the-art algorithms on nine out of 12 benchmark supervised domain adaptation experiments defined on two object recognition data sets. In cancer biology experiments, we use our algorithm to predict mutation status of important cancer genes from gene expression profiles using two distinct cancer populations, namely, patient-derived primary tumor data and in-vitro-derived cancer cell line data. We show that we can increase our generalization performance on primary tumors using cell lines as an auxiliary data source. Mehmet Gönen, Adam A. Margolin |
AAAI | 1 |
| 2014 | Embedding Heterogeneous Data by Preserving Multiple KernelsabstractHeterogeneous data may arise in many real-life applications under different scenarios. In this paper, we formulate a general framework to address the problem of modeling heterogeneous data. Our main contribution is a novel embedding method, called multiple kernel preserving embedding (MKPE), which projects heterogeneous data into a unified embedding space by preserving crossdomain interactions and within-domain similarities simultaneously. These interactions and similarities between data points are approximated with Gaussian kernels to transfer local neighborhood information to the projected subspace. We also extend our method for out-of-sample embedding using a parametric formulation in the projection step. The performance of MKPE is illustrated on two tasks: (i) modeling biological interaction networks and (ii) cross-domain information retrieval. Empirical results of these two tasks validate the predictive performance of our algorithm. Mehmet Gönen |
ECAI | 1 |
| 2014 | Bayesian Multiview Dimensionality Reduction for Learning Predictive SubspacesabstractMultiview learning basically tries to exploit different feature representations to obtain better learners. For example, in video and image recognition problems, there are many possible feature representations such as color- and texture-based features. There are two common ways of exploiting multiple views: forcing similarity (i) in predictions and (ii) in latent subspace. In this paper, we introduce a novel Bayesian multiview dimensionality reduction method coupled with supervised learning to find predictive subspaces and its inference details. Experiments show that our proposed method obtains very good results on image recognition tasks in terms of classification and retrieval performances. Mehmet Gönen, Gülefsan Bozkurt Gönen, Fikret S. Gürgen |
ECAI | 1 |
| 2014 | Localized Data Fusion for Kernel k-Means Clustering with Application to Cancer Biology
Mehmet Gönen, Adam A. Margolin |
NIPS | 1 |
| 2014 | Drug susceptibility prediction against a panel of drugs using kernelized Bayesian multitask learningabstractMOTIVATION: Human immunodeficiency virus (HIV) and cancer require personalized therapies owing to their inherent heterogeneous nature. For both diseases, large-scale pharmacogenomic screens of molecularly characterized samples have been generated with the hope of identifying genetic predictors of drug susceptibility. Thus, computational algorithms capable of inferring robust predictors of drug responses from genomic information are of great practical importance. Most of the existing computational studies that consider drug susceptibility prediction against a panel of drugs formulate a separate learning problem for each drug, which cannot make use of commonalities between subsets of drugs. RESULTS: In this study, we propose to solve the problem of drug susceptibility prediction against a panel of drugs in a multitask learning framework by formulating a novel Bayesian algorithm that combines kernel-based non-linear dimensionality reduction and binary classification (or regression). The main novelty of our method is the joint Bayesian formulation of projecting data points into a shared subspace and learning predictive models for all drugs in this subspace, which helps us to eliminate off-target effects and drug-specific experimental noise. Another novelty of our method is the ability of handling missing phenotype values owing to experimental conditions and quality control reasons. We demonstrate the performance of our algorithm via cross-validation experiments on two benchmark drug susceptibility datasets of HIV and cancer. Our method obtains statistically significantly better predictive performance on most of the drugs compared with baseline single-task algorithms that learn drug-specific models. These results show that predicting drug susceptibility against a panel of drugs simultaneously within a multitask learning framework improves overall predictive performance over single-task learning approaches. AVAILABILITY AND IMPLEMENTATION: Our Matlab implementations for binary classification and regression are available at https://github.com/mehmetgonen/kbmtl. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Mehmet Gönen, Adam A. Margolin |
Bioinform. | 1 |
| 2014 | Multi-task and multi-view learning of user state
Melih Kandemir, Akos Vetek, Mehmet Gönen, Arto Klami, Samuel Kaski |
Neurocomputing | 3 |
| 2014 | Kernelized Bayesian Matrix FactorizationabstractWe extend kernelized matrix factorization with a full-Bayesian treatment and with an ability to work with multiple side information sources expressed as different kernels. Kernels have been introduced to integrate side information about the rows and columns, which is necessary for making out-of-matrix predictions. We discuss specifically binary output matrices but extensions to realvalued matrices are straightforward. We extend the state of the art in two key aspects: (i) A full-conjugate probabilistic formulation of the kernelized matrix factorization enables an efficient variational approximation, whereas full-Bayesian treatments are not computationally feasible in the earlier approaches. (ii) Multiple side information sources are included, treated as different kernels in multiple kernel learning which additionally reveals which side sources are informative. We then show that the framework can also be used for supervised and semi-supervised multilabel classification and multi-output regression, by considering samples and outputs as the domains where matrix factorization operates. Our method outperforms alternatives in predicting drug-protein interactions on two data sets. On multilabel classification, our algorithm obtains the lowest Hamming losses on 10 out of 14 data sets compared to five state-of-the-art multilabel classification algorithms. We finally show that the proposed approach outperforms alternatives in multi-output regression experiments on a yeast cell cycle data set. Mehmet Gönen, Samuel Kaski |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2014 | Coupled dimensionality reduction and classification for supervised and semi-supervised multilabel learning
Mehmet Gönen |
Pattern Recognit. Lett. | 1 |
| 2013 | Kernelized Bayesian Matrix FactorizationabstractWe extend kernelized matrix factorization with a fully Bayesian treatment and with an ability to work with multiple side information sources expressed as different kernels. Kernel functions have been introduced to matrix factorization to integrate side information about the rows and columns (e.g., objects and users in recommender systems), which is necessary for making out-of-matrix (i.e., cold start) predictions. We discuss specifically bipartite graph inference, where the output matrix is binary, but extensions to more general matrices are straightforward. We extend the state of the art in two key aspects: (i) A fully conjugate probabilistic formulation of the kernelized matrix factorization problem enables an efficient variational approximation, whereas fully Bayesian treatments are not computationally feasible in the earlier approaches. (ii) Multiple side information sources are included, treated as different kernels in multiple kernel learning that additionally reveals which side information sources are informative. Our method outperforms alternatives in predicting drug-protein interactions on two data sets. We then show that our framework can also be used for solving multilabel learning problems by considering samples and labels as the two domains where matrix factorization operates on. Our algorithm obtains the lowest Hamming loss values on 10 out of 14 multilabel classification data sets compared to five state-of-the-art multilabel learning algorithms. Mehmet Gönen, Suleiman A. Khan, Samuel Kaski |
ICML (3) | 1 |
| 2013 | Predicting Emotional States of Images Using Bayesian Multiple Kernel Learning
He Zhang 0009, Mehmet Gönen, Zhirong Yang, Erkki Oja |
ICONIP (3) | 2 |
| 2013 | Affective Abstract Image Classification and Retrieval Using Multiple Kernel Learning
He Zhang 0009, Zhirong Yang, Mehmet Gönen, Markus Koskela, Jorma Laaksonen, Timo Honkela, Erkki Oja |
ICONIP (3) | 3 |
| 2013 | Localized algorithms for multiple kernel learning
Mehmet Gönen, Ethem Alpaydin |
Pattern Recognit. | 1 |
| 2013 | Bayesian Supervised Dimensionality ReductionabstractDimensionality reduction is commonly used as a preprocessing step before training a supervised learner. However, coupled training of dimensionality reduction and supervised learning steps may improve the prediction performance. In this paper, we introduce a simple and novel Bayesian supervised dimensionality reduction method that combines linear dimensionality reduction and linear supervised learning in a principled way. We present both Gibbs sampling and variational approximation approaches to learn the proposed probabilistic model for multiclass classification. We also extend our formulation toward model selection using automatic relevance determination in order to find the intrinsic dimensionality. Classification experiments on three benchmark data sets show that the new model significantly outperforms seven baseline linear dimensionality reduction algorithms on very low dimensions in terms of generalization performance on test data. The proposed model also obtains the best results on an image recognition task in terms of classification and retrieval performances. Mehmet Gönen |
IEEE Trans. Cybern. | 1 |
| 2013 | Supervised Multiple Kernel Embedding for Learning Predictive SubspacesabstractFor supervised learning problems, dimensionality reduction is generally applied as a preprocessing step. However, coupled training of dimensionality reduction and supervised learning steps may improve the prediction performance. In this paper, we propose a novel dimensionality reduction algorithm coupled with a supervised kernel-based learner, called supervised multiple kernel embedding, that integrates multiple kernel learning to dimensionality reduction and performs prediction on the projected subspace with a joint optimization framework. Combining multiple kernels allows us to combine different feature representations and/or similarity measures toward a unified subspace. We perform experiments on one digit recognition and two bioinformatics data sets. Our proposed method significantly outperforms multiple kernel Fisher discriminant analysis followed by a standard kernel-based learner, especially on low dimensions. Mehmet Gönen |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2012 | Bayesian Efficient Multiple Kernel Learning
Mehmet Gönen |
ICML | 1 |
| 2012 | Bayesian Supervised Multilabel Learning with Coupled Embedding and ClassificationabstractCoupled training of dimensionality reduction and classification is proposed previously to improve the prediction performance for single-label problems. Following this line of research, in this paper, we introduce a novel Bayesian supervised multilabel learning method that combines linear dimensionality reduction with linear binary classification. We present a deterministic variational approximation approach to learn the proposed probabilistic model for multilabel classification. We perform experiments on four benchmark multilabel learning data sets by comparing our method with four baseline linear dimensionality reduction algorithms. Experiments show that the proposed approach achieves good performance values in terms of hamming loss, macro F1, and micro F1 on held-out test data. The low-dimensional embeddings obtained by our method are also very useful for exploratory data analysis. Mehmet Gönen |
SDM | 1 |
| 2012 | Predicting drug-target interactions from chemical and genomic kernels using Bayesian matrix factorizationabstractMOTIVATION: Identifying interactions between drug compounds and target proteins has a great practical importance in the drug discovery process for known diseases. Existing databases contain very few experimentally validated drug-target interactions and formulating successful computational methods for predicting interactions remains challenging. RESULTS: In this study, we consider four different drug-target interaction networks from humans involving enzymes, ion channels, G-protein-coupled receptors and nuclear receptors. We then propose a novel Bayesian formulation that combines dimensionality reduction, matrix factorization and binary classification for predicting drug-target interaction networks using only chemical similarity between drug compounds and genomic similarity between target proteins. The novelty of our approach comes from the joint Bayesian formulation of projecting drug compounds and target proteins into a unified subspace using the similarities and estimating the interaction network in that subspace. We propose using a variational approximation in order to obtain an efficient inference scheme and give its detailed derivations. Finally, we demonstrate the performance of our proposed method in three different scenarios: (i) exploratory data analysis using low-dimensional projections, (ii) predicting interactions for the out-of-sample drug compounds and (iii) predicting unknown interactions of the given network. AVAILABILITY: Software and Supplementary Material are available at http://users.ics.aalto.fi/gonen/kbmf2k. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Mehmet Gönen |
Bioinform. | 1 |
| 2012 | Probabilistic and discriminative group-wise feature selection methods for credit risk analysis
Gülefsan Bozkurt Gönen, Mehmet Gönen, Fikret S. Gürgen |
Expert Syst. Appl. | 2 |
| 2011 | Multitask Learning Using Regularized Multiple Kernel Learning
Mehmet Gönen, Melih Kandemir, Samuel Kaski |
ICONIP (2) | 1 |
| 2011 | Multiple Kernel Learning Algorithms
Mehmet Gönen, Ethem Alpaydin |
J. Mach. Learn. Res. | 1 |
| 2011 | Regularizing multiple kernel learning using response surface methodology
Mehmet Gönen, Ethem Alpaydin |
Pattern Recognit. | 1 |
| 2010 | Localized Multiple Kernel RegressionabstractMultiple kernel learning (MKL) uses a weighted combination of kernels where the weight of each kernel is optimized during training. However, MKL assigns the same weight to a kernel over the whole input space. Our main objective is the formulation of the localized multiple kernel learning (LMKL) framework that allows kernels to be combined with different weights in different regions of the input space by using a gating model. In this paper, we apply the LMKL framework to regression estimation and derive a learning algorithm for this extension. Canonical support vector regression may over fit unless the kernel parameters are selected appropriately; we see that even if provide more kernels than necessary, LMKL uses only as many as needed and does not overfit due to its inherent regularization. Mehmet Gönen, Ethem Alpaydin |
ICPR | 1 |
| 2010 | Supervised learning of local projection kernels
Mehmet Gönen, Ethem Alpaydin |
Neurocomputing | 1 |
| 2010 | Cost-conscious multiple kernel learning
Mehmet Gönen, Ethem Alpaydin |
Pattern Recognit. Lett. | 1 |
| 2008 | Localized multiple kernel learningabstractRecently, instead of selecting a single kernel, multiple kernel learning (MKL) has been proposed which uses a convex combination of kernels, where the weight of each kernel is optimized during training. However, MKL assigns the same weight to a kernel over the whole input space. In this paper, we develop a localized multiple kernel learning (LMKL) algorithm using a gating model for selecting the appropriate kernel function locally. The localizing gating model and the kernel-based classifier are coupled and their optimization is done in a joint manner. Empirical results on ten benchmark and two bioinformatics data sets validate the applicability of our approach. LMKL achieves statistically similar accuracy results compared with MKL by storing fewer support vectors. LMKL can also combine multiple copies of the same kernel function localized in different parts. For example, LMKL with multiple linear kernels gives better accuracy results than using a single linear kernel on bioinformatics data sets. Mehmet Gönen, Ethem Alpaydin |
ICML | 1 |
| 2008 | Multiclass Posterior Probability Support Vector MachinesabstractTao, et al have recently proposed the posterior probability support vector machine (PPSVM) which uses soft labels derived from estimated posterior probabilities to be more robust to noise and outliers. Tao, et al's model uses a window-based density estimator to calculate the posterior probabilities and is a binary classifier. We propose a neighbor-based density estimator and also extend the model to the multiclass case. Our bias-variance analysis shows that the decrease in error by PPSVM is due to a decrease in bias. On 20 benchmark data sets, we observe that PPSVM obtains accuracy results that are higher or comparable to those of canonical SVM using significantly fewer support vectors. Mehmet Gönen, Ayse Gönül Tanugur, Ethem Alpaydin |
IEEE Trans. Neural Networks | 1 |