Mehmet Tan

dblp:72/4720 · DBLP profile ↗
← Back
22ranked-venue papers
10as first author
5since 2021 · last 2021
0000-0002-1741-0570ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 15 · 7 first-author · 4 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author
YearPublicationVenuePosition
2021 Automated Molecule Generation using Deep Q-Learning and Graph Neural Networks
abstract
The concept of generating molecular structures with specific desirable characteristics underlies some of the crucial problems in drug discovery. In this paper, we present a model which constructs new molecules for specific desired properties. This model uses graph neural networks to generate molecular representations, and combining these representations with Reinforcement Learning architecture (Deep Q-Learning), builds new molecules. We used two different graph neural network architectures: Graph Convolutional Network and Graph Attention Network. We compared the molecular representations obtained from these two models with the Morgan Fingerprint representation in three separate experiments using the same Reinforcement Learning design. These experiments are single-objective optimization (drug-likeness), optimizing the penalized LogP with similarity constraint, and multi-objective optimization (drug-likeness with similarity constraint). The results show that the Reinforcement Learning models trained with molecular representations obtained from graph neural networks are more successful than the model trained with Morgan Fingerprint representation.
Riza Isik, Mehmet Tan
BIBM2
2021 Using Clinical Drug Representations for Improving Mortality and Length of Stay Predictions
abstract
Drug representations have played an important role in cheminformatics. However, in the healthcare domain, drug representations have been underused relative to the rest of Electronic Health Record (EHR) data, due to the complexity of high dimensional drug representations and the lack of proper pipeline that will allow to convert clinical drugs to their representations. Time-varying vital signs, laboratory measurements, and related time-series signals are commonly used to predict clinical outcomes. In this work, we demonstrated that using clinical drug representations in addition to other clinical features has significant potential to increase the performance of mortality and length of stay (LOS) models. We evaluate the two different drug representation methods (Extended -Connectivity Fingerprint- ECFP and SMILES-Transformer embedding) on clinical outcome predictions. The results have shown that the proposed multimodal approach achieves substantial enhancement on clinical tasks over baseline models. U sing clinical drug representations as additional features improve the LOS prediction for Area Under the Receiver Operating Characteristics (AUROC) around %6 and for Area Under Precision-Recall Curve (AUPRC) by around % 5. Furthermore, for the mortality prediction task, there is an improvement of around % 2 over the time series baseline in terms of AUROC and %3.5 in terms of AUPRC. The code for the proposed method is available at https://github.com/tanlab/MIMIC-III-Clinical-Drug-Representations.
Batuhan Bardak, Mehmet Tan
CIBCB2
2021 DeepGREP: A deep convolutional neural network for predicting gene-regulating effects of small molecules
abstract
Accurately predicting desired gene expression effects by using the representations of drugs and genes in silico is a key task in chemogenomics. This paper proposes DeepGREP, a deep learning model that can predict small molecules' gene regulation effects. The main motivation of this work is improving chemical-induced differential gene expression prediction by using a convolutional-based architecture to represent drugs and genes more effectively. To evaluate the performance of the DeepGREP, we conducted several experiments and compared them with DeepCop, the baseline model. The results show that DeepGREP outperforms the baseline model and significantly improves the gene expression prediction for AUC by around 4%, F-Score by around 15%, and Enrichment Factor by around 22%. We also demonstrate that the proposed method mostly outperforms the baseline in more difficulties setting of generalization to unseen molecules by using cold-drug splitting.
Benan Bardak, Mehmet Tan
CIBCB2
2021 Improving clinical outcome predictions using convolution over medical entities with multimodal learning
Batuhan Bardak, Mehmet Tan
Artif. Intell. Medicine2
2021 A Convolutional Deep Clustering Framework for Gene Expression Time Series
abstract
The functional or regulatory processes within the cell are explicitly governed by the expression levels of a subset of its genes. Gene expression time series captures activities of individual genes over time and aids revealing underlying cellular dynamics. An important step in high-throughput gene expression time series experiment is clustering genes based on their temporal expression patterns and is conventionally achieved by unsupervised machine learning techniques. However, most of the clustering techniques either suffer from the short length of gene expression time series or ignore temporal structure of the data. In this work, we propose DeepTrust, a novel deep learning-based framework for gene expression time series clustering which can overcome these issues. DeepTrust initially transforms time series data into images to obtain richer data representations. Afterwards, a deep convolutional clustering algorithm is applied on the constructed images. Analyses on both simulated and biological data sets exhibit the efficiency of this new framework, compared to widely used clustering techniques. We also utilize enrichment analyses to illustrate the biological plausibility of the clusters detected by DeepTrust. Our code and data are available from http://github.com/tanlab/DeepTrust.
Ozan Firat Özgül, Batuhan Bardak, Mehmet Tan
IEEE ACM Trans. Comput. Biol. Bioinform.3
2020 Chemical Induced Differential Gene Expression Prediction on LINCS Database
abstract
Understanding the mechanism of action for drugs is vital for drug discovery. Identifying the effect of drugs on gene expression can shed light on the system-side influence of the chemical compounds in biological organisms. In this paper, we propose to use multi-task neural networks to predict chemical induced differential gene expression on cancer cell lines based solely on features of chemicals. Our model predicts differential gene expression identified by a method called Characteristic Direction on a large scale chemical induced gene expression database (LINCS L1000). The results show that the multi-task networks outperform the other single task baselines. We also compare different representations of chemicals and report effect of clustering genes on the prediction performance.
Riza Isik, Isiksu Eksioglu, Bahattin Can Maral, Benan Bardak, Mehmet Tan
BIBE5
2017 Disease outbreak prediction by data integration and multi-task learning
abstract
The requirements for treatments vary for different diseases. These have to be considered in order to plan ahead the expenditures for the health care system. In this sense, disease surveillance has a significant impact on resource planning. To this end, we study the problem of predicting the number of incidences for a given disease based on the internet search and access log statistics. A number of papers appear in the literature that study this problem of predicting outbreaks, especially for Influenza. In this paper, in addition to investigating disease incidences other than Influenza, we propose to use the statistics for different diseases together for achieving transfer learning. We argue that we can increase prediction performance by considering diseases together in a multi-task learning setting due to our assumption of structure sharing. The results we obtained are promising as we achieved performance improvements in this setting. The code and data-sets used in the study are available from http://mtan.etu.edu.tr/Supplementary/Outbreak-prediction/.
Batuhan Bardak, Mehmet Tan
CIBCB2
2016 Effective gene expression data generation framework based on multi-model approach
Utku Sirin, Utku Erdogdu, Faruk Polat, Mehmet Tan, Reda Alhajj
Artif. Intell. Medicine4
2016 Prediction of anti-cancer drug response by kernelized multi-task learning
Mehmet Tan
Artif. Intell. Medicine1
2015 Prediction of influenza outbreaks by integrating Wikipedia article access logs and Google flu trend data
abstract
Prediction of influenza outbreaks is of utmost importance for health practitioners, officers and people. After the increasing usage of internet, it became easier and more valuable to fetch and process internet search query data. There are two significant platforms that people widely use, Google and Wikipedia. In both platforms, access logs are available which means that we can see how often any query/article was searched. Google has its own web service for monitoring and forecasting influenza-illness which is called the Google Flu Trends. It provides estimates of influenza activity for some countries. The second alternative is Wikipedia access logs which provide the number of visits for the articles on Wikipedia. There are papers which work with these platforms separately. In this paper, we propose a new technique to use these two sources together to improve the prediction of influenza outbreaks. We achieved promising results for both nowcasting and forecasting with linear regression models.
Batuhan Bardak, Mehmet Tan
BIBE2
2014 Drug sensitivity prediction for cancer cell lines based on pairwise kernels and miRNA profiles
abstract
Cancer cell lines comprise an important tool to design and evaluate new drug candidates. Prediction of in vivo drug response for cancer cell lines has become attractive due to recently issued large scale drug screen databases. The data provided by these databases can be the key to model drug sensitivity for cancer cell lines. The data provided by these databases is in the form of drug cell line pairs where a natural method for prediction of drug response, therefore is pairwise support vector machines. This paper presents results on the application of pairwise kernels for drug response prediction, where the results are promising compared to some previously well-performed methods on this task. In addition, effect of exploiting microRNA profiles of cancer cell lines together with mRNA profiles is given.
Mehmet Tan
BIBM1
2012 Effective Enrichment of Gene Expression Data Sets
abstract
The ever-growing need for gene-expression data analysis motivates studies in sample generation due to the lack of enough gene-expression data. It is common that there are thousands of genes but only tens or rarely hundreds of samples available. In this paper, we attempt to formulate the sample generation task as follows: first, building alternative Gene Regulatory Network (GRN) models, second, sampling data from each of them, and then filtering the generated samples using metrics that measure compatibility, diversity and coverage with respect to the original dataset. We constructed two alternative GRN models using Probabilistic Boolean Networks and Ordinary Differential Equations. We developed a multi-objective filtering mechanism based on the three metrics to assess the quality of the newly generated data. We presented a number of experiments to show effectiveness and applicability of the proposed multi-model framework.
Utku Sirin, Utku Erdogdu, Mehmet Tan, Faruk Polat, Reda Alhajj
ICMLA (1)3
2012 Revealing miRNA Regulation and miRNA Target Prediction Using Constraint-Based Learning
abstract
The past decades have witnessed advances in genomic technology; and this has allowed laboratories to generate vast amount of biological data, including microarray gene expression data. Effective analysis of the data helps in better understanding the mechanisms behind the complex behavior of the cell. Actually, a huge body of research focuses on the role of gene regulatory networks (GRNs) in controlling the cell. However, studying the heterogeneous interactions between mRNA and miRNA has received less attention. Fortunately, revealing the targets of miRNAs started to gain some consideration from the research community. Further, integrating mRNA gene expression and miRNA expression data is receiving more attention; the target is to understand the role of miRNA in regulating mRNA in different cell contexts; this could lead to predicting miRNA targets and constructing miRNA-mRNA interaction networks. On the other hand, we have already demonstrated the power of constraint-based learning as a promising technique to learn the structure of GRN , which are homogeneous in the sense that they contain one type of nodes, namely, genes. In this study, we extend our previous work to show how constraint-based learning can be effectively applied to tackle a more challenging problem, namely, to learn the structure of heterogeneous networks, like mRNA-miRNA network. In other words, to build the whole picture of the heterogeneous interactions, we used constraint-based learning algorithms which usually perform well on sparse graphs to predict the interactions within heterogeneous networks, namely, miRNA-mRNA interactions. We are able to achieve this by extending our PCPDPr algorithm, which works on homogeneous networks. The extended version named htrPCPDPr is capable of handling networks connecting two heterogeneous sets of nodes into a bipartite graph. This way, we propose a new learning mechanism to predict miRNA targets from expression profiles of both mRNA and miRNA, in addition to sequence-based prior knowledge about the interactions. The method has been applied to different set of genes related to the Alzheimer disease; the results reported in this paper demonstrate the novelty, applicability, and effectiveness of the proposed approach.
Mohammed Al-Shalalfa, Mehmet Tan, Ghada Naji, Reda Alhajj, Faruk Polat, Jon G. Rokne
IEEE Trans. Syst. Man Cybern. Part C2
2011 Employing Machine Learning Techniques for Data Enrichment: Increasing the Number of Samples for Effective Gene Expression Data Analysis
abstract
For certain domains, e.g. bioinformatics, producing more real samples is costly, error prone and time consuming. Therefore, there is a need for an intelligent automated process capable of substituting the real samples by artificial samples that carry the same characteristics as the real samples and hence could be used for running comprehensive testing of new methodologies. Motivated by this need, we describe a novel approach that integrates Probabilistic Boolean Network and genetic algorithm based techniques into a framework that uses some existing real samples as input and successfully produces new samples as output. The new samples will inspire the characteristics of the existing samples without duplicating them. This leads to diversity in the samples and hence a more rich set of samples to be used in testing. The developed framework incorporates two models (perspectives) for sample generation. We illustrate its applicability for producing new gene expression data samples, a high demanding area that has not received attention. The two perspectives employed in the process are based on models that are not closely related, the independence eliminates the bias of having the produced approach covering only certain characteristics of the domain and leading to samples skewed towards one direction. The produced results are very promising in showing the effectiveness, usefulness and applicability of the proposed multi-model framework.
Utku Erdogdu, Mehmet Tan, Reda Alhajj, Faruk Polat, Douglas J. Demetrick, Jon G. Rokne
BIBM2
2011 Influence of Prior Knowledge in Constraint-Based Learning of Gene Regulatory Networks
abstract
Constraint-based structure learning algorithms generally perform well on sparse graphs. Although sparsity is not uncommon, there are some domains where the underlying graph can have some dense regions; one of these domains is gene regulatory networks, which is the main motivation to undertake the study described in this paper. We propose a new constraint-based algorithm that can both increase the quality of output and decrease the computational requirements for learning the structure of gene regulatory networks. The algorithm is based on and extends the PC algorithm. Two different types of information are derived from the prior knowledge; one is the probability of existence of edges, and the other is the nodes that seem to be dependent on a large number of nodes compared to other nodes in the graph. Also a new method based on Gene Ontology for gene regulatory network validation is proposed. We demonstrate the applicability and effectiveness of the proposed algorithms on both synthetic and real data sets.
Mehmet Tan, Mohammed Al-Shalalfa, Reda Alhajj, Faruk Polat
IEEE ACM Trans. Comput. Biol. Bioinform.1
2010 Feature selection for graph kernels
abstract
Graph classification is important for different scientific applications; it can be exploited in various problems related to bioinformatics and cheminformatics. Given their graphs, there is increasing need for classifying small molecules to predict their properties such as activity, toxicity or mutagenicity. Using subtrees as feature set for graph classification in kernel methods has been shown to perform well in classifying small molecules. It is also well-known that feature selection can improve the performance of classifiers. However, most of the graph kernels are not selective in choosing which subtrees to include in the set of features. Instead, they use all subtrees of a certain property as their feature set. We argue that not all the latter features are needed for effective classification. In this paper, we investigate the effect of selecting subset of the subtrees as features for graph kernels, i.e., we try to identify and keep useful features; all the remaining subtrees are eliminated. A masking procedure, which boils down to feature selection, is proposed for classifying graphs. We conducted experiments on several molecule classification datasets; the results demonstrate the applicability and effectiveness of the proposed feature selection process.
Mehmet Tan, Faruk Polat, Reda Alhajj
BIBM1
2010 Scalable approach for effective control of gene regulatory networks
Mehmet Tan, Reda Alhajj, Faruk Polat
Artif. Intell. Medicine1
2010 Automated Large-Scale Control of Gene Regulatory Networks
abstract
Controlling gene regulatory networks (GRNs) is an important and hard problem. As it is the case in all control problems, the curse of dimensionality is the main issue in real applications. It is possible that hundreds of genes may regulate one biological activity in an organism; this implies a huge state space, even in the case of Boolean models. This is also evident in the literature that shows that only models of small portions of the genome could be used in control applications. In this paper, we empower our framework for controlling GRNs by eliminating the need for expert knowledge to specify some crucial threshold that is necessary for producing effective results. Our framework is characterized by applying the factored Markov decision problem (FMDP) method to the control problem of GRNs. The FMDP is a suitable framework for large state spaces as it represents the probability distribution of state transitions using compact models so that more space and time efficient algorithms could be devised for solving control problems. We successfully mapped the GRN control problem to an FMDP and propose a model reduction algorithm that helps find approximate solutions for large networks by using existing FMDP solvers. The test results reported in this paper demonstrate the efficiency and effectiveness of the proposed approach.
Mehmet Tan, Reda Alhajj, Faruk Polat
IEEE Trans. Syst. Man Cybern. Part B1
2009 Derivation of Transcriptional Regulatory Relationships by Partial Least Squares Regression
abstract
As the number of genes in a transcriptional regulatory network is large and the number of samples in biological data types is usually small, there is a need for integrating multiple data types for reverse engineering these networks. In this paper, we propose a method to integrate microarray gene expression, ChIP-chip and transcription factor binding motif data sets in a partial least squares regression model to derive transcription factors (TFs) -gene interactions. Both single and synergistic effects of TFs on the promoters are considered in the model. A method that dynamically updates the significance level based on ChIP-chip and binding motif data is proposed. The results evaluated by methods based on gene ontology demonstrate the effectiveness of the proposed approach.
Mehmet Tan, Faruk Polat, Reda Alhajj
BIBM1
2008 Large-scale approximate intervention strategies for Probabilistic Boolean Networks as models of gene regulation
abstract
Control of Probabilistic Boolean Networks as models of gene regulation is an important problem; the solution may help researchers in various different areas. But as generally applies to control problems, the size of the state space in gene regulatory networks is too large to be considered for comprehensive solution to the problem; this is evident from the work done in the field, where only very small portions of the whole genome of an organism could be used in control applications. The Factored Markov Decision Problem (FMDP) framework avoids enumerating the whole state space by representing the probability distribution of state transitions using compact models like dynamic bayesian networks. In this paper, we successfully applied FMDP to gene regulatory network control, and proposed a model minimization method that helps finding better approximate policies by using existing FMDP solvers. The results reported on gene expression data demonstrate the applicability and effectiveness of the proposed approach.
Mehmet Tan, Reda Alhajj, Faruk Polat
BIBE1
2008 Combining multiple types of biological data in constraint-based learning of gene regulatory networks
abstract
Due to the complex structure and scale of gene regulatory networks, we support the argument that combination of multiple types of biological data to derive satisfactory network structures is necessary to understand the regulatory mechanisms of cellular systems. In this paper, we propose a simple but effective method of combining two types of biological data, namely microarray and transcription factor (TF) binding data, to construct gene regulatory networks. The proposed algorithm is based on and extends the well-known PC algorithm. Further, we developed a method for measuring the significance of the interactions between the genes and the TFs. The reported test results on both synthetic and real data sets demonstrate the applicability and effectiveness of the proposed approach; we also report the results of some comparative analysis that highlights the power of the proposed approach.
Mehmet Tan, Mohammed Al-Shalalfa, Reda Alhajj, Faruk Polat
CIBCB1
2007 Feature Reduction for Gene Regulatory Network Control
abstract
Scalability is one of the most important issues in control problems, including the control of gene regulatory networks. In this paper, we argue that it is possible to improve scalability of gene regulatory networks control by reducing the number of genes to be considered by the control policy; and consequently propose a novel method to estimate genes that are less important for control. The reported test results on real and synthetic data demonstrate the applicability and effectiveness of the proposed approach.
Mehmet Tan, Faruk Polat, Reda Alhajj
BIBE1