Rengül Çetin-Atalay

dblp:33/3424 · DBLP profile ↗
← Back
17ranked-venue papers
0as first author
5since 2021 · last 2026
0000-0003-2408-6606ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 14 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 Molecular contrastive learning with graph attention network (MoCL-GAT) for enhanced molecular representation
abstract
BACKGROUND: Learning the representation of molecules is crucial for drug discovery but is often hindered by the scarcity of labeled experimental data, which limits the performance of supervised machine learning models. While self-supervised learning (SSL) offers a solution by leveraging vast unlabeled chemical databases, many existing methods focus on learning from either local structural information or global molecular properties, but not both simultaneously. We introduce MoCL-GAT, a novel contrastive and transfer learning-based SSL framework that addresses this gap by simultaneously learning from two complementary objectives. It combines a local contrastive task on molecular subgraphs to capture fine-grained chemical environments with a global predictive task to learn holistic molecular descriptors. This dual-objective approach, powered by a Graph Attention Network, is designed to create more robust, versatile, and transferable molecular representations. RESULTS: Pre-trained on 1.9 million compounds, MoCL-GAT was fine-tuned on diverse benchmarks. It achieved state-of-the-art performance on molecular property prediction tasks, with an AUROC of 0.928 on BBBP and 0.768 on SIDER, and top-ranking RMSEs of 0.570 for ESOL and 1.818 for FreeSolv. Critically, fine-tuned models consistently and significantly outperformed models trained from scratch, confirming the value of pre-training. CONCLUSIONS: These results validate that MoCL-GAT's dual-objective approach learns highly effective and transferable representations, enabling more accurate and data-efficient predictions for key cheminformatics challenges. The source code for MoCL-GAT is publicly available on Zenodo at https://doi.org/10.5281/zenodo.16927285 .
Alperen Dalkiran, Ahmet Süreyya Rifaioglu, Rengül Çetin-Atalay, Aybar C. Acar, Tunca Dogan, M. Volkan Atalay
BMC Bioinform.3
2023 Transfer learning for drug-target interaction prediction
abstract
MOTIVATION: Utilizing AI-driven approaches for drug-target interaction (DTI) prediction require large volumes of training data which are not available for the majority of target proteins. In this study, we investigate the use of deep transfer learning for the prediction of interactions between drug candidate compounds and understudied target proteins with scarce training data. The idea here is to first train a deep neural network classifier with a generalized source training dataset of large size and then to reuse this pre-trained neural network as an initial configuration for re-training/fine-tuning purposes with a small-sized specialized target training dataset. To explore this idea, we selected six protein families that have critical importance in biomedicine: kinases, G-protein-coupled receptors (GPCRs), ion channels, nuclear receptors, proteases, and transporters. In two independent experiments, the protein families of transporters and nuclear receptors were individually set as the target datasets, while the remaining five families were used as the source datasets. Several size-based target family training datasets were formed in a controlled manner to assess the benefit provided by the transfer learning approach. RESULTS: Here, we present a systematic evaluation of our approach by pre-training a feed-forward neural network with source training datasets and applying different modes of transfer learning from the pre-trained source network to a target dataset. The performance of deep transfer learning is evaluated and compared with that of training the same deep neural network from scratch. We found that when the training dataset contains fewer than 100 compounds, transfer learning outperforms the conventional strategy of training the system from scratch, suggesting that transfer learning is advantageous for predicting binders to under-studied targets. AVAILABILITY AND IMPLEMENTATION: The source code and datasets are available at https://github.com/cansyl/TransferLearning4DTI. Our web-based service containing the ready-to-use pre-trained models is accessible at https://tl4dti.kansil.org.
Alperen Dalkiran, Ahmet Atakan, Ahmet Süreyya Rifaioglu, Maria Jesus Martin, Rengül Çetin-Atalay, Aybar C. Acar, Tunca Dogan, Volkan Atalay
Bioinform.5
2022 SLPred: a multi-view subcellular localization prediction tool for multi-location human proteins
abstract
SUMMARY: Accurate prediction of the subcellular locations (SLs) of proteins is a critical topic in protein science. In this study, we present SLPred, an ensemble-based multi-view and multi-label protein subcellular localization prediction tool. For a query protein sequence, SLPred provides predictions for nine main SLs using independent machine-learning models trained for each location. We used UniProtKB/Swiss-Prot human protein entries and their curated SL annotations as our source data. We connected all disjoint terms in the UniProt SL hierarchy based on the corresponding term relationships in the cellular component category of Gene Ontology and constructed a training dataset that is both reliable and large scale using the re-organized hierarchy. We tested SLPred on multiple benchmarking datasets including our-in house sets and compared its performance against six state-of-the-art methods. Results indicated that SLPred outperforms other tools in the majority of cases. AVAILABILITY AND IMPLEMENTATION: SLPred is available both as an open-access and user-friendly web-server (https://slpred.kansil.org) and a stand-alone tool (https://github.com/kansil/SLPred). All datasets used in this study are also available at https://slpred.kansil.org. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Gökhan Özsari, Ahmet Süreyya Rifaioglu, Ahmet Atakan, Tunca Dogan, Maria Jesus Martin, Rengül Çetin-Atalay, M. Volkan Atalay
Bioinform.6
2021 MDeePred: novel multi-channel protein featurization for deep learning-based binding affinity prediction in drug discovery
abstract
MOTIVATION: Identification of interactions between bioactive small molecules and target proteins is crucial for novel drug discovery, drug repurposing and uncovering off-target effects. Due to the tremendous size of the chemical space, experimental bioactivity screening efforts require the aid of computational approaches. Although deep learning models have been successful in predicting bioactive compounds, effective and comprehensive featurization of proteins, to be given as input to deep neural networks, remains a challenge. RESULTS: Here, we present a novel protein featurization approach to be used in deep learning-based compound-target protein binding affinity prediction. In the proposed method, multiple types of protein features such as sequence, structural, evolutionary and physicochemical properties are incorporated within multiple 2D vectors, which is then fed to state-of-the-art pairwise input hybrid deep neural networks to predict the real-valued compound-target protein interactions. The method adopts the proteochemometric approach, where both the compound and target protein features are used at the input level to model their interaction. The whole system is called MDeePred and it is a new method to be used for the purposes of computational drug discovery and repositioning. We evaluated MDeePred on well-known benchmark datasets and compared its performance with the state-of-the-art methods. We also performed in vitro comparative analysis of MDeePred predictions with selected kinase inhibitors' action on cancer cells. MDeePred is a scalable method with sufficiently high predictive performance. The featurization approach proposed here can also be utilized for other protein-related predictive tasks. AVAILABILITY AND IMPLEMENTATION: The source code, datasets, additional information and user instructions of MDeePred are available at https://github.com/cansyl/MDeePred. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ahmet Süreyya Rifaioglu, Rengül Çetin-Atalay, Deniz Cansen Kahraman, Tunca Dogan, Maria Jesus Martin, Volkan Atalay
Bioinform.2
2021 Protein domain-based prediction of drug/compound-target interactions and experimental validation on LIM kinases
abstract
Predictive approaches such as virtual screening have been used in drug discovery with the objective of reducing developmental time and costs. Current machine learning and network-based approaches have issues related to generalization, usability, or model interpretability, especially due to the complexity of target proteins' structure/function, and bias in system training datasets. Here, we propose a new method "DRUIDom" (DRUg Interacting Domain prediction) to identify bio-interactions between drug candidate compounds and targets by utilizing the domain modularity of proteins, to overcome problems associated with current approaches. DRUIDom is composed of two methodological steps. First, ligands/compounds are statistically mapped to structural domains of their target proteins, with the aim of identifying their interactions. As such, other proteins containing the same mapped domain or domain pair become new candidate targets for the corresponding compounds. Next, a million-scale dataset of small molecule compounds, including those mapped to domains in the previous step, are clustered based on their molecular similarities, and their domain associations are propagated to other compounds within the same clusters. Experimentally verified bioactivity data points, obtained from public databases, are meticulously filtered to construct datasets of active/interacting and inactive/non-interacting drug/compound-target pairs (~2.9M data points), and used as training data for calculating parameters of compound-domain mappings, which led to 27,032 high-confidence associations between 250 domains and 8,165 compounds, and a finalized output of ~5 million new compound-protein interactions. DRUIDom is experimentally validated by syntheses and bioactivity analyses of compounds predicted to target LIM-kinase proteins, which play critical roles in the regulation of cell motility, cell cycle progression, and differentiation through actin filament dynamics. We showed that LIMK-inhibitor-2 and its derivatives significantly block the cancer cell migration through inhibition of LIMK phosphorylation and the downstream protein cofilin. One of the derivative compounds (LIMKi-2d) was identified as a promising candidate due to its action on resistant Mahlavu liver cancer cells. The results demonstrated that DRUIDom can be exploited to identify drug candidate compounds for intended targets and to predict new target proteins based on the defined compound-domain relationships. Datasets, results, and the source code of DRUIDom are fully-available at: https://github.com/cansyl/DRUIDom.
Tunca Dogan, Ece Akhan, Marcus Baumann, Altay Koyas, Heval Atas, Ian R. Baxendale, Maria Jesus Martin, Rengül Çetin-Atalay
PLoS Comput. Biol.8
2020 iBioProVis: interactive visualization and analysis of compound bioactivity space
abstract
Bioinformatics (2010) doi: 10.1093/bioinformatics/btaa496 Co-author Tunca Doğan’s name was initially misspelled as Tunca Doğana. This has been corrected online.
Ataberk Donmez, Ahmet Süreyya Rifaioglu, Aybar C. Acar, Tunca Dogan, Rengül Çetin-Atalay, Volkan Atalay
Bioinform.5
2020 iBioProVis: interactive visualization and analysis of compound bioactivity space
abstract
SUMMARY: iBioProVis is an interactive tool for visual analysis of the compound bioactivity space in the context of target proteins, drugs and drug candidate compounds. iBioProVis tool takes target protein identifiers and, optionally, compound SMILES as input, and uses the state-of-the-art non-linear dimensionality reduction method t-Distributed Stochastic Neighbor Embedding (t-SNE) to plot the distribution of compounds embedded in a 2D map, based on the similarity of structural properties of compounds and in the context of compounds' cognate targets. Similar compounds, which are embedded to proximate points on the 2D map, may bind the same or similar target proteins. Thus, iBioProVis can be used to easily observe the structural distribution of one or two target proteins' known ligands on the 2D compound space, and to infer new binders to the same protein, or to infer new potential target(s) for a compound of interest, based on this distribution. Principal component analysis (PCA) projection of the input compounds is also provided, Hence the user can interactively observe the same compound or a group of selected compounds which is projected by both PCA and embedded by t-SNE. iBioProVis also provides detailed information about drugs and drug candidate compounds through cross-references to widely used and well-known databases, in the form of linked table views. Two use-case studies were demonstrated, one being on angiotensin-converting enzyme 2 (ACE2) protein which is Severe Acute Respiratory Syndrome Coronavirus 2 (SARS-CoV-2) Spike protein receptor. ACE2 binding compounds and seven antiviral drugs were closely embedded in which two of them have been under clinical trial for Coronavirus disease 19 (COVID-19). AVAILABILITY AND IMPLEMENTATION: iBioProVis and its carefully filtered dataset are available at https://ibpv.kansil.org/ for public use. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ataberk Donmez, Ahmet Süreyya Rifaioglu, Aybar C. Acar, Tunca Dogan, Rengül Çetin-Atalay, Volkan Atalay, Wren Jonathan
Bioinform.5
2020 DeepDistance: A multi-task deep regression model for cell detection in inverted microscopy images
Can Koyuncu 0001, Gozde Nur Gunesli, Rengül Çetin-Atalay, Cigdem Demir
Medical Image Anal.3
2019 Recent applications of deep learning and machine intelligence on in silico drug discovery: methods, tools and databases
abstract
The identification of interactions between drugs/compounds and their targets is crucial for the development of new drugs. In vitro screening experiments (i.e. bioassays) are frequently used for this purpose; however, experimental approaches are insufficient to explore novel drug-target interactions, mainly because of feasibility problems, as they are labour intensive, costly and time consuming. A computational field known as 'virtual screening' (VS) has emerged in the past decades to aid experimental drug discovery studies by statistically estimating unknown bio-interactions between compounds and biological targets. These methods use the physico-chemical and structural properties of compounds and/or target proteins along with the experimentally verified bio-interaction information to generate predictive models. Lately, sophisticated machine learning techniques are applied in VS to elevate the predictive performance. The objective of this study is to examine and discuss the recent applications of machine learning techniques in VS, including deep learning, which became highly popular after giving rise to epochal developments in the fields of computer vision and natural language processing. The past 3 years have witnessed an unprecedented amount of research studies considering the application of deep learning in biomedicine, including computational drug discovery. In this review, we first describe the main instruments of VS methods, including compound and protein features (i.e. representations and descriptors), frequently used libraries and toolkits for VS, bioactivity databases and gold-standard data sets for system training and benchmarking. We subsequently review recent VS studies with a strong emphasis on deep learning applications. Finally, we discuss the present state of the field, including the current challenges and suggest future directions. We believe that this survey will provide insight to the researchers working in the field of computational drug discovery in terms of comprehending and developing novel bio-prediction methods.
Ahmet Süreyya Rifaioglu, Heval Atas, Maria Jesus Martin, Rengül Çetin-Atalay, Volkan Atalay, Tunca Dogan
Briefings Bioinform.4
2018 ECPred: a tool for the prediction of the enzymatic functions of protein sequences based on the EC nomenclature
abstract
BACKGROUND: The automated prediction of the enzymatic functions of uncharacterized proteins is a crucial topic in bioinformatics. Although several methods and tools have been proposed to classify enzymes, most of these studies are limited to specific functional classes and levels of the Enzyme Commission (EC) number hierarchy. Besides, most of the previous methods incorporated only a single input feature type, which limits the applicability to the wide functional space. Here, we proposed a novel enzymatic function prediction tool, ECPred, based on ensemble of machine learning classifiers. RESULTS: In ECPred, each EC number constituted an individual class and therefore, had an independent learning model. Enzyme vs. non-enzyme classification is incorporated into ECPred along with a hierarchical prediction approach exploiting the tree structure of the EC nomenclature. ECPred provides predictions for 858 EC numbers in total including 6 main classes, 55 subclass classes, 163 sub-subclass classes and 634 substrate classes. The proposed method is tested and compared with the state-of-the-art enzyme function prediction tools by using independent temporal hold-out and no-Pfam datasets constructed during this study. CONCLUSIONS: ECPred is presented both as a stand-alone and a web based tool to provide probabilistic enzymatic function predictions (at all five levels of EC) for uncharacterized protein sequences. Also, the datasets of this study will be a valuable resource for future benchmarking studies. ECPred is available for download, together with all of the datasets used in this study, at: https://github.com/cansyl/ECPred . ECPred webserver can be accessed through http://cansyl.metu.edu.tr/ECPred.html .
Alperen Dalkiran, Ahmet Süreyya Rifaioglu, Maria Jesus Martin, Rengül Çetin-Atalay, Volkan Atalay, Tunca Dogan
BMC Bioinform.4
2015 Multi-resolution super-pixels and their applications on fluorescent mesenchymal stem cells images using 1-D SIFT merging
abstract
A new multi-resolution super-pixel based algorithm is proposed to track cell size, count and motion in Mesenchymal Stem Cells (MSCs) images. Multi-resolution super-pixels are obtained by placing varying density seeds on the image. The density of the seeds are determined according to the local high frequency components of the MSCs image. In this way a multi-resolution super-pixels decomposition of the image is obtained. A second contribution of the paper is novel decision rule for merging similar neighboring super-pixels. One-dimensional version of the well known scale invariant feature transform (SIFT) is developed and applied to the histograms of the neighboring super-pixels to determine similar regions. The proposed algorithm is experimentally shown to be successful in segmenting and tracking cells in MSCs images.
Onur Yorulmaz, Oguzhan Oguz, Ece Akhan, Donus Tuncel, Rengül Çetin-Atalay, A. Enis Çetin
ICIP5
2013 A multiplication-free framework for signal processing and applications in biomedical image analysis
abstract
A new framework for signal processing is introduced based on a novel vector product definition that permits a multiplier-free implementation. First a new product of two real numbers is defined as the sum of their absolute values, with the sign determined by product of the hard-limited numbers. This new product of real numbers is used to define a similar product of vectors in RN. The new vector product of two identical vectors reduces to a scaled version of the l1norm of the vector. The main advantage of this framework is that it yields multiplication-free computationally efficient algorithms for performing some important tasks in signal processing. An application to the problem of cancer cell line image classification is presented that uses the notion of a co-difference matrix that is analogous to a covariance matrix except that the vector products are based on our new proposed framework. Results show the effectiveness of this approach when the proposed co-difference matrix is compared with a covariance matrix.
Alexander Suhre, Musa Furkan Keskin, Tulin Ersahin, Rengül Çetin-Atalay, Rashid Ansari, A. Enis Çetin
ICASSP4
2013 Attributed Relational Graphs for Cell Nucleus Segmentation in Fluorescence Microscopy Images
abstract
More rapid and accurate high-throughput screening in molecular cellular biology research has become possible with the development of automated microscopy imaging, for which cell nucleus segmentation commonly constitutes the core step. Although several promising methods exist for segmenting the nuclei of monolayer isolated and less-confluent cells, it still remains an open problem to segment the nuclei of more-confluent cells, which tend to grow in overlayers. To address this problem, we propose a new model-based nucleus segmentation algorithm. This algorithm models how a human locates a nucleus by identifying the nucleus boundaries and piecing them together. In this algorithm, we define four types of primitives to represent nucleus boundaries at different orientations and construct an attributed relational graph on the primitives to represent their spatial relations. Then, we reduce the nucleus identification problem to finding predefined structural patterns in the constructed graph and also use the primitives in region growing to delineate the nucleus borders. Working with fluorescence microscopy images, our experiments demonstrate that the proposed algorithm identifies nuclei better than previous nucleus segmentation algorithms.
Salim Arslan, Tulin Ersahin, Rengül Çetin-Atalay, Cigdem Demir
IEEE Trans. Medical Imaging3
2012 Microscopic image classification via ℂWT-based covariance descriptors using Kullback-Leibler distance
abstract
In this paper, we present a novel method for classification of cancer cell line images using complex wavelet-based region covariance matrix descriptors. Microscopic images containing irregular carcinoma cell patterns are represented by randomly selected subwindows which possibly correspond to foreground pixels. For each subwindow, a new region descriptor utilizing the dual-tree complex wavelet transform coefficients as pixel features is computed. ℂWT as a feature extraction tool is preferred primarily because of its ability to characterize singularities at multiple orientations, which often arise in carcinoma cell lines, and approximate shift invariance property. We propose new dissimilarity measures between covariance matrices based on Kullback-Leibler (KL) divergence and L2-norm, which turn out to be as successful as the classical KL divergence, but with much less computational complexity. Experimental results demonstrate the effectiveness of the proposed image classification framework. The proposed algorithm outperforms the recently published eigenvalue-based Bayesian classification method.
Musa Furkan Keskin, A. Enis Çetin, Tulin Ersahin, Rengül Çetin-Atalay
ISCAS4
2005 Implicit motif distribution based hybrid computational kernel for sequence classification
abstract
MOTIVATION: We designed a general computational kernel for classification problems that require specific motif extraction and search from sequences. Instead of searching for explicit motifs, our approach finds the distribution of implicit motifs and uses as a feature for classification. Implicit motif distribution approach may be used as modus operandi for bioinformatics problems that require specific motif extraction and search, which is otherwise computationally prohibitive. RESULTS: A system named P2SL that infer protein subcellular targeting was developed through this computational kernel. Targeting-signal was modeled by the distribution of subsequence occurrences (implicit motifs) using self-organizing maps. The boundaries among the classes were then determined with a set of support vector machines. P2SL hybrid computational system achieved approximately 81% of prediction accuracy rate over ER targeted, cytosolic, mitochondrial and nuclear protein localization classes. P2SL additionally offers the distribution potential of proteins among localization classes, which is particularly important for proteins, shuttle between nucleus and cytosol. AVAILABILITY: http://staff.vbi.vt.edu/volkan/p2sl and http://www.i-cancer.fen.bilkent.edu.tr/p2sl CONTACT: [email protected].
Volkan Atalay, Rengül Çetin-Atalay
Bioinform.2
2004 An ontology for collaborative construction and analysis of cellular pathways
abstract
MOTIVATION: As the scientific curiosity in genome studies shifts toward identification of functions of the genomes in large scale, data produced about cellular processes at molecular level has been accumulating with an accelerating rate. In this regard, it is essential to be able to store, integrate, access and analyze this data effectively with the help of software tools. Clearly this requires a strong ontology that is intuitive, comprehensive and uncomplicated. RESULTS: We define an ontology for an intuitive, comprehensive and uncomplicated representation of cellular events. The ontology presented here enables integration of fragmented or incomplete pathway information via collaboration, and supports manipulation of the stored data. In addition, it facilitates concurrent modifications to the data while maintaining its validity and consistency. Furthermore, novel structures for representation of multiple levels of abstraction for pathways and homologies is provided. Lastly, our ontology supports efficient querying of large amounts of data. We have also developed a software tool named pathway analysis tool for integration and knowledge acquisition (PATIKA) providing an integrated, multi-user environment for visualizing and manipulating network of cellular events. PATIKA implements the basics of our ontology.
Emek Demir, Ozgun Babur, Ugur Dogrusoz, Attila Gürsoy, A. Ayaz, Gürcan Gülesir, Gurkan Nisanci, Rengül Çetin-Atalay
Bioinform.8
2002 PATIKA: an integrated visual environment for collaborative construction and analysis of cellular pathways
abstract
MOTIVATION: Availability of the sequences of entire genomes shifts the scientific curiosity towards the identification of function of the genomes in large scale as in genome studies. In the near future, data produced about cellular processes at molecular level will accumulate with an accelerating rate as a result of proteomics studies. In this regard, it is essential to develop tools for storing, integrating, accessing, and analyzing this data effectively. RESULTS: We define an ontology for a comprehensive representation of cellular events. The ontology presented here enables integration of fragmented or incomplete pathway information and supports manipulation and incorporation of the stored data, as well as multiple levels of abstraction. Based on this ontology, we present the architecture of an integrated environment named Patika (Pathway Analysis Tool for Integration and Knowledge Acquisition). Patika is composed of a server-side, scalable, object-oriented database and client-side editors to provide an integrated, multi-user environment for visualizing and manipulating network of cellular events. This tool features automated pathway layout, functional computation support, advanced querying and a user-friendly graphical interface. We expect that Patika will be a valuable tool for rapid knowledge acquisition, microarray generated large-scale data interpretation, disease gene identification, and drug development. AVAILABILITY: A prototype of Patika is available upon request from the authors.
Emek Demir, Ozgun Babur, Ugur Dogrusoz, Attila Gürsoy, Gurkan Nisanci, Rengül Çetin-Atalay, Mehmet Ozturk
Bioinform.6