Attila Gürsoy

dblp:g/AttilaGursoy · DBLP profile ↗
← Back
21ranked-venue papers
4as first author
5since 2021 · last 2025
0000-0002-2297-2113ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 16 · 5 since 2021Systems, architecture and hardware · 5 · 4 first-author
YearPublicationVenuePosition
2025 DeepAllo: allosteric site prediction using protein language model (pLM) with multitask learning
abstract
MOTIVATION: Allostery, the process by which binding at one site perturbs a distant site, is being rendered as a key focus in the field of drug development with its substantial impact on protein function. The identification of allosteric pockets (sites) is a challenging task and several techniques have been developed, including Machine Learning to predict allosteric pockets that utilize both static and pocket features. RESULTS: Our work, DeepAllo, is the first study that combines fine-tuned protein language model (pLM) with FPocket features and shows an increase in prediction performance of allosteric sites over previous studies. The pLM model was fine-tuned on AlloSteric Database (ASD) in Multitask Learning setting and was further used as a feature extractor to train XGBoost and AutoML models. The best model predicts allosteric pockets with 89.66% F1 score and 90.5% of allosteric pockets in the top 3 positions, outperforming previous results. A case study has been performed on proteins with known allosteric pockets, which shows the proof of our approach. Moreover, an effort was made to explain the pLM by visualizing its attention mechanism among allosteric and non-allosteric residues. AVAILABILITY AND IMPLEMENTATION: The source code is available on GitHub (https://github.com/MoaazK/deepallo) and archived on Zenodo (DOI: 10.5281/zenodo.15255379). The trained model is hosted on Hugging Face (DOI: 10.57967/hf/5198). The dataset used for training and evaluation is archived on Zenodo (DOI: 10.5281/zenodo.15255437).
Moaaz Khokhar, Ozlem Keskin, Attila Gürsoy
Bioinform.3
2024 Structural coverage of the human interactome
abstract
Complex biological processes in cells are embedded in the interactome, representing the complete set of protein-protein interactions. Mapping and analyzing the protein structures are essential to fully comprehending these processes' molecular details. Therefore, knowing the structural coverage of the interactome is important to show the current limitations. Structural modeling of protein-protein interactions requires accurate protein structures. In this study, we mapped all experimental structures to the reference human proteome. Later, we found the enrichment in structural coverage when complementary methods such as homology modeling and deep learning (AlphaFold) were included. We then collected the interactions from the literature and databases to form the reference human interactome, resulting in 117 897 non-redundant interactions. When we analyzed the structural coverage of the interactome, we found that the number of experimentally determined protein complex structures is scarce, corresponding to 3.95% of all binary interactions. We also analyzed known and modeled structures to potentially construct the structural interactome with a docking method. Our analysis showed that 12.97% of the interactions from HuRI and 73.62% and 32.94% from the filtered versions of STRING and HIPPIE could potentially be modeled with high structural coverage or accuracy, respectively. Overall, this paper provides an overview of the current state of structural coverage of the human proteome and interactome.
Kayra Kosoglu, Zeynep Aydin, Nurcan Tuncbag, Attila Gürsoy, Ozlem Keskin
Briefings Bioinform.4
2023 The interaction between mutated DYRK1A and APP suggests a potential target for vascular cognitive impairment
abstract
Vascular cognitive impairment, being a cerebrovascular condition that is more common in the aging population, results in a deterioration in cognitive capacities. Regarding these aspects, the identification of associated proteins, pathways, and treatment targets is crucial. The interactions that result in Vascular Cognitive Impairment can be better understood by looking into the interactors of the well-researched APP, which has been linked to dementia. Therefore, our study focused on identifying novel APP interactions by protein-protein interaction networks. After constructing the APP-centered dynamic-structural PPI network we investigated these interactions with PRISM. We suggested a number of interactions that could be relevant to VCI and found that APP binds more selectively to SORT1, MAPK8, and UBQLN1 proteins. Additionally, we propose that three mutations in the kinase DYRK1A (V165I, S337P, and D401G) may be essential for VCI through its interaction with APP.
Melisa Ece Zeylan, Attila Gürsoy, Simge Senyuz, Ozlem Keskin
BIBM2
2022 HMI-PRED 2.0: a biologist-oriented web application for prediction of host-microbe protein-protein interaction by interface mimicry
abstract
SUMMARY: HMI-PRED 2.0 is a publicly available web service for the prediction of host-microbe protein-protein interaction by interface mimicry that is intended to be used without extensive computational experience. A microbial protein structure is screened against a database covering the entire available structural space of complexes of known human proteins. AVAILABILITY AND IMPLEMENTATION: HMI-PRED 2.0 provides user-friendly graphic interfaces for predicting, visualizing and analyzing host-microbe interactions. HMI-PRED 2.0 is available at https://hmipred.org/.
Hansaim Lim, Chung-Jung Tsai, Ozlem Keskin, Ruth Nussinov, Attila Gürsoy
Bioinform.5
2022 SARS-CoV-2 Interactome 3D: A Web interface for 3D visualization and analysis of SARS-CoV-2-human mimicry and interactions
abstract
SUMMARY: We present a web-based server for navigating and visualizing possible interactions between SARS-CoV-2 and human host proteins. The interactions are obtained from HMI_Pred which relies on the rationale that virus proteins mimic host proteins. The structural alignment of the viral protein with one side of the human protein-protein interface determines the mimicry. The mimicked human proteins and predicted interactions, and the binding sites are presented. The user can choose one of the 18 SARS-CoV-2 protein structures and visualize the potential 3D complexes it forms with human proteins. The mimicked interface is also provided. The user can superimpose two interacting human proteins in order to see whether they bind to the same site or different sites on the viral protein. The server also tabulates all available mimicked interactions together with their match scores and number of aligned residues. This is the first server listing and cataloging all interactions between SARS-CoV-2 and human protein structures, enabled by our innovative interface mimicry strategy. AVAILABILITY AND IMPLEMENTATION: The server is available at https://interactome.ku.edu.tr/sars/.
Damla Ovek, Ameer Taweel, Zeynep Abali, Ece Tezsezen, Yunus Emre Koroglu, Chung-Jung Tsai, Ruth Nussinov, Ozlem Keskin, Attila Gürsoy
Bioinform.9
2019 3D spatial organization and network-guided comparison of mutation profiles in Glioblastoma reveals similarities across patients
abstract
Glioblastoma multiforme (GBM) is the most aggressive type of brain tumor. Molecular heterogeneity is a hallmark of GBM tumors that is a barrier in developing treatment strategies. In this study, we used the nonsynonymous mutations of GBM tumors deposited in The Cancer Genome Atlas (TCGA) and applied a systems level approach based on biophysical characteristics of mutations and their organization in patient-specific subnetworks to reduce inter-patient heterogeneity and to gain potential clinically relevant insights. Approximately 10% of the mutations are located in "patches" which are defined as the set of residues spatially in close proximity that are mutated across multiple patients. Grouping mutations as 3D patches reduces the heterogeneity across patients. There are multiple patches that are relatively small in oncogenes, whereas there are a small number of very large patches in tumor suppressors. Additionally, different patches in the same protein are often located at different domains that can mediate different functions. We stratified the patients into five groups based on their potentially affected pathways that are revealed from the patient-specific subnetworks. These subnetworks were constructed by integrating mutation profiles of the patients with the interactome data. Network-guided clustering showed significant association between the groups and patient survival (P-value = 0.0408). Also, each group carries a set of signature 3D mutation patches that affect predominant pathways. We integrated drug sensitivity data of GBM cell lines with the mutation patches and the patient groups to analyze the possible therapeutic outcome of these patches. We found that Pazopanib might be effective in Group 3 by targeting CSF1R. Additionally, inhibiting ATM that is a mediator of PTEN phosphorylation may be ineffective in Group 2. We believe that from mutations to networks and eventually to clinical and therapeutic data, this study provides a novel perspective in the network-guided precision medicine.
Cansu Dincer, Tugba Turkoglu Kaya, Ozlem Keskin, Attila Gürsoy, Nurcan Tuncbag
PLoS Comput. Biol.4
2018 Analysis of single amino acid variations in singlet hot spots of protein-protein interfaces
abstract
Motivation: Single amino acid variations (SAVs) in protein-protein interaction (PPI) sites play critical roles in diseases. PPI sites (interfaces) have a small subset of residues called hot spots that contribute significantly to the binding energy, and they may form clusters called hot regions. Singlet hot spots are the single amino acid hot spots outside of the hot regions. The distribution of SAVs on the interface residues may be related to their disease association. Results: We performed statistical and structural analyses of SAVs with literature curated experimental thermodynamics data, and demonstrated that SAVs which destabilize PPIs are more likely to be found in singlet hot spots rather than hot regions and energetically less important interface residues. In contrast, non-hot spot residues are significantly enriched in neutral SAVs, which do not affect PPI stability. Surprisingly, we observed that singlet hot spots tend to be enriched in disease-causing SAVs, while benign SAVs significantly occur in non-hot spot residues. Our work demonstrates that SAVs in singlet hot spot residues have significant effect on protein stability and function. Availability and implementation: The dataset used in this paper is available as Supplementary Material. The data can be found at http://prism.ccbb.ku.edu.tr/data/sav/ as well. Supplementary information: Supplementary data are available at Bioinformatics online.
E. Sila Ozdemir, Attila Gürsoy, Ozlem Keskin
Bioinform.2
2017 Topological, functional, and structural analyses of protein-protein interaction networks of breast cancer lung and brain metastases
abstract
Breast cancer is the second most common cause of death among women. However, it is not deadly if the cancerous cells remain in the breast. The life threat starts when cancerous cells travel to other parts of body like lung, liver, bone and brain. So, most breast cancer deaths derive from metastasis to other organs. In this study, we introduce novel proteins and cellular pathways that play important roles in brain and lung metastases of breast cancer using Protein-Protein Interaction (PPI) networks. Our topological analysis identified genes such as RPL5, MMP2 and DPP4 which are already known to be associated with lung or brain metastasis. Additionally, we found four and nine novel candidate genes that are specific to lung and brain metastases, respectively. The functional enrichment analysis showed that KEGG pathways associated with the immune system and infectious diseases, particularly the chemokine signaling pathway, are important for lung metastasis. On the other hand, pathways related to genetic information processing were more involved in brain metastasis. By enriching the traditional PPI network with protein structural data, we show the effects of mutations on specific protein-protein interactions. By using the different conformations of protein CXCL12, we show the effect of H25R mutation on CXCL12 dimerization.
Farideh Halakou, Attila Gürsoy, Emel Sen Kilic, Ozlem Keskin
CIBCB2
2014 The Structural Pathway of Interleukin 1 (IL-1) Initiated Signaling Reveals Mechanisms of Oncogenic Mutations and SNPs in Inflammation and Cancer
abstract
Interleukin-1 (IL-1) is a large cytokine family closely related to innate immunity and inflammation. IL-1 proteins are key players in signaling pathways such as apoptosis, TLR, MAPK, NLR and NF-κB. The IL-1 pathway is also associated with cancer, and chronic inflammation increases the risk of tumor development via oncogenic mutations. Here we illustrate that the structures of interfaces between proteins in this pathway bearing the mutations may reveal how. Proteins are frequently regulated via their interactions, which can turn them ON or OFF. We show that oncogenic mutations are significantly at or adjoining interface regions, and can abolish (or enhance) the protein-protein interaction, making the protein constitutively active (or inactive, if it is a repressor). We combine known structures of protein-protein complexes and those that we have predicted for the IL-1 pathway, and integrate them with literature information. In the reconstructed pathway there are 104 interactions between proteins whose three dimensional structures are experimentally identified; only 15 have experimentally-determined structures of the interacting complexes. By predicting the protein-protein complexes throughout the pathway via the PRISM algorithm, the structural coverage increases from 15% to 71%. In silico mutagenesis and comparison of the predicted binding energies reveal the mechanisms of how oncogenic and single nucleotide polymorphism (SNP) mutations can abrogate the interactions or increase the binding affinity of the mutant to the native partner. Computational mapping of mutations on the interface of the predicted complexes may constitute a powerful strategy to explain the mechanisms of activation/inhibition. It can also help explain how an oncogenic mutation or SNP works.
Saliha Ece Acuner, Attila Gürsoy, Ruth Nussinov, Ozlem Keskin
PLoS Comput. Biol.2
2010 Interaction prediction and classification of PDZ domains
abstract
BACKGROUND: PDZ domain is a well-conserved, structural protein domain found in hundreds of signaling proteins that are otherwise unrelated. PDZ domains can bind to the C-terminal peptides of different proteins and act as glue, clustering different protein complexes together, targeting specific proteins and routing these proteins in signaling pathways. These domains are classified into classes I, II and III, depending on their binding partners and the nature of bonds formed. Binding specificities of PDZ domains are very crucial in order to understand the complexity of signaling pathways. It is still an open question how these domains recognize and bind their partners. RESULTS: The focus of the current study is two folds: 1) predicting to which peptides a PDZ domain will bind and 2) classification of PDZ domains, as Class I, II or I-II, given the primary sequences of the PDZ domains. Trigram and bigram amino acid frequencies are used as features in machine learning methods. Using 85 PDZ domains and 181 peptides, our model reaches high prediction accuracy (91.4%) for binary interaction prediction which outperforms previously investigated similar methods. Also, we can predict classes of PDZ domains with an accuracy of 90.7%. We propose three critical amino acid sequence motifs that could have important roles on specificity pattern of PDZ domains. CONCLUSIONS: Our model on PDZ interaction dataset shows that our approach produces encouraging results. The method can be further used as a virtual screening technique to reduce the search space for putative candidate target proteins and drug-like molecules of PDZ domains.
Sibel Kalyoncu, Ozlem Keskin, Attila Gürsoy
BMC Bioinform.3
2009 A survey of available tools and web servers for analysis of protein-protein interactions and interfaces
abstract
The unanimous agreement that cellular processes are (largely) governed by interactions between proteins has led to enormous community efforts culminating in overwhelming information relating to these proteins; to the regulation of their interactions, to the way in which they interact and to the function which is determined by these interactions. These data have been organized in databases and servers. However, to make these really useful, it is essential not only to be aware of these, but in particular to have a working knowledge of which tools to use for a given problem; what are the tool advantages and drawbacks; and no less important how to combine these for a particular goal since usually it is not one tool, but some combination of tool-modules that is needed. This is the goal of this review.
Nurcan Tuncbag, Gozde Kar, Ozlem Keskin, Attila Gürsoy, Ruth Nussinov
Briefings Bioinform.4
2009 Identification of computational hot spots in protein interfaces: combining solvent accessibility and inter-residue potentials improves the accuracy
abstract
MOTIVATION: Hot spots are residues comprising only a small fraction of interfaces yet accounting for the majority of the binding energy. These residues are critical in understanding the principles of protein interactions. Experimental studies like alanine scanning mutagenesis require significant effort; therefore, there is a need for computational methods to predict hot spots in protein interfaces. RESULTS: We present a new intuitive efficient method to determine computational hot spots based on conservation (C), solvent accessibility [accessible surface area (ASA)] and statistical pairwise residue potentials (PP) of the interface residues. Combination of these features is examined in a comprehensive way to study their effect in hot spot detection. The predicted hot spots are observed to match with the experimental hot spots with an accuracy of 70% and a precision of 64% in Alanine Scanning Energetics Database (ASEdb), and accuracy of 70% and a precision of 73% in Binding Interface Database (BID). Several machine learning methods are also applied to predict hot spots. Performance of our empirical approach exceeds learning-based methods and other existing hot spot prediction methods. Residue occlusion from solvent in the complexes and pairwise potentials are found to be the main discriminative features in hot spot prediction. CONCLUSION: Our empirical method is a simple approach in hot spot prediction yet with its high accuracy and computational effectiveness. We believe that this method provides insights for the researchers working on characterization of protein binding sites and design of specific therapeutic agents for protein interactions. AVAILABILITY: The list of training and test sets are available as Supplementary Data at http://prism.ccbb.ku.edu.tr/hotpoint/supplement.doc. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Nurcan Tuncbag, Attila Gürsoy, Ozlem Keskin
Bioinform.2
2009 Human Cancer Protein-Protein Interaction Network: A Structural Perspective
abstract
Protein-protein interaction networks provide a global picture of cellular function and biological processes. Some proteins act as hub proteins, highly connected to others, whereas some others have few interactions. The dysfunction of some interactions causes many diseases, including cancer. Proteins interact through their interfaces. Therefore, studying the interface properties of cancer-related proteins will help explain their role in the interaction networks. Similar or overlapping binding sites should be used repeatedly in single interface hub proteins, making them promiscuous. Alternatively, multi-interface hub proteins make use of several distinct binding sites to bind to different partners. We propose a methodology to integrate protein interfaces into cancer interaction networks (ciSPIN, cancer structural protein interface network). The interactions in the human protein interaction network are replaced by interfaces, coming from either known or predicted complexes. We provide a detailed analysis of cancer related human protein-protein interfaces and the topological properties of the cancer network. The results reveal that cancer-related proteins have smaller, more planar, more charged and less hydrophobic binding sites than non-cancer proteins, which may indicate low affinity and high specificity of the cancer-related interactions. We also classified the genes in ciSPIN according to phenotypes. Within phenotypes, for breast cancer, colorectal cancer and leukemia, interface properties were found to be discriminating from non-cancer interfaces with an accuracy of 71%, 67%, 61%, respectively. In addition, cancer-related proteins tend to interact with their partners through distinct interfaces, corresponding mostly to multi-interface hubs, which comprise 56% of cancer-related proteins, and constituting the nodes with higher essentiality in the network (76%). We illustrate the interface related affinity properties of two cancer-related hub proteins: Erbb3, a multi interface, and Raf1, a single interface hub. The results reveal that affinity of interactions of the multi-interface hub tends to be higher than that of the single-interface hub. These findings might be important in obtaining new targets in cancer as well as finding the details of specific binding regions of putative cancer drug candidates.
Gozde Kar, Attila Gürsoy, Ozlem Keskin
PLoS Comput. Biol.2
2005 Prediction of protein-protein interactions by combining structure and sequence conservation in protein interfaces
abstract
MOTIVATION: Elucidation of the full network of protein-protein interactions is crucial for understanding of the principles of biological systems and processes. Thus, there is a need for in silico methods for predicting interactions. We present a novel algorithm for automated prediction of protein-protein interactions that employs a unique bottom-up approach combining structure and sequence conservation in protein interfaces. RESULTS: Running the algorithm on a template dataset of 67 interfaces and a sequentially non-redundant dataset of 6170 protein structures, 62 616 potential interactions are predicted. These interactions are compared with the ones in two publicly available interaction databases (Database of Interacting Proteins and Biomolecular Interaction Network Database) and also the Protein Data Bank. A significant number of predictions are verified in these databases. The unverified ones may correspond to (1) interactions that are not covered in these databases but known in literature, (2) unknown interactions that actually occur in nature and (3) interactions that do not occur naturally but may possibly be realized synthetically in laboratory conditions. Some unverified interactions, supported significantly with studies found in the literature, are discussed. AVAILABILITY: http://gordion.hpc.eng.ku.edu.tr/prism CONTACT: [email protected]; [email protected].
A. Selim Aytuna, Attila Gürsoy, Ozlem Keskin
Bioinform.2
2005 PHR: A Parallel Hierarchical Radiosity System with Dynamic Load Balancing
Ali Kemal Sinop, Tolga Abaci, Ümit Akkus, Attila Gürsoy, Ugur Güdükbay
J. Supercomput.4
2004 An ontology for collaborative construction and analysis of cellular pathways
abstract
MOTIVATION: As the scientific curiosity in genome studies shifts toward identification of functions of the genomes in large scale, data produced about cellular processes at molecular level has been accumulating with an accelerating rate. In this regard, it is essential to be able to store, integrate, access and analyze this data effectively with the help of software tools. Clearly this requires a strong ontology that is intuitive, comprehensive and uncomplicated. RESULTS: We define an ontology for an intuitive, comprehensive and uncomplicated representation of cellular events. The ontology presented here enables integration of fragmented or incomplete pathway information via collaboration, and supports manipulation of the stored data. In addition, it facilitates concurrent modifications to the data while maintaining its validity and consistency. Furthermore, novel structures for representation of multiple levels of abstraction for pathways and homologies is provided. Lastly, our ontology supports efficient querying of large amounts of data. We have also developed a software tool named pathway analysis tool for integration and knowledge acquisition (PATIKA) providing an integrated, multi-user environment for visualizing and manipulating network of cellular events. PATIKA implements the basics of our ontology.
Emek Demir, Ozgun Babur, Ugur Dogrusoz, Attila Gürsoy, A. Ayaz, Gürcan Gülesir, Gurkan Nisanci, Rengül Çetin-Atalay
Bioinform.4
2004 Performance and modularity benefits of message-driven execution
Attila Gürsoy, Laxmikant V. Kalé
J. Parallel Distributed Comput.1
2002 PATIKA: an integrated visual environment for collaborative construction and analysis of cellular pathways
abstract
MOTIVATION: Availability of the sequences of entire genomes shifts the scientific curiosity towards the identification of function of the genomes in large scale as in genome studies. In the near future, data produced about cellular processes at molecular level will accumulate with an accelerating rate as a result of proteomics studies. In this regard, it is essential to develop tools for storing, integrating, accessing, and analyzing this data effectively. RESULTS: We define an ontology for a comprehensive representation of cellular events. The ontology presented here enables integration of fragmented or incomplete pathway information and supports manipulation and incorporation of the stored data, as well as multiple levels of abstraction. Based on this ontology, we present the architecture of an integrated environment named Patika (Pathway Analysis Tool for Integration and Knowledge Acquisition). Patika is composed of a server-side, scalable, object-oriented database and client-side editors to provide an integrated, multi-user environment for visualizing and manipulating network of cellular events. This tool features automated pathway layout, functional computation support, advanced querying and a user-friendly graphical interface. We expect that Patika will be a valuable tool for rapid knowledge acquisition, microarray generated large-scale data interpretation, disease gene identification, and drug development. AVAILABILITY: A prototype of Patika is available upon request from the authors.
Emek Demir, Ozgun Babur, Ugur Dogrusoz, Attila Gürsoy, Gurkan Nisanci, Rengül Çetin-Atalay, Mehmet Ozturk
Bioinform.4
2001 Parallel Pruning for K-Means Clustering on Shared Memory Architectures
Attila Gürsoy, Ilker Cengiz
Euro-Par1
2000 Neighbourhood Preserving Load Balancing: A Self-Organizing Approach
Attila Gürsoy, Murat Atun
Euro-Par1
1991 High level support for divide-and-conquer parallelism
abstract
In thts paper we present a simple language based on C for expressing dtvide-and-conquer computations.The "language" conswts of a few stmple extensions to C. It allows for many varzatzons tn the standard divzde-and-conquer paradigm.It as tmptemented using the Chare Kernel parallel programming system.The Chare Kernel supports dynamic creation of work with dynamtc load balanctng strategies, and machine independent executton.As a result, implementation of languages and systems such as that described in thts paper as stmp!ified significantly.A translator translates divide-and-conquer programs to Chare Kernel programs, handling details of synch roni~ata'on and communication automatically.The design of the language is presented, followed by a description of its implementation, and performance results on many parallel machines, including NC UBE/two, iPSC/2, and the Sequent symmetry.User programs do not have to be changed to run on any of these machines.computation-intensive problems will be routinely speeded up using parallel processing.Although many commercial systems have appeared in the market, programming them to meet this expectation is still a challenging task.Parallel programming is obviously more difficult than sequential programming.It is necessary to simplify and support the task of writing parallel applications, and also to ensure that the investment in parallel software is protected through architectural advances and new generation of parallel machines.The Chare Kernel [11] is a machine-independent MIMD parallel programming system that is aimed at this objective.The system provides an explicitly parallel language -which uses C [12] as its base lan-
Attila Gürsoy, Laxmikant V. Kalé
SC1