Mustafa Coskun

dblp:155/0035 · DBLP profile ↗
← Back
10ranked-venue papers
7as first author
6since 2021 · last 2025
0000-0003-4805-1416ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Computer networks · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 MultiScale Spectral GNN for Fraud Detection
Melike Yildiz Aktas, Mustafa Coskun, Chang-Tien Lu
ASONAM (2)2
2025 ArnoldiGCL: Graph Contrastive Learning via Learnable Arnoldi-Based Guided Spectral Chebyshev Polynomial Filters
abstract
Graph Contrastive Learning (GCL) emerged as a powerful paradigm in self-supervised graph representation learning. While earlier applications of GCL rely on homophily assumptions, spectral graph neural networks (GNNs) enhance the effectiveness of GCL on heterophilic graphs by incorporating both low-pass and high-pass filters. However, due to numerical considerations, existing approaches oversimplify low-pass and high-pass filters by modeling them as basic linear operations, failing to capture complex topological relationships.
Mustafa Coskun, Abdelkader Baggag, Mehmet Koyutürk
KDD (2)1
2025 Multiplex Embedding of Biological Networks Using Cross-Network Node Similarities
abstract
Network embedding techniques, which provide low-dimensional representations of the nodes in a network, have been commonly applied to many machine learning problems in computational biology. In most of these applications, multiple networks (e.g., different types of interactions/associations or semantically identical networks that come from different sources) are available. Multiplex network embedding aims to derive strength from these data sources by integrating multiple networks that share a common set of nodes. Existing approaches to this problem treat all layers of the multiplex network separately while performing integration, ignoring the differences in the topology and sparsity patterns of different networks. Here, we formulate an optimization problem that accounts for inner-network smoothness and topological similarity of networks to compute diffusion states for each network. To quantify the topological similarity of nodes in different networks, we utilize shared neighborhood across networks. To compute the diffusion states of integrated networks, we propose an efficient algorithm for accelerating iterations, which yields two-fold improvement over the runtime of the state-of-the-art power iteration techniques. Finally, we integrate the resulting diffusion states and apply dimensionality reduction (singular value decomposition after log transformation) to compute node embeddings. We evaluate the performance of the resulting algorithm, Crossim, in the context of protein function prediction. Our experimental results show that the embeddings computed by Crossim consistently improve predictive accuracy over algorithms that do not take into account the topological similarity of different networks, suggesting that accounting for topological similarity across multiple network layers can improve network integration.
Mustafa Coskun, Mehmet Koyutürk
IEEE Trans. Comput. Biol. Bioinform.1
2025 Topological-Similarity Based Canonical Representations for Biological Link Prediction
abstract
Graph machine learning algorithms are being commonly applied to a broad range of prediction tasks in systems biology. An important design criterion in this regard is the definition of "topological similarity" between two nodes in a network, which is used to design convolution matrices for graph convolution or loss functions to evaluate node embeddings. Many measures of topological similarity exist in network science literature (e.g., random walk based proximity, shared neighborhood) and recent comparative studies show that the choice of topological similarity can have a significant effect on the performance and reliability of graph machine learning models. We propose GraphCan, a framework for computing canonical representations for biological networks using a similarity-based Graph Convolutional Network (GCN). GraphCan integrates multiple node similarity measures (Common Neighbor, Adamic Adar, Random Walk with Restart, Von Neumann, Resource Allocation, Hub-Depressed Index, Hub-Promoted Index, and adjacency matrix) to compute canonical node embeddings for a given network. The resulting embeddings can be utilized directly for downstream machine learning tasks. We comprehensively evaluate GraphCan in the context of various link prediction tasks in systems biology. Our results show that the integration of multiple similarity measures improves the robustness of the framework, especially when the input networks are sparse.
Mustafa Coskun, Mehmet Koyutürk
IEEE Trans. Comput. Biol. Bioinform.2
2021 Node similarity-based graph convolution for link prediction in biological networks
abstract
BACKGROUND: Link prediction is an important and well-studied problem in network biology. Recently, graph representation learning methods, including Graph Convolutional Network (GCN)-based node embedding have drawn increasing attention in link prediction. MOTIVATION: An important component of GCN-based network embedding is the convolution matrix, which is used to propagate features across the network. Existing algorithms use the degree-normalized adjacency matrix for this purpose, as this matrix is closely related to the graph Laplacian, capturing the spectral properties of the network. In parallel, it has been shown that GCNs with a single layer can generate more robust embeddings by reducing the number of parameters. Laplacian-based convolution is not well suited to single-layered GCNs, as it limits the propagation of information to immediate neighbors of a node. RESULTS: Capitalizing on the rich literature on unsupervised link prediction, we propose using node similarity-based convolution matrices in GCNs to compute node embeddings for link prediction. We consider eight representative node-similarity measures (Common Neighbors, Jaccard Index, Adamic-Adar, Resource Allocation, Hub- Depressed Index, Hub-Promoted Index, Sorenson Index and Salton Index) for this purpose. We systematically compare the performance of the resulting algorithms against GCNs that use the degree-normalized adjacency matrix for convolution, as well as other link prediction algorithms. In our experiments, we use three-link prediction tasks involving biomedical networks: drug-disease association prediction, drug-drug interaction prediction and protein-protein interaction prediction. Our results show that node similarity-based convolution matrices significantly improve the link prediction performance of GCN-based embeddings. CONCLUSION: As sophisticated machine-learning frameworks are increasingly employed in biological applications, historically well-established methods can be useful in making a head-start. AVAILABILITY AND IMPLEMENTATION: Our method, SiGraC, is implemented as a Python library and is freely available at https://github.com/mustafaCoskunAgu/SiGraC.
Mustafa Coskun, Mehmet Koyutürk
Bioinform.1
2021 Fast computation of Katz index for efficient processing of link prediction queries
Mustafa Coskun, Abdelkader Baggag, Mehmet Koyutürk
Data Min. Knowl. Discov.1
2019 OFFER Referees Suggester for the Journal Editors
abstract
Assigning appropriate referees to a journal or conference paper is a vital task for many reasons, including enhancing the journal venue quality and reliance, fair judgement of the papers, and among many others. While assigning the referees to the papers, the editors of a journal venue need to find suitable referees who are both related to field of the given paper and have no conflict of interest with the authors of the paper. Editorial-wise this referee assignment process is implemented in a hand-crafted manner, i.e., the editor needs to find the most suitable referees to the paper via a search engine and manually refines the all unrelated and having conflict of interest authors to the paper authors. Clearly, such a manual referee searching process is tedious and time consuming for the editors. In this paper, we present an alternate automated approach for assigning referees problem using intrinsic random walk with restart proximity measure. In our experiments based on a comprehensive DBLP networks, we show that our approach, called OFFER, significantly outperforms state-of-the-art the random walk with restart based method.
Mustafa Coskun, Hilal Hacilar, Cengiz Gezer, Vehbi C. Gungor
ISCC1
2019 Hidden Smile Correlation Discovery Across Subjects Using Random Walk with Restart
abstract
Fine-grained smile analysis is a complicated and challenging process. Understanding other party's smiles is one of the key tasks associated with realizing the implicit messages transmitted by the human. Considering this kind of message transmission is a major feature of human communication, understanding smiling has great potential value to promote the development of humanoid robots and animated software agents. Therefore, a fine-grained smile analysis system is proposed to uncover the hidden smile correlation across subjects. The system incorporates head pose as prior knowledge and employs conditional random forest to detect fiducial points on face. After that, a steady-state probability defined by a succession of Markov random steps is used to indicate the relevance score between smiles across subjects. We demonstrate performance of the proposed system on both constrained and unconstrained face datasets. The experimental results show that the proposed system is able to classify 4 smile levels and uncover hidden smile correlations across subjects successfully.
Mustafa Coskun, Alaa Badokhon, Menghan Liu, Ming-Chun Huang
IEEE Trans. Affect. Comput.2
2018 Indexed Fast Network Proximity Querying
abstract
Node proximity queries are among the most common operations on network databases. A common measure of node proximity is random walk based proximity, which has been shown to be less susceptible to noise and missing data. Real-time processing of random-walk based proximity queries poses significant computational challenges for larger graphs with over billions of nodes and edges, since it involves solution of large linear systems of equations. Due to the importance of this operation, significant effort has been devoted to developing efficient methods for random-walk based node proximity computations. These methods either aim to speed up iterative computations by exploiting numerical properties of random walks, or rely on computation and storage of matrix inverses to avoid computation during query processing. Although both approaches have been well studied, the speedup achieved by iterative approaches does not translate to real-time query processing, and the storage requirements of inversion-based approaches prohibit their use on very large graph databases. We present a novel approach to significantly reducing the computational cost of random walk based node proximity queries with scalable indexing. Our approach combines domain graph-partitioning based indexing with fast iterative computations during query processing using Chebyshev polynomials over the complex elliptic plane. This approach combines the query processing benefits of inversion techniques with the memory and storage benefits of iterative approache. Using real-world networks with billions of nodes and edges, and top- k proximity queries as the benchmark problem, we show that our algorithm, I-C hopper , significantly outperforms existing methods. Specifically, it drastically reduces convergence time of the iterative procedure, while also reducing storage requirements for indexing.
Mustafa Coskun, Ananth Grama, Mehmet Koyutürk
Proc. VLDB Endow.1
2016 Efficient Processing of Network Proximity Queries via Chebyshev Acceleration
abstract
Network proximity is at the heart of a large class of network analytics and information retrieval techniques, including node/ edge rankings, network alignment, and randomwalk based proximity queries, among many others. Owing to its importance, significant effort has been devoted to accelerating iterative processes underlying network proximity computations. These techniques rely on numerical properties of power iterations, as well as structural properties of the networks to reduce the run time of iterative algorithms.
Mustafa Coskun, Ananth Grama, Mehmet Koyutürk
KDD1