Emre Sefer

dblp:70/7563 · DBLP profile ↗
← Back
21ranked-venue papers
10as first author
11since 2021 · last 2026
0000-0002-9186-0270ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 12 · 6 first-author · 6 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-author · 2 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Diagnosing LLM Reranker Behavior Under Fixed Evidence Pools
Baris Arat, Emre Sefer
SIGIR2
2026 A comparative analysis of topological domain callers over RNA-associated interactome
abstract
BACKGROUND: Topological domains (TADs) are consecutive genomic locations with denser local interactions to a certain extent, and they are important for cellular gene expression control and modulation. TADs were first identified when studying three-dimensional genomic structures over Hi-C interaction datasets. Many studies have focused on developing approaches in inferring TADs, which led to multiple TAD-caller approaches development. On the other hand, the number of RNA interactome datasets, such as RNA-RNA and RNA-DNA interactions, has recently been increasing. Even though TADs have been extensively studied in Hi-C datasets, they have not been studied across these RNA interactomes. RESULT: We conducted a systematized comparison of 28 TAD-callers across mammalians over RNA-associated interactomes (RAIs), especially RNA-RNA and RNA-DNA interaction datasets at a high resolution. Our findings highlighted the significant enrichment of Cohesin/CTCF proteins at RNA-TAD boundaries for both RNA-RNA and RNA-DNA interactomes, especially those with corner dots, which are similar to the results for original TADs. The sizes and numbers of RNA-TADs vary significantly between different TAD caller approaches and RNA interactome resolution, suggesting the importance of considering RNA-TADs as hierarchical domains rather than distinct intervals. CONCLUSION: We examined the core principles and assumptions behind TAD-callers over RNA interactomes. To our best knowledge, this is the first time that many TAD inference methods are adapted to infer TAD-like domains on RNAs. Our results provide valuable guidance in selecting the most suitable methods for TAD inference over RNA interactomes.
Tuna Alaygut, Emre Sefer
BMC Bioinform.2
2026 Deep time-series methods to enhance company earnings prediction performance
Deniz Ozbakir, Arda Erdogan, Uygar Gun, Emre Sefer
Eng. Appl. Artif. Intell.4
2026 Diffusion-based valuable NFT generation
abstract
Abstract Non-fungible tokens (NFTs) have revolutionized digital ownership, offering unique provenance and value to digital assets. Existing text-to-image models do not have the incentive mechanisms to generate statistically rare features, even when they optimize for visual fidelity. This paper introduces DiffNFTGen, a new generative framework that is the first to combine a customized RarityReward measure derived from a Vision Transformer (ViT) with reinforcement learning. The suggested method ensures fidelity to NFT styles while explicitly maximizing the generation of rare features by fine-tuning Stable Diffusion using Proximal Policy Optimization (PPO) and Kullback-Leibler (KL) divergence regularization. DiNFTGen achieves a 2.4x greater rarity score than baseline models while keeping competitive visual quality, according to quantitative evaluation utilizing Frédechet Inception Distance (FID) and Rarity Score. In order to examine the trade-off between fidelity and rarity, we also perform ablation studies regarding reward weighting. The model’s capacity to generalize NFT styles to new domains is confirmed by qualitative evaluations. The datasets, analysis code, and suggested approach are accessible on https://github.com/seferlab/diffnftgen .
Emir Ulurak, Beyza Kaya, Emre Sefer
Multim. Tools Appl.3
2026 GAT-HiC: Efficient Reconstruction of 3D Chromosome Structure via Residual Graph Attention Neural Networks
abstract
Hi-C is an experimental technique to measure the genome-wide topological dynamics and three-dimensional (3D) shape of chromosomes indirectly via counting the number of interactions between distinct sets of loci. One can estimate the 3D shape of a chromosome over these indirect interaction datasets. Here, we come up with graph attention and residual network-based GAT-HiC to predict three-dimensional chromosome structure from Hi-C interactions. GAT-HiC is distinct from the existing 3D chromosome shape prediction approaches in a way that it can generalize to data that is different than train data. So, we can train GAT-HiC on one type of Hi-C interaction matrix and infer on a completely dissimilar interaction matrix. GAT-HiC combines the unsupervised vertex embedding method Node2vec with an attention-based graph neural network when predicting each genomic loci's three-dimensional coordinates from Hi-C interaction matrix. We test the performance of our method across multiple Hi-C interaction datasets, where a trained model can be generalized across distinct cell populations, distinct restriction enzymes, and distinct Hi-C resolutions over human and mouse. GAT-HiC can reconstruct accurately in all these scenarios. Our method outperforms the existing approaches in terms of the accuracy of three-dimensional chromosome shape inference over interaction datasets.
Beyza Kaya, Emre Sefer
IEEE Trans. Comput. Biol. Bioinform.2
2026 Optimal Reconstruction of Graph Evolution History Under Preferential Attachment Model
abstract
Research into the evolution of biological networks enhances our understanding of the functional roles of various biomolecular properties. Graph growth models, such as the Preferential Attachment (PA) model, help characterize the evolutionary dynamics of protein interaction networks by modeling the preferential attachment of duplicated proteins' new interactions to existing ones. This approach generates realistic, scale-free PPI networks in a systematic manner. However, existing methods for reconstructing ancestral graphs based on the PA model have predominantly relied on greedy algorithms, often leading to suboptimal results. In this study, we introduce ILP-PA, a novel approach based on Integer Linear Programming (ILP), designed to reconstruct historical PPI graphs by maximizing likelihood within the Preferential Attachment model. Our ILP formulation also leverages systematic heuristics from general-purpose ILP solvers, allowing the analysis of near-optimal and multiple optimal solutions, which can be valuable for various applications across different fields. We evaluate the effectiveness of our approach on both synthetic data and three real protein-protein interaction graphs, specifically the Commander complex, bZIP transcription factor family, and herpesvirus interaction network. Compared to existing techniques, our ILP-PA solutions achieve higher likelihoods and demonstrate greater robustness to model mismatches and noise in the data. Furthermore, our solutions align more closely with biological findings from various studies across real datasets.
Elena Kudret, Emre Sefer
IEEE Trans. Comput. Biol. Bioinform.2
2025 Anomaly Detection via Graph Contrastive Learning
abstract
Graph contrastive learning (GCL) techniques have shown superior performance in many tasks, such as social networks and recommendation systems, which makes them good solution candidates for detecting anomalies more accurately. The existing noncontrastive learning-based approaches do not fully take into account dynamic agents that can camouflage themselves. These agents either establish frequent associations with regular objects or intentionally skip forming relationships with the remaining objects, which are called head and tail anomalies respectively. In terms of graph topology, such agent behavior makes the graph more imbalanced. To handle both types of anomalies, we come up with a novel ensemble graph contrastive learning-based GCAD (Graph Contrasted Anomaly Detection) which is an ensemble of two approaches: 1- We learn representations and embeddings by leveraging Siamese architecture, which learns to minimize/maximize similarity between graph pairs at different scales while capturing their hierarchical structure. 2- We integrate a self-supervised learning framework using graph augmentations (like node and edge dropout) and contrastive learning to learn robust graph embeddings. In this case, the main idea is to generate multiple views of a graph using augmentations and then maximize the agreement between these views using contrastive loss. We show our approach outperforms the competing approaches in detecting both tail and head anomalies across 6 different datasets from citation and finance domains. The ablation studies also show the importance of GCAD components as well as its robustness.
Emre Sefer
SDM1
2025 Financial asset price prediction with graph neural network-based temporal deep learning models
Yasin Uygun, Emre Sefer
Neural Comput. Appl.2
2023 BERT2OME: Prediction of 2′-O-Methylation Modifications From RNA Sequence by Transformer Architecture Based on BERT
abstract
Recent work on language models has resulted in state-of-the-art performance on various language tasks. Among these, Bidirectional Encoder Representations from Transformers (BERT) has focused on contextualizing word embeddings to extract context and semantics of the words. On the other hand, post-transcriptional 2'-O-methylation (Nm) RNA modification is important in various cellular tasks and related to a number of diseases. The existing high-throughput experimental techniques take longer time to detect these modifications, and costly in exploring these functional processes. Here, to deeply understand the associated biological processes faster, we come up with an efficient method Bert2Ome to infer 2'-O-methylation RNA modification sites from RNA sequences. Bert2Ome combines BERT-based model with convolutional neural networks (CNN) to infer the relationship between the modification sites and RNA sequence content. Unlike the methods proposed so far, Bert2Ome assumes each given RNA sequence as a text and focuses on improving the modification prediction performance by integrating the pretrained deep learning-based language model BERT. Additionally, our transformer-based approach could infer modification sites across multiple species. According to 5-fold cross-validation, human and mouse accuracies were 99.15% and 94.35% respectively. Similarly, ROC AUC scores were 0.99, 0.94 for the same species. Detailed results show that Bert2Ome reduces the time consumed in biological experiments and outperforms the existing approaches across different datasets and species over multiple metrics. Additionally, deep learning approaches such as 2D CNNs are more promising in learning BERT attributes than more conventional machine learning methods.
Necla Nisa Soylu, Emre Sefer
IEEE ACM Trans. Comput. Biol. Bioinform.2
2022 A comparison of topologically associating domain callers over mammals at high resolution
abstract
BACKGROUND: Topologically associating domains (TADs) are locally highly-interacting genome regions, which also play a critical role in regulating gene expression in the cell. TADs have been first identified while investigating the 3D genome structure over High-throughput Chromosome Conformation Capture (Hi-C) interaction dataset. Substantial degree of efforts have been devoted to develop techniques for inferring TADs from Hi-C interaction dataset. Many TAD-calling methods have been developed which differ in their criteria and assumptions in TAD inference. Correspondingly, TADs inferred via these callers vary in terms of both similarities and biological features they are enriched in. RESULT: We have carried out a systematic comparison of 27 TAD-calling methods over mammals. We use Micro-C, a recent high-resolution variant of Hi-C, to compare TADs at a very high resolution, and classify the methods into 3 categories: feature-based methods, Clustering methods, Graph-partitioning methods. We have evaluated TAD boundaries, gaps between adjacent TADs, and quality of TADs across various criteria. We also found particularly CTCF and Cohesin proteins to be effective in formation of TADs with corner dots. We have also assessed the callers performance on simulated datasets since a gold standard for TADs is missing. TAD sizes and numbers change remarkably between TAD callers and dataset resolutions, indicating that TADs are hierarchically-organized domains, instead of disjoint regions. A core subset of feature-based TAD callers regularly perform the best while inferring reproducible domains, which are also enriched for TAD related biological properties. CONCLUSION: We have analyzed the fundamental principles of TAD-calling methods, and identified the existing situation in TAD inference across high resolution Micro-C interaction datasets over mammals. We come up with a systematic, comprehensive, and concise framework to evaluate the TAD-calling methods performance across Micro-C datasets. Our research will be useful in selecting appropriate methods for TAD inference and evaluation based on available data, experimental design, and biological question of interest. We also introduce our analysis as a benchmarking tool with publicly available source code.
Emre Sefer
BMC Bioinform.1
2022 BioCode: A Data-Driven Procedure to Learn the Growth of Biological Networks
abstract
Probabilistic biological network growth models have been utilized for many tasks including but not limited to capturing mechanism and dynamics of biological growth activities, null model representation, capturing anomalies, etc. Well-known examples of these probabilistic models are Kronecker model, preferential attachment model, and duplication-based model. However, we should frequently keep developing new models to better fit and explain the observed network features while new networks are being observed. Additionally, it is difficult to develop a growth model each time we study a new network. In this paper, we propose BioCode, a framework to automatically discover novel biological growth models matching user-specified graph attributes in directed and undirected biological graphs. BioCode designs a basic set of instructions which are common enough to model a number of well-known biological graph growth models. We combine such instruction-wise representation with a genetic algorithm based optimization procedure to encode models for various biological networks. We mainly evaluate the performance of BioCode in discovering models for biological collaboration networks, gene regulatory networks, and protein interaction networks which features such as assortativity, clustering coefficient, degree distribution closely match with the true ones in the corresponding real biological networks. As shown by the tests on the simulated graphs, the variance of the distributions of biological networks generated by BioCode is similar to the known models' variance for these biological network types.
Emre Sefer
IEEE ACM Trans. Comput. Biol. Bioinform.1
2016 Shall We Dense? Comparing Design Strategies for Time Series Expression Experiments
Emre Sefer, Ziv Bar-Joseph
RECOMB1
2016 Diffusion archeology for diffusion progression history reconstruction
Emre Sefer, Carl Kingsford
Knowl. Inf. Syst.1
2015 Convex Risk Minimization to Infer Networks from probabilistic diffusion data at multiple scales
abstract
SEIR (Susceptible-Exposed-Infected-Recovered) is a general and widely-used diffusion model that can model the diffusion in different contexts such as idea spreading and disease propagation. Here, we tackle the problem of inferring graph edges if we can only observe a SEIR diffusion process spreading over the nodes of a graph. This problem is of importance in the common case where node states can be estimated with less cost than the edges can be found. Some applications include inferring a contact network from disease spread data, inferring a reference network from idea spreading, or estimating influenza diffusion rates between U.S. states. We improve upon the existing approaches for this problem in three ways: (1) we assume we are provided only with the probabilistic information about the state of each node which may also be undersampled or incomplete; (2) we present a more general framework that better uses trace data to model edge non-existence under SEIR model; (3) we can infer the network at both micro and macro scales. Experiments on both real and synthetic data show that our method is accurate under these challenging cases at multiple scales, and it performs consistently better than the existing methods. For instance, we can infer a high school human contact network at microscale by tracking influenza diffusion almost 10% better than the existing methods as well as the estimated networks closely mimick the full range of properties of the true network. We also estimated the strength of the influenza diffusion between and inside the U.S. states from Google Flu Trends data at macroscale. Estimated rates are correlated with the human transportation rates between the states to a certain degree, and we gain interesting insight into the influenza diffusion in U.S. such as the importance of the less populous states in epidemics as well as the asymmetric influenza diffusion between U.S. states.
Emre Sefer, Carl Kingsford
ICDE1
2015 Deconvolution of Ensemble Chromatin Interaction Data Reveals the Latent Mixing Structures in Cell Subpopulations
Emre Sefer, Geet Duggal, Carl Kingsford
RECOMB1
2015 Semi-nonparametric Modeling of Topological Domain Formation from Epigenetic Data
Emre Sefer, Carl Kingsford
WABI1
2014 Diffusion Archaeology for Diffusion Progression History Reconstruction
abstract
Diffusion through graphs can be used to model many real-world process, such as the spread of diseases, social network memes, computer viruses, or water contaminants. Often, a real-world diffusion cannot be directly observed while it is occurring -- perhaps it is not noticed until some time has passed, continuous monitoring is too costly, or privacy concerns limit data access. This leads to the need to reconstruct how the present state of the diffusion came to be from partial diffusion data. Here, we tackle the problem of reconstructing a diffusion history from one or more snapshots of the diffusion state. This ability can be invaluable to learn when certain computer nodes are infected or which people are the initial disease spreaders to control future diffusions. We formulate this problem over discrete-time SEIRS-type diffusion models in terms of maximum likelihood. We design methods that are based on sub modularity and a novel prize-collecting dominating-set vertex cover (PCDSVC) relaxation that can identify likely diffusion steps with some provable performance guarantees. Our methods are the first to be able to reconstruct complete diffusion histories accurately in real and simulated situations. As a special case, they can also identify the initial spreaders better than existing methods for that problem. Our results for both meme and contaminant diffusion show that the partial diffusion data problem can be overcome with proper modeling and methods, and that hidden temporal characteristics of diffusion can be predicted from limited data.
Emre Sefer, Carl Kingsford
ICDM1
2012 The missing models: a data-driven approach for learning how networks grow
abstract
Probabilistic models of network growth have been extensively studied as idealized representations of network evolution. Models, such as the Kronecker model, duplication-based models, and preferential attachment models, have been used for tasks such as representing null models, detecting anomalies, algorithm testing, and developing an understanding of various mechanistic growth processes. However, developing a new growth model to fit observed properties of a network is a difficult task, and as new networks are studied, new models must constantly be developed. Here, we present a framework, called GrowCode, for the automatic discovery of novel growth models that match user-specified topological features in undirected graphs. GrowCode introduces a set of basic commands that are general enough to encode several previously developed models. Coupling this formal representation with an optimization approach, we show that GrowCode is able to discover models for protein interaction networks, autonomous systems networks, and scientific collaboration networks that closely match properties such as the degree distribution, the clustering coefficient, and assortativity that are observed in real networks of these classes. Additional tests on simulated networks show that the models learned by GrowCode generate distributions of graphs with similar variance as existing models for these classes.
Rob Patro, Geet Duggal, Emre Sefer, Hao Wang 0024, Darya Filippova, Carl Kingsford
KDD3
2012 Resolving Spatial Inconsistencies in Chromosome Conformation Data
Geet Duggal, Rob Patro, Emre Sefer, Hao Wang 0024, Darya Filippova, Samir Khuller, Carl Kingsford
WABI3
2011 Metric Labeling and Semi-metric Embedding for Protein Annotation Prediction
Emre Sefer, Carl Kingsford
RECOMB1
2011 Parsimonious Reconstruction of Network Evolution
Rob Patro, Emre Sefer, Justin Malin, Guillaume Marçais, Saket Navlakha, Carl Kingsford
WABI2