Samet Tenekeci

dblp:304/9329 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0001-8875-4111ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 7 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Automating software size measurement from python code using language models
Samet Tenekeci, Hüseyin Ünlü, Bedir Arda Gül, Damla Keles, Murat Küük, Onur Demirörs
Autom. Softw. Eng.1
2026 Correction: Automating software size measurement from python code using language models
Samet Tenekeci, Hüseyin Ünlü, Bedir Arda Gül, Damla Keles, Murat Küçük, Onur Demirörs
Autom. Softw. Eng.1
2026 Automating software size measurement with language models: Insights from industrial case studies
Hüseyin Ünlü, Samet Tenekeci, Dhia Eddine Kennouche, Onur Demirörs
J. Syst. Softw.2
2026 A Contrastive Learning Framework for Efficient Viral Escape Prediction
abstract
Understanding the complex rules and mechanisms behind viral evolution is crucial for developing better preventive treatments, yet predicting immune-evading mutations remains challenging. Recent advances in protein language models have led to novel approaches for in silico analysis of viral escape. We introduce CoV-SNN, unifying variant classification and escape prediction within an efficient contrastive learning framework. CoV-SNN, built on a Siamese neural network architecture, classifies previously unseen variants by modeling sequence-level similarities and differences through embeddings from CoV-RoBERTa, a lightweight protein language model trained on high-quality SARS-CoV-2 Spike sequences. It prioritizes escape sequences using an enhanced Constrained Semantic Change Search (CSCS) function that maps antigenic variation to semantic change and viral fitness to sequence probability. We evaluate CoV-SNN on novel sequences containing both wet-lab-verified and computationally generated escape mutations. CoV-SNN achieves 98.8% accuracy in multi-class variant classification, an AUC of 0.909 in zero-shot variant classification, and 97.7% escape precision at Top-10%, while providing up to a 125-fold inference speedup. These results suggest that contrastive learning can support scalable in silico surveillance and prioritization of immune-evasive mutations.
Samet Tenekeci, Efe Sezgin, Selma Tekir
IEEE Trans. Comput. Biol. Bioinform.1
2024 Predicting Software Functional Size Using Natural Language Processing: An Exploratory Case Study
abstract
Software Size Measurement (SSM) plays an essential role in software project management as it enables the acquisition of software size, which is the primary input for development effort and schedule estimation. However, many small and medium-sized companies cannot perform objective SSM and Software Effort Estimation (SEE) due to the lack of resources and an expert workforce. This results in inadequate estimates and projects exceeding the planned time and budget. Therefore, organizations need to perform objective SSM and SEE using minimal resources without an expert workforce. In this research, we conducted an exploratory case study to predict the functional size of software project requirements using state-of-the-art large language models (LLMs). For this aim, we fine-tuned BERT and BERT_SE with a set of user stories and their respective functional size in COSMIC Function Points (CFP). We gathered the user stories included in different project requirement documents. In total size prediction, we achieved 72.8% accuracy with BERT and 74.4% accuracy with BERT_SE. In data movement-based size prediction, we achieved 87.5% average accuracy with BERT and 88.1% average accuracy with BERT_SE. Although we use relatively small datasets in model training, these results are promising and hold significant value as they demonstrate the practical utility of language models in SSM.
Hüseyin Ünlü, Samet Tenekeci, Can Çiftçi, Ibrahim Baran Oral, Tunahan Atalay, Tuna Hacaloglu, Burcu Musaoglu, Onur Demirörs
SEAA2
2024 Predicting Software Size and Effort from Code Using Natural Language Processing
Samet Tenekeci, Hüseyin Ünlü, Emre Dikenelli, Ugurcan Selçuk, Görkem Kilinç Soylu, Onur Demirörs
IWSM-Mensura1
2022 Integrative Biological Network Analysis to Identify Shared Genes in Metabolic Disorders
abstract
Identification of common molecular mechanisms in interrelated diseases is essential for better prognoses and targeted therapies. However, complexity of metabolic pathways makes it difficult to discover common disease genes underlying metabolic disorders; and it requires more sophisticated bioinformatics models that combine different types of biological data and computational methods. Accordingly, we built an integrative network analysis model to identify shared disease genes in metabolic syndrome (MS), type 2 diabetes (T2D), and coronary artery disease (CAD). We constructed weighted gene co-expression networks by combining gene expression, protein-protein interaction, and gene ontology data from multiple sources. For 90 different configurations of disease networks, we detected the significant modules by using MCL, SPICi, and Linkcomm graph clustering algorithms. We also performed a comparative evaluation on disease modules to determine the best method providing the highest biological validity. By overlapping the disease modules, we identified 22 shared genes for MS-CAD and T2D-CAD. Moreover, 19 out of these genes were directly or indirectly associated with relevant diseases in the previous medical studies. This study does not only demonstrate the performance of different biological data sources and computational methods in disease-gene discovery, but also offers potential insights into common genetic mechanisms of the metabolic disorders.
Samet Tenekeci, Zerrin Isik
IEEE ACM Trans. Comput. Biol. Bioinform.1
2021 Event Oriented vs Object Oriented Analysis for Microservice Architecture: An Exploratory Case Study
abstract
The rapidly developing internet infrastructure together with the advances in software technology has enabled the development of cloud-based modern web applications that are much more responsive, flexible, and reliable compared to traditional monolithic applications. Such modern applications require new software design paradigms and architectures. Microservice-based architecture (MSbA), which aims to create small, isolated, loosely-coupled applications that work in cohesion, becoming widespread as one of these approaches. MSbA allows the developed applications to be deployed and maintained separately, as well as scaled on demand. However, there is no de facto method for the analysis and design of systems for these architectures. In this paper, we compared the usefulness of the object-oriented (OO) and event-oriented (EO) approaches for analyzing and designing MS-based systems. More specifically, we performed an exploratory case study to analyze, design, and implement a software application dealing with the ‘application and evaluation process of graduate students at IzTech’. This paper discusses the results of this case study. We observe that the EO approaches have significant advantages with respect to the OO approaches.
Hüseyin Ünlü, Samet Tenekeci, Ali Yildiz, Onur Demirörs
SEAA2
2021 Author Reputation Measurement on Question and Answer Sites by the Classification of Author-Generated Content
abstract
In the field of software engineering, practitioners’ share in the constructed knowledge cannot be underestimated and is mostly in the form of grey literature (GL). GL is a valuable resource though it is subjective and lacks an objective quality assurance methodology. In this paper, a quality assessment scheme is proposed for question and answer (Q&A) sites. In particular, we target stack overflow (SO) and stack exchange (SE) sites. We model the problem of author reputation measurement as a classification task on the author-provided answers. The authors’ mean, median, and total answer scores are used as inputs for class labeling. State-of-the-art language models (BERT and DistilBERT) with a softmax layer on top are utilized as classifiers and compared to SVM and random baselines. Our best model achieves [Formula: see text] accuracy in binary classification in SO design patterns tag and [Formula: see text] accuracy in SE software engineering category. Superior performance in SE software engineering can be explained by its larger dataset size. In addition to quantitative evaluation, we provide qualitative evidence, which supports that the system’s predicted reputation labels match the quality of provided answers.
Erhan Sezerer, Samet Tenekeci, Ali Acar, Bora Baloglu, Selma Tekir
Int. J. Softw. Eng. Knowl. Eng.2