EDBT 2026 Demo / reviewers in the wild / expert
Gökhan Kul
dblp:145/7676
· DBLP profile ↗
11ranked-venue papers
1as first author
10since 2021 · last 2026
0000-0001-6467-1979ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 3 · 3 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VulStyle: A Multi-Modal Pre-Training for Code Stylometry-Augmented Vulnerability Detection
Chidera Biringa, Ajmal Abbas, Vishnu Selvaraj, Gökhan Kul |
DSN | 4 |
| 2026 | MITRE ATT&CK-based Attack Chain Prediction Using Hybrid LSTM-Markov Models for Cyber Risk Assessment
Mayank Raj, Nathaniel D. Bastian, Lance Fiondella, Gökhan Kul |
SECRYPT (2) | 4 |
| 2025 | Regression and Time Series Mixture Approaches to Predict System Performance and Assess ResilienceabstractResilience engineering is the ability to design, build, and sustain systems that can deal effectively with disruptive events. Previous research focused on resilience models that were not designed to predict multiple disruptions and recoveries, and resilience metrics, which are typically calculated after disruptions. Therefore, this article introduces a new approach combining regression and time series methods to track and predict system performance under multiple shocks, offering a framework for planning resilience tests and guiding data collection applicable to various systems and processes. To illustrate, subsets ranging from 50% to 80% of a historical job loss dataset from the 1980 U.S. recession were used for model fitting to assess generalization and stability. Goodness-of-fit measures, confidence intervals, and resilience metrics validated this approach against established statistical methods and a neural network model. The results indicate that traditional statistical models fail to capture minor changes when fitted with small datasets, and neural networks are overly sensitive to the size of the training data. In contrast, the novel mixture approach considering immediate and delayed disruptions exhibits superior long-term predictive performance and greater accuracy in forecasting resilience metrics, even when only 50% of the data is used for model fitting. Priscila Silva, Gaspard Baye, Mindy Hotchkiss, Gökhan Kul, Nathaniel D. Bastian, Lance Fiondella |
IEEE Trans. Reliab. | 4 |
| 2024 | Predicting F1-Scores of Classifiers in Network Intrusion Detection SystemsabstractWith the evolution of the Internet of Things, network intrusion detection systems (NIDS) are vital for protecting networks by monitoring and analyzing network traffic to detect potential cyber threats. Deep neural networks (DNNs) are widely used in NIDS for their accurate classification and response capabilities against threats. However, there is a lack of detailed discussion in the literature about evaluating DNN performance in real-time scenarios for monitoring and assurance of NIDS. This paper fills this gap by applying multiple linear regression models to predict the F1-score of a DNN attack classifier. The predictive models are evaluated using a pre-trained DNN on a NIDS benchmark dataset, where three different distance metrics computed between real-time instances and known attack patterns stored in historical data are considered as model covariates. Our findings show that the multiple linear regression model with interaction between covariates confidently forecasts the F1-score for future periods, demonstrating its ability to anticipate future observations with an empirical coverage of 94.4%. Priscila Silva, Gaspard Baye, Alexandre Broggi, Nathaniel D. Bastian, Gökhan Kul, Lance Fiondella |
ICCCN | 5 |
| 2024 | A Reproducible Tutorial on Reproducibility in Database Systems ResearchabstractReproducibility is a key aspect of the scientific method, and it is essential for building trust in the results of research. This tutorial aims to provide concrete guidance on how to leverage containerized reproducibility using Docker for database systems research. In this tutorial, we present a step-by-step guide on how to prepare a Docker-based artifact for an experiment. We will cover topics such as Dockerfiles, Docker images, Docker Compose, automation using Python, Bash, and Make, and also artifact documentation and packaging best practices. The tutorial itself is a reproducible artifact, and we provide a public GitHub repository with all the code and examples used in the tutorial. This repository can serve as a starting point to prepare artifacts for experiments and publications. Tim Fischer 0003, Denis Hirn, Gökhan Kul |
Proc. VLDB Endow. | 3 |
| 2024 | PACE: A Program Analysis Framework for Continuous Performance PredictionabstractSoftware development teams establish elaborate continuous integration pipelines containing automated test cases to accelerate the development process of software. Automated tests help to verify the correctness of code modifications decreasing the response time to changing requirements. However, when the software teams do not track the performance impact of pending modifications, they may need to spend considerable time refactoring existing code. This article presents PACE , a program analysis framework that provides continuous feedback on the performance impact of pending code updates. We design performance microbenchmarks by mapping the execution time of functional test cases given a code update. We map microbenchmarks to code stylometry features and feed them to predictors for performance predictions. Our experiments achieved significant performance in predicting code performance, outperforming current state-of-the-art by 75% on neural-represented code stylometry features. Chidera Biringa, Gökhan Kul |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2023 | Performance Analysis of Deep-Learning Based Open Set Recognition Algorithms for Network Intrusion Detection SystemsabstractOpen Set Recognition (OSR) is the ability of a machine learning (ML) algorithm to classify the known and recognize the unknown. In other words, OSR enables novelty detection in classification algorithms. This broader approach is critical to detect new types of attacks, including zero-days, thereby improving the effectiveness and efficiency of various MLenabled mission-critical systems, such as cyber-physical, facial recognition, spam filtering, and cyber defense systems such as intrusion detection systems (IDS). In ML algorithms, like deep learning (DL) classifiers, hyperparameters control the learning process; their values affect other model parameters, such as weights and biases, which affect the performance of OSR algorithms. Moreover, OSR introduces additional parameters, making DL classifiers bigger and training them more computationally intensive. Determining the suitable set of hyperparameters and parameters is a computationally expensive task. Alternative OSR algorithms have demonstrated promising results on image datasets, but only limited studies have been performed in the context of IDS. This paper proposes OpenSetPerf, an empirical investigation of three prominent OSR algorithms using a current, real-world network intrusion detection systems (NIDS) benchmark dataset to discover the relationship between the DL-based OSR algorithm’s hyperparameter values and their performance. OpenSetperf evaluates these algorithms using quantitative studies with widely used ML performance evaluation metrics. Gaspard Baye, Priscila Silva, Alexandre Broggi, Lance Fiondella, Nathaniel D. Bastian, Gökhan Kul |
NOMS | 6 |
| 2022 | Covariate Software Vulnerability Discovery Model to Support Cybersecurity Test & Evaluation (Practical Experience Report)abstractVulnerability discovery models (VDM) have been proposed as an application of software reliability growth models (SRGM) to software security related defects. VDM model the number of vulnerabilities discovered as a function of testing time, enabling quantitative measures of security. Despite their obvious utility, past VDM have been limited to parametric forms that do not consider the multiple activities software testers undertake in order to identify vulnerabilities. In contrast, covariate SRGM characterize the software defect discovery process in terms of one or more test activities. However, data sets documenting multiple security testing activities suitable for application of covariate models are not readily available in the open literature. To demonstrate the applicability of covariate SRGM to vul-nerability discovery, this research identified a web application to target as well as multiple tools and techniques to test for vulnerabilities. The time dedicated to each test activity and the corresponding number of unique vulnerabilities discovered were documented and prepared in a format suitable for application of covariate SRGM. Analysis and prediction were then performed and compared with a flexible VDM without covariates, namely the Alhazmi-Malaiya Logistic Model (AML). Our results indicate that covariate VDM significantly outperformed the AML model on predictive and information theoretic measures of goodness of fit, suggesting that covariate VDM are a suitable and effective method to predict the impact of applying specific vulnerability discovery tools and techniques. Julia Sorrentino, Priscila Silva, Gaspard Baye, Gökhan Kul, Lance Fiondella |
ISSRE | 4 |
| 2022 | A Secure Design Pattern Approach Toward Tackling Lateral-Injection AttacksabstractSoftware weaknesses that create attack surfaces for adversarial exploits, such as lateral SQL injection (LSQLi) attacks, are usually introduced during the design phase of software development. Security design patterns are sometimes applied to tackle these weaknesses. However, due to the stealthy nature of lateral-based attacks, employing traditional security patterns to address these threats is insufficient. Hence, we present SEAL, a secure design that extrapolates architectural, design, and implementation abstraction levels to delegate security strategies toward tackling LSQLi attacks. We evaluated SEAL using case study software, where we assumed the role of an adversary and injected several attack vectors tasked with compromising the confidentiality and integrity of its database. Our evaluation of SEAL demonstrated its capacity to address LSQLi attacks. Chidera Biringa, Gökhan Kul |
SIN | 2 |
| 2021 | Automated User Experience Testing through Multi-Dimensional Performance Impact AnalysisabstractAlthough there are many automated software testing suites, they usually focus on unit, system, and interface testing. However, especially software updates such as new security features have the potential to diminish user experience. In this paper, we propose a novel automated user experience testing methodology that learns how code changes impact the time unit and system tests take, and extrapolate user experience changes based on this information. Such a tool can be integrated into existing continuous integration pipelines, and it provides software teams immediate user experience feedback. We construct a feature set from lexical, layout, and syntactic characteristics of the code, and using Abstract Syntax Tree-Based Embeddings, we can calculate the approximate semantic distance to feed into a machine learning algorithm. In our experiments, we use several regression methods to estimate the time impact of software updates. Our open-source tool achieved a 3.7% mean absolute error rate with a random forest regressor. Chidera Biringa, Gökhan Kul |
AST | 2 |
| 2018 | Similarity Metrics for SQL Query ClusteringabstractDatabase access logs are the starting point for many forms of database administration, from database performance tuning, to security auditing, to benchmark design, and many more. Unfortunately, query logs are also large and unwieldy, and it can be difficult for an analyst to extract broad patterns from the set of queries found therein. Clustering is a natural first step towards understanding the massive query logs. However, many clustering methods rely on the notion of pairwise similarity, which is challenging to compute for SQL queries, especially when the underlying data and database schema is unavailable. We investigate the problem of computing similarity between queries, relying only on the query structure. We conduct a rigorous evaluation of three query similarity heuristics proposed in the literature applied to query clustering on multiple query log datasets, representing different types of query workloads. To improve the accuracy of the three heuristics, we propose a generic feature engineering strategy, using classical query rewrites to standardize query structure. The proposed strategy results in a significant improvement in the performance of all three similarity heuristics. Gökhan Kul, Duc Thanh Anh Luong, Ting Xie 0002, Varun Chandola, Oliver Kennedy, Shambhu J. Upadhyaya |
IEEE Trans. Knowl. Data Eng. | 1 |