VLDB 2026 Research / reviewers in the wild / expert
Bora Caglayan
dblp:60/7093 · also Bora Çaglayan
· DBLP profile ↗
19ranked-venue papers
5as first author
3since 2021 · last 2025
0000-0002-8491-4453ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 16 · 5 first-author · 3 since 2021Artificial intelligence and machine learning · 3Databases, data management, data science and information retrieval · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RADICE: Causal Graph Based Root Cause Analysis for System Performance DiagnosticabstractRadice: (Italian noun) root Root cause analysis is one of the most crucial operations in software reliability regarding system performance diagnostic. It aims to identify the root causes of system performance anomalies, allowing the resolution or the future prevention of issues that can cause millions of dollars in losses. Common existing approaches relying on data correlation or full domain expert knowledge are inaccurate or infeasible in most industrial cases, since correlation does not imply causation, and domain experts may not have full knowledge of complex and real-time systems. In this work, we define a novel causal domain knowledge model representing causal relations about the underlying system components to allow domain experts to contribute partial domain knowledge for root cause analysis. We then introduce RADICE, an algorithm that through the causal graph discovery, enhancement, refinement, and subtraction processes is able to output a root cause causal sub-graph showing the causal relations between the system components affected by the anomaly. We evaluated RADICE with simulated data and reported a real data use case, sharing the lessons we learned. The experiments show that RADICE provides better results than other baseline methods, including causal discovery algorithms and correlation based approaches for root cause analysis. Andrea Tonon, Bora Caglayan, Tong Gui, Mingxue Wang |
SANER | 3 |
| 2024 | BIS: NL2SQL Service Evaluation Benchmark for Business Intelligence Scenarios
Bora Caglayan, Mingxue Wang, John D. Kelleher, Shen Fei, Gui Tong, Jiandong Ding, Puchao Zhang |
ICSOC (2) | 1 |
| 2024 | Shrec: a Sre Behaviour Knowledge Graph Model for Shell Command RecommendationsabstractIn IT system operations, shell commands are common command line tools used by site reliability engineers (SREs) for daily tasks, such as system configuration, package deployment, and performance optimization. The efficiency in their execution has a crucial business impact since shell commands very often aim to execute critical operations, such as the resolution of system faults. However, many shell commands involve long parameters that make them hard to remember and type. Additionally, the experience and knowledge of SREs using these commands are almost always not preserved. In this work, we propose SHREC, a SRE behaviour knowledge graph model for shell command recommendations. We model the SRE shell behaviour knowledge as a knowledge graph and propose a strategy to directly extract such a knowledge from SRE historical shell operations. The knowledge graph is then used to provide shell command recommendations in real-time to improve the SRE operation efficiency. Our empirical study based on real shell commands executed in our company demonstrates that Shrec can improve the SRE operation efficiency, allowing to share and re-utilize the SRE knowledge. Andrea Tonon, Bora Caglayan, Mingxue Wang, Puchao Zhang |
SANER | 2 |
| 2019 | Competition-Based Crowdsourcing Software Development: A Multi-Method Study from a Customer PerspectiveabstractCrowdsourcing is emerging as an alternative outsourcing strategy which is gaining increasing attention in the software engineering community. However, crowdsourcing software development involves complex tasks which differ significantly from the micro-tasks that can be found on crowdsourcing platforms such as Amazon Mechanical Turk which are much shorter in duration, are typically very simple, and do not involve any task interdependencies. To achieve the potential benefits of crowdsourcing in the software development context, companies need to understand how this strategy works, and what factors might affect crowd participation. We present a multi-method qualitative and quantitative theory-building research study. First, we derive a set of key concerns from the crowdsourcing literature as an initial analytical framework for an exploratory case study in a Fortune 500 company. We complement the case study findings with an analysis of 13,602 crowdsourcing competitions over a ten-year period on the very popular Topcoder crowdsourcing platform. Drawing from our empirical findings and the crowdsourcing literature, we propose a theoretical model of crowd interest and actual participation in crowdsourcing competitions. We evaluate this model using Structural Equation Modeling. Among the findings are that the level of prize and duration of competitions do not significantly increase crowd interest in competitions. Klaas-Jan Stol, Bora Caglayan, Brian Fitzgerald 0001 |
IEEE Trans. Software Eng. | 2 |
| 2018 | DeepAD: A Generic Framework Based on Deep Learning for Time Series Anomaly Detection
Teodora Sandra Buda, Bora Caglayan, Haytham Assem |
PAKDD (1) | 2 |
| 2018 | ST-DenNetFus: A New Deep Learning Approach for Network Demand Prediction
Haytham Assem, Bora Caglayan, Teodora Sandra Buda, Declan O'Sullivan |
ECML/PKDD (3) | 2 |
| 2018 | The relationship between evolutionary coupling and defects in large industrial software (journal-first abstract)abstractIn this study, we investigate the effect of EC on the defect-proneness of large industrial software systems and explain why the effects vary. Serkan Kirbas, Bora Caglayan, Tracy Hall, Steve Counsell, David Bowes, Alper Sen 0001, Ayse Basar Bener |
SANER | 2 |
| 2018 | Predicting bug-fixing time: A replication study using an open source software project
Shirin Akbarinasaji, Bora Caglayan, Ayse Basar Bener |
J. Syst. Softw. | 2 |
| 2017 | The relationship between evolutionary coupling and defects in large industrial softwareabstractAbstract Evolutionary coupling (EC) is defined as the implicit relationship between 2 or more software artifacts that are frequently changed together. Changing software is widely reported to be defect‐prone. In this study, we investigate the effect of EC on the defect proneness of large industrial software systems and explain why the effects vary. We analysed 2 large industrial systems: a legacy financial system and a modern telecommunications system. We collected historical data for 7 years from 5 different software repositories containing 176 thousand files. We applied correlation and regression analysis to explore the relationship between EC and software defects, and we analysed defect types, size, and process metrics to explain different effects of EC on defects through correlation. Our results indicate that there is generally a positive correlation between EC and defects, but the correlation strength varies. Evolutionary coupling is less likely to have a relationship to software defects for parts of the software with fewer files and where fewer developers contributed. Evolutionary coupling measures showed higher correlation with some types of defects (based on root causes) such as code implementation and acceptance criteria. Although EC measures may be useful to explain defects, the explanatory power of such measures depends on defect types, size, and process metrics. Serkan Kirbas, Bora Caglayan, Tracy Hall, Steve Counsell, David Bowes, Alper Sen 0001, Ayse Basar Bener |
J. Softw. Evol. Process. | 2 |
| 2016 | Software Analytics in Practice: A Defect Prediction Model Using Code SmellsabstractIn software engineering, maintainability is related to investigating the defects and their causes, correcting the defects and modifying the system to meet customer requirements. Maintenance is a time consuming activity within the software life cycle. Therefore, there is a need for efficiently organizing the software resources in terms of time, cost and personnel for maintenance activity. One way of efficiently managing maintenance resources is to predict defects that may occur after the deployment. Many researchers so far have built defect prediction models using different sets of metrics such as churn and static code metrics. However, hidden causes of defects such as code smells have not been investigated thoroughly. In this study we propose using data science and analytics techniques on software data to build defect prediction models. In order to build the prediction model we used code smells metrics, churn metrics and combination of churn and code smells metrics. The results of our experiments on two different software companies show that code smells is a good indicator of defect proneness of the software product. Therefore, we recommend that code smells metrics should be used to train a defect prediction model to guide the software maintenance team. Behjat Soltanifar, Shirin Akbarinasaji, Bora Caglayan, Ayse Basar Bener, Asli Filiz, Bryan M. Kramer |
IDEAS | 3 |
| 2016 | Effect of developer collaboration activity on software quality in two large scale projects
Bora Caglayan, Ayse Basar Bener |
J. Syst. Softw. | 1 |
| 2015 | Merits of Organizational Metrics in Defect Prediction: An Industrial ReplicationabstractDefect prediction models presented in the literature lack generalization unless the original study can be replicated using new datasets and in different organizational settings. Practitioners can also benefit from replicating studies in their own environment by gaining insights and comparing their findings with those reported. In this work, we replicated an earlier study in order to investigate the merits of organizational metrics in building defect prediction models for large-scale enterprise software. We mined the organizational, code complexity, code churn and pre-release bug metrics of that large scale software and built defect prediction models for each metric set. In the original study, organizational metrics were found to achieve the highest performance. In our case, models based on organizational metrics performed better than models based on churn metrics but were outperformed by pre-release metric models. Further, we verified four individual organizational metrics as indicators for defects. We conclude that the performance of different metric sets in building defect prediction models depends on the project's characteristics and the targeted prediction level. Our replication of earlier research enabled assessing the validity and limitations of organizational metrics in a different context. Bora Caglayan, Burak Turhan, Ayse Basar Bener, Mayy Habayeb, Andriy V. Miranskyy, Enzo Cialini |
ICSE (2) | 1 |
| 2015 | Predicting defective modules in different test phases
Bora Caglayan, Ayse Tosun Misirli, Ayse Basar Bener, Andriy V. Miranskyy |
Softw. Qual. J. | 1 |
| 2014 | The effect of evolutionary coupling on software defects: an industrial case study on a legacy systemabstractEvolutionary coupling is defined as the implicit relationship between two or more software artifacts that are frequently changed together. In this study we investigate the effect of evolutionary coupling on defect proneness of a large financial legacy software in an industrial software development environment. We collected historical data for 5 years from 3 different software repositories containing 150 thousand files on 274 modules. Our results indicate that there is a positive correlation between evolutionary coupling and defect measures. Furthermore, we built linear and logistic regression models by using evolutionary coupling measures in order to explain defects. Although regression analysis results show that evolutionary coupling measures can be useful to explain defects, especially for modules in which high correlation is detected, explanatory power decreases dramatically with the decreasing correlation. Serkan Kirbas, Alper Sen 0001, Bora Caglayan, Ayse Basar Bener, Rasim Mahmutogullari |
ESEM | 3 |
| 2014 | Effect of temporal collaboration network, maintenance activity, and experience on defect exposureabstractContext: Number of defects fixed in a given month is used as an input for several project management decisions such as release time, maintenance effort estimation and software quality assessment. Past activity of developers and testers may help us understand the future number of reported defects. Goal: To find a simple and easy to implement solution, predicting defect exposure. Method: We propose a temporal collaboration network model that uses the history of collaboration among developers, testers, and other issue originators to estimate the defect exposure for the next month. Results: Our empirical results show that temporal collaboration model could be used to predict the number of exposed defects in the next month with R2 values of 0.73. We also show that temporality gives a more realistic picture of collaboration network compared to a static one. Conclusions: We believe that our novel approach may be used to better plan for the upcoming releases, helping managers to make evidence based decisions. Andriy V. Miranskyy, Bora Caglayan, Ayse Basar Bener, Enzo Cialini |
ESEM | 2 |
| 2012 | Dione: an integrated measurement and defect prediction solutionabstractWe present an integrated measurement and defect prediction tool: Dione. Our tool enables organizations to measure, monitor, and control product quality through learning based defect prediction. Similar existing tools either provide data collection and analytics, or work just as a prediction engine. Therefore, companies need to deal with multiple tools with incompatible interfaces in order to deploy a complete measurement and prediction solution. Dione provides a fully integrated solution where data extraction, defect prediction and reporting steps fit seamlessly. In this paper, we present the major functionality and architectural elements of Dione followed by an overview of our demonstration. Bora Caglayan, Ayse Tosun Misirli, Gül Çalikli, Ayse Basar Bener, Turgay Aytac, Burak Turhan |
SIGSOFT FSE | 1 |
| 2011 | Defect prediction using social network analysis on issue repositoriesabstractPeople are the most important pillar of software development process. It is critical to understand how they interact with each other and how these interactions affect the quality of the end product in terms of defects. In this research we propose to include a new set of metrics, a.k.a. social network metrics on issue repositories in predicting defects. Social network metrics on issue repositories has not been used before to predict defect proneness of a software product. To validate our hypotheses we used two datasets, development data of IBM1 Rational ® Team Concert™ (RTC) and Drupal, to conduct our experiments. The results of the experiments revealed that compared to other set of metrics such as churn metrics using social network metrics on issue repositories either considerably decreases high false alarm rates without compromising the detection rates or considerably increases low prediction rates without compromising low false alarm rates. Therefore we recommend practitioners to collect social network metrics on issue repositories since people related information is a strong indicator of past patterns in a given team. Serdar Biçer, Ayse Basar Bener, Bora Caglayan |
ICSSP | 3 |
| 2010 | Do More People Make the Code More Defect Prone?: Social Network Analysis in OSS Projects
Salifu Alhassan, Bora Caglayan, Ayse Basar Bener |
SEKE | 2 |
| 2009 | Prest: An Intelligent Software Metrics Extraction, Analysis and Defect Prediction Tool
Ekrem Kocaguneli, Ayse Tosun Misirli, Ayse Basar Bener, Burak Turhan, Bora Caglayan |
SEKE | 5 |