VLDB 2026 Research / reviewers in the wild / expert
Gül Çalikli
dblp:11/7657 · also Gul Calikli, Handan Gül Çalikli
· DBLP profile ↗
27ranked-venue papers
7as first author
11since 2021 · last 2026
0000-0003-4578-1747ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 22 · 7 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 3Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Automated Test Suite Enhancement Using Large Language Models with Few-shot PromptingabstractUnit testing is essential for verifying the functional correctness of code modules (e.g., classes, methods), but manually writing unit tests is often labor-intensive and time-consuming. Unit tests generated by tools that employ traditional approaches, such as search-based software testing (SBST), lack readability, naturalness, and practical usability. LLMs have recently provided promising results and become integral to developers’ daily practices. Consequently, software repositories now include a mix of human-written tests, LLM-generated tests, and those from tools employing traditional approaches such as SBST. While LLMs’ zero-shot capabilities have been widely studied, their few-shot learning potential for unit test generation remains underexplored. Few-shot prompting enables LLMs to learn from examples in the prompt and automatically retrieving such examples could enhance test suites. This paper empirically investigates how few-shot prompting with different test artifact sources, comprising human, SBST, or LLM, affects the quality of LLM-generated unit tests as program comprehension artifacts and their contribution to improving existing test suites by evaluating not only correctness and coverage but also readability, cognitive complexity, and maintainability in hybrid human–AI codebases. We conducted experiments on HumanEval and ClassEval datasets using GPT-4.o, which is integrated into GitHub Copilot and widely used among developers. We also assessed retrieval-based methods for selecting relevant examples. Our results show that LLMs can generate high-quality tests via few-shot prompting, with human-written examples producing the best coverage and correctness. Additionally, selecting examples based on the combined similarity of problem description and code consistently yields the most effective few-shot prompts. Data and Materials: https://doi.org/10.5281/zenodo.15561007 Alex Chudic, Gül Çalikli |
ICPC | 2 |
| 2025 | Perspectives, Needs and Challenges for Sustainable Software Engineering Teams: A FinServ Case StudyabstractBackground: Sustainable Software Engineering (SSE) is slowly becoming an industry need. While general perspectives of SSE have been studied, little is known about the context-specific sustainability needs and constraints of sectors like financial services, which are highly regulated, data-intensive, and handle millions of transactions daily. Aim: To address this gap, our research focuses on a financial services company (FinServCo) that invited us to investigate perspectives on sustainability in their IT function: how it could be put into practice, who is responsible for it, and what the challenges are. Method: We conducted an exploratory qualitative case study using interviews and a focus group with six higher management employees and 16 software engineers comprising various experience levels from junior developers to team leaders. Results: We found a clear divergence in how sustainability is perceived across organisational levels. Higher management focused on technical and economic aspects, such as cloud migration and business continuity through data availability, while developers emphasised human-centric concerns like workload and stress. Some developers were sceptical of organisational motives, viewing initiatives as PR-driven. Key challenges included knowledge gaps, cultural resistance, legacy systems, and limited client demand. Many participants favoured a dedicated sustainability team, often drawing comparisons to established security governance structures. Conclusions: SSE is shaped by organisational roles, cultural dynamics, and sector-specific constraints. The disconnect between organisational goals and individual developer needs highlights the importance of context-sensitive, co-designed interventions. Study material containing codebook and the interview questions file: https://doi.org/10.5281/zenodo.16359596 Satwik Ghanta, Peggy Gregory, Gül Çalikli |
ESEM | 3 |
| 2025 | In-Context Learning as an Effective Estimator of Functional Correctness of LLM-Generated CodeabstractWhen applying LLM-based code generation to software development projects that follow a feature-driven or rapid application development approach, it becomes necessary to estimate the functional correctness of the generated code in the absence of test cases. Just as a user selects a relevant document from a ranked list of retrieved ones, a software generation workflow requires a developer to choose (and potentially refine) a generated solution from a ranked list of alternative solutions, ordered by their posterior likelihoods. This implies that estimating the quality of a ranked list - akin to estimating ''relevance'' for query performance prediction (QPP) in IR - is also crucial for generative software development, where quality is defined in terms of ''functional correctness''. In this paper, we propose an in-context learning (ICL) based approach for code quality estimation. Our findings demonstrate that providing few-shot examples of functionally correct code from a training set enhances the performance of existing QPP approaches as well as a zero-shot-based approach for code quality estimation. Susmita Das 0002, Madhusudan Ghosh, Priyanka Swami, Debasis Ganguly, Gül Çalikli |
SIGIR | 5 |
| 2024 | Guidelines for using financial incentives in software-engineering experimentationabstractAbstract Context: Empirical studies with human participants (e.g., controlled experiments) are established methods in Software Engineering (SE) research to understand developers’ activities or the pros and cons of a technique, tool, or practice. Various guidelines and recommendations on designing and conducting different types of empirical studies in SE exist. However, the use of financial incentives (i.e., paying participants to compensate for their effort and improve the validity of a study) is rarely mentioned Objective: In this article, we analyze and discuss the use of financial incentives for human-oriented SE experimentation to derive corresponding guidelines and recommendations for researchers. Specifically, we propose how to extend the current state-of-the-art and provide a better understanding of when and how to incentivize. Method: We captured the state-of-the-art in SE by performing a Systematic Literature Review (SLR) involving 105 publications from six conferences and five journals published in 2020 and 2021. Then, we conducted an interdisciplinary analysis based on guidelines from experimental economics and behavioral psychology, two disciplines that research and use financial incentives. Results: Our results show that financial incentives are sparsely used in SE experimentation, mostly as completion fees. Especially performance-based and task-related financial incentives (i.e., payoff functions) are not used, even though we identified studies for which the validity may benefit from tailored payoff functions. To tackle this issue, we contribute an overview of how experiments in SE may benefit from financial incentivisation, a guideline for deciding on their use, and 11 recommendations on how to design them. Conclusions: We hope that our contributions get incorporated into standards (e.g., the ACM SIGSOFT Empirical Standards), helping researchers understand whether the use of financial incentives is useful for their experiments and how to define a suitable incentivisation strategy. Jacob Krüger, Gül Çalikli, Dmitri Bershadskyy, Siegmar Otto, Sarah Zabel, Robert Heyer |
Empir. Softw. Eng. | 2 |
| 2024 | Virtual Platform: Effective and Seamless Variability Management for Software SystemsabstractCustomization is a general trend in software engineering, demanding systems that support variable stakeholder requirements. Two opposing strategies are commonly used to create variants: software clone & own and software configuration with an integrated platform. Organizations often start with the former, which is cheap and agile, but does not scale. The latter scales by establishing an integrated platform that shares software assets between variants, but requires high up-front investments or risky migration processes. So, could we have a method that allows an easy transition or even combine the benefits of both strategies? We propose a method and tool that supports a truly incremental development of variant-rich systems, exploiting a spectrum between the opposing strategies. We design, formalize, and prototype a variability-management framework: the virtual platform. Virtual platform bridges clone & own and platform-oriented development. Relying on programming-language independent conceptual structures representing software assets, it offers operators for engineering and evolving a system, comprising: traditional, asset-oriented operators and novel, feature-oriented operators for incrementally adopting concepts of an integrated platform. The operators record meta-data that is exploited by other operators to support the transition. Among others, they eliminate expensive feature-location effort or the need to trace clones. A cost-and-benefit analysis of using the virtual platform to simulate the development of a real-world variant-rich system shows that it leads to benefits in terms of saved effort and time for clone detection and feature location. Furthermore, we present a user study indicating that the virtual platform effectively supports exploratory and hands-on tasks, outperforming manual development concerning correctness. We also observed that participants were significantly faster when performing typical variability management tasks using the virtual platform. Furthermore, participants perceived manual development to be significantly more difficult than using the virtual platform, preferring virtual platform for all our tasks. We supplement our findings with recommendations on when to use virtual platform and on incorporating the virtual platform in practice. Wardah Mahmood, Gül Çalikli, Daniel Strüber 0001, Ralf Lämmel, Mukelabai Mukelabai, Thorsten Berger |
IEEE Trans. Software Eng. | 2 |
| 2023 | Competencies for Code ReviewabstractPeer code review is a widely practiced software engineering process in which software developers collaboratively evaluate and improve source code quality. Whether developers can perform good reviews depends on whether they have sufficient competence and experience. However, the knowledge of what competencies developers need to execute code review is currently limited, thus hindering, for example, the creation of effective support tools and training strategies. To address this gap, we firstly identified 27 competencies relevant to performing code review through expert validation. Later, we conducted an online survey with 105 reviewers to rank these competencies along four dimensions: frequency of usage, importance, proficiency, and desire of reviewers to improve in that competency. The survey shows that technical competencies are considered essential to performing reviews and that respondents feel generally confident in their technical proficiency. Moreover, reviewers feel less confident in how to communicate clearly and give constructive feedback - competencies they consider like-wise an essential part of reviewing. Therefore, research and education should focus in more detail on how to support and develop reviewers' potential to communicate effectively during reviews. In the paper, we also discuss further implications for training, code review performance assessment, and reviewers of different experience level. Data and materials: https://doi.org/10.5281/zenodo.7401313 Pavlína Wurzel Gonçalves, Gül Çalikli, Alexander Serebrenik, Alberto Bacchelli |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2022 | Less is More: Supporting Developers in Vulnerability Detection during Code ReviewabstractReviewing source code from a security perspective has proven to be a difficult task. Indeed, previous research has shown that developers often miss even popular and easy-to-detect vulnerabilities during code review. Initial evidence suggests that a significant cause may lie in the reviewers' mental attitude and common practices. Larissa Braz, Christian Aeberhard, Gül Çalikli, Alberto Bacchelli |
ICSE | 3 |
| 2022 | First come first served: the impact of file position on code reviewabstractThe most popular code review tools (e.g., Gerrit and GitHub) present the files to review sorted in alphabetical order. Could this choice or, more generally, the relative position in which a file is presented bias the outcome of code reviews? We investigate this hypothesis by triangulating complementary evidence in a two-step study. First, we observe developers’ code review activity. We analyze the review comments pertaining to 219,476 Pull Requests (PRs) from 138 popular Java projects on GitHub. We found files shown earlier in a PR to receive more comments than files shown later, also when controlling for possible confounding factors: e.g., the presence of discussion threads or the lines added in a file. Second, we measure the impact of file position on defect finding in code review. Recruit- ing 106 participants, we conduct an online controlled experiment in which we measure participants’ performance in detecting two unrelated defects seeded into two different files. Participants are assigned to one of two treatments in which the position of the defective files is switched. For one type of defect, participants are not affected by its file’s position; for the other, they have 64% lower odds to identify it when its file is last as opposed to first. Overall, our findings provide evidence that the relative position in which files are presented has an impact on code reviews’ outcome; we discuss these results and implications for tool design and code review. Enrico Fregnan, Larissa Braz, Marco D'Ambros, Gül Çalikli, Alberto Bacchelli |
ESEC/SIGSOFT FSE | 4 |
| 2022 | Interpersonal Conflicts During Code Review: Developers' Experience and PracticesabstractCode review consists of manual inspection, discussion, and judgment of source code by developers other than the code's author. Due to discussions around competing ideas and group decision-making processes, interpersonal conflicts during code reviews are expected. This study systematically investigates how developers perceive code review conflicts and addresses interpersonal conflicts during code reviews as a theoretical construct. Through the thematic analysis of interviews conducted with 22 developers, we confirm that conflicts during code reviews are commonplace, anticipated and seen as normal by developers. Even though conflicts do happen and carry a negative impact for the review, conflicts-if resolved constructively-can also create value and bring improvement. Moreover, the analysis provided insights on how strongly conflicts during code review and its context (i.e., code, developer, team, organization) are intertwined. Finally, there are aspects specific to code review conflicts that call for the research and application of customized conflict resolution and management techniques, some of which are discussed in this paper. Preprint: https://arxiv.org/abs/2201.05425 Data and material: https://doi.org/10.5281/zenodo.5848794 Pavlína Wurzel Gonçalves, Gül Çalikli, Alberto Bacchelli |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2021 | Why Don't Developers Detect Improper Input Validation? '; DROP TABLE Papers; -abstractImproper Input Validation (IIV) is a software vulnerability that occurs when a system does not safely handle input data. Even though IIV is easy to detect and fix, it still commonly happens in practice. In this paper, we study to what extent developers can detect IIV and investigate underlying reasons. This knowledge is essential to better understand how to support developers in creating secure software systems. We conduct an online experiment with 146 participants, of which 105 report at least three years of professional software development experience. Our results show that the existence of a visible attack scenario facilitates the detection of IIV vulnerabilities and that a significant portion of developers who did not find the vulnerability initially could identify it when warned about its existence. Yet, a total of 60 participants could not detect the vulnerability even after the warning. Other factors, such as the frequency with which the participants perform code reviews, influence the detection of IIV. Preprint: https://arxiv.org/abs/2102.06251. Data and materials: https://doi.org/10.5281/zenodo.3996696. Larissa Braz, Enrico Fregnan, Gül Çalikli, Alberto Bacchelli |
ICSE | 3 |
| 2021 | How Explicit Feature Traces Did Not Impact Developers' MemoryabstractSoftware features are intuitive entities used to abstract and manage the functionalities of a software system, for instance, in product-line engineering and agile software development. Nonetheless, developers rarely make features explicit in code, which is why they have to perform costly program comprehension and particularly feature location to (re-)gain knowledge about the code. In a previous paper, we conducted an experiment on how explicit feature traces impact developers' program comprehension by facilitating feature location. We found that annotating features in code improved program comprehension, while decomposing them into classes had a negative impact. Additionally, but not reported in that paper, we were concerned with understanding whether the different traces would impact developers' memory regarding the code and its features. To this end, we repeatedly asked our participants questions about the code on different levels of detail within time periods of two weeks. Since developers' memory decays over time, we expected that our participants would provide fewer correct answers over time, with differences depending on the feature traces in their code. Unfortunately, the actual results were inconclusive and up for interpretation, particularly due to challenges in designing an experiment on developers' memory. In this paper, we discuss our experimental design, the null results, and challenges for improving the methodology of future studies in this direction. Jacob Krüger, Gül Çalikli, Thorsten Berger, Thomas Leich |
SANER | 2 |
| 2020 | Primers or reminders?: the effects of existing review comments on code reviewabstractIn contemporary code review, the comments put by reviewers on a specific code change are immediately visible to the other reviewers involved. Could this visibility prime new reviewers' attention (due to the human's proneness to availability bias), thus biasing the code review outcome? In this study, we investigate this topic by conducting a controlled experiment with 85 developers who perform a code review and a psychological experiment. With the psychological experiment, we find that ≈70% of participants are prone to availability bias. However, when it comes to the code review, our experiment results show that participants are primed only when the existing code review comment is about a type of bug that is not normally considered; when this comment is visible, participants are more likely to find another occurrence of this type of bug. Moreover, this priming effect does not influence reviewers' likelihood of detecting other types of bugs. Our findings suggest that the current code review practice is effective because existing review comments about bugs in code changes are not negative primers, rather positive reminders for bugs that would otherwise be overlooked during code review. Data and materials: https://doi.org/10.5281/zenodo.3653856 Davide Spadini, Gül Çalikli, Alberto Bacchelli |
ICSE | 2 |
| 2019 | Empirical Analysis of Hidden Technical Debt Patterns in Machine Learning Software
Mohannad Alahdab, Gül Çalikli |
PROFES | 2 |
| 2019 | Effects of explicit feature traceability on program comprehensionabstractDevelopers spend a substantial amount of their time with program comprehension. To improve their comprehension and refresh their memory, developers need to communicate with other developers, read the documentation, and analyze the source code. Many studies show that developers focus primarily on the source code and that small improvements can have a strong impact. As such, it is crucial to bring the code itself into a more comprehensible form. A particular technique for this purpose are explicit feature traces to easily identify a program’s functionalities. To improve our empirical understanding about the effects of feature traces, we report an online experiment with 49 professional software developers. We studied the impact of explicit feature traces, namely annotations and decomposition, on program comprehension and compared them to the same code without traces. Besides this experiment, we also asked our participants about their opinions in order to combine quantitative and qualitative data. Our results indicate that, as opposed to purely object-oriented code: (1) annotations can have positive effects on program comprehension; (2) decomposition can have a negative impact on bug localization; and (3) our participants perceive both techniques as beneficial. Moreover, none of the three code versions yields significant improvements on task completion time. Overall, our results indicate that lightweight traceability, such as using annotations, provides immediate benefits to developers during software development and maintenance without extensive training or tooling; and can improve current industrial practices that rely on heavyweight traceability tools (e.g., DOORS) and retroactive fulfillment of standards (e.g., ISO-26262, DO-178B). Jacob Krüger, Gül Çalikli, Thorsten Berger, Thomas Leich, Gunter Saake |
ESEC/SIGSOFT FSE | 2 |
| 2018 | Safety-Critical Systems and Agile Development: A Mapping StudyabstractIn the last decades, agile methods had a huge impact on how software is developed. In many cases, this has led to significant benefits, such as quality and speed of software deliveries to customers. However, safety-critical systems have widely been dismissed from benefiting from agile methods. Products that include safety critical aspects are therefore faced with a situation in which the development of safety-critical parts can significantly limit the potential speed-up through agile methods, for the full product, but also in the non-safety critical parts. For such products, the ability to develop safety-critical software in an agile way will generate a competitive advantage. In order to enable future research in this important area, we present in this paper a mapping of the current state of practice based on a mixed method approach. Starting from a workshop with experts from six large Swedish product development companies we develop a lens for our analysis. We then present a systematic mapping study on safety-critical systems and agile development through this lens in order to map potential benefits, challenges, and solution candidates for guiding future research. Rashidah Kasauli, Eric Knauss, Benjamin Kanagwa, Agneta Nilsson, Gül Çalikli |
SEAA | 5 |
| 2018 | Measure early and decide fast: transforming quality management and measurement to continuous deploymentabstractContinuous deployment has become software companies' inevitable response to the economic pressures of the market. At the same time, software quality is crucial in order to meet customers' expectations and hence succeed in the market. Therefore, current quality management processes require transformation in order to keep up with the fast pace of the market while at the same time meeting customers' expectations. In order to figure out how the current quality management process should be transformed to keep up with the fast pace of the market while ensuring both product quality and continuous deployment, we conducted a qualitative study at a large infrastructure provider company. During the interviews we conducted with the quality manager, developer and test architect, we used a metrics portfolio consisting of 59 candidate metrics that can be used in the transformed quality management process. Our findings show that, out of these candidate metrics, 9 metrics should be used in the internal quality measurement dashboard for quality check at the end of the software development life-cycle (SDLC) before the software is released to customer site, while 3 metrics should be used by quality manager to monitor earlier phases of SDLC and 5 metrics should also be delegated to earlier phases of SDLC but without the involvement of the quality manager. To summarize, our study support the claim that quality managers should not be only gatekeepers, but also proactive controllers of quality by monitoring earlier phases of the SDLC. Gül Çalikli, Miroslaw Staron, Wilhelm Meding |
ICSSP | 1 |
| 2018 | Involving External Stakeholders in Project CoursesabstractProblem: The involvement of external stakeholders in capstone projects and project courses is desirable due to its potential positive effects on the students. Capstone projects particularly profit from the inclusion of an industrial partner to make the project relevant and help students acquire professional skills. In addition, an increasing push towards education that is aligned with industry and incorporates industrial partners can be observed. However, the involvement of external stakeholders in teaching moments can create friction and could, in the worst case, lead to frustration of all involved parties. Contribution: We developed a model that allows analysing the involvement of external stakeholders in university courses both in a retrospective fashion, to gain insights from past course instances, and in a constructive fashion, to plan the involvement of external stakeholders. Key Concepts: The conceptual model and the accompanying guideline guide the teachers in their analysis of stakeholder involvement. The model is comprised of several activities (define, execute, and evaluate the collaboration). The guideline provides questions that the teachers should answer for each of these activities. In the constructive use, the model allows teachers to define an action plan based on an analysis of potential stakeholders and the pedagogical objectives. In the retrospective use, the model allows teachers to identify issues that appeared during the project and their underlying causes. Drawing from ideas of the reflective practitioner, the model contains an emphasis on reflection and interpretation of the observations made by the teacher and other groups involved in the courses. Key Lessons: Applying the model retrospectively to a total of eight courses shows that it is possible to reveal hitherto implicit risks and assumptions and to gain a better insight into the interaction between external stakeholders and students. Our empirical data reveals seven recurring risk themes that categorise the different risks appearing in the analysed courses. These themes can also be used to categorise mitigation strategies to address these risks proactively. Additionally, aspects not related to external stakeholders, e.g., about the interaction of the project with other courses in the study programme, have been revealed. The constructive use of the model for one course has proved helpful in identifying action alternatives and finally deciding to not include external stakeholders in the project due to the perceived cost-benefit-ratio. Implications to Practice: Our evaluation shows that the model is a viable and useful tool that allows teachers to reason about and plan the involvement of external stakeholders in a variety of course settings, and in particular in capstone projects. Jan-Philipp Steghöfer, Håkan Burden, Regina Hebig, Gül Çalikli, Robert Feldt, Imed Hammouda, Jennifer Horkoff, Eric Knauss, Grischa Liebel |
ACM Trans. Comput. Educ. | 4 |
| 2018 | Threat analysis of software systems: A systematic literature review
Katja Tuma, Gül Çalikli, Riccardo Scandariato |
J. Syst. Softw. | 2 |
| 2018 | Guest Editorial Special Section on Engineering Industrial Big Data Analytics Platforms for Internet of ThingsabstractOver the last few years, a large number of Internet of Things (IoT) solutions have come to the IoT marketplace. Typically, each of these IoT solutions are designed to perform a single or minimal number of tasks (primary usage). We believe a significant amount of knowledge and insights are hidden in these data silos that can be used to improve our lives; such data include our behaviors, habits, preferences, life patterns, and resource consumption. To discover such knowledge, we need to acquire and analyze this data together in a large scale. To discover useful information and deriving conclusions toward supporting efficient and effective decision making, industrial IoT platform needs to support variety of different data analytics processes such as inspecting, cleaning, transforming, and modeling data, especially in big data context. IoT middleware platforms have been developed in both academic and industrial settings in order to facilitate IoT data management tasks including data analytics. However, engineering these general-purpose industrial-grade big data analytics platforms need to address many challenges. We have accepted six manuscripts out of 24 submissions for this special section (25% acceptance rate) after the strict peerreview processes. Each manuscript has been blindly reviewed by at least three external reviewers before the decisions were made. The papers are briefly summarized. Charith Perera, Athanasios V. Vasilakos, Gül Çalikli, Quan Z. Sheng, Kuanching Li |
IEEE Trans. Ind. Informatics | 3 |
| 2017 | Learning to share: engineering adaptive decision-support for online social networksabstractSome online social networks (OSNs) allow users to define friendship-groups as reusable shortcuts for sharing information with multiple contacts. Posting exclusively to a friendship-group gives some privacy control, while supporting communication with (and within) this group. However, recipients of such posts may want to reuse content for their own social advantage, and can bypass existing controls by copy-pasting into a new post; this cross-posting poses privacy risks. This paper presents a learning to share approach that enables the incorporation of more nuanced privacy controls into OSNs. Specifically, we propose a reusable, adaptive software architecture that uses rigorous runtime analysis to help OSN users to make informed decisions about suitable audiences for their posts. This is achieved by supporting dynamic formation of recipient-groups that benefit social interactions while reducing privacy risks. We exemplify the use of our approach in the context of Facebook. Yasmin Rafiq, Luke Dickens, Alessandra Russo, Arosha K. Bandara, Mu Yang, Avelie Stuart, Mark Levine, Gül Çalikli, Blaine A. Price, Bashar Nuseibeh |
ASE | 8 |
| 2015 | Empirical analysis of factors affecting confirmation bias levels of software engineers
Gül Çalikli, Ayse Basar Bener |
Softw. Qual. J. | 1 |
| 2013 | Towards a Metric Suite Proposal to Quantify Confirmation Biases of DevelopersabstractThe goal of software metrics is the identification and measurement of the essential parameters that affect software development. Metrics can be used to improve software quality and productivity. Existing metrics in the literature are mostly product or process related. However, thought processes of people have a significant impact on software quality as software is designed, implemented and tested by people. Therefore, in defining new metrics, we need to take into account human cognitive aspects. Our research aims to address this need through the proposal of a new metric scheme to quantify a specific human cognitive aspect, namely "confirmation bias". In our previous research, in order to quantify confirmation bias, we defined a methodology to measure confirmation biases of people. In this research, we propose a metric suite that would be used by practitioners during daily decision making. Our proposed metric set consists of six metrics with a theoretical basis in cognitive psychology and measurement theory. Empirical sample of these metrics are collected from two software companies that are specialized in two different domains in order to demonstrate their feasibility. We suggest ways in which practitioners may use these metrics to improve software development process. Gül Çalikli, Ayse Basar Bener, Turgay Aytac, Övünç Bozcan |
ESEM | 1 |
| 2013 | The Impact of Confirmation Bias on the Release-based Defect Prediction of Developer Groups
Gül Çalikli, Ayse Basar Bener |
SEKE | 1 |
| 2013 | Influence of confirmation biases of developers on software quality: an empirical study
Gül Çalikli, Ayse Basar Bener |
Softw. Qual. J. | 1 |
| 2012 | Dione: an integrated measurement and defect prediction solutionabstractWe present an integrated measurement and defect prediction tool: Dione. Our tool enables organizations to measure, monitor, and control product quality through learning based defect prediction. Similar existing tools either provide data collection and analytics, or work just as a prediction engine. Therefore, companies need to deal with multiple tools with incompatible interfaces in order to deploy a complete measurement and prediction solution. Dione provides a fully integrated solution where data extraction, defect prediction and reporting steps fit seamlessly. In this paper, we present the major functionality and architectural elements of Dione followed by an overview of our demonstration. Bora Caglayan, Ayse Tosun Misirli, Gül Çalikli, Ayse Basar Bener, Turgay Aytac, Burak Turhan |
SIGSOFT FSE | 3 |
| 2010 | Preliminary analysis of the effects of confirmation bias on software defect densityabstractIn cognitive psychology, confirmation bias is defined as the tendency of people to verify hypotheses rather than refuting them. During unit testing software developers should aim to fail their code. However, due to confirmation bias, most defects might be overlooked leading to an increase in software defect density. In this research, we empirically analyze the effect of confirmation bias of software developers on software defect density. Gül Çalikli, Ayse Basar Bener |
ESEM | 1 |
| 2010 | An analysis of the effects of company culture, education and experience on confirmation bias levels of software developers and testersabstractIn this paper, we present a preliminary analysis of factors such as company culture, education and experience, on confirmation bias levels of software developers and testers. Confirmation bias is defined as the tendency of people to verify their hypotheses rather than refuting them and thus it has an effect on all software testing. Gül Çalikli, Ayse Basar Bener, Berna Arslan |
ICSE (2) | 1 |