Burak Turhan

dblp:62/1974 · DBLP profile ↗
← Back
90ranked-venue papers
12as first author
23since 2021 · last 2025
0000-0003-1511-2163ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 85 · 10 first-author · 23 since 2021Artificial intelligence and machine learning · 7 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 4Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Cognitive Biases in Software Engineering: Debiasing Through Reconception
abstract
Background: Cognitive biases are systematic errors in reasoning that can lead to inaccurate decision-making across all areas of software production, regardless of the domain, programming language, or development method. Given the central role of software across all sectors, such biases can result in largescale inefficiencies, delays, and increased costs. Goal: While prior software engineering (SE) research predominantly focuses on specific tasks and quantitative methods, the vision of this paper is to study qualitatively how cognitive biases emerge, persist, and can be mitigated to improve decision-making throughout the Software Development Life Cycle (SDLC). Method: This vision uniquely applies the concept of reconception as a theoretical lens to explore how software professionals' cognitive models influence their management of dialectical oppositions, such as project velocity vs product quality and open-source vs proprietary control. Utilising Socio-Technical Grounded Theory (STGT) and dialectical inquiry, it examines how reconceiving these oppositions affects cognitive models and decision-making processes. Expected Outcome: A broad and high-level theoretical framework that explains how selected cognitive biases influence SE decision-making across the SDLC. The framework is developed through a two-phase STGT process: identifying salient bias categories in the first phase and focusing on the role of specific biases in shaping decision patterns in the second. This highlevel theory will open new research opportunities for future investigations into bias-informed decision-making in SE. Conclusion: This theoretical framework will support the development of empirically grounded debiasing strategies for more reliable decision-making in software engineering.
Heidi Hietala, Burak Turhan
ESEM2
2025 Navigating fairness: practitioners' understanding, challenges, and strategies in AI/ML development
abstract
Abstract The rise in the use of AI/ML applications across industries has sparked more discussions about the fairness of AI/ML in recent times. While prior research on the fairness of AI/ML exists, there is a lack of empirical studies focused on understanding perspectives and experiences of AI practitioners in developing a fair AI/ML system. Understanding AI practitioners’ perspectives and experiences on the fairness of AI/ML systems is important because they are directly involved in its development and deployment and their insights can offer valuable real-world perspectives on the challenges associated with ensuring fairness in AI/ML systems. We conducted semi-structured interviews with 22 AI practitioners to investigate their understanding of what a ‘fair AI/ML’ is, the challenges they face in developing a fair AI/ML system, the consequences of developing an unfair AI/ML system, and the strategies they employ to ensure AI/ML system fairness. By exploring AI practitioners’ perspectives and experiences, this study provides actionable insights to enhance AI/ML fairness, which may promote fairer systems, reduce bias, and foster public trust in AI technologies. Additionally, we also identify areas for further investigation and offer recommendations to aid AI practitioners and AI companies in navigating fairness.
Aastha Pant, Rashina Hoda, Chakkrit Tantithamthavorn, Burak Turhan
Empir. Softw. Eng.4
2025 Instance Space Analysis of Testing of Autonomous Vehicles in Critical Scenarios
abstract
Before being deployed on roads, Autonomous Vehicles (AVs) must undergo comprehensive testing. Safety-critical situations, however, are infrequent in usual driving conditions, so simulated scenarios are used to create them. A test scenario comprises static and dynamic features related to the AV and the test environment; the representation of these features is complex and makes testing a heavy process. A test scenario is effective if it identifies incorrect behaviors of the AV. In this article, we present a technique for identifying key features of test scenarios associated with their effectiveness using Instance Space Analysis (ISA). ISA generates a ( \(2D\) ) representation of test scenarios and their features. This visualization helps to identify combinations of features that make a test scenario effective. We present a graphical representation of each feature that helps identify how well each testing technique explores the search space. While identifying key features is a primary goal, this study specifically seeks to determine the critical features that differentiate the performance of algorithms. Finally, we present metrics to assess the robustness of testing algorithms and the scenarios generated. Collecting essential features in combination with their values associated with effectiveness can be used for selection and prioritization of effective test cases.
Victor Crespo-Rodriguez, Neelofar, Aldeida Aleti, Burak Turhan
ACM Trans. Softw. Eng. Methodol.4
2025 Does Treatment Adherence Impact Experiment Results in TDD?
abstract
Context:In software engineering (SE) experiments, the way in which a treatment is applied could affect results. Different interpretations of how to apply the treatment and decisions on treatment adherence could lead to different results when data are analysed.Objective:This paper aims to study whether treatment adherence has an impact on the results of an SE experiment.Method:The experiment used as test case for our research uses Test-Driven Development (TDD) and Incremental Test-Last Development, (ITLD) as treatments. We reported elsewhere the design and results of such an experiment where 24 participants were recruited from industry. Here, we compare experiment results depending on the use of data from adherent participants or data from all the participants irrespective of their adherence to treatments.Results:Only 40% of the participants adhere to both TDD protocol and to the ITLD protocol; 27% never followed TDD; 20% used TDD even in the control group; 13% are defiers (used TDD in ITLD session but not in TDD session). Considering that both TDD and ITLD are less complex than other SE methods, we can hypothesize that more complex SE techniques could get even lower adherence to the treatment.Conclusion:Both TDD and ITLD are applied differently across participants. Training participants could not be enough to ensure a medium to large adherence of experiment participants. Adherence to treatments impacts results and should not be taken for granted in SE experiments.
Itir Karac, José Ignacio Panach, Burak Turhan, Natalia Juristo Juzgado
IEEE Trans. Software Eng.3
2024 Ethics in AI through the practitioner's view: a grounded theory literature review
abstract
Abstract The term ethics is widely used, explored, and debated in the context of developing Artificial Intelligence (AI) based software systems. In recent years, numerous incidents have raised the profile of ethical issues in AI development and led to public concerns about the proliferation of AI technology in our everyday lives. But what do we know about the views and experiences of those who develop these systems – the AI practitioners? We conducted a grounded theory literature review (GTLR) of 38 primary empirical studies that included AI practitioners’ views on ethics in AI and analysed them to derive five categories: practitioner awareness, perception, need, challenge, and approach. These are underpinned by multiple codes and concepts that we explain with evidence from the included studies. We present a taxonomy of ethics in AI from practitioners’ viewpoints to assist AI practitioners in identifying and understanding the different aspects of AI ethics. The taxonomy provides a landscape view of the key aspects that concern AI practitioners when it comes to ethics in AI. We also share an agenda for future research studies and recommendations for practitioners, managers, and organisations to help in their efforts to better consider and implement ethics in AI.
Aastha Pant, Rashina Hoda, Chakkrit Tantithamthavorn, Burak Turhan
Empir. Softw. Eng.4
2024 Ethics in the Age of AI: An Analysis of AI Practitioners' Awareness and Challenges
abstract
Ethics in AI has become a debated topic of public and expert discourse in recent years. But what do people who build AI—AI practitioners—have to say about their understanding of AI ethics and the challenges associated with incorporating it into the AI-based systems they develop? Understanding AI practitioners’ views on AI ethics is important as they are the ones closest to the AI systems and can bring about changes and improvements. We conducted a survey aimed at understanding AI practitioners’ awareness of AI ethics and their challenges in incorporating ethics. Based on 100 AI practitioners’ responses, our findings indicate that the majority of AI practitioners had a reasonable familiarity with the concept of AI ethics, primarily due to workplace rules and policies . Privacy protection and security was the ethical principle that the majority of them were aware of. Formal education/training was considered somewhat helpful in preparing practitioners to incorporate AI ethics. The challenges that AI practitioners faced in the development of ethical AI-based systems included (i) general challenges, (ii) technology-related challenges, and (iii) human-related challenges. We also identified areas needing further investigation and provided recommendations to assist AI practitioners and companies in incorporating ethics into AI development.
Aastha Pant, Rashina Hoda, Simone V. Spiegler, Chakkrit Tantithamthavorn, Burak Turhan
ACM Trans. Softw. Eng. Methodol.5
2024 On the Impact of Lower Recall and Precision in Defect Prediction for Guiding Search-based Software Testing
abstract
Defect predictors, static bug detectors, and humans inspecting the code can propose locations in the program that are more likely to be buggy before they are discovered through testing. Automated test generators such as search-based software testing (SBST) techniques can use this information to direct their search for test cases to likely buggy code, thus speeding up the process of detecting existing bugs in those locations. Often the predictions given by these tools or humans are imprecise, which can misguide the SBST technique and may deteriorate its performance. In this article, we study the impact of imprecision in defect prediction on the bug detection effectiveness of SBST. Our study finds that the recall of the defect predictor, i.e., the proportion of correctly identified buggy code, has a significant impact on bug detection effectiveness of SBST with a large effect size. More precisely, the SBST technique detects 7.5 fewer bugs on average (out of 420 bugs) for every 5% decrements of the recall. However, the effect of precision, a measure for false alarms, is not of meaningful practical significance, as indicated by a very small effect size. In the context of combining defect prediction and SBST, our recommendation is to increase the recall of defect predictors as a primary objective and precision as a secondary objective. In our experiments, we find that 75% precision is as good as 100% precision. To account for the imprecision of defect predictors, in particular low recall values, SBST techniques should be designed to search for test cases that also cover the predicted non-buggy parts of the program, while prioritising the parts that have been predicted as buggy.
Anjana Perera, Burak Turhan, Aldeida Aleti, Marcel Böhme
ACM Trans. Softw. Eng. Methodol.2
2024 Beyond Accuracy: An Empirical Study on Unit Testing in Open-source Deep Learning Projects
abstract
Deep Learning (DL) models have rapidly advanced, focusing on achieving high performance through testing model accuracy and robustness. However, it is unclear whether DL projects, as software systems, are tested thoroughly or functionally correct when there is a need to treat and test them like other software systems. Therefore, we empirically study the unit tests in open-source DL projects, analyzing 9,129 projects from GitHub. We find that: (1) unit tested DL projects have positive correlation with the open-source project metrics and have a higher acceptance rate of pull requests; (2) 68% of the sampled DL projects are not unit tested at all; (3) the layer and utilities (utils) of DL models have the most unit tests. Based on these findings and previous research outcomes, we built a mapping taxonomy between unit tests and faults in DL projects. We discuss the implications of our findings for developers and researchers and highlight the need for unit testing in open-source DL projects to ensure their reliability and stability. The study contributes to this community by raising awareness of the importance of unit testing in DL projects and encouraging further research in this area.
Han Wang 0023, Sijia Yu, Chunyang Chen 0001, Burak Turhan, Xiaodong Zhu 0001
ACM Trans. Softw. Eng. Methodol.4
2023 Automated detection, categorisation and developers' experience with the violations of honesty in mobile apps
abstract
Abstract Human values such as honesty, social responsibility, fairness, privacy, and the like are things considered important by individuals and society. Software systems, including mobile software applications (apps), may ignore or violate such values, leading to negative effects in various ways for individuals and society. While some works have investigated different aspects of human values in software engineering, this mixed-methods study focuses on honesty as a critical human value. In particular, we studied (i) how to detect honesty violations in mobile apps, (ii) the types of honesty violations in mobile apps, and (iii) the perspectives of app developers on these detected honesty violations. We first develop and evaluate 7 machine learning (ML) models to automatically detect violations of the value of honesty in app reviews from an end-user perspective. The most promising was a Deep Neural Network model with F1 score of 0.921. We then conducted a manual analysis of 401 reviews containing honesty violations and characterised honesty violations in mobile apps into 10 categories: unfair cancellation and refund policies; false advertisements; delusive subscriptions; cheating systems; inaccurate information; unfair fees; no service; deletion of reviews; impersonation; and fraudulent-looking apps. A developer survey and interview study with mobile developers then identified 7 key causes behind honesty violations in mobile apps and 8 strategies to avoid or fix such violations. The findings of our developer study also articulate the negative consequences that honesty violations might bring for businesses, developers, and users. Finally, the app developers’ feedback shows that our prototype ML-based models can have promising benefits in practice.
Humphrey O. Obie, Hung Du, Kashumi Madampe, Mojtaba Shahin, Idowu Ilekura, John C. Grundy, Li Li 0029, Jon Whittle 0001, Burak Turhan, Hourieh Khalajzadeh
Empir. Softw. Eng.9
2023 An Experimental Assessment of Using Theoretical Defect Predictors to Guide Search-Based Software Testing
abstract
Automated test generators, such as search-based software testing (SBST) techniques are primarily guided by coverage information. As a result, they are very effective at achieving high code coverage. However, is high code coverage alone sufficient to detect bugs effectively? In this paper, we propose a new SBST technique, predictive many objective sorting algorithm (PreMOSA), which augments coverage information with defect prediction information to decide where to increase the test coverage in the class under test (CUT). Through an experimental evaluation using 420 labelled bugs on the Defects4J benchmark and using theoretical defect predictors, we demonstrate the improved effectiveness and efficiency of PreMOSA in detecting bugs when using any acceptable defect predictor, i.e., a defect predictor with recall and precision$\geq$75%, compared to the state-of-the-art dynamic many objective sorting algorithm (DynaMOSA). PreMOSA detects up to 8.3% more labelled bugs on average than DynaMOSA when given a time budget of 2 minutes for test generation per CUT.
Anjana Perera, Aldeida Aleti, Burak Turhan, Marcel Böhme
IEEE Trans. Software Eng.3
2023 Confirmation Bias and Time Pressure: A Family of Experiments in Software Testing
abstract
Background: Software testers manifest confirmation bias (the cognitive tendency) when they design relatively more specification consistent test cases than specification inconsistent test cases. Time pressure may influence confirmation bias of testers per the research in the psychology discipline.Objective: We examine the manifestation of confirmation bias of software testers while designing functional test cases, and the effect of time pressure on confirmation bias in the same context.Method: We executed one internal and two external experimental replications concerning the original experimentation in Oulu. We analyse individual replications and meta-analyse our family of experiments (the original and replications) for joint results on the phenomena. Results: Our findings indicate a significant manifestation of confirmation bias by software testers during the designing of functional test cases. Time pressure significantly promoted confirmation bias among testers per the joint results of the family. The different experimental sites affected the results; however, we did not detect any effects of site-specific variables.Conclusion: Software testers should develop an outside-of-the-box thinking attitude to counter the manifestation of confirmation bias. Time pressure can be manoeuvred by centring manual suites on the designing and consequently the execution of inconsistent test cases, while automated testing focuses on consistent ones.
Iflaah Salman, Burak Turhan, Robert Ramac, Vladimir Mandic
IEEE Trans. Software Eng.2
2022 On the Violation of Honesty in Mobile Apps: Automated Detection and Categories
abstract
Human values such as integrity, privacy, curiosity, security, and honesty are guiding principles for what people consider important in life. Such human values may be violated by mobile software applications (apps), and the negative effects of such human value violations can be seen in various ways in society. In this work, we focus on the human value of honesty. We present a model to support the automatic identification of violations of the value of honesty from app reviews from an end-user perspective. Beyond the automatic detection of honesty violations by apps, we also aim to better understand different categories of honesty violations expressed by users in their app reviews. The result of our manual analysis of our honesty violations dataset shows that honesty violations can be characterised into ten categories: unfair cancellation and refund policies; false advertisements; delusive subscriptions; cheating systems; inaccurate information; unfair fees; no service; deletion of reviews; impersonation; and fraudulent-looking apps. Based on these results, we argue for a conscious effort in developing more honest software artefacts including mobile apps, and the promotion of honesty as a key value in software development practices. Furthermore, we discuss the role of app distribution platforms as enforcers of ethical systems supporting human values, and highlight some proposed next steps for human values in software engineering (SE) research.
Humphrey O. Obie, Idowu Ilekura, Hung Du, Mojtaba Shahin, John C. Grundy, Li Li 0029, Jon Whittle 0001, Burak Turhan
MSR8
2022 A fine-grained data set and analysis of tangling in bug fixing commits
abstract
Abstract Context Tangled commits are changes to software that address multiple concerns at once. For researchers interested in bugs, tangled commits mean that they actually study not only bugs, but also other concerns irrelevant for the study of bugs. Objective We want to improve our understanding of the prevalence of tangling and the types of changes that are tangled within bug fixing commits. Methods We use a crowd sourcing approach for manual labeling to validate which changes contribute to bug fixes for each line in bug fixing commits. Each line is labeled by four participants. If at least three participants agree on the same label, we have consensus. Results We estimate that between 17% and 32% of all changes in bug fixing commits modify the source code to fix the underlying problem. However, when we only consider changes to the production code files this ratio increases to 66% to 87%. We find that about 11% of lines are hard to label leading to active disagreements between participants. Due to confirmed tangling and the uncertainty in our data, we estimate that 3% to 47% of data is noisy without manual untangling, depending on the use case. Conclusion Tangled commits have a high prevalence in bug fixes and can lead to a large amount of noise in the data. Prior research indicates that this noise may alter results. As researchers, we should be skeptics and assume that unvalidated data is likely very noisy, until proven otherwise.
Steffen Herbold, Alexander Trautsch, Benjamin Ledel, Alireza Aghamohammadi, Taher Ahmed Ghaleb, Kuljit Kaur Chahal, Tim Bossenmaier, Bhaveet Nagaria, Philip Makedonski, Matin Nili Ahmadabadi, Kristóf Szabados, Helge Spieker, Matej Madeja, Nathaniel Hoy, Valentina Lenarduzzi, Shangwen Wang, Gema Rodríguez-Pérez, Ricardo Colomo-Palacios, Roberto Verdecchia, Paramvir Singh, Yihao Qin, Debasish Chakroborti, Willard Davis, Vijay Walunj, Diego Marcilio, Omar Alam, Abdullah Aldaeej, Idan Amit, Burak Turhan, Simon Eismann, Anna-Katharina Wickert, Ivano Malavolta, Matús Sulír, Fatemeh Hendijani Fard, Austin Z. Henley, Stratos Kourtzanidis, Eray Tüzün, Christoph Treude, Simin Maleki Shamasbi, Ivan Pashchenko, Marvin Wyrich, James C. Davis 0001, Alexander Serebrenik, Ella Albrecht, Ethem Utku Aktas, Daniel Strüber 0001, Johannes Erbel
Empir. Softw. Eng.30
2022 Search-based fairness testing for regression-based machine learning systems
abstract
Abstract Context Machine learning (ML) software systems are permeating many aspects of our life, such as healthcare, transportation, banking, and recruitment. These systems are trained with data that is often biased, resulting in biased behaviour. To address this issue, fairness testing approaches have been proposed to test ML systems for fairness, which predominantly focus on assessing classification-based ML systems. These methods are not applicable to regression-based systems, for example, they do not quantify the magnitude of the disparity in predicted outcomes, which we identify as important in the context of regression-based ML systems. Method: We conduct this study as design science research. We identify the problem instance in the context of emergency department (ED) wait-time prediction. In this paper, we develop an effective and efficient fairness testing approach to evaluate the fairness of regression-based ML systems. We propose fairness degree, which is a new fairness measure for regression-based ML systems, and a novel search-based fairness testing (SBFT) approach for testing regression-based machine learning systems. We apply the proposed solutions to ED wait-time prediction software. Results: We experimentally evaluate the effectiveness and efficiency of the proposed approach with ML systems trained on real observational data from the healthcare domain. We demonstrate that SBFT significantly outperforms existing fairness testing approaches, with up to 111% and 190% increase in effectiveness and efficiency of SBFT compared to the best performing existing approaches. Conclusion: These findings indicate that our novel fairness measure and the new approach for fairness testing of regression-based ML systems can identify the degree of fairness in predictions, which can help software teams to make data-informed decisions about whether such software systems are ready to deploy. The scientific knowledge gained from our work can be phrased as a technological rule; to measure the fairness of the regression-based ML systems in the context of emergency department wait-time prediction use fairness degree and search-based techniques to approximate it.
Anjana Perera, Aldeida Aleti, Chakkrit Tantithamthavorn, Jirayus Jiarpakdee, Burak Turhan, Lisa Kuhn, Katie Walker
Empir. Softw. Eng.5
2022 Use and Misuse of the Term "Experiment" in Mining Software Repositories Research
abstract
The significant momentum and importance of Mining Software Repositories (MSR) in Software Engineering (SE) has fostered new opportunities and challenges for extensive empirical research. However, MSR researchers seem to struggle to characterize the empirical methods they use into the existing empirical SE body of knowledge. This is especially the case of MSR experiments. To provide evidence on the special characteristics of MSR experiments and their differences with experiments traditionally acknowledged in SE so far, we elicited the hallmarks that differentiate an experiment from other types of empirical studies and characterized the hallmarks and types of experiments in MSR. We analyzed MSR literature obtained from a small-scale systematic mapping study to assess the use of the term experiment in MSR. We found that 19% of the papers claiming to be an experiment are indeed not an experiment at all but also observational studies, so they use the term in a misleading way. From the remaining 81% of the papers, only one of them refers to a genuine controlled experiment while the others stand for experiments with limited control. MSR researchers tend to overlook such limitations, compromising the interpretation of the results of their studies. We provide recommendations and insights to support the improvement of MSR experiments.
Claudia P. Ayala, Burak Turhan, Xavier Franch, Natalia Juristo Juzgado
IEEE Trans. Software Eng.2
2022 How Templated Requirements Specifications Inhibit Creativity in Software Engineering
abstract
Desiderata is a general term for stakeholder needs, desires or preferences. Recent experiments demonstrate that presenting desiderata as templated requirements specifications leads to less creative solutions. However, these experiments do not establish how the presentation of desiderata affects design creativity. This study, therefore, aims to explore the cognitive mechanisms by which presenting desiderata as templated requirements specifications reduces creativity during software design. Forty-two software designers, organized into 21 pairs, participated in a dialog-based protocol study. Their interactions were transcribed and the transcripts were analyzed in two ways: (1) using inductive process coding and (2) using an a-priori coding scheme focusing on fixation and critical thinking. Process coding shows that participants exhibited seven categories of behavior: making design moves, uncritically accepting, rejecting, grouping, questioning, assuming and considering quality criteria. Closed coding shows that participants tend to accept given requirements and priority levels while rejecting newer, more innovative design ideas. Overall, the results suggest that designers fixate on desiderata presented as templated requirements specifications, hindering critical thinking. More precisely, requirements fixation mediates the negative relationship between specification formality and creativity.
Rahul Mohanani, Paul Ralph, Burak Turhan, Vladimir Mandic
IEEE Trans. Software Eng.3
2022 What Leads to a Confirmatory or Disconfirmatory Behavior of Software Testers?
abstract
Background:The existing literature in software engineering reports adverse effects of confirmation bias on software testing. Confirmation bias among software testers leads to confirmatory behavior, which is designing or executing relatively more specification consistent test cases (confirmatory behavior) than specification inconsistent test cases (disconfirmatory behavior).Objective:We aim to explore the antecedents to confirmatory and disconfirmatory behavior of software testers. Furthermore, we aim to understand why and how those antecedents lead to (dis)confirmatory behavior.Method:We follow grounded theory method for the analyses of the data collected through semi-structured interviews with twelve software testers.Results:We identified twenty antecedents to (dis)confirmatory behavior, and classified them in nine categories. Experience and Time are the two major categories. Experience is a disconfirmatory category, which also determines which behavior (confirmatory or disconfirmatory) occurs first among software testers, as an effect of other antecedents. Time Pressure is a confirmatory antecedent of the Time category. It also contributes to the confirmatory effects of antecedents of other categories.Conclusion:The disconfirmatory antecedents, especially that belong to the testing process, e.g., test suite reviews by project team members, may help circumvent the deleterious effects of confirmation bias in software testing. If a team’s resources permit, the designing and execution of a test suite could be divided among the test team members, as different perspectives of testers may help to detect more errors. The results of our study are based on a single context where dedicated testing teams focus on higher levels of testing. The study’s scope does not account for the testing performed by developers. Future work includes exploring other contexts to extend our results.
Iflaah Salman, Pilar Rodríguez 0002, Burak Turhan, Ayse Tosun Misirli, Arda Gureller
IEEE Trans. Software Eng.3
2021 Does Domain Change the Opinion of Individuals on Human Values? A Preliminary Investigation on eHealth Apps End-users
abstract
The elicitation of end-users& human values - such as freedom, honesty, transparency, etc - is important in the development of software systems. We carried out two preliminary Q-studies to understand (a) the general human value opinion types of eHealth applications (apps) end-users (b) the eHealth domain human value opinion types of eHealth apps end-users (c) whether there are differences between the general and eHealth domain opinion types. Our early results show three value opinion types using generic value instruments: (1) fun-loving, success-driven and independent end-user, (2) security-conscious, socially-concerned, and success-driven end-user, and (3) benevolent, success-driven, and conformist end-user. Our results also show two value opinion types using domain-specific value instruments: (1) security-conscious, reputable, and honest end-user, and (2) success-driven, reputable and pain-avoiding end-user. Given these results, consideration should be given to domain context in the design and application of values elicitation instruments.
Humphrey O. Obie, Mojtaba Shahin, John C. Grundy, Burak Turhan, Li Li 0029, Jon Whittle 0001
APSEC4
2021 A family of experiments on test-driven development
Adrián Santos, Sira Vegas, Óscar Dieste Tubío, Fernando Uyaguari, Ayse Tosun Misirli, Davide Fucci, Burak Turhan, Giuseppe Scanniello, Simone Romano 0001, Itir Karac, Marco Kuhrmann, Vladimir Mandic, Robert Ramac, Dietmar Pfahl, Christian Engblom, Jarno Kyykka, Kerli Rungi, Carolina Palomeque, Jaroslav Spisak, Markku Oivo, Natalia Juristo Juzgado
Empir. Softw. Eng.7
2021 Studying test-driven development and its retainment over a six-month time span
Maria Teresa Baldassarre, Danilo Caivano, Davide Fucci, Natalia Juristo Juzgado, Simone Romano 0001, Giuseppe Scanniello, Burak Turhan
J. Syst. Softw.7
2021 On researcher bias in Software Engineering experiments
Simone Romano 0001, Davide Fucci, Giuseppe Scanniello, Maria Teresa Baldassarre, Burak Turhan, Natalia Juristo Juzgado
J. Syst. Softw.5
2021 A Controlled Experiment with Novice Developers on the Impact of Task Description Granularity on Software Quality in Test-Driven Development
abstract
Background: Test-Driven Development (TDD) is an iterative software development process characterized by test-code-refactor cycle. TDD recommends that developers work on small and manageable tasks at each iteration. However, the ability to break tasks into small work items effectively is a learned skill that improves with experience. In experimental studies of TDD, the granularity of task descriptions is an overlooked factor. In particular, providing a more granular task description in terms of a set of sub-tasks versus providing a coarser-grained, generic description. Objective: We aim to investigate the impact of task description granularity on the outcome of TDD, as implemented by novice developers, with respect to software quality, as measured by functional correctness and functional completeness. Method: We conducted a one-factor crossover experiment with 48 graduate students in an academic environment. Each participant applied TDD and implemented two tasks, where one of the tasks was presented using a more granular task description. Resulting artifacts were evaluated with acceptance tests to assess functional correctness and functional completeness. Linear mixed-effects models (LMM) were used for analysis. Results: Software quality improved significantly when participants applied TDD using more granular task descriptions. The effect of task description granularity is statistically significant and had a medium to large effect size. Moreover, the task was found to be a significant predictor of software quality which is an interesting result (because two tasks used in the experiment were considered to be of similar complexity). Conclusion: For novice TDD practitioners, the outcome of TDD is highly coupled with the ability to break down the task into smaller parts. For researchers, task selection and task description granularity requires more attention in the design of TDD experiments. Task description granularity should be taken into account in secondary studies. Further comparative studies are needed to investigate whether task descriptions affect other development processes similarly.
Itir Karac, Burak Turhan, Natalia Juristo Juzgado
IEEE Trans. Software Eng.2
2021 Requirements Framing Affects Design Creativity
abstract
Design creativity, the originality and practicality of a solution concept, is critical for the success of many software projects. However, little research has investigated the relationship between the way desiderata are presented and design creativity. This study therefore investigates the impact of presenting desiderata as ideas, requirements or prioritized requirements on design creativity. Two between-subjects randomized controlled experiments were conducted with 42 and 34 participants. Participants were asked to create design concepts from a list of desiderata. Participants who received desiderata framed as requirements or prioritized requirements created designs that are, on average, less original but more practical than the designs created by participants who received desiderata framed as ideas. This suggests that more formal, structured presentations of desiderata are less appropriate where more innovative solutions are desired. The results also show that design performance is highly susceptible to minor changes in the vernacular used to communicate desiderata.
Rahul Mohanani, Burak Turhan, Paul Ralph
IEEE Trans. Software Eng.2
2020 Researcher Bias in Software Engineering Experiments: a Qualitative Investigation
abstract
Researcher Bias (RB) occurs when researchers influence the results of an empirical study based on their expectations. RB might be due to the use of Questionable Research Practices (QRPs). In research fields like medicine, blinding techniques have been applied to counteract RB. We conducted an explorative qualitative survey to investigate RB in Software Engineering (SE) experiments, with respect to: (i) QRPs potentially leading to RB, (ii) causes behind RB, and (iii) possible actions to counteract RB including blinding techniques. Data collection was based on semi-structured interviews. We interviewed nine active experts in the empirical SE community. We then analyzed the transcripts of these interviews through thematic analysis. We found that some QRPs are acceptable in certain cases. Also, it appears that the presence of RB is perceived in SE and, to counteract RB, a number of solutions have been highlighted: some are intended for SE researchers and others for the boards of SE research outlets.
Simone Romano 0001, Davide Fucci, Giuseppe Scanniello, Maria Teresa Baldassarre, Burak Turhan, Natalia Juristo Juzgado
SEAA5
2020 Defect Prediction Guided Search-Based Software Testing
abstract
Today, most automated test generators, such as search-based software testing (SBST) techniques focus on achieving high code coverage. However, high code coverage is not sufficient to maximise the number of bugs found, especially when given a limited testing budget. In this paper, we propose an automated test generation technique that is also guided by the estimated degree of defectiveness of the source code. Parts of the code that are likely to be more defective receive more testing budget than the less defective parts. To measure the degree of defectiveness, we leverage Schwa, a notable defect prediction technique.
Anjana Perera, Aldeida Aleti, Marcel Böhme, Burak Turhan
ASE4
2020 On the need of preserving order of data when validating within-project defect classifiers
abstract
Abstract We are in the shoes of a practitioner who uses previous project releases’ data to predict which classes of the current release are defect-prone. In this scenario, the practitioner would like to use the most accurate classifier among the many available ones. A validation technique, hereinafter “technique”, defines how to measure the prediction accuracy of a classifier. Several previous research efforts analyzed several techniques. However, no previous study compared validation techniques in the within-project across-release class-level context or considered techniques that preserve the order of data. In this paper, we investigate which technique recommends the most accurate classifier. We use the last release of a project as the ground truth to evaluate the classifier’s accuracy and hence the ability of a technique to recommend an accurate classifier. We consider nine classifiers, two industry and 13 open projects, and three validation techniques: namely 10-fold cross-validation (i.e., the most used technique), bootstrap (i.e., the recommended technique), and walk-forward (i.e., a technique preserving the order of data). Our results show that: 1) classifiers differ in accuracy in all datasets regardless of their entity per value, 2) walk-forward outperforms both 10-fold cross-validation and bootstrap statistically in all three accuracy metrics: AUC of the selected classifier, bias and absolute bias, 3) surprisingly, all techniques resulted to be more prone to overestimate than to underestimate the performances of classifiers, and 3) the defect rate resulted in changing between the second and first half in both industry projects and 83% of open-source datasets. This study recommends the use of techniques that preserve the order of data such as walk-forward over 10-fold cross-validation and bootstrap in the within-project across-release class-level context given the above empirical results and that walk-forward is by nature more simple, inexpensive, and stable than the other two techniques.
Davide Falessi, Jacky Huang, Likhita Narayana, Jennifer Fong Thai, Burak Turhan
Empir. Softw. Eng.5
2020 Correction to: On the need of preserving order of data when validating within-project defect classifiers
abstract
To fulfill the contractual requirement of the Compact agreement, the following funding note has to be added and placed in the Funding section of the original article: Open access funding provided by Università degli Studi di Roma Tor Vergata within the CRUI-CARE Agreement.
Davide Falessi, Jacky Huang, Likhita Narayana, Jennifer Fong Thai, Burak Turhan
Empir. Softw. Eng.5
2020 Pandemic programming
abstract
Abstract Context As a novel coronavirus swept the world in early 2020, thousands of software developers began working from home. Many did so on short notice, under difficult and stressful conditions. Objective This study investigates the effects of the pandemic on developers’ wellbeing and productivity. Method A questionnaire survey was created mainly from existing, validated scales and translated into 12 languages. The data was analyzed using non-parametric inferential statistics and structural equation modeling. Results The questionnaire received 2225 usable responses from 53 countries. Factor analysis supported the validity of the scales and the structural model achieved a good fit (CFI = 0.961, RMSEA = 0.051, SRMR = 0.067). Confirmatory results include: (1) the pandemic has had a negative effect on developers’ wellbeing and productivity; (2) productivity and wellbeing are closely related; (3) disaster preparedness, fear related to the pandemic and home office ergonomics all affect wellbeing or productivity. Exploratory analysis suggests that: (1) women, parents and people with disabilities may be disproportionately affected; (2) different people need different kinds of support. Conclusions To improve employee productivity, software companies should focus on maximizing employee wellbeing and improving the ergonomics of employees’ home offices. Women, parents and disabled persons may require extra support.
Paul Ralph, Sebastian Baltes, Gianisa Adisaputri, Richard Torkar, Vladimir Kovalenko, Marcos Kalinowski, Nicole Novielli, Shin Yoo, Xavier Devroey, Xin Tan 0003, Minghui Zhou 0001, Burak Turhan, Rashina Hoda, Hideaki Hata, Gregorio Robles, Amin Milani Fard, Rana Alkadhi
Empir. Softw. Eng.12
2020 Guest Editorial: Special Issue on Predictive Models and Data Analytics in Software Engineering
Ayse Tosun Misirli, Shane McIntosh, Leandro L. Minku, Burak Turhan
Empir. Softw. Eng.4
2020 Increasing validity through replication: an illustrative TDD case
abstract
Abstract Software engineering (SE) experiments suffer from threats to validity that may impact their results. Replication allows researchers building on top of previous experiments’ weaknesses and increasing the reliability of the findings. Illustrating the benefits of replication to increase the reliability of the findings and uncover moderator variables. We replicate an experiment on test-driven development (TDD) and address some of its threats to validity and those of a previous replication. We compare the replications’ results and hypothesize on plausible moderators impacting results. Differences across TDD replications’ results might be due to the operationalization of the response variables, the allocation of subjects to treatments, the allowance to work outside the laboratory, the provision of stubs, or the task. Replications allow examining the robustness of the findings, hypothesizing on plausible moderators influencing results, and strengthening the evidence obtained.
Adrián Santos, Sira Vegas, Fernando Uyaguari, Óscar Dieste Tubío, Burak Turhan, Natalia Juristo Juzgado
Softw. Qual. J.5
2020 Cognitive Biases in Software Engineering: A Systematic Mapping Study
abstract
One source of software project challenges and failures is the systematic errors introduced by human cognitive biases. Although extensively explored in cognitive psychology, investigations concerning cognitive biases have only recently gained popularity in software engineering research. This paper therefore systematically maps, aggregates and synthesizes the literature on cognitive biases in software engineering to generate a comprehensive body of knowledge, understand state-of-the-art research and provide guidelines for future research and practise. Focusing on bias antecedents, effects and mitigation techniques, we identified 65 articles (published between 1990 and 2016), which investigate 37 cognitive biases. Despite strong and increasing interest, the results reveal a scarcity of research on mitigation techniques and poor theoretical foundations in understanding and interpreting cognitive biases. Although bias-related research has generated many new insights in the software engineering community, specific bias mitigation techniques are still needed for software professionals to overcome the deleterious effects of cognitive biases on their work.
Rahul Mohanani, Iflaah Salman, Burak Turhan, Pilar Rodríguez 0002, Paul Ralph
IEEE Trans. Software Eng.3
2020 Key Stakeholders' Value Propositions for Feature Selection in Software-Intensive Products: An Industrial Case Study
abstract
Numerous software companies are adopting value-based decision making. However, what does value mean for key stakeholders making decisions? How do different stakeholder groups understand value? Without an explicit understanding of what value means, decisions are subject to ambiguity and vagueness, which are likely to bias them. This case study provides an in-depth analysis of key stakeholders' value propositions when selecting features for a large telecommunications company's software-intensive product. Stakeholders' value propositions were elicited via interviews, which were analyzed using Grounded Theory coding techniques (open and selective coding). Thirty-six value propositions were identified and classified into six dimensions: customer value, market competitiveness, economic value/profitability, cost efficiency, technology & architecture, and company strategy. Our results show that although propositions in the customer value dimension were those mentioned the most, the concept of value for feature selection encompasses a wide range of value propositions. Moreover, stakeholder groups focused on different and complementary value dimensions, calling to the importance of involving all key stakeholders in the decision making process. Although our results are particularly relevant to companies similar to the one described herein, they aim to generate a learning process on value-based feature selection for practitioners and researchers in general.
Pilar Rodríguez 0002, Emilia Mendes, Burak Turhan
IEEE Trans. Software Eng.3
2019 Guest editorial: special section on predictive models and data analytics in software engineering
David Bowes, Emad Shihab, Burak Turhan
Empir. Softw. Eng.3
2019 A controlled experiment on time pressure and confirmation bias in functional software testing
abstract
Confirmation bias is a person’s tendency to look for evidence that strengthens his/her prior beliefs rather than refutes them. Manifestation of confirmation bias in software testing may have adverse effects on software quality. Psychology research suggests that time pressure could trigger confirmation bias. In the software industry, this phenomenon may deteriorate software quality. In this study, we investigate whether testers manifest confirmation bias and how it is affected by time pressure in functional software testing. We performed a controlled experiment with 42 graduate students to assess manifestation of confirmation bias in terms of the conformity of their designed test cases to the provided requirements specification. We employed a one factor with two treatments between-subjects experimental design. We observed, overall, participants designed significantly more confirmatory test cases as compared to disconfirmatory ones, which is in line with previous research. However, we did not observe time pressure as an antecedent to an increased rate of confirmatory testing behaviour. People tend to design confirmatory test cases regardless of time pressure. For practice, we find it necessary that testers develop self-awareness of confirmation bias and counter its potential adverse effects with a disconfirmatory attitude. We recommend further replications to investigate the effect of time pressure as a potential contributor to the manifestation of confirmation bias.
Iflaah Salman, Burak Turhan, Sira Vegas
Empir. Softw. Eng.2
2019 A Systematic Literature Review and Meta-Analysis on Cross Project Defect Prediction
abstract
Background: Cross project defect prediction (CPDP) recently gained considerable attention, yet there are no systematic efforts to analyse existing empirical evidence. Objective: To synthesise literature to understand the state-of-the-art in CPDP with respect to metrics, models, data approaches, datasets and associated performances. Further, we aim to assess the performance of CPDP versus within project DP models. Method: We conducted a systematic literature review. Results from primary studies are synthesised (thematic, meta-analysis) to answer research questions. Results: We identified 30 primary studies passing quality assessment. Performance measures, except precision, vary with the choice of metrics. Recall, precision, f-measure, and AUC are the most common measures. Models based on Nearest-Neighbour and Decision Tree tend to perform well in CPDP, whereas the popular naïve Bayes yields average performance. Performance of ensembles varies greatly across f-measure and AUC. Data approaches address CPDP challenges using row/column processing, which improve CPDP in terms of recall at the cost of precision. This is observed in multiple occasions including the meta-analysis of CPDP versus WPDP. NASA and Jureczko datasets seem to favour CPDP over WPDP more frequently. Conclusion: CPDP is still a challenge and requires more research before trustworthy applications can take place. We provide guidelines for further research.
Seyedrebvar Hosseini, Burak Turhan, Dimuthu Gunarathna
IEEE Trans. Software Eng.2
2018 The effect of noise on software engineers' performance
abstract
Background: Noise, defined as an unwanted sound, is one of the commonest factors that could affect people's performance in their daily work activities. The software engineering research community has marginally investigated the effects of noise on software engineers' performance.
Simone Romano 0001, Giuseppe Scanniello, Davide Fucci, Natalia Juristo Juzgado, Burak Turhan
ESEM5
2018 A longitudinal cohort study on the retainment of test-driven development
abstract
Background: Test-Driven Development (TDD) is an agile software development practice, which is claimed to boost both external quality of software products and developers' productivity.
Davide Fucci, Simone Romano 0001, Maria Teresa Baldassarre, Danilo Caivano, Giuseppe Scanniello, Burak Turhan, Natalia Juristo Juzgado
ESEM6
2018 An Exploratory Study of Search Based Training Data Selection for Cross Project Defect Prediction
abstract
Context: Search based approaches are gaining attention in cross project defect prediction (CPDP). The complexity of such approaches and existence of various design decisions are important issues to consider. Objective: We aim at investigating factors that can affect the performance of search based selection (SBS) approaches. We study a genetic instance selection approach (GIS) and present an evaluation of design options for search based CPDP. Method: Using an exploratory approach, data from different options of models are gathered and analyzed through ANOVA tests and effect sizes. Results: Both feature sets and validation dataset selection options show small or insignificant impacts on F-measure and precision, unlike the more affected false positive and true negative rates. Size of training data does not seem to be related to significant changes in F-measure and precision and high variability in performance are discouraging evidence for using larger datasets. Fitness function is one of the major factors that impact performance with much larger effect than the choice of validation dataset. Finally, while showing slight impacts, data label changes do not seem to be the top contributor to performance. Conclusions: We conclude that exploratory approaches can be effective for making design decisions in constructing search based CPDP models. Effect of individual tuned learners and their interaction with other affecting parameters and more in depth study of quality affecting factors guided by label changes are directions to investigate.
Seyedrebvar Hosseini, Burak Turhan
SEAA2
2018 Empirical evaluation of the effects of experience on code quality and programmer productivity: an exploratory study
abstract
This extended abstract summarizes an article, which has been published in the Empirical Software Engineering Journal and was selected for the Journal-First presentations at the International Conference on Software and System Process (ICSSP 2018).
Óscar Dieste Tubío, Alejandrina Aranda, Fernando Uyaguari, Burak Turhan, Ayse Tosun Misirli, Davide Fucci, Markku Oivo, Natalia Juristo Juzgado
ICSSP4
2018 Effect of time-pressure on perceived and actual performance in functional software testing
abstract
Background: Time-pressure is an inevitable reality of software industry that influences the performance of software engineers. It may result in adverse effects on software quality or distort the perception of performance on executed tasks to differ from actual performance. Objective: We aim to investigate the effect of time-pressure on perceived and actual performance of software testers in the context of functional software testing. Method: We performed two controlled experiments with 87 graduate students in two academic terms. We assessed actual performance in terms of coverage (i.e. percentage of test cases correctly identified) and perceived performance using NASA-TLX. We have an independent factorial design for our experimental study. Results: The results reveal a significant effect of time-pressure on actual performance. However, we could not observe a significant effect of time-pressure on the perceived performance of the participants for the task undertaken. We also observed a significant negative correlation between actual and perceived performance when controlled for time-pressure and experimental session factors. Conclusion: Time-pressure affects the actual performance in a testing task but the perception of accomplishment by the testers is sustained irrespective of time-pressure, indicating an over-estimation issue. Perception of performance should be adjusted to align with reality to account for the effect of time pressure. This will lead to better self estimates of performance.
Iflaah Salman, Burak Turhan
ICSSP2
2018 On the effectiveness of unit tests in test-driven development
abstract
Background: Writing unit tests is one of the primary activities in test-driven development. Yet, the existing reviews report few evidence supporting or refuting the effect of this development approach on test case quality. Lack of ability and skills of developers to produce sufficiently good test cases are also reported as limitations of applying test-driven development in industrial practice. Objective: We investigate the impact of test-driven development on the effectiveness of unit test cases compared to an incremental test last development in an industrial context. Method: We conducted an experiment in an industrial setting with 24 professionals. Professionals followed the two development approaches to implement the tasks. We measure unit test effectiveness in terms of mutation score. We also measure branch and method coverage of test suites to compare our results with the literature. Results: In terms of mutation score, we have found that the test cases written for a test-driven development task have a higher defect detection ability than test cases written for an incremental test-last development task. Subjects wrote test cases that cover more branches on a test-driven development task compared to the other task. However, test cases written for an incremental test-last development task cover more methods than those written for the second task. Conclusion: Our findings are different from previous studies conducted at academic settings. Professionals were able to perform more effective unit testing with test-driven development. Furthermore, we observe that the coverage measure preferred in academic studies reveal different aspects of a development approach. Our results need to be validated in larger industrial contexts.
Ayse Tosun Misirli, Muzamil Ahmed, Burak Turhan, Natalia Juristo Juzgado
ICSSP3
2018 Empirical software engineering experts on the use of students and professionals in experiments
Davide Falessi, Natalia Juristo Juzgado, Claes Wohlin, Burak Turhan, Jürgen Münch, Andreas Jedlitschka, Markku Oivo
Empir. Softw. Eng.4
2018 Four commentaries on the use of students and professionals in empirical software engineering experiments
Robert Feldt, Thomas Zimmermann 0001, Gunnar R. Bergersen, Davide Falessi, Andreas Jedlitschka, Natalia Juristo Juzgado, Jürgen Münch, Markku Oivo, Per Runeson, Martin J. Shepperd, Dag I. K. Sjøberg, Burak Turhan
Empir. Softw. Eng.12
2018 A benchmark study on the effectiveness of search-based data selection and feature selection for cross project defect prediction
Seyedrebvar Hosseini, Burak Turhan, Mika Mäntylä
Inf. Softw. Technol.2
2018 Guest editorial: special issue on predictive models for software quality
Leandro L. Minku, Ayse Basar Bener, Burak Turhan
Softw. Qual. J.3
2017 Perceptions of Creativity in Software Engineering Research and Practice
abstract
Software engineering, especially design and requirements engineering, is intensely creative. However, practitioners and researchers appear to perceive creativity differently, hindering knowledge transfer. To explore and understand these perceptual differences, this paper combines a systematic mapping study of SE research literature with an interview study of practitioners. The subsequent analysis of 84 primary studies and 17 semi-structured interviews reveal some agreement (e.g. creativity is a process that produces novel and useful ideas). However, it also reveals important differences in the way creativity is conceptualized, measured and improved. These differences undermine evidence-based techniques to enhance and measure creativity in SE research and practice.
Rahul Mohanani, Prabhat Ram, Ahmed Lasisi, Paul Ralph, Burak Turhan
SEAA5
2017 Guest editorial: special issue on realising artificial intelligence synergies in software engineering
Rachel Harrison, Ayse Basar Bener, Çetin Meriçli, Burak Turhan
Autom. Softw. Eng.4
2017 Empirical evaluation of the effects of experience on code quality and programmer productivity: an exploratory study
Óscar Dieste Tubío, Alejandrina Aranda, Fernando Uyaguari, Burak Turhan, Ayse Tosun Misirli, Davide Fucci, Markku Oivo, Natalia Juristo Juzgado
Empir. Softw. Eng.4
2017 An industry experiment on the effects of test-driven development on external quality and productivity
Ayse Tosun Misirli, Óscar Dieste Tubío, Davide Fucci, Sira Vegas, Burak Turhan, Hakan Erdogmus, Adrián Santos, Markku Oivo, Kimmo Toro, Janne Järvinen, Natalia Juristo Juzgado
Empir. Softw. Eng.5
2017 Findings from a multi-method study on test-driven development
Simone Romano 0001, Davide Fucci, Giuseppe Scanniello, Burak Turhan, Natalia Juristo Juzgado
Inf. Softw. Technol.4
2017 Special section on realizing artificial intelligence synergies in software engineering
Çetin Meriçli, Burak Turhan
Softw. Qual. J.2
2017 A Dissection of the Test-Driven Development Process: Does It Really Matter to Test-First or to Test-Last?
abstract
Background: Test-driven development (TDD) is a technique that repeats short coding cycles interleaved with testing. The developer first writes a unit test for the desired functionality, followed by the necessary production code, and refactors the code. Many empirical studies neglect unique process characteristics related to TDD iterative nature. Aim: We formulate four process characteristic: sequencing, granularity, uniformity, and refactoring effort. We investigate how these characteristics impact quality and productivity in TDD and related variations. Method: We analyzed 82 data points collected from 39 professionals, each capturing the process used while performing a specific development task. We built regression models to assess the impact of process characteristics on quality and productivity. Quality was measured by functional correctness. Result: Quality and productivity improvements were primarily positively associated with the granularity and uniformity. Sequencing, the order in which test and production code are written, had no important influence. Refactoring effort was negatively associated with both outcomes. We explain the unexpected negative correlation with quality by possible prevalence of mixed refactoring. Conclusion: The claimed benefits of TDD may not be due to its distinctive test-first dynamic, but rather due to the fact that TDD-like processes encourage fine-grained, steady steps that improve focus and flow.
Davide Fucci, Hakan Erdogmus, Burak Turhan, Markku Oivo, Natalia Juristo Juzgado
IEEE Trans. Software Eng.3
2016 An External Replication on the Effects of Test-driven Development Using a Multi-site Blind Analysis Approach
abstract
Context: Test-driven development (TDD) is an agile practice claimed to improve the quality of a software product, as well as the productivity of its developers. A previous study (i.e., baseline experiment) at the University of Oulu (Finland) compared TDD to a test-last development (TLD) approach through a randomized controlled trial. The results failed to support the claims. Goal: We want to validate the original study results by replicating it at the University of Basilicata (Italy), using a different design. Method: We replicated the baseline experiment, using a crossover design, with 21 graduate students. We kept the settings and context as close as possible to the baseline experiment. In order to limit researchers bias, we involved two other sites (UPM, Spain, and Brunel, UK) to conduct blind analysis of the data. Results: The Kruskal-Wallis tests did not show any significant difference between TDD and TLD in terms of testing effort (p-value = .27), external code quality (p-value = .82), and developers' productivity (p-value = .83). Nevertheless, our data revealed a difference based on the order in which TDD and TLD were applied, though no carry over effect. Conclusions: We verify the baseline study results, yet our results raises concerns regarding the selection of experimental objects, particularly with respect to their interaction with the order in which of treatments are applied.
Davide Fucci, Giuseppe Scanniello, Simone Romano 0001, Martin J. Shepperd, Boyce Sigweni, Fernando Uyaguari, Burak Turhan, Natalia Juristo Juzgado, Markku Oivo
ESEM7
2016 Providing Tool-Support for Value-Based Decision-Making: A Usability Assessment
abstract
Numerous companies worldwide make their decisions related to software projects/products in a value neutral way, using only earned value systems, which represent short-term goals. Better decisions can be made using a value-based approach, achieving cost-effective results and reliable construction and maintenance of products. However, moving from a value-neutral to a value-based paradigm can be a challenge. We provide tool-support, which was co-created in collaboration with three software companies, to ease the paradigm shift. Our tool supports both individual and group-based decisions using several visualization mechanisms. Despite the co-creation process employed while developing the VALUE tool, there are specific issues relating to its usability that must also be assessed in order to reduce any possible drawbacks for its adoption by industry. This paper details three usability studies that were carried out to assess the VALUE tool's usability. The results also suggest that the tool is ready to be taken into use in the industry.
Vitor Freitas, Emilia Mendes, Burak Turhan
SEAA3
2015 On the effects of programming and testing skills on external quality and productivity in a test-driven development context
abstract
Background: In previous studies, a model was proposed that investigated how the developers' unit testing effort impacted their productivity as well as the external quality of the software they developed.
Davide Fucci, Burak Turhan, Markku Oivo
EASE2
2015 Merits of Organizational Metrics in Defect Prediction: An Industrial Replication
abstract
Defect prediction models presented in the literature lack generalization unless the original study can be replicated using new datasets and in different organizational settings. Practitioners can also benefit from replicating studies in their own environment by gaining insights and comparing their findings with those reported. In this work, we replicated an earlier study in order to investigate the merits of organizational metrics in building defect prediction models for large-scale enterprise software. We mined the organizational, code complexity, code churn and pre-release bug metrics of that large scale software and built defect prediction models for each metric set. In the original study, organizational metrics were found to achieve the highest performance. In our case, models based on organizational metrics performed better than models based on churn metrics but were outperformed by pre-release metric models. Further, we verified four individual organizational metrics as indicators for defects. We conclude that the performance of different metric sets in building defect prediction models depends on the project's characteristics and the targeted prediction level. Our replication of earlier research enabled assessing the validity and limitations of organizational metrics in a different context.
Bora Caglayan, Burak Turhan, Ayse Basar Bener, Mayy Habayeb, Andriy V. Miranskyy, Enzo Cialini
ICSE (2)2
2015 4th International Workshop on Realizing AI Synergies in Software Engineering (RAISE 2015)
abstract
This workshop is the fourth in the series and continued to build upon the work carried out at the previous iterations of the International Workshop on Realizing Artificial Intelligence Synergies in Software Engineering, which were held at ICSE in 2012, 2013 and 2014. RAISE 2015 brought together researchers and practitioners from the artificial intelligence (AI) and software engineering (SE) disciplines to build on the interdis- ciplinary synergies that exist and to stimulate further interaction across these disciplines. Mutually beneficial characteristics have appeared in the past few decades and are still evolving due to new challenges and technological advances. Hence, the question that motivates and drives the RAISE Workshop series is: "Are SE and AI researchers ignoring important insights from AI and SE?". To pursue this question, RAISE'15 explored not only the application of AI techniques to SE problems but also the application of SE techniques to AI problems. RAISE not only strengthens the AI- and-SE community but also continues to develop a roadmap of strategic research directions for AI and SE.
Burak Turhan, Ayse Basar Bener, Rachel Harrison, Andriy V. Miranskyy, Çetin Meriçli, Leandro L. Minku
ICSE (2)1
2015 Towards an operationalization of test-driven development skills: An industrial empirical study
Davide Fucci, Burak Turhan, Natalia Juristo Juzgado, Óscar Dieste Tubío, Ayse Tosun Misirli, Markku Oivo
Inf. Softw. Technol.2
2014 Conformance factor in test-driven development: initial results from an enhanced replication
abstract
Test-driven development (TDD) is an iterative software development technique where unit-tests are defined before production code. The proponents of TDD claim that it improves both external quality and developers' productivity. In particular, Erdogmus et al. (i.e., original study) proposed a two-stage model to investigate these claims regarding TDD's effects. Our aim is to enhance the model proposed in the original study by investigating an additional factor: TDD process conformance. We conducted a close, external replication of the original study accompanied by a correlation analysis to check whether process conformance is related to improvements for the subjects using TDD. We partially confirmed the results of the original study. Moreover, we observed a correlation between process conformance and quality, but not productivity. We found no evidence to support the claim that external quality and productivity are improved by the adoption of TDD compared to test-last development. Finally, conformance to TDD process improves the quality and does not affect productivity. We conclude that the role of process conformance is relevant in studying the quality and productivity-related effects of TDD.
Davide Fucci, Burak Turhan, Markku Oivo
EASE2
2014 Impact of process conformance on the effects of test-driven development
abstract
Context: One limitation of the empirical studies about test-driven development (TDD) is knowing whether the developers followed the advocated test-code-refactor cycle. Research dealt with the issue of process conformance only in terms of internal validity, while investigating the role of other confounding variables that might explain the controversial effects of TDD. None of the research included process conformance as a fundamental part of the analysis.
Davide Fucci, Burak Turhan, Markku Oivo
ESEM2
2014 On the role of tests in test-driven development: a differentiated and partial replication
Davide Fucci, Burak Turhan
Empir. Softw. Eng.2
2013 A Replicated Experiment on the Effectiveness of Test-First Development
abstract
Background: Test-first development (TF) is regarded as a development practice that can lead to better quality of software products, as well as improved developer productivity. By implementing unit tests before the corresponding production code, the tests themselves are the main driver to such improvements. The role of tests on the effectiveness of TF has been studied in a controlled experiment by Erdogmus et al. (i.e. original study). Aim: Our goal is to examine the impact of test-first (TF) development on product quality and developer productivity, specifically the role that tests play in it. Method: We replicated the original study's controlled experiment by comparing an experimental group applying TF to a control group applying a test-last approach. We then carried out a correlation study in order to understand whether the number of tests is a good predictor for external quality and/or productivity. Results: Mann-Whitney tests did not show any significant difference between the two groups in terms of number of tests written (W=114.5, p=0.38), developers' productivity (W=90, p=0.82) and external quality (W=81.55, p=0.53). In addition, while a significant correlation exists between the number of tests and productivity (Spearman's ρ = 0.57, p<;0.001), none was found in the case of external quality (Spearman's ρ = 0.17, p=0.18). Conclusions: We conclude that TF neither improves nor deteriorates the external quality or the productivity when compared to the test-last approach, leaving room for other variables to impact the effects of TF. This replication has partially confirmed the findings of the original study.
Davide Fucci, Burak Turhan
ESEM2
2013 Constructing Defect Predictors and Communicating the Outcomes to Practitioners
abstract
Background: An alternative to expert-based decisions is to take data-driven decisions and software analytics is the key enabler for this evidence-based management approach. Defect prediction is one popular application area of software analytics, however with serious challenges to deploy into practice. Goal: We aim at developing and deploying a defect prediction model for guiding practitioners to focus their activities on the most problematic parts of the software and improve the efficiency of the testing process. Method: We present a pilot study, where we developed a defect prediction model and different modes of information representation of the data and the model outcomes, namely: commit hotness ranking, error probability mapping to the source and visualization of interactions among teams through errors. We also share the challenges and lessons learned in the process. Result: In terms of standard performance measures, the constructed defect prediction model performs similar to those reported in earlier studies, e.g. 80% of errors can be detected by inspecting 30% of the source. However, the feedback from practitioners indicates that such performance figures are not useful to have an impact in their daily work. Pointing out most problematic source files, even isolating error-prone sections within files are regarded as stating the obvious by the practitioners, though the latter is found to be helpful for activities such as refactoring. On the other hand, visualizing the interactions among teams, based on the errors introduced and fixed, turns out to be the most helpful representation as it helps pinpointing communication related issues within and across teams. Conclusion: The constructed predictor can give accurate information about the most error prone parts. Creating practical representations from this data is possible, but takes effort. The error prediction research done in Elektrobit Wireless Ltd is concluded to be useful and we will further improve the presentations made from the error prediction data.
Taneli Taipale, Mika Qvist, Burak Turhan
ESEM3
2013 Message from the PROMISE 2013 Chairs
abstract
PROMISE conference is an annual forum for researchers and practitioners to present, discuss and exchange ideas, results, expertise and experiences in construction and/or application of prediction models in software engineering. Such models could be targeted at: planning, design, implementation, testing, maintenance, quality assurance, evaluation, process improvement, management, decision making, and risk assessment in software and systems development. PROMISE is distinguished from similar forums with its public data repository and focus on methodological details, providing a unique interdisciplinary venue for software engineering and machine learning communities, and seeking for verifiable and repeatable prediction models that are useful in practice.
Burak Turhan, Stefan Wagner 0001, Ayse Basar Bener, Massimiliano Di Penta
ESEM1
2013 Data science for software engineering
abstract
Target audience: Software practitioners and researchers wanting to understand the state of the art in using data science for software engineering (SE). Content: In the age of big data, data science (the knowledge of deriving meaningful outcomes from data) is an essential skill that should be equipped by software engineers. It can be used to predict useful information on new projects based on completed projects. This tutorial offers core insights about the state-of-the-art in this important field. What participants will learn: Before data science: this tutorial discusses the tasks needed to deploy machine-learning algorithms to organizations (Part 1: Organization Issues). During data science: from discretization to clustering to dichotomization and statistical analysis. And the rest: When local data is scarce, we show how to adapt data from other organizations to local problems. When privacy concerns block access, we show how to privatize data while still being able to mine it. When working with data of dubious quality, we show how to prune spurious information. When data or models seem too complex, we show how to simplify data mining results. When data is too scarce to support intricate models, we show methods for generating predictions. When the world changes, and old models need to be updated, we show how to handle those updates. When the effect is too complex for one model, we show how to reason across ensembles of models. Pre-requisites: This tutorial makes minimal use of maths of advanced algorithms and would be understandable by developers and technical managers.
Tim Menzies, Ekrem Kocaguneli, Fayola Peters, Burak Turhan, Leandro L. Minku
ICSE4
2013 Empirical evaluation of the effects of mixed project data on learning defect predictors
Burak Turhan, Ayse Tosun Misirli, Ayse Basar Bener
Inf. Softw. Technol.1
2013 Local versus Global Lessons for Defect Prediction and Effort Estimation
abstract
Existing research is unclear on how to generate lessons learned for defect prediction and effort estimation. Should we seek lessons that are global to multiple projects or just local to particular projects? This paper aims to comparatively evaluate local versus global lessons learned for effort estimation and defect prediction. We applied automated clustering tools to effort and defect datasets from the PROMISE repository. Rule learners generated lessons learned from all the data, from local projects, or just from each cluster. The results indicate that the lessons learned after combining small parts of different data sources (i.e., the clusters) were superior to either generalizations formed over all the data or local lessons formed from particular projects. We conclude that when researchers attempt to draw lessons from some historical data source, they should 1) ignore any existing local divisions into multiple sources, 2) cluster across all available data, then 3) restrict the learning of lessons to the clusters from other sources that are nearest to the test data.
Tim Menzies, Andrew Butcher, David R. Cok, Andrian Marcus, Lucas Layman, Forrest Shull, Burak Turhan, Thomas Zimmermann 0001
IEEE Trans. Software Eng.7
2012 Dione: an integrated measurement and defect prediction solution
abstract
We present an integrated measurement and defect prediction tool: Dione. Our tool enables organizations to measure, monitor, and control product quality through learning based defect prediction. Similar existing tools either provide data collection and analytics, or work just as a prediction engine. Therefore, companies need to deal with multiple tools with incompatible interfaces in order to deploy a complete measurement and prediction solution. Dione provides a fully integrated solution where data extraction, defect prediction and reporting steps fit seamlessly. In this paper, we present the major functionality and architectural elements of Dione followed by an overview of our demonstration.
Bora Caglayan, Ayse Tosun Misirli, Gül Çalikli, Ayse Basar Bener, Turgay Aytac, Burak Turhan
SIGSOFT FSE6
2012 On the dataset shift problem in software engineering prediction models
Burak Turhan
Empir. Softw. Eng.1
2012 Learning Better Inspection Optimization Policies
abstract
Recent research has shown the value of social metrics for defect prediction. Yet many repositories lack the information required for a social analysis. So, what other means exist to infer how developers interact around their code? One option is static code metrics that have already demonstrated their usefulness in analyzing change in evolving software systems. But do they also help in defect prediction? To address this question we selected a set of static code metrics to determine what classes are most "active" (i.e., the classes where the developers spend much time interacting with each other's design and implementation decisions) in 33 open-source Java systems that lack details about individual developers. In particular, we assessed the merit of these activity-centric measures in the context of "inspection optimization" — a technique that allows for reading the fewest lines of code in order to find the most defects. For the task of inspection optimization these activity measures perform as well as (usually, within 4%) a theoretical upper bound on the performance of any set of measures. As a result, we argue that activity-centric static code metrics are an excellent predictor for defects.
Markus Lumpe, Rajesh Vasa, Tim Menzies, Rebecca Rush, Burak Turhan
Int. J. Softw. Eng. Knowl. Eng.5
2011 A comparative study for estimating software development effort intervals
Ayse Bakir, Burak Turhan, Ayse Basar Bener
Softw. Qual. J.2
2011 An industrial case study of classifier ensembles for locating software defects
Ayse Tosun Misirli, Ayse Basar Bener, Burak Turhan
Softw. Qual. J.3
2010 Regularities in Learning Defect Predictors
Burak Turhan, Ayse Basar Bener, Tim Menzies
PROFES1
2010 Agile Adoption Strategies in the Context of Agile in the Large: FLEXI Agile Adoption Industrial Inventory
Anna Rohunen, Pilar Rodríguez 0002, Pasi Kuvaja, Lech Krzanik, Jouni Markkula, Burak Turhan
XP6
2010 A Quantitative Comparison of Test-First and Test-Last Code in an Industrial Project
Burak Turhan, Ayse Basar Bener, Pasi Kuvaja, Markku Oivo
XP1
2010 Defect prediction from static code features: current results, limitations, new approaches
Tim Menzies, Zach Milton, Burak Turhan, Bojan Cukic, Yue Jiang 0001, Ayse Basar Bener
Autom. Softw. Eng.3
2010 Practical considerations in deploying statistical methods for defect prediction: A case study within the Turkish telecommunications industry
Ayse Tosun Misirli, Ayse Basar Bener, Burak Turhan, Tim Menzies
Inf. Softw. Technol.3
2010 A new perspective on data homogeneity in software cost estimation: a study in the embedded systems domain
Ayse Bakir, Burak Turhan, Ayse Basar Bener
Softw. Qual. J.2
2009 Prest: An Intelligent Software Metrics Extraction, Analysis and Defect Prediction Tool
Ekrem Kocaguneli, Ayse Tosun Misirli, Ayse Basar Bener, Burak Turhan, Bora Caglayan
SEKE4
2009 Analysis of Naive Bayes' assumptions on software fault data: An empirical study
Burak Turhan, Ayse Basar Bener
Data Knowl. Eng.1
2009 On the relative value of cross-company and within-company data for defect prediction
Burak Turhan, Tim Menzies, Ayse Basar Bener, Justin S. Di Stefano
Empir. Softw. Eng.1
2009 An expert system for determining candidate software classes for refactoring
Yasemin Kösker, Burak Turhan, Ayse Basar Bener
Expert Syst. Appl.2
2009 Feature weighting heuristics for analogy-based effort estimation models
Ayse Tosun Misirli, Burak Turhan, Ayse Basar Bener
Expert Syst. Appl.2
2009 Data mining source code for locating software bugs: A case study in telecommunication industry
Burak Turhan, Gözde Koçak, Ayse Basar Bener
Expert Syst. Appl.1
2009 Ensemble of neural networks with associative memory (ENNA) for estimating software development costs
Yigit Kultur, Burak Turhan, Ayse Basar Bener
Knowl. Based Syst.2
2008 Ensemble of software defect predictors: a case study
abstract
In this paper, we present a defect prediction model based on ensemble of classifiers, which has not been fully explored so far in this type of research. We have conducted several experiments on public datasets. Our results reveal that ensemble of classifiers considerably improve the defect detection capability compared to Naive Bayes algorithm. We also conduct a cost-benefit analysis for our ensemble, where it turns out that it is enough to inspect 32% of the code on the average, for detecting 76% of the defects.
Ayse Tosun Misirli, Burak Turhan, Ayse Basar Bener
ESEM2
2008 Weighted Static Code Attributes for Software Defect Prediction
Burak Turhan, Ayse Basar Bener
SEKE1
2008 ENNA: software effort estimation using ensemble of neural networks with associative memory
abstract
Companies usually have limited amount of data for effort estimation. Machine learning methods have been preferred over parametric models due to their flexibility to calibrate the model for the available data. On the other hand, as machine learning methods become more complex they need more data to learn from. Therefore the challenge is to increase the performance of the algorithm when there is limited data. In this research we used a relatively complex machine learning algorithm, neural networks, and showed that stable and accurate estimations are achievable with an ensemble using associative memory. Our experimental results revealed that our proposed algorithm (ENNA) achieves on the average PRED(25) = 36.4 which is a significant increase compared to Neural Network (NN) PRED(25) = 8.
Yigit Kultur, Burak Turhan, Ayse Basar Bener
SIGSOFT FSE2
2007 Evaluation of Feature Extraction Methods on Software Cost Estimation
abstract
This research investigates the effects of linear and non-linear feature extraction methods on the cost estimation performance. We use principal component analysis (PCA) and Isomap for extracting new features from observed ones and evaluate these methods with support vector regression (SVR) on publicly available datasets. Our results for these datasets indicate there is no significant difference between the performances of these linear and non-linear feature extraction methods.
Burak Turhan, F. Onur Kutlubay, Ayse Basar Bener
ESEM1
2007 A Template for Real World Team Projects for Highly Populated Software Engineering Classes
abstract
Assigning projects of group work in the context of software engineering courses has become a commonly used practice in several educational institutions. Previously reported results examined different aspects of this approach. The problem is that most studies are based on relatively small group sizes. In this article a large scale project template for a class-wide project that is currently in use in the Department of Computer Engineering, Bogazici University, will be presented.
Burak Turhan, Ayse Basar Bener
ICSE1