EDBT 2026 Demo / reviewers in the wild / expert
Richard Torkar
dblp:52/972
· DBLP profile ↗
59ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0002-0118-8143ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 56 · 3 first-author · 12 since 2021Artificial intelligence and machine learning · 6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mitigating omitted variable bias in empirical software engineeringabstractOmitted variable bias occurs when a statistical model leaves out variables that are relevant determinants of the studied effects. This results in the model attributing the missing variables’ effect to some of the included variables—hence over- or under-estimating the latter’s true effect. Omitted variable bias presents a significant threat to the validity of empirical research, particularly in non-experimental studies such as those common in empirical software engineering. This paper illustrates the impact of omitted variable bias on two illustrative examples in the software engineering domain, and uses them to present methods to investigate the possible presence of omitted variable bias, to estimate its impact, and to mitigate its drawbacks. The analysis techniques we present are based on causal structural models of the variables of interest, which provide a practical, intuitive summary of the key relations among variables. This paper demonstrates a sequence of analysis steps that inform the design and execution of similar empirical studies in software engineering. An important observation is that it pays off to invest effort investigating omitted variable bias before actually executing an empirical study, because this effort can lead to a more solid study design, and to a reduction in its threats to validity. Carlo A. Furia, Richard Torkar |
Empir. Softw. Eng. | 2 |
| 2025 | Applying bayesian data analysis for causal inference about requirements quality: a controlled experimentabstractAbstract It is commonly accepted that the quality of requirements specifications impacts subsequent software engineering activities. However, we still lack empirical evidence to support organizations in deciding whether their requirements are good enough or impede subsequent activities. We aim to contribute empirical evidence to the effect that requirements quality defects have on a software engineering activity that depends on this requirement. We conduct a controlled experiment in which 25 participants from industry and university generate domain models from four natural language requirements containing different quality defects. We evaluate the resulting models using both frequentist and Bayesian data analysis. Contrary to our expectations, our results show that the use of passive voice only has a minor impact on the resulting domain models. The use of ambiguous pronouns, however, shows a strong effect on various properties of the resulting domain models. Most notably, ambiguous pronouns lead to incorrect associations in domain models. Despite being equally advised against by literature and frequentist methods, the Bayesian data analysis shows that the two investigated quality defects have vastly different impacts on software engineering activities and, hence, deserve different levels of attention. Our employed method can be further utilized by researchers to improve reliable, detailed empirical evidence on requirements quality. Julian Frattini, Davide Fucci, Richard Torkar, Lloyd Montgomery, Michael Unterkalmsteiner, Jannik Fischbach, Daniel Méndez 0001 |
Empir. Softw. Eng. | 3 |
| 2025 | Governing the commons: code ownership and code-clones in large-scale software developmentabstractAbstract Context In software development organizations employing weak or collective ownership, different teams are allowed and expected to autonomously perform changes in various components. This creates diversity both in the knowledge of, and in the responsibility for, individual components. Objective Our objective is to understand how and why different teams introduce technical debt in the form of code clones as they change different components. Method We collected data about change size and clone introductions made by ten teams in eight components which was part of a large industrial software system. We then designed a Multi-Level Generalized Linear Model (MLGLM), to illustrate the teams’ differing behavior. Finally, we discussed the results with three development teams, plus line manager and the architect team, evaluating whether the model inferences aligned with what they expected. Responses were recorded and thematically coded. Results The results show that teams do behave differently in different components, and the feedback from the teams indicates that this method of illustrating team behavior can be useful as a complement to traditional summary statistics of ownership. Conclusions We find that our model-based approach produces useful visualizations of team introductions of code clones as they change different components. Practitioners stated that the visualizations gave them insights that were useful, and by comparing with an average team, inter-team comparisons can be avoided. Thus, this has the potential to be a useful feedback tool for teams in software development organizations that employ weak or collective ownership. Anders Sundelin, Javier Gonzalez-Huerta, Richard Torkar, Krzysztof Wnuk |
Empir. Softw. Eng. | 3 |
| 2024 | The broken windows theory applies to technical debtabstractAbstract Context: The term technical debt (TD) describes the aggregation of sub-optimal solutions that serve to impede the evolution and maintenance of a system. Some claim that the broken windows theory (BWT), a concept borrowed from criminology, also applies to software development projects. The theory states that the presence of indications of previous crime (such as a broken window) will increase the likelihood of further criminal activity; TD could be considered the broken windows of software systems. Objective: To empirically investigate the causal relationship between the TD density of a system and the propensity of developers to introduce new TD during the extension of that system. Method: The study used a mixed-methods research strategy consisting of a controlled experiment with an accompanying survey and follow-up interviews. The experiment had a total of 29 developers of varying experience levels completing system extension tasks in already existing systems with high or low TD density. Results: The analysis revealed significant effects of TD level on the subjects’ tendency to re-implement (rather than reuse) functionality, choose non-descriptive variable names, and introduce other code smells identified by the software tool , all with at least $$95\%$$ 95 % credible intervals. Coclusions: Three separate significant results along with a validating qualitative result combine to form substantial evidence of the BWT’s existence in software engineering contexts. This study finds that existing TD can have a major impact on developers propensity to introduce new TD of various types during development. William Levén, Hampus Broman, Terese Besker, Richard Torkar |
Empir. Softw. Eng. | 4 |
| 2024 | Not all requirements prioritization criteria are equal at all times: A quantitative analysisabstractRequirement prioritization is recognized as an important decision-making activity in requirements engineering . Requirement prioritization is applied to determine which requirements should be implemented and released. In order to prioritize requirements, there are several approaches/techniques/tools that use different requirements prioritization criteria, which are often identified by gut feeling instead of an in-depth analysis of which criteria are most important to use. Therefore, in this study we investigate which requirements prioritization criteria are most important to use in industry when determining which requirements are implemented and released, and if the importance of the criteria change depending on how far a requirement has reached in the development process. We conducted a quantitative study where quantitative data was collected through a case study of one completed project from one software developing company by extracting 32,139 requirements prioritization decisions based on eight requirements prioritization criteria for 11,110 requirements. The results show that not all requirements prioritization criteria are equally important, and this change depending on how far a requirement has reached in the development process. For example, for requirements prioritization decisions before iteration/sprint planning, having high Business value had an impact on the decisions, but after iteration/sprint planning, having high Business value had no impact. Editor’s note: Open Science material was validated by the Journal of Systems and Software Open Science Board. Richard Berntsson-Svensson, Richard Torkar |
J. Syst. Softw. | 2 |
| 2024 | Towards Causal Analysis of Empirical Software Engineering Data: The Impact of Programming Languages on Coding CompetitionsabstractThere is abundant observational data in the software engineering domain, whereas running large-scale controlled experiments is often practically impossible. Thus, most empirical studies can only report statistical correlations —instead of potentially more insightful and robust causal relations. To support analyzing purely observational data for causal relations and to assess any differences between purely predictive and causal models of the same data, this article discusses some novel techniques based on structural causal models (such as directed acyclic graphs of causal Bayesian networks). Using these techniques, one can rigorously express, and partially validate, causal hypotheses and then use the causal information to guide the construction of a statistical model that captures genuine causal relations—such that correlation does imply causation. We apply these ideas to analyzing public data about programmer performance in Code Jam, a large world-wide coding contest organized by Google every year. Specifically, we look at the impact of different programming languages on a participant’s performance in the contest. While the overall effect associated with programming languages is weak compared to other variables—regardless of whether we consider correlational or causal links—we found considerable differences between a purely associational and a causal analysis of the very same data. The takeaway message is that even an imperfect causal analysis of observational data can help answer the salient research questions more precisely and more robustly than with just purely predictive techniques—where genuine causal effects may be confounded. Carlo A. Furia, Richard Torkar, Robert Feldt |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2022 | Take a deep breath: Benefits of neuroplasticity practices for software developers and computer workers in a family of experimentsabstractAbstract Context Computer workers in general, and software developers specifically, are under a high amount of stress due to continuous deadlines and, often, over-commitment. Objective This study investigates the effects of a neuroplasticity practice, a specific breathing practice, on the attention awareness, well-being, perceived productivity, and self-efficacy of computer workers. Method The intervention was a 12-week program with a weekly live session that included a talk on a well-being topic and a facilitated group breathing session. During the intervention period, we solicited one daily journal note and one weekly well-being rating. We created a questionnaire mainly from existing, validated scales as entry and exit survey for data points for comparison before and after the intervention. We replicated the intervention in a similarly structured 8-week program. The data was analyzed using Bayesian multi-level models for the quantitative part and thematic analysis for the qualitative part. Results The intervention showed improvements in participants’ experienced inner states despite an ongoing pandemic and intense outer circumstances for most. Over the course of the study, we found an improvement in the participants’ ratings of how often they found themselves in good spirits as well as in a calm and relaxed state. We also aggregate a large number of deep inner reflections and growth processes that may not have surfaced for the participants without deliberate engagement in such a program. Conclusion The data indicates usefulness and effectiveness of an intervention for computer workers in terms of increasing well-being and resilience. Everyone needs a way to deliberately relax, unplug, and recover. A breathing practice is a simple way to do so, and the results call for establishing a larger body of work to make this common practice. Birgit Penzenstadler, Richard Torkar, Cristina Martinez Montes |
Empir. Softw. Eng. | 2 |
| 2022 | Applying Bayesian Analysis Guidelines to Empirical Software Engineering Data: The Case of Programming Languages and Code QualityabstractStatistical analysis is the tool of choice to turn data into information and then information into empirical knowledge. However, the process that goes from data to knowledge is long, uncertain, and riddled with pitfalls. To be valid, it should be supported by detailed, rigorous guidelines that help ferret out issues with the data or model and lead to qualified results that strike a reasonable balance between generality and practical relevance. Such guidelines are being developed by statisticians to support the latest techniques for Bayesian data analysis. In this article, we frame these guidelines in a way that is apt to empirical research in software engineering. To demonstrate the guidelines in practice, we apply them to reanalyze a GitHub dataset about code quality in different programming languages. The dataset’s original analysis [Ray et al. 55 ] and a critical reanalysis [Berger et al. 6 ] have attracted considerable attention—in no small part because they target a topic (the impact of different programming languages) on which strong opinions abound. The goals of our reanalysis are largely orthogonal to this previous work, as we are concerned with demonstrating, on data in an interesting domain, how to build a principled Bayesian data analysis and to showcase its benefits. In the process, we will also shed light on some critical aspects of the analyzed data and of the relationship between programming languages and code quality—such as the impact of project-specific characteristics other than the used programming language. The high-level conclusions of our exercise will be that Bayesian statistical techniques can be applied to analyze software engineering data in a way that is principled, flexible, and leads to convincing results that inform the state-of-the-art while highlighting the boundaries of its validity. The guidelines can support building solid statistical analyses and connecting their results. Thus, they can help buttress continued progress in empirical software engineering research. Carlo A. Furia, Richard Torkar, Robert Feldt |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2022 | A Method to Assess and Argue for Practical Significance in Software EngineeringabstractA key goal of empirical research in software engineering is to assess practical significance, which answers the question whether the observed effects of some compared treatments show a relevant difference in practice in realistic scenarios. Even though plenty of standard techniques exist to assess statistical significance, connecting it to practical significance is not straightforward or routinely done; indeed, only a few empirical studies in software engineering assess practical significance in a principled and systematic way. In this paper, we argue that Bayesian data analysis provides suitable tools to assess practical significance rigorously. We demonstrate our claims in a case study comparing different test techniques. The case study's data was previously analyzed (Afzalet al., 2015) using standard techniques focusing on statistical significance. Here, we build a multilevel model of the same data, which we fit and validate using Bayesian techniques. Our method is to apply cumulative prospect theory on top of the statistical model to quantitatively connect our statistical analysis output to a practically meaningful context. This is then the basis both for assessing and arguing for practical significance. Our study demonstrates that Bayesian analysis provides a technically rigorous yet practical framework for empirical software engineering. A substantial side effect is that any uncertainty in the underlying data will be propagated through the statistical model, and its effects on practical significance are made clear. Thus, in combination with cumulative prospect theory, Bayesian analysis supports seamlessly assessing practical significance in an empirical software engineering context, thus potentially clarifying and extending the relevance of research for practitioners. Richard Torkar, Carlo A. Furia, Robert Feldt, Francisco Gomes de Oliveira Neto, Lucas Gren, Per Lenberg, Neil A. Ernst |
IEEE Trans. Software Eng. | 1 |
| 2021 | Measuring affective states from technical debtabstractAbstract Context Software engineering is a human activity. Despite this, human aspects are under-represented in technical debt research, perhaps because they are challenging to evaluate. Objective This study’s objective was to investigate the relationship between technical debt and affective states (feelings, emotions, and moods) from software practitioners. Method Forty participants ( N = 40) from twelve companies took part in a mixed-methods approach, consisting of a repeated-measures ( r = 5) experiment ( n = 200), a survey, and semi-structured interviews. From the qualitative data, it is clear that technical debt activates a substantial portion of the emotional spectrum and is psychologically taxing. Further, the practitioners’ reactions to technical debt appear to fall in different levels of maturity. Results The statistical analysis shows that different design smells (strong indicators of technical debt) negatively or positively impact affective states. Conclusions We argue that human aspects in technical debt are important factors to consider, as they may result in, e.g., procrastination, apprehension, and burnout. Jesper Olsson, Erik Risfelt, Terese Besker, Antonio Martini 0001, Richard Torkar |
Empir. Softw. Eng. | 5 |
| 2021 | An empirical study of Linespots: A novel past-fault algorithmabstractSummary This paper proposes the novel past‐faults fault prediction algorithm Linespots, based on the Bugspots algorithm. We analyse the predictive performance and runtime of Linespots compared with Bugspots with an empirical study using the most significant self‐built dataset as of now, including high‐quality samples for validation. As a novelty in fault prediction, we use Bayesian data analysis and Directed Acyclic Graphs to model the effects. We found consistent improvements in the predictive performance of Linespots over Bugspots for all seven evaluation metrics. We conclude that Linespots should be used over Bugspots in all cases where no real‐time performance is necessary. Maximilian Scholz, Richard Torkar |
Softw. Test. Verification Reliab. | 2 |
| 2021 | Bayesian Data Analysis in Empirical Software Engineering ResearchabstractStatistics comes in two main flavors: frequentist and Bayesian. For historical and technical reasons, frequentist statistics have traditionally dominated empirical data analysis, and certainly remain prevalent in empirical software engineering. This situation is unfortunate because frequentist statistics suffer from a number of shortcomings-such as lack of flexibility and results that are unintuitive and hard to interpret-that curtail their effectiveness when dealing with the heterogeneous data that is increasingly available for empirical analysis of software engineering practice. In this paper, we pinpoint these shortcomings, and present Bayesian data analysis techniques that provide tangible benefits-as they can provide clearer results that are simultaneously robust and nuanced. After a short, high-level introduction to the basic tools of Bayesian statistics, we present the reanalysis of two empirical studies on the effectiveness of automatically generated tests and the performance of programming languages. By contrasting the original frequentist analyses with our new Bayesian analyses, we demonstrate the concrete advantages of the latter. To conclude we advocate a more prominent role for Bayesian statistical techniques in empirical software engineering research and practice. Carlo A. Furia, Robert Feldt, Richard Torkar |
IEEE Trans. Software Eng. | 3 |
| 2020 | The Impact of a Proposal for Innovation Measurement in the Software IndustryabstractBackground: Measuring an organization's capability to innovate and assessing its innovation output and performance is a challenging task. Previously, a comprehensive model and a suite of measurements to support this task were proposed. Aims: In the current paper, seven years since the publication of the paper titled Towards innovation measurement in the software industry, we have reflected on the impact of the work. Method: We have mainly relied on quantitative and qualitative analysis of the citations of the paper using an established classification schema. Results: We found that the article has had a significant scientific impact (indicated by the number of citations), i.e., (1) cited in literature from both software engineering and other fields, (2) cited in grey literature and peer-reviewed literature, and (3) substantial citations in literature not published in the English language. However, we consider a majority of the citations in the peer-reviewed literature (75 out of 116) as neutral, i.e., they have not used the innovation measurement paper in any substantial way. All in all, 38 out of 116 have used, modified or based their work on the definitions, measurements or the model proposed in the article. This analysis revealed a significant weakness of the citing work, i.e., among the citing papers, we found only two explicit comparisons to the innovation measurement proposal, and we found no papers that identify weaknesses of said proposal. Conclusions: This work highlights the need for being cautious of relying solely on the number of citations for understanding impact, and the need for further improving and supporting the peer-review process to identify unwarranted citations in papers. Nauman Bin Ali, Henry Edison, Richard Torkar |
ESEM | 3 |
| 2020 | Pandemic programmingabstractAbstract Context As a novel coronavirus swept the world in early 2020, thousands of software developers began working from home. Many did so on short notice, under difficult and stressful conditions. Objective This study investigates the effects of the pandemic on developers’ wellbeing and productivity. Method A questionnaire survey was created mainly from existing, validated scales and translated into 12 languages. The data was analyzed using non-parametric inferential statistics and structural equation modeling. Results The questionnaire received 2225 usable responses from 53 countries. Factor analysis supported the validity of the scales and the structural model achieved a good fit (CFI = 0.961, RMSEA = 0.051, SRMR = 0.067). Confirmatory results include: (1) the pandemic has had a negative effect on developers’ wellbeing and productivity; (2) productivity and wellbeing are closely related; (3) disaster preparedness, fear related to the pandemic and home office ergonomics all affect wellbeing or productivity. Exploratory analysis suggests that: (1) women, parents and people with disabilities may be disproportionately affected; (2) different people need different kinds of support. Conclusions To improve employee productivity, software companies should focus on maximizing employee wellbeing and improving the ergonomics of employees’ home offices. Women, parents and disabled persons may require extra support. Paul Ralph, Sebastian Baltes, Gianisa Adisaputri, Richard Torkar, Vladimir Kovalenko, Marcos Kalinowski, Nicole Novielli, Shin Yoo, Xavier Devroey, Xin Tan 0003, Minghui Zhou 0001, Burak Turhan, Rashina Hoda, Hideaki Hata, Gregorio Robles, Amin Milani Fard, Rana Alkadhi |
Empir. Softw. Eng. | 4 |
| 2019 | Estimating Return on Investment for GUI Test Automation FrameworksabstractAutomated graphical user interface (GUI) tests can reduce manual testing activities and increase test frequency. This motivates the conversion of manual test cases into automated GUI tests. However, it is not clear whether such automation is cost-effective given that GUI automation scripts add to the code base and demand maintenance as a system evolves. In this paper, we introduce a method for estimating maintenance cost and Return on Investment (ROI) for Automated GUI Testing (AGT). The method utilizes the existing source code change history and has the potential to be used for the evaluation of other testing or quality assurance automation technologies. We evaluate the method for a real-world, industrial software system and compare two fundamentally different AGT frameworks, namely Selenium and EyeAutomate, to estimate and compare their ROI. We also report on their defect-finding capabilities and usability. The quantitative data is complemented by interviews with employees at the company the study has been conducted at. The method was successfully applied, and estimated maintenance cost and ROI for both frameworks are reported. Overall, the study supports earlier results showing that implementation time is the leading cost for introducing AGT. The findings further suggest that, while EyeAutomate tests are significantly faster to implement, Selenium tests require more of a programming background but less maintenance. Felix Dobslaw, Robert Feldt, David Michaelsson, Patrick Haar, Francisco Gomes de Oliveira Neto, Richard Torkar |
ISSRE | 6 |
| 2019 | The Unfulfilled Potential of Data-Driven Decision Making in Agile Software DevelopmentabstractAbstract With the general trend towards data-driven decision making (DDDM), organizations are looking for ways to use DDDM to improve their decisions. However, few studies have looked into the practitioners view of DDDM, in particular for agile organizations. In this paper we investigated the experiences of using DDDM, and how data can improve decision making. An emailed questionnaire was sent out to 124 industry practitioners in agile software developing companies, of which 84 answered. The results show that few practitioners indicated a wide-spread use of DDDM in their current decision making practices. The practitioners were more positive to its future use for higher-level and more general decision making, fairly positive to its use for requirements elicitation and prioritization decisions, while being less positive to its future use at the team level. The practitioners do see a lot of potential for DDDM in an agile context; however, currently unfulfilled. Richard Berntsson-Svensson, Robert Feldt, Richard Torkar |
XP | 3 |
| 2019 | Evolution of statistical analysis in empirical software engineering research: Current state and steps forward
Francisco Gomes de Oliveira Neto, Richard Torkar, Robert Feldt, Lucas Gren, Carlo A. Furia |
J. Syst. Softw. | 2 |
| 2018 | Transferring interactive search-based software testing to industry
Bogdan Marculescu, Robert Feldt, Richard Torkar, Simon M. Poulding |
J. Syst. Softw. | 3 |
| 2017 | A Controlled Experiment on Coverage Maximization of Automated Model-Based Software Test Cases in the Automotive IndustryabstractIn the automotive industry, as the complexity of electronic control units (ECUs) increase, there is a need for the creation of models that facilitate early tests to ensure functionality, but there is little guidance on how to write these tests in order to achieve maximum coverage. Our prototype CANoe+, which builds on the CANoe and GraphWalker tools, was evaluated against CANoe with regard to coverage maximization of generated test cases from the viewpoint of both software developers and software testers. Rashid Darwish, Lynnie Nakyanzi Gwosuta, Richard Torkar |
ICST | 3 |
| 2017 | Group development and group maturity when building agile teams: A qualitative and quantitative investigation at eight large companies
Lucas Gren, Richard Torkar, Robert Feldt |
J. Syst. Softw. | 2 |
| 2016 | Capturing cost avoidance through reuse: systematic literature review and industrial evaluationabstractBackground: Cost avoidance through reuse shows the benefits gained by the software organisations when reusing an artefact. Cost avoidance captures benefits that are not captured by cost savings e.g. spending that would have increased in the absence of the cost avoidance activity. This type of benefit can be combined with quality aspects of the product e.g. costs avoided because of defect prevention. Cost avoidance is a key driver for software reuse. Objectives: The main objectives of this study are: (1) To assess the status of capturing cost avoidance through reuse in the academia; (2) Based on the first objective, propose improvements in capturing of reuse cost avoidance, integrate these into an instrument, and evaluate the instrument in the software industry. Method: The study starts with a systematic literature review (SLR) on capturing of cost avoidance through reuse. Later, a solution is proposed and evaluated in the industry to address the shortcomings identified during the systematic literature review. Results: The results of a systematic literature review describe three previous studies on reuse cost avoidance and show that no solution, to capture reuse cost avoidance, was validated in industry. Afterwards, an instrument and a data collection form are proposed that can be used to capture the cost avoided by reusing any type of reuse artefact. The instrument and data collection form (describing guidelines) were demonstrated to a focus group, as part of static evaluation. Based on the feedback, the instrument was updated and evaluated in industry at 6 development sites, in 3 different countries, covering 24 projects in total. Conclusion: The proposed solution performed well in industrial evaluation. With this solution, practitioners were able to do calculations for reuse costs avoidance and use the results as decision support for identifying potential artefacts to reuse. Mohsin Irshad, Richard Torkar, Kai Petersen, Wasif Afzal |
EASE | 2 |
| 2016 | Using Exploration Focused Techniques to Augment Search-Based Software Testing: An Experimental EvaluationabstractSearch-based software testing (SBST) often uses objective-based approaches to solve testing problems. There are, however, situations where the validity and completeness of objectives cannot be ascertained, or where there is insufficient information to define objectives at all. Incomplete or incorrect objectives may steer the search away from interesting behavior of the software under test (SUT) and from potentially useful test cases. This papers investigates the degree to which exploration-based algorithms can be used to complement an objective-based tool we have previously developed and evaluated in industry. In particular, we would like to assess how exploration-based algorithms perform in situations where little information on the behavior space is available a priori. We have conducted an experiment comparing the performance of an exploration-based algorithm with an objective-based one on a problem with a high-dimensional behavior space. In addition, we evaluate to what extent that performance degrades in situations where computational resources are limited. Our experiment shows that exploration-based algorithms are useful in covering a larger area of the behavior space and result in a more diverse solution population. Typically, of the candidate solutions that exploration-based algorithms propose, more than 80% were not covered by their objective-based counterpart. This increased diversity is present in the resulting population even when computational resources are limited. We conclude that exploration-focused algorithms are a useful means of investigating high-dimensional spaces, even in situations where limited information and limited resources are available. Bogdan Marculescu, Robert Feldt, Richard Torkar |
ICST | 3 |
| 2016 | Tester interactivity makes a difference in search-based software testing: A controlled experiment
Bogdan Marculescu, Simon M. Poulding, Robert Feldt, Kai Petersen, Richard Torkar |
Inf. Softw. Technol. | 5 |
| 2016 | Full modification coverage through automatic similarity-based test case selection
Francisco Gomes de Oliveira Neto, Richard Torkar, Patrícia Duarte de Lima Machado |
Inf. Softw. Technol. | 2 |
| 2016 | Software test process improvement approaches: A systematic literature review and an industrial case study
Wasif Afzal, Snehal Alone, Kerstin Glocksien, Richard Torkar |
J. Syst. Softw. | 4 |
| 2015 | An Initiative to Improve Reproducibility and Empirical Evaluation of Software Testing TechniquesabstractThe current concern regarding quality of evaluation performed in existing studies reveals the need for methods and tools to assist in the definition and execution of empirical studies and experiments. However, when trying to apply general methods from empirical software engineering in specific fields, such as evaluation of software testing techniques, new obstacles and threats to validity appears, hindering researchers' use of empirical methods. This paper discusses those issues specific for evaluation of software testing techniques and proposes an initiative for a collaborative effort to encourage reproducibility of experiments evaluating software testing techniques (STT). We also propose the development of a tool that enables automatic execution and analysis of experiments producing a reproducible research compendia as output that is, in turn, shared among researchers. There are many expected benefits from this Endeavour, such as providing a foundation for evaluation of existing and upcoming STT, and allowing researchers to devise and publish better experiments. Francisco Gomes de Oliveira Neto, Richard Torkar, Patrícia Duarte de Lima Machado |
ICSE (2) | 2 |
| 2015 | An experiment on the effectiveness and efficiency of exploratory testing
Wasif Afzal, Ahmad Nauman Ghazi, Juha Itkonen, Richard Torkar, Anneliese Amschler Andrews, Khurram Bhatti |
Empir. Softw. Eng. | 4 |
| 2015 | Guest Editors' Introduction: Special section on Software Engineering and Advanced Applications
Rick Rabiser, Richard Torkar |
Inf. Softw. Technol. | 2 |
| 2015 | The prospects of a quantitative measurement of agility: A validation study on an agile maturity model
Lucas Gren, Richard Torkar, Robert Feldt |
J. Syst. Softw. | 2 |
| 2014 | Outliers and Replication in Software EngineeringabstractEmpirical software engineering is a research field of growing interest. Studies within this field handles an increasing amount of data. In order to replicate a study the data needs to be accessible and all processing of this data needs to be reproducible. Specifically, the handling of deviating data points, also known as outliers, needs to be documented in order for a study to be replicated. This study investigated the data availability for recently published studies within empirical software engineering. Furthermore, it also investigated if outliers are documented in the same research field. Papers were reviewed using a literature review and the presence of outliers was investigated using an unsupervised outlier detection method. Only 37% of the papers reviewed had their data accessible. Furthermore, in many cases outliers were present in the reviewed studies but 63% of the papers studies did not mention how outliers were handled. The data availability within empirical software engineering research is low and is hindering replication of studies. Additionally, the lack of documentation regarding how outliers are handled is hindering replication. Henrik Larsson, Erik Lindqvist, Richard Torkar |
APSEC (1) | 3 |
| 2014 | Information Sources and Their Importance to Prioritize Test Cases in the Heterogeneous Systems Context
Ahmad Nauman Ghazi, Jesper Andersson, Richard Torkar, Kai Petersen, Jürgen Börstler |
EuroSPI | 3 |
| 2014 | Empirical evaluations on the cost-effectiveness of state-based testing: An industrial case study
Nina Elisabeth Holt, Lionel C. Briand, Richard Torkar |
Inf. Softw. Technol. | 3 |
| 2014 | Prediction of faults-slip-through in large software projects: an empirical evaluation
Wasif Afzal, Richard Torkar, Robert Feldt, Tony Gorschek |
Softw. Qual. J. | 2 |
| 2014 | Overcoming the Equivalent Mutant Problem: A Systematic Literature Review and a Comparative Experiment of Second Order MutationabstractContext. The equivalent mutant problem (EMP) is one of the crucial problems in mutation testing widely studied over decades. Objectives. The objectives are: to present a systematic literature review (SLR) in the field of EMP; to identify, classify and improve the existing, or implement new, methods which try to overcome EMP and evaluate them. Method. We performed SLR based on the search of digital libraries. We implemented four second order mutation (SOM) strategies, in addition to first order mutation (FOM), and compared them from different perspectives. Results. Our SLR identified 17 relevant techniques (in 22 articles) and three categories of techniques: detecting (DEM); suggesting (SEM); and avoiding equivalent mutant generation (AEMG). The experiment indicated that SOM in general and JudyDiffOp strategy in particular provide the best results in the following areas: total number of mutants generated; the association between the type of mutation strategy and whether the generated mutants were equivalent or not; the number of not killed mutants; mutation testing time; time needed for manual classification. Conclusions . The results in the DEM category are still far from perfect. Thus, the SEM and AEMG categories have been developed. The JudyDiffOp algorithm achieved good results in many areas. Lech Madeyski, Wojciech Orzeszyna, Richard Torkar, Mariusz Jozala |
IEEE Trans. Software Eng. | 3 |
| 2013 | Practitioner-Oriented Visualization in an Interactive Search-Based Software Test Creation ToolabstractSearch-based software testing uses meta-heuristic search techniques to automate or partially automate testing tasks, such as test case generation or test data generation. It uses a fitness function to encode the quality characteristics that are relevant, for a given problem, and guides the search to acceptable solutions in a potentially vast search space. From an industrial perspective, this opens up the possibility of generating and evaluating lots of test cases without raising costs to unacceptable levels. First, however, the applicability of search-based software engineering in an industrial setting must be evaluated. In practice, it is difficult to develop a priori a fitness function that covers all practical aspects of a problem. Interaction with human experts offers access to experience that is otherwise unavailable and allows the creation of a more informed and accurate fitness function. Moreover, our industrial partner has already expressed a view that the knowledge and experience of domain specialists are more important to the overall quality of the systems they develop than software engineering expertise. In this paper we describe our application of Interactive Search Based Software Testing (ISBST) in an industrial setting. We used SBST to search for test cases for an industrial software module and based, in part, on interaction with a human domain specialist. Our evaluation showed that such an approach is feasible, though it also identified potential difficulties relating to the interaction between the domain specialist and the system. Bogdan Marculescu, Robert Feldt, Richard Torkar |
APSEC (2) | 3 |
| 2013 | Objective Re-weighting to Guide an Interactive Search Based Software Testing SystemabstractEven hardware-focused industries today develop products where software is both a large and important component. Engineers tasked with developing and integrating these products do not always have a software engineering background. To ensure quality, tools are needed that automate and support software testing while allowing these domain specialists to leverage their knowledge and experience. Search-based testing could be a key aspect in creating an automated tool for supporting testing activities. However, domain specific quality criteria and trade-offs make it difficult to develop a general fitness function a priori, so interaction between domain specialists and such a tool would be critical to its success. In this paper we present a system for interactive search-based software testing and investigate a way for domain specialists to guide the search by dynamically re-weighting quality goals. Our empirical investigation shows that objective re-weighting can help a human domain specialist interactively guide the search, without requiring specialized knowledge of the system and without sacrificing population diversity. Bogdan Marculescu, Robert Feldt, Richard Torkar |
ICMLA (2) | 3 |
| 2013 | Software fault prediction metrics: A systematic literature review
Danijel Radjenovic, Marjan Hericko, Richard Torkar, Ales Zivkovic |
Inf. Softw. Technol. | 3 |
| 2013 | Equality in cumulative voting: A systematic review with an improvement proposal
K. Rinkevics, Richard Torkar |
Inf. Softw. Technol. | 2 |
| 2013 | Test case selection for black-box regression testing of database applications
Erik Rogstad, Lionel C. Briand, Richard Torkar |
Inf. Softw. Technol. | 3 |
| 2013 | Towards innovation measurement in the software industry
Henry Edison, Nauman Bin Ali, Richard Torkar |
J. Syst. Softw. | 3 |
| 2012 | State-Based Testing: Industrial Evaluation of the Cost-Effectiveness of Round-Trip Path and Sneak-Path StrategiesabstractIn the context of safety-critical software development, one important step in ensuring safe behavior is conformance testing, i.e., checking compliance between expected behavior and implementation. Round-trip path testing (RTP) is one example of conformance testing. Another essential step, however, is sneak-path testing, that is testing of how software reacts to unexpected events for a particular system state. Despite the importance of being systematic while testing, all testing activities take place, even for safety-critical software, under resource constraints. In this paper, we present an empirical evaluation of the cost-effectiveness of RTP when combined with sneak-path testing in the context of an industrial control system. Results highlight the importance of sneak-path testing since unexpected behavior is shown to be difficult to detect by other common, state-based test strategies. Results also suggest that sneak-path testing is a cost-effective supplement to RTP. Nina Elisabeth Holt, Richard Torkar, Lionel C. Briand, Kai Hansen |
ISSRE | 2 |
| 2012 | A Unified Model for Server Usage and Operational Costs Based on User Profiles: An Industrial Evaluation
Johannes Pelto-Piri, Peter Molin, Richard Torkar |
SEKE | 3 |
| 2012 | A Concept for an Interactive Search-Based Software Testing System
Bogdan Marculescu, Robert Feldt, Richard Torkar |
SSBSE | 3 |
| 2012 | Resampling Methods in Software Quality ClassificationabstractIn the presence of a number of algorithms for classification and prediction in software engineering, there is a need to have a systematic way of assessing their performances. The performance assessment is typically done by some form of partitioning or resampling of the original data to alleviate biased estimation. For predictive and classification studies in software engineering, there is a lack of a definitive advice on the most appropriate resampling method to use. This is seen as one of the contributing factors for not being able to draw general conclusions on what modeling technique or set of predictor variables are the most appropriate. Furthermore, the use of a variety of resampling methods make it impossible to perform any formal meta-analysis of the primary study results. Therefore, it is desirable to examine the influence of various resampling methods and to quantify possible differences. Objective and method: This study empirically compares five common resampling methods (hold-out validation, repeated random sub-sampling, 10-fold cross-validation, leave-one-out cross-validation and non-parametric bootstrapping) using 8 publicly available data sets with genetic programming (GP) and multiple linear regression (MLR) as software quality classification approaches. Location of (PF, PD) pairs in the ROC (receiver operating characteristics) space and area under an ROC curve (AUC) are used as accuracy indicators. Results: The results show that in terms of the location of (PF, PD) pairs in the ROC space, bootstrapping results are in the preferred region for 3 of the 8 data sets for GP and for 4 of the 8 data sets for MLR. Based on the AUC measure, there are no significant differences between the different resampling methods using GP and MLR. Conclusion: There can be certain data set properties responsible for insignificant differences between the resampling methods based on AUC. These include imbalanced data sets, insignificant predictor variables and high-dimensional data sets. With the current selection of data sets and classification techniques, bootstrapping is a preferred method based on the location of (PF, PD) pair data in the ROC space. Hold-out validation is not a good choice for comparatively smaller data sets, where leave-one-out cross-validation (LOOCV) performs better. For comparatively larger data sets, 10-fold cross-validation performs better than LOOCV. Wasif Afzal, Richard Torkar, Robert Feldt |
Int. J. Softw. Eng. Knowl. Eng. | 2 |
| 2012 | Requirements Traceability: a Systematic Review and Industry Case StudyabstractRequirements traceability enables software engineers to trace a requirement from its emergence to its fulfillment. In this paper we examine requirements traceability definitions, challenges, tools and techniques, by the use of a systematic review performing an exhaustive search through the years 1997–2007. We present a number of common definitions, challenges, available tools and techniques (presenting empirical evidence when found), while complementing the results and analysis with a static validation in industry through a series of interviews. Richard Torkar, Tony Gorschek, Robert Feldt, Mikael Svahnberg, Uzair Akbar Raja, Kashif Kamran |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2012 | Introduction of a process maturity model for market-driven product management and requirements engineeringabstractSUMMARY The area of software product development of software intensive products has received much attention, especially in the area of requirements engineering and product management. Many companies are faced with new challenges when operating in an environment where potential requirements number in thousands or even tens of thousands, and where a product does not have a customer, but any number of customers or markets. The development organization carries not only all the costs of development, but also takes all the risks. In this environment traditional bespoke requirements engineering, together with traditional process assessment and improvement models fall short as they do not address the unique challenges of a market‐driven environment. This paper introduces the Market‐driven Requirements Engineering Process Model, aimed at enabling process improvement and process assurance for organizations faced with these new challenges. The model is also validated in the industry through three case studies where the model is used for process assessment and improvement suggestion. Initial results show that the model is appropriate for process improvement for organizations operating in a market‐driven environment. In addition, the model was designed to be light weight in terms of low cost and thus adapted not only for large organizations but suitable for small and medium enterprises as well. Copyright © 2011 John Wiley & Sons, Ltd. Tony Gorschek, Andrigo Gomes, Andreas Pettersson, Richard Torkar |
J. Softw. Maintenance Res. Pract. | 4 |
| 2012 | Quality Requirements in Industrial Practice - An Extended Interview Study at Eleven CompaniesabstractIn order to create a successful software product and assure its quality, it is not enough to fulfill the functional requirements, it is also crucial to find the right balance among competing quality requirements (QR). An extended, previously piloted, interview study was performed to identify specific challenges associated with the selection, tradeoff, and management of QR in industrial practice. Data were collected through semistructured interviews with 11 product managers and 11 project leaders from 11 software companies. The contribution of this study is fourfold: First, it compares how QR are handled in two cases, companies working in business-to-business markets and companies that are working in business-to-consumer markets. These two are also compared in terms of impact on the handling of QR. Second, it compares the perceptions and priorities of QR by product and project management, respectively. Third, it includes an examination of the interdependencies among quality requirements perceived as most important by the practitioners. Fourth, it characterizes the selection and management of QR in downstream development activities. Richard Berntsson-Svensson, Tony Gorschek, Björn Regnell, Richard Torkar, Ali Shahrokni, Robert Feldt |
IEEE Trans. Software Eng. | 4 |
| 2011 | Search-based software testing and test data generation for a dynamic programming languageabstractManually creating test cases is time consuming and error prone. Search-based software testing can help automate this process and thus reduce time and effort and increase quality by automatically generating relevant test cases. Previous research has mainly focused on static programming languages and simple test data inputs such as numbers. This is not practical for dynamic programming languages that are increasingly used by software developers. Here we present an approach for search-based software testing for dynamically typed programming languages that can generate test scenarios and both simple and more complex test data. The approach is implemented as a tool, RuTeG, in and for the dynamic programming language Ruby. It combines an evolutionary search for test cases that give structural code coverage with a learning component to restrict the space of possible types of inputs. The latter is called for in dynamic languages since we cannot always know statically which types of objects are valid inputs. Experiments on 14 cases taken from real-world Ruby projects show that RuTeG achieves full or higher statement coverage on more cases and does so faster than randomly generated test cases. Stefan Mairhofer, Robert Feldt, Richard Torkar |
GECCO | 3 |
| 2011 | Prioritization of quality requirements: State of practice in eleven companiesabstractRequirements prioritization is recognized as an important but challenging activity in software product development. For a product to be successful, it is crucial to find the right balance among competing quality requirements. Although literature offers many methods for requirements prioritization, the research on prioritization of quality requirements is limited. This study identifies how quality requirements are prioritized in practice at 11 successful companies developing software intensive systems. We found that ad-hoc prioritization and priority grouping of requirements are the dominant methods for prioritizing quality requirements. The results also show that it is common to use customer input as criteria for prioritization but absence of any criteria was also common. The results suggests that quality requirements by default have a lower priority than functional requirements, and that they only get attention in the prioritizing process if decision-makers are dedicated to invest specific time and resources on QR prioritization. The results of this study may help future research on quality requirements to focus investigations on industry-relevant issues. Richard Berntsson-Svensson, Tony Gorschek, Björn Regnell, Richard Torkar, Ali Shahrokni, Robert Feldt, Aybüke Aurum |
RE | 4 |
| 2011 | On the application of genetic programming for software engineering predictive modeling: A systematic review
Wasif Afzal, Richard Torkar |
Expert Syst. Appl. | 2 |
| 2010 | Challenges with Software Verification and Validation Activities in the Space IndustryabstractDeveloping software for high-dependable space applications and systems is a formidable task. With new political and market pressures on the space industry to deliver more software at a lower cost, optimization of their methods and standards need to be investigated. The industry has to follow standards that strictly set quality goals and prescribes engineering processes and methods to fulfill them. The overall goal of this study is to evaluate if current use of the standards from the European Cooperation for Space Standardization (ECSS) is cost efficient and if there are ways to make the process leaner while still maintaining quality and to analyze if their verification and validation (V&V) activities can be optimized. This paper presents results from two industrial case studies of companies in the European space industry that are following ECSS standards in various V&V activities. The case studies reported here focus on how ECSS standards are used by the companies, how that affects their processes and, in the end, how their V&V activities can be further optimized. Robert Feldt, Richard Torkar, Ehsan Ahmad, Bilal Raza |
ICST | 2 |
| 2010 | Links between the personalities, views and attitudes of software engineers
Robert Feldt, Lefteris Angelis, Richard Torkar, Maria Samuelsson |
Inf. Softw. Technol. | 3 |
| 2010 | A systematic review on strategic release planning models
Mikael Svahnberg, Tony Gorschek, Robert Feldt, Richard Torkar, Saad Bin Saleem, Muhammad Usman Shafique |
Inf. Softw. Technol. | 4 |
| 2009 | A systematic review of search-based testing for non-functional system properties
Wasif Afzal, Richard Torkar, Robert Feldt |
Inf. Softw. Technol. | 2 |
| 2008 | A Comparative Evaluation of Using Genetic Programming for Predicting Fault Count DataabstractThere have been a number of software reliability growth models (SRGMs) proposed in literature. Due to several reasons, such as violation of models' assumptions and complexity of models, the practitioners face difficulties in knowing which models to apply in practice. This paper presents a comparative evaluation of traditional models and use of genetic programming (GP) for modeling software reliability growth based on weekly fault count data of three different industrial projects. The motivation of using a GP approach is its ability to evolve a model based entirely on prior data without the need of making underlying assumptions. The results show the strengths of using GP for predicting fault count data. Wasif Afzal, Richard Torkar |
ICSEA | 2 |
| 2008 | Pitfalls in Remote Team Coordination: Lessons Learned from a Case Study
Darja Smite, Nils Brede Moe, Richard Torkar |
PROFES | 3 |
| 2008 | A Systematic Mapping Study on Non-Functional Search-based Software Testing
Wasif Afzal, Richard Torkar, Robert Feldt |
SEKE | 2 |
| 2003 | New Quality Estimations in Random TestingabstractBy reformulating the issue of random testing into an equivalent problem we are able to introduce a new kind of quality estimations based on Monte Carlo integration and the central limit theorem. This method also provides a limited but working "success theory" in the case of no detected failures. In an empirical evaluation using hundreds of billions of simulated tests we furthermore find a very good match between the quality estimations presented in this article and the true failure frequencies. Both simple modulus defects as well as seeded defects in two extensively employed numerical routines were subject to investigation in the empirical work. Stefan Mankefors-Christiernin, Richard Torkar, Andreas Boklund |
ISSRE | 2 |
| 2003 | An Exploratory Study of Component Reliability Using Unit TestingabstractUsing basic unit testing techniques we found 25 faults in a core component within a larger component oriented framework after the component had already started to he reused. We found that, even though this particular component had been subject to subsystem and system testing and used for some time, several faults were discovered which seriously would have affected applications using it, especially in terms of reliability. This study clearly indicates the need of a new approach to testing and verification within component-based development and reuse. Richard Torkar, Stefan Mankefors-Christiernin, Krister Hansson, Andreas Jonsson |
ISSRE | 1 |