Per Runeson

dblp:24/24 · DBLP profile ↗
← Back
122ranked-venue papers
19as first author
22since 2021 · last 2025
0000-0003-2795-4851ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 119 · 17 first-author · 21 since 2021Artificial intelligence and machine learning · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author
YearPublicationVenuePosition
2025 FuseRank: Filtered Vector Search in Multimodal Structured Data
abstract
Single-stage filtering in vector search offers a significant advancement over conventional two-stage metadata filtering, which tends to suffer from high latency or low recall. We introduce FuseRank – a new multimodal filtered retrieval framework based on the extended vector space model, which unifies retrieval and filtering into a single approximate nearest neighbor query. FuseRank supports numerical, categorical, binary, and spatial tabular modalities through dedicated modality vectorizers. We implement FuseRank on a well-known filterless vector database as a reproducible pre-production prototype. Experiments on two real-world datasets yield retrieval results comparable to traditional two-stage filtered search, demonstrating the feasibility of platform-independent single-stage retrieval. This enables native modality filtering across any vector database backend via dot product computation, effectively avoiding vendor lock-in and reliance on platform-specific filtering logic and syntax. FuseRank also allows flexible weighting of each modality’s contribution to the final ranking score, while remaining easy to implement and extend to other modalities.
Dimitris Paraschakis, Rasmus Ros, Adam Asaad, Markus Borg, Per Runeson
KES5
2025 Exploring the Performance of ML Model Size for Classification in Relation to Energy Consumption
Andreas Bexell, Lo Gullstrand Heander, Emma Söderberg, Sigrid Eldh, Per Runeson
PROFES5
2025 AI Alignment for Ethical Compliance and Risk Mitigation in Industrial Applications
Rushali Gupta, Qunying Song, Matthias Wagner 0008, Emelie Engström, Emma Söderberg, Markus Borg, Per Runeson
PROFES7
2025 Assertions in software testing: survey, landscape, and trends
abstract
Abstract Assertions are one of the most useful automated techniques for checking program’s behaviour and hence have been used for different verification and validation tasks. We provide an overview of the last two decades of research involving “assertions” in software testing. Based on a term-based search, we filtered the inclusion of relevant papers and synthesised them with respect to the problem addressed, the solution designed, and the evaluation conducted. The survey rendered 145 papers on assertions in software testing. After test oracle, the dominant problem is test generation, followed by engineering aspects of assertions. Solutions are typically embedded in tool prototypes and evaluated throughout a limited number of cases, whereas using large-scale industrial settings is still a noticeable method. We conclude that assertions would be worth more attention in future research, particularly regarding the new and emerging demands (e.g., verification of programs with uncertainty), for effective, applicable, and domain-specific solutions.
Masoumeh Taromirad, Per Runeson
Int. J. Softw. Tools Technol. Transf.2
2024 Threats to Validity in Software Engineering - hypocritical paper section or essential analysis?
abstract
Background: In recent years, a discourse on how to systematically consider and report threats to validity started to gain momentum within the empirical software engineering community. Aims: With this study, we aim to systematically underpin the current state of threats to validity practices in software engineering research. Method: We conduct a literature review comprising 91 papers awarded with the ACM SIGSOFT Distinguished Paper Award at the ACM/IEEE International Conference on Software Engineering. Data is extracted and analyzed by considering six main facets of threats to validity, e.g., their explicit documentation, categorization, discussion of limitations, and trade-offs. Results: Results corroborate current critiques to the threats management state of the art. Threats result to be seldom discussed in depth, and are mostly considered as an enforced afterthought rather than an active concern of the research design and execution. Conclusions: To improve the observed practice, we derived items to consider for researchers, reviewers and readers, and call for a community action to increase the understanding of knowledge creation in empirical software engineering research.
Patricia Lago, Per Runeson, Qunying Song, Roberto Verdecchia
ESEM2
2024 Advancing Software Monitoring: An Industry Survey on ML-Driven Alert Management Strategies
abstract
With the dynamic nature of modern software development and operations environments and the increasing complexity of cloud-based software systems, traditional monitoring practices are often insufficient to timely identify and handle unexpected operational failures. To address these challenges, this paper presents the findings from a quantitative industry survey focused on the application of Machine Learning (ML) to enhance software monitoring and alert management strategies. The survey targets industry professionals, aiming to understand the current challenges and future trends in ML-driven software monitoring. We analyze 25 responses from 11 different software companies to conclude if and how ML is being integrated into their monitoring systems. Key findings revealed a growing but still limited reliance on ML to intelligently filter raw monitoring data, prioritize issues, and respond to system alerts, thereby improving operational efficiency and system reliability. The paper also discusses the barriers to adopting ML-based solutions and provides insights into the future direction of software monitoring.
Adha Hrusto, Per Runeson, Emelie Engström, Magnus C. Ohlsson
SEAA2
2024 Inter-Organizational Data Sharing Processes - An Exploratory Analysis of Incentives and Challenges
abstract
Businesses across different areas of interest are increasingly depending on data, particularly for machine learning (ML) applications. To ensure data provisioning, inter-organizational data sharing is proposed, e.g. in the form of data ecosystems. The aim of this study was to perform an exploratory investigation into the data sharing practices that exist in business-to-business (B2B) and business-to-customers (B2C) relations, in order to shape a knowledge foundation for future research. We launched a qualitative survey, using interviews as data collection method. We conducted and analyzed eleven interviews with representatives from seven different companies across several industries with the aim of finding key practices, differences and similarities between approaches, so we could formulate the future research goals and questions. We grouped the core findings of this study into three categories: organizational aspects of data sharing, where we noticed the importance of data sharing and data ownership as business driver; technical aspects of data sharing, related to data types, formats, maintenance and infrastructures; and challenges, with privacy being the highest concern along with the data volumes and cost of data.
Konstantin Malysh, Johan Linåker, Per Runeson
SEAA4
2024 A theory of factors affecting continuous experimentation (FACE)
abstract
Abstract Context Continuous experimentation (CE) is used by many companies with internet-facing products to improve their business models and software solutions based on user data. Some companies deliberately adopt a systematic experiment-driven approach to software development while some companies use CE in a more ad-hoc fashion. Objective The goal of this study is to identify factors for success in CE that explain the variations in the utility and efficacy of CE between different companies. Method We conducted a multi-case study of 12 companies involved with CE and performed 27 interviews with practitioners at these companies. Based on that empirical data, we then built a theory of factors at play in CE. Results We introduce a theory of Factors Affecting Continuous Experimentation (FACE). The theory includes three factors, namely 1) processes and infrastructure for CE, 2) the user problem complexity of the product offering, and 3) incentive structures for CE. The theory explains how these factors affect the effectiveness of CE and its ability to achieve problem-solution and product-market fit. Conclusions Our theory may inspire practitioners to assess an organisation’s potential for adopting CE and to identify factors that pose challenges in gaining value from CE practices. Our results also provide a basis for defining practitioner guidelines and a starting point for further research on how contextual factors affect CE and how these may be mitigated.
Rasmus Ros, Elizabeth Bjarnason, Per Runeson
Empir. Softw. Eng.3
2024 Industry Practices for Challenging Autonomous Driving Systems with Critical Scenarios
abstract
Testing autonomous driving systems for safety and reliability is essential, yet complex. A primary challenge is identifying relevant test scenarios, especially the critical ones that may expose hazards or harm to autonomous vehicles and other road users. Although numerous approaches and tools for critical scenario identification are proposed, the industry practices for selection, implementation, and evaluation of approaches, are not well understood. Therefore, we aim at exploring practical aspects of how autonomous driving systems are tested, particularly the identification and use of critical scenarios. We interviewed 13 practitioners from 7 companies in autonomous driving in Sweden. We used thematic modeling to analyse and synthesize the interview data. As a result, we present 9 themes of practices and 4 themes of challenges related to critical scenarios. Our analysis indicates there is little joint effort in the industry, despite every approach has its own limitations, and tools and platforms are lacking. To that end, we recommend the industry and academia combine different approaches, collaborate among different stakeholders, and continuously learn the field. The contributions of our study are exploration and synthesis of industry practices and related challenges for critical scenario identification and testing, and potential increase of industry relevance for future studies.
Qunying Song, Emelie Engström, Per Runeson
ACM Trans. Softw. Eng. Methodol.3
2023 Towards optimization of anomaly detection in DevOps
abstract
DevOps has recently become a mainstream solution for bridging the gaps between development (Dev) and operations (Ops) enabling cross-functional collaboration. The DevOps concept of continuous monitoring may bring a lot of benefits to development teams such as early detection of run-time errors and various performance anomalies. We aim to explore deep learning (DL) solutions for detection of anomalous systems behavior based on collected monitoring data that consists of applications’ and systems’ performance metrics. Moreover, we specifically address a shortage of approaches for evaluating DL models without any ground truth data. We perform a case study in a real DevOps environment, following the principles of the design science paradigm. The research activities span from practice to theory and from problem to solution domain, including problem conceptualization, solution design, instantiation, and empirical validation. We proposed and implemented a cloud solution for DL model deployment and evaluation empowered by feedback from the development team. The labeled data generated through the feedback was used for evaluation of current and training of new DL models in several iterations. The overall results showed that reconstruction-based models such as autoencoders, are quite robust to any parameter modification and are among the preferred for anomaly detection in multivariate monitoring data. Leveraging raw monitoring data and DL-inspired solutions, DevOps teams may get critical insights into the software and its operation. In our case, this proved to be an efficient way of discovering early signs of production failures.
Adha Hrusto, Emelie Engström, Per Runeson
Inf. Softw. Technol.3
2023 Industry-academia collaboration for realism in software engineering research: Insights and recommendations
abstract
Effective industry-academia collaboration may increase software engineering research relevance by increased realism, yet very challenging for reasons like confidentiality concerns, different objectives and priorities. We analyse industry-academia collaboration scenarios based on our own experiences as Ph.D. student and supervisor, and provide insights and recommendations to facilitate future collaborations with industry. We first present our industry-academia collaboration experiences that span over two and a half years with different companies. Then, we analyse both facilitators and problems from those scenarios and synthesize recommendations based on that. Five different scenarios are analysed, including both success and failure scenarios. Reflections and insights into these experiences as well as some general recommendations are presented. We believe such experiences and insights are helpful for academic researchers to pursue industry-academia collaboration. We plan to continuously report our experience and provide our suggestions for effective collaboration with industry.
Qunying Song, Per Runeson
Inf. Softw. Technol.2
2023 Threats to validity in software engineering research: A critical reflection
Roberto Verdecchia, Emelie Engström, Patricia Lago, Per Runeson, Qunying Song
Inf. Softw. Technol.4
2023 Critical scenario identification for realistic testing of autonomous driving systems
abstract
Abstract Autonomous driving has become an important research area for road traffic, whereas testing of autonomous driving systems to ensure a safe and reliable operation remains an open challenge. Substantial real-world testing or massive driving data collection does not scale since the potential test scenarios in real-world traffic are infinite, and covering large shares of them in the test is impractical. Thus, critical ones have to be prioritized. We have developed an approach for critical test scenario identification and in this study, we implement the approach and validate it on two real autonomous driving systems from industry by integrating it into their tool-chain. Our main contribution in this work is the demonstration and validation of our approach for critical scenario identification for testing real autonomous driving systems.
Qunying Song, Kaige Tan, Per Runeson, Stefan Persson
Softw. Qual. J.3
2022 A Scenario Distribution Model for Effective and Efficient Testing of Autonomous Driving Systems
abstract
While autonomous driving systems are expected to change future means of mobility and reduce road accidents, understanding intensive and complex traffic situations is essential to enable testing of such systems under realistic traffic conditions. Particularly, we need to cover more relevant driving scenarios in the test. However, we do not want to spend time and resources testing useless scenarios that never happen in the real road traffic. In this work, we propose a new model that defines the distribution of scenarios using TTC (Time-to-Collision) for the vehicle–pedestrian interactions at unsignalized crossings based on the traffic density. The scenario distribution can be used as an input for test scenario generation and selection. We validate the model using real traffic data collected in Sweden and the result indicates that the model is effective and consistently upholds the real distribution, especially for critical scenarios with TTC less than 3 seconds. We also demonstrate the use of the model by connecting it to the testing of an auto-braking function from the industry. As a first step, our contribution is a model that predicts the worst-case distribution of scenarios using TTC and provides a mandatory input for testing autonomous driving systems.
Qunying Song, Per Runeson, Stefan Persson
ASE2
2022 A/B Testing in the Small: An Empirical Exploration of Controlled Experimentation on Internal Tools
Amalia Paulsson, Per Runeson, Rasmus Ros
PROFES2
2022 Near Failure Analysis Using Dynamic Behavioural Data
Masoumeh Taromirad, Per Runeson
PROFES2
2022 Sustaining Open Data as a Digital Common - Design principles for Common Pool Resources applied to Open Data Ecosystems
abstract
Motivation. Digital commons is an emerging phenomenon and of increasing importance, as we enter a digital society. Open data is one example that makes up a pivotal input and foundation for many of today’s digital services and applications. Ensuring sustainable provisioning and maintenance of the data, therefore, becomes even more important. Aim. We aim to investigate how such provisioning and maintenance can be collaboratively performed in the community surrounding a common. Specifically, we look at Open Data Ecosystems (ODEs), a type of community of actors, openly sharing and evolving data on a technological platform. Method. We use Elinor Ostrom’s design principles for Common Pool Resources as a lens to systematically analyze the governance of earlier reported cases of ODEs using a theory-oriented software engineering framework. Results. We find that, while natural commons must regulate consumption, digital commons such as open data maintained by an ODE must stimulate both use and data provisioning. Governance needs to enable such stimulus while also ensuring that the collective action can still be coordinated and managed within the frame of available maintenance resources of a community. Subtractability is, in this sense, a concern regarding the resources required to maintain the quality and value of the data, rather than the availability of data. Further, we derive empirically-based recommended practices for ODEs based on the design principles by Ostrom for how to design a governance structure in a way that enables a sustainable and collaborative provisioning and maintenance of the data. Conclusion. ODEs are expected to play a role in data provisioning which democratize the digital society and enables innovation from smaller commercial actors. Our empirically based guidelines intend to support this development.
Johan Linåker, Per Runeson
OpenSym2
2021 Controlled experimentation in continuous experimentation: Knowledge and challenges
abstract
Continuous experimentation and A/B testing is an established industry practice that has been researched for more than 10 years. Our aim is to synthesize the conducted research. We wanted to find the core constituents of a framework for continuous experimentation and the solutions that are applied within the field. Finally, we were interested in the challenges and benefits reported of continuous experimentation. We applied forward snowballing on a known set of papers and identified a total of 128 relevant papers. Based on this set of papers we performed two qualitative narrative syntheses and a thematic synthesis to answer the research questions. The framework constituents for continuous experimentation include experimentation processes as well as supportive technical and organizational infrastructure. The solutions found in the literature were synthesized to nine themes, e.g. experiment design, automated experiments, or metric specification. Concerning the challenges of continuous experimentation, the analysis identified cultural, organizational, business, technical, statistical, ethical, and domain-specific challenges. Further, the study concludes that the benefits of experimentation are mostly implicit in the studies. The research on continuous experimentation has yielded a large body of knowledge on experimentation. The synthesis of published research presented within include recommended infrastructure and experimentation process models, guidelines to mitigate the identified challenges, and what problems the various published solutions solve.
Florian Auer, Rasmus Ros, Lukas Kaltenbrunner, Per Runeson, Michael Felderer
Inf. Softw. Technol.4
2021 Governance and Management of Green IT: A Multi-Case Study
J. David Patón-Romero, Maria Teresa Baldassarre, Moisés Rodríguez 0001, Per Runeson, Martin Höst, Mario Piattini
Inf. Softw. Technol.4
2021 Guiding the selection of research methodology in industry-academia collaboration in software engineering
abstract
The literature concerning research methodologies and methods has increased in software engineering in the last decade. However, there is limited guidance on selecting an appropriate research methodology for a given research study or project. Based on a selection of research methodologies suitable for software engineering research in collaboration between industry and academia, we present, discuss and compare the methodologies aiming to provide guidance on which research methodology to choose in a given situation to ensure successful industry–academia collaboration in research. Three research methodologies were chosen for two main reasons. Design Science and Action Research were selected for their usage in software engineering. We also chose a model emanating from software engineering, i.e., the Technology Transfer Model. An overview of each methodology is provided. It is followed by a discussion and an illustration concerning their use in industry–academia collaborative research. The three methodologies are then compared using a set of criteria as a basis for our guidance. The discussion and comparison of the three research methodologies revealed general similarities and distinct differences. All three research methodologies are easily mapped to the general research process describe–solve–practice, while the main driver behind the formulation of the research methodologies is different. Thus, we guide in selecting a research methodology given the primary research objective for a given research study or project in collaboration between industry and academia. We observe that the three research methodologies have different main objectives and differ in some characteristics, although still having a lot in common. We conclude that it is vital to make an informed decision concerning which research methodology to use. The presentation and comparison aim to guide selecting an appropriate research methodology when conducting research in collaboration between industry and academia.
Claes Wohlin, Per Runeson
Inf. Softw. Technol.2
2021 Open Data Ecosystems - An empirical investigation into an emerging industry collaboration concept
abstract
Software systems are increasingly depending on data, particularly with the rising use of machine learning, and developers are looking for new sources of data. Open Data Ecosystems (ODE) is an emerging concept for data sharing under public licenses in software ecosystems, similar to Open Source Software (OSS). It has certain similarities to Open Government Data (OGD), where public agencies share data for innovation and transparency. We aimed to explore open data ecosystems involving commercial actors. Thus, we organized five focus groups with 27 practitioners from 22 companies, public organizations, and research institutes. Based on the outcomes, we surveyed three cases of emerging ODE practice to further understand the concepts and to validate the initial findings. The main outcome is an initial conceptual model of ODEs’ value, intrinsics, governance, and evolution, and propositions for practice and further research. We found that ODE must be value driven. Regarding the intrinsics of data, we found their type, meta-data, and legal frameworks influential for their openness. We also found the characteristics of ecosystem initiation, organization, data acquisition and openness be differentiating, which we advise research and practice to take into consideration.
Per Runeson, Thomas Olsson 0001, Johan Linåker
J. Syst. Softw.1
2021 A case study of industry-academia communication in a joint software engineering research project
abstract
Abstract Empirical software engineering research relies on good communication with industrial partners. Conducting joint research both requires and contributes to bridging the communication gap between industry and academia (IA) in software engineering. This study aims to explore communication between the two parties in such a setting. To better understand what facilitates good IA communication and what project outcomes such communication promotes, we performed a case study, in the context of a long‐term IA joint project, followed by a validating survey among practitioners and researchers with experience of working in similar settings. We identified five facilitators of IA communication and nine project outcomes related to this communication. The facilitators concern the relevance of the research, practitioners' attitude and involvement in research, frequency of communication and longevity of the collaboration. The project outcomes promoted by this communication include, for researchers, changes in teaching and new scientific venues, and for practitioners, increased awareness, changes to practice, and new tools and source code. Besides, both parties gain new knowledge and develop social‐networks through IA communication. Our study presents empirically based insights that can provide advise on how to improve communication in IA research projects and thus the co‐creation of software engineering knowledge that is anchored in both practice and research.
Sergio Rico, Elizabeth Bjarnason, Emelie Engström, Martin Höst, Per Runeson
J. Softw. Evol. Process.5
2020 Getting Started with Chaos Engineering - design of an implementation framework in practice
abstract
Background. Chaos Engineering is proposed as a practice to verify a system's resilience under real, operational conditions. It employs fault injection, is originally developed at Netflix, and supported by several tools from there and other sources. Aims. We aim to introduce Chaos Engineering at ICA Gruppen AB, a group of companies whose core business is grocery retail, to improve their systems' resilience, and to capture our knowledge gained from literature and interviews in a process framework for the introduction of Chaos Engineering. Method. The research is conducted under the design science paradigm, where the problem is conceptualized through a literature study of Chaos Engineering and exploratory interviews in the company. The solution framework is designed based on the literature and a tool survey, and validated by letting software engineers at ICA apply parts of it to the software systems of ica.se website, including its e-shop. Results. The main contributions are a synthesis of Chaos Engineering literature and tools, in depth understanding of the needs of the case company, and guidelines for introducing Chaos Engineering. Conclusions. The applied parts were concluded to be feasible and they successfully discovered a set of initial improvement opportunities for the system's resilience, as well as a suitable Chaos Engineering practice for future resilience testing of the system. We recommend companies using the framework as a guide for the implementation of Chaos Engineering.
Hugo Jernberg, Per Runeson, Emelie Engström
ESEM2
2020 Challenges and Opportunities in Open Data Collaboration - a focus group study
abstract
Data-driven software is becoming prevalent, especially with the advent of machine learning and artificial intelligence. With data-driven systems come both challenges - to keep collecting and maintaining high quality data - and opportunities - open innovation by sharing data with others. We propose Open Data Collaboration (ODC) to describe pecuniary and non-pecuniary sharing of open data, similar to Open Source Software (OSS) and in contrast to Open Government Data (OGD), where public authorities share data. To understand challenges and opportunities with ODC, we organized five focus groups with in total 27 practitioners from 22 companies, public organizations, and research institutes. In the discussions, we observed a general interest in the subject, both from private companies and public authorities. We also noticed similarities in attitudes to open innovation practices, i.e. initial resistance which gradually turned into interest. While several of the participants were experienced in open source software, no had shared data openly. Based on the findings, we identify challenges which we set out to continue addressing in future research.
Per Runeson, Thomas Olsson 0001
SEAA1
2020 Public Sector Platforms going Open: Creating and Growing an Ecosystem with Open Collaborative Development
abstract
Background: By creating ecosystems around platforms of Open Source Software (OSS) and Open Data (OD), and adopting open collaborative development practices, platform providers may exploit open innovation benefits. However, adopting such practices in a traditionally closed organization is a maturity process that we hypothesize cannot be undergone without friction.
Johan Linåker, Per Runeson
OpenSym2
2020 How software engineering research aligns with design science: a review
abstract
Abstract Background Assessing and communicating software engineering research can be challenging. Design science is recognized as an appropriate research paradigm for applied research, but is rarely explicitly used as a way to present planned or achieved research contributions in software engineering. Applying the design science lens to software engineering research may improve the assessment and communication of research contributions. Aim The aim of this study is 1) to understand whether the design science lens helps summarize and assess software engineering research contributions, and 2) to characterize different types of design science contributions in the software engineering literature. Method In previous research, we developed a visual abstract template, summarizing the core constructs of the design science paradigm. In this study, we use this template in a review of a set of 38 award winning software engineering publications to extract, analyze and characterize their design science contributions. Results We identified five clusters of papers, classifying them according to their different types of design science contributions. Conclusions The design science lens helps emphasize the theoretical contribution of research output—in terms of technological rules—and reflect on the practical relevance, novelty and rigor of the rules proposed by the research.
Emelie Engström, Margaret-Anne D. Storey, Per Runeson, Martin Höst, Maria Teresa Baldassarre
Empir. Softw. Eng.3
2019 Open Tools for Software Engineering: Validation of a Theory of Openness in the Automotive Industry
abstract
Context: Open tools (e.g., Jenkins, Gerrit and Git) offer a lucrative alternative to commercial tools. Many companies and developers from OSS communities make a collaborative effort to improve the tools. Prior to this study, we developed an empirically based theory for companies' strategic choices on the development of these tools, based on empirical observations in the telecom domain. Aim: The aim of this study is to validate the theory of openness for tools in software engineering, in another domain, automotive. Specifically, we validated the theory propositions and mapped the case companies onto the model of openness. Method: We run focus groups in two automotive companies, collecting data in a survey and follow-up discussions. We used the repertory grid technique to analyze the survey responses, in combination with qualitative data from the focus group, to validate the propositions. Results: Openness of tools has the potential to reduce development costs and time, and may lead to process and product innovation. This study confirms three out of five theory propositions, on cost and time reduction, and the complementary role of open tools. One propositions was not possible to validate due to lack of investment in OSS tools communities by both companies. However, our findings extend the fifth proposition to require management being involved for both the proactive and reactive strategy. Further, we observe that the move towards open tools happen with a paradigm shift towards openness in the automotive domain, and lead to standardization of tools. Both companies confirm that they need legal procedures for the contribution, as well as an internal champion, driving the open tools strategy. Conclusion: We validated the theory, originating from the telecom domain, partially using two automotive companies. Both case companies are classified as laggards (reactive, cost saving) in the model of openness presented in the theory. Furthermore, we would like to have more validations studies to validate the remaining quadrants (e.g., leverage, lucrativeness and leaders).
Hussan Munir, Per Runeson, Krzysztof Wnuk
EASE2
2019 Open data collaborations: a snapshot of an emerging practice
abstract
Data defined software is becoming more and more prevalent, especially with the advent of machine learning and artificial intelligence. With data defined systems come both challenges - to continue to collect and maintain quality data - and opportunities - open innovation by sharing with others. We propose Open Data Collaboration (ODC) to describe pecuniary and non-pecuniary sharing of open data, similar to Open Source Software. To understand challenges and opportunities with ODC, we ran focus groups with 22 companies and organizations. We observed an interest in the subject, but we conclude that the overall maturity is low and ODC is rare.
Thomas Olsson 0001, Per Runeson
OpenSym2
2018 Four commentaries on the use of students and professionals in empirical software engineering experiments
Robert Feldt, Thomas Zimmermann 0001, Gunnar R. Bergersen, Davide Falessi, Andreas Jedlitschka, Natalia Juristo Juzgado, Jürgen Münch, Markku Oivo, Per Runeson, Martin J. Shepperd, Dag I. K. Sjøberg, Burak Turhan
Empir. Softw. Eng.9
2018 Open innovation using open source tools: a case study at Sony Mobile
abstract
Despite growing interest of Open Innovation (OI) in Software Engineering (SE), little is known about what triggers software organizations to adopt it and how this affects SE practices. OI can be realized in numerous of ways, including Open Source Software (OSS) involvement. Outcomes from OI are not restricted to product innovation but also include process innovation, e.g. improved SE practices and methods. This study explores the involvement of a software organization (Sony Mobile) in OSS communities from an OI perspective and what SE practices (requirements engineering and testing) have been adapted in relation to OI. It also highlights the innovative outcomes resulting from OI. An exploratory embedded case study investigates how Sony Mobile use and contribute to Jenkins and Gerrit; the two central OSS tools in their continuous integration tool chain. Quantitative analysis was performed on change log data from source code repositories in order to identify the top contributors and triangulated with the results from five semi-structured interviews to explore the nature of the commits. The findings of the case study include five major themes: i) The process of opening up towards the tool communities correlates in time with a general adoption of OSS in the organization. ii) Assets not seen as competitive advantage nor a source of revenue are made open to OSS communities, and gradually, the organization turns more open. iii) The requirements engineering process towards the community is informal and based on engagement. iv) The need for systematic and automated testing is still in its infancy, but the needs are identified. v) The innovation outcomes included free features and maintenance, and were believed to increase speed and quality in development. Adopting OI was a result of a paradigm shift of moving from Windows to Linux. This shift enabled Sony Mobile to utilize the Jenkins and Gerrit communities to make their internal development process better for its software developers and testers.
Hussan Munir, Johan Linåker, Krzysztof Wnuk, Per Runeson, Björn Regnell
Empir. Softw. Eng.4
2018 A theory of openness for software engineering tools in software organizations
Hussan Munir, Per Runeson, Krzysztof Wnuk
Inf. Softw. Technol.2
2017 A Machine Learning Approach for Semi-Automated Search and Selection in Literature Studies
abstract
Background. Search and selection of primary studies in Systematic Literature Reviews (SLR) is labour intensive, and hard to replicate and update. Aims. We explore a machine learning approach to support semi-automated search and selection in SLRs to address these weaknesses. Method. We 1) train a classifier on an initial set of papers, 2) extend this set of papers by automated search and snowballing, 3) have the researcher validate the top paper, selected by the classifier, and 4) update the set of papers and iterate the process until a stopping criterion is met. Results. We demonstrate with a proof-of-concept tool that the proposed automated search and selection approach generates valid search strings and that the performance for subsets of primary studies can reduce the manual work by half. Conclusions. The approach is promising and the demonstrated advantages include cost savings and replicability. The next steps include further tool development and evaluate the approach on a complete SLR.
Rasmus Ros, Elizabeth Bjarnason, Per Runeson
EASE3
2017 Using a Visual Abstract as a Lens for Communicating and Promoting Design Science Research in Software Engineering
abstract
Empirical software engineering research aims to generate prescriptive knowledge that can help software engineers improve their work and overcome their challenges, but deriving these insights from real-world problems can be challenging. In this paper, we promote design science as an effective way to produce and communicate prescriptive knowledge. We propose using a visual abstract template to communicate design science contributions and highlight the main problem/solution constructs of this area of research, as well as to present the validity aspects of design knowledge. Our conceptualization of design science is derived from existing literature and we illustrate its use by applying the visual abstract to an example use case. This is work in progress and further evaluation by practitioners and researchers will be forthcoming.
Margaret-Anne D. Storey, Emelie Engström, Martin Höst, Per Runeson, Elizabeth Bjarnason
ESEM4
2017 Unit Verification Effects on Reused Components in Sequential Project Releases
abstract
Background. The effects of different practices on fault distributions in evolving complex software systems is not fully understood. Software reuse and unit verification are practices used to improve system reliability by minimising the number of late faults. Reused software benefits from already being verified while unit verification aims to find faults early.Aims. We want to study effects of software reuse and unit verification on future modifications, fault densities of software units, and fault distributions.Method. We applied statistical analysis to a sample of 520 units that were reused and modified within four sequential projects from one product line in the telecommunication domain.Results. In reused units, the results of unit verification are correlated to a smaller degree of modifications and decreased fault densities.Conclusion. Unit verification in complex systems may improve system evolution in terms of smaller modifications and decrease of fault densities. The unit verification faults in reused components may be used as predictors of component modification and fault density.
Tihana Galinac Grbac, Per Runeson, Darko Huljenic
SEAA2
2017 Software engineers' information seeking behavior in change impact analysis: an interview study
abstract
Software engineers working in large projects must navigate complex information landscapes. Change Impact Analysis (CIA) is a task that relies on engineers' successful information seeking in databases storing, e.g., source code, requirements, design descriptions, and test case specifications. Several previous approaches to support information seeking are task-specific, thus understanding engineers' seeking behavior in specific tasks is fundamental. We present an industrial case study on how engineers seek information in CIA, with a particular focus on traceability and development artifacts that are not source code. We show that engineers have different information seeking behavior, and that some do not consider traceability particularly useful when conducting CIA. Furthermore, we observe a tendency for engineers to prefer less rigid types of support rather than formal approaches, i.e., engineers value support that allows flexibility in how to practically conduct CIA. Finally, due to diverse information seeking behavior, we argue that future CIA support should embrace individual preferences to identify change impact by empowering several seeking alternatives, including searching, browsing, and tracing.
Markus Borg, Emil Alégroth, Per Runeson
ICPC3
2017 Automated Controlled Experimentation on Software by Evolutionary Bandit Optimization
Rasmus Ros, Elizabeth Bjarnason, Per Runeson
SSBSE3
2017 Supporting Change Impact Analysis Using a Recommendation System: An Industrial Case Study in a Safety-Critical Context
abstract
Change Impact Analysis (CIA) during software evolution of safety-critical systems is a labor-intensive task. Several authors have proposed tool support for CIA, but very few tools were evaluated in industry. We present a case study on ImpRec, a recommendation System for Software Engineering (RSSE), tailored for CIA at a process automation company. ImpRec builds on assisted tracing, using information retrieval solutions and mining software repositories to recommend development artifacts, potentially impacted when resolving incoming issue reports. In contrast to the majority of tools for automated CIA, ImpRec explicitly targets development artifacts that are not source code. We evaluate ImpRec in a two-phase study. First, we measure the correctness of ImpRec's recommendations by a simulation based on 12 years' worth of issue reports in the company. Second, we assess the utility of working with ImpRec by deploying the RSSE in two development teams on different continents. The results suggest that ImpRec presents about 40 percent of the true impact among the top-10 recommendations. Furthermore, user log analysis indicates that ImpRec can support CIA in industry, and developers acknowledge the value of ImpRec in interviews. In conclusion, our findings show the potential of reusing traceability associated with developers' past activities in an RSSE.
Markus Borg, Krzysztof Wnuk, Björn Regnell, Per Runeson
IEEE Trans. Software Eng.4
2016 Foreword to the special issue on empirical evidence on software product line engineering
Ebrahim Bagheri, David Benavides 0001, Klaus Schmid, Per Runeson
Empir. Softw. Eng.4
2016 Automated bug assignment: Ensemble-based machine learning in large scale industrial contexts
Leif Jonsson, Markus Borg, David Broman, Kristian Sandahl, Sigrid Eldh, Per Runeson
Empir. Softw. Eng.6
2016 Open innovation in software engineering: a systematic mapping study
Hussan Munir, Krzysztof Wnuk, Per Runeson
Empir. Softw. Eng.3
2016 A theory of distances in software engineering
Elizabeth Bjarnason, Kari Smolander, Emelie Engström, Per Runeson
Inf. Softw. Technol.4
2016 A quantitative analysis of the unit verification perspective on fault distributions in complex software systems: an operational replication
Tihana Galinac Grbac, Per Runeson, Darko Huljenic
Softw. Qual. J.2
2015 Navigating Information Overload Caused by Automated Testing - a Clustering Approach in Multi-Branch Development
abstract
Background. Test automation is a widely used technique to increase the efficiency of software testing. However, executing more test cases increases the effort required to analyze test results. At Qlik, automated tests run nightly for up to 20 development branches, each containing thousands of test cases, resulting in information overload. Aim. We therefore develop a tool that supports the analysis of test results. Method. We create NIOCAT, a tool that clusters similar test case failures, to help the analyst identify underlying causes. To evaluate the tool, experiments on manually created subsets of failed test cases representing different use cases are conducted, and a focus group meeting is held with test analysts at Qlik. Results. The case study shows that NIOCAT creates accurate clusters, in line with analyses performed by human analysts. Further, the potential time-savings of our approach is confirmed by the participants in the focus group. Conclusions. NIOCAT provides a feasible complement to current automated testing practices at Qlik by reducing information overload.
Nicklas Erman, Vanja Tufvesson, Markus Borg, Per Runeson, Anders Ardö
ICST4
2015 Software testing in open innovation: an exploratory case study of the acceptance test harness for jenkins
abstract
Open Innovation (OI) has gained significant attention since the term was introduced in 2003. However, little is known whether general software testing processes are well suited for OI. An exploratory case study on the Acceptance Test Harness (ATH) is conducted to investigate OI testing activities of Jenkins. As far as the research methodology is concerned, we extracted the change log data of ATH followed by five interviews with key contributors in the development of ATH. The findings of the study are threefold. First, it highlights the key stakeholders involved in the development of ATH. Second, the study compares the ATH testing activities with ISO/IEC/IEEE testing process and presents a tailored process for software testing in OI. Finally, the study underlines some key challenges that software intensive organizations face while working with the testing in OI.
Hussan Munir, Per Runeson
ICSSP2
2015 Case studies synthesis: a thematic, cross-case, and narrative synthesis worked example
Daniela S. Cruzes, Tore Dybå, Per Runeson, Martin Höst
Empir. Softw. Eng.3
2014 A replicated study on duplicate detection: using apache lucene to search among Android defects
abstract
Context: Duplicate detection is a fundamental part of issue management. Systems able to predict whether a new defect report will be closed as a duplicate, may decrease costs by limiting rework and collecting related pieces of information. Goal: Our work explores using Apache Lucene for large-scale duplicate detection based on textual content. Also, we evaluate the previous claim that results are improved if the title is weighted as more important than the description. Method: We conduct a conceptual replication of a well-cited study conducted at Sony Ericsson, using Lucene for searching in the public Android defect repository. In line with the original study, we explore how varying the weighting of the title and the description affects the accuracy. Results: We show that Lucene obtains the best results when the defect report title is weighted three times higher than the description, a bigger difference than has been previously acknowledged. Conclusions: Our work shows the potential of using Lucene as a scalable solution for duplicate detection.
Markus Borg, Per Runeson, Jens Johansson, Mika Mäntylä
ESEM2
2014 Towards a framework to support large scale sampling in software engineering surveys
abstract
Context: The low quality and small size of samples in empirical studies in software engineering hamper the interpretation and generalization of their results. Therefore, enlarging sample sizes and improving their quality represent an important research challenge. Goal: We aim to define a conceptual framework, including requirements for establishing adequate sources for sampling subjects in software engineering surveys. Method: We use previous experience on applying systematic sampling strategies combined with contemporary web technologies in previously executed surveys, to organize the conceptual framework. We analyze its application to different sources of sampling. Results: The framework was observed to be feasible after its application to nine different large-scale sources of sampling. Conclusions: The analyzed crowdsourcing tools do not support essential requirements to be considered sources of sampling, while free-lancing tools and professional social network do.
Rafael Maiani de Mello, Pedro Correa da Silva, Per Runeson, Guilherme Horta Travassos
ESEM3
2014 Supporting Regression Test Scoping with Visual Analytics
abstract
Background: Test managers have to repeatedly select test cases for test activities during evolution of large software systems. Researchers have widely studied automated test scoping, but have not fully investigated decision support with human interaction. We previously proposed the introduction of visual analytics for this purpose. Aim: In this empirical study we investigate how to design such decision support. Method: We explored the use of visual analytics using heat maps of historical test data for test scoping support by letting test managers evaluate prototype visualizations in three focus groups with in total nine industrial test experts. Results: All test managers in the study found the visual analytics useful for supporting test planning. However, our results show that different tasks and contexts require different types of visualizations. Conclusion: Important properties for test planning support are: ability to overview testing from different perspectives, ability to filter and zoom to compare subsets of the testing with respect to various attributes and the ability to manipulate the subset under analysis by selecting and deselecting test cases. Our results may be used to support the introduction of visual test analytics in practice.
Emelie Engström, Mika Mäntylä, Per Runeson, Markus Borg
ICST3
2014 Challenges and practices in aligning requirements with verification and validation: a case study of six companies
Elizabeth Bjarnason, Per Runeson, Markus Borg, Michael Unterkalmsteiner, Emelie Engström, Björn Regnell, Giedre Sabaliauskaite, Annabella Loconsole, Tony Gorschek, Robert Feldt
Empir. Softw. Eng.2
2014 Recovering from a decade: a systematic mapping of information retrieval approaches to software traceability
Markus Borg, Per Runeson, Anders Ardö
Empir. Softw. Eng.2
2014 Variation factors in the design and analysis of replicated controlled experiments - Three (dis)similar studies on inspections versus unit testing
Per Runeson, Andreas Stefik, Anneliese Amschler Andrews
Empir. Softw. Eng.1
2014 Bridges and barriers to hardware-dependent software ecosystem participation - A case study
Krzysztof Wnuk, Per Runeson, Matilda Lantz, Oskar Weijden
Inf. Softw. Technol.2
2014 Early identification of bottlenecks in very large scale system of systems software development
abstract
System of systems are of high complexity, and for each system, many different requirements are implemented in parallel. Systems are developed with some degree of managerial independence but later on have to work together. In this situation, many requirements are written, implemented, and tested in parallel for different systems that are to be integrated. This makes identifying bottlenecks challenging, and visualizations often used on project level (such as Kanban boards or burndown charts) have to be extended/complemented to cope with the increased complexity. In response to these challenges, the contributions of this study are to propose the following: (i) a visualization for early identification and proactive removal of bottlenecks; (ii) a visualization to check on the success of bottleneck resolution; and (iii) to provide an industry evaluation of the visualizations in a case study of a system of systems developed at Ericsson AB in Sweden. The feedback by the practitioners showed that the visualizations were perceived as useful in improving throughput and lead time. The quantitative analysis showed that the visualizations were able in identifying bottlenecks and showing improvements or the lack thereof. On the basis of the qualitative and quantitative data collected, we conclude that the visualizations are useful in bottleneck identification and resolution. Copyright © 2014 John Wiley & Sons, Ltd.
Kai Petersen, Peter Roos, Staffan Nyström, Per Runeson
J. Softw. Evol. Process.4
2014 Guest editorial: special section on regression testing
Shin Yoo, Per Runeson
Softw. Qual. J.2
2013 IR in Software Traceability: From a Bird's Eye View
abstract
Background. Several researchers have proposed creating after-the-fact structure among software artifacts using trace recovery based on Information Retrieval (IR). Due to significant variation points in previous studies, results are not easily aggregated. Aim. We aim at an overview picture of the outcome of previous evaluations. Method. Based on a systematic mapping study, we perform a synthesis of published research. Results. Our synthesis shows that there are no empirical evidence that any IR model outperforms another model consistently. We also display a strong dependency between the Precision and Recall (P-R) values and the input datasets. Finally, our mapping of P-R values on the possible output space highlights the difficulty of recovering accurate trace links using naïve cut-off strategies. Conclusion. Based on our findings, we stress the need for empirical evaluations beyond the basic P-R 'race'.
Markus Borg, Per Runeson
ESEM2
2013 Challenges in Flexible Safety-Critical Software Development - An Industrial Qualitative Survey
Jesper Pedersen Notander, Martin Höst, Per Runeson
PROFES3
2013 Test overlay in an emerging software product line - An industrial case study
Emelie Engström, Per Runeson
Inf. Softw. Technol.2
2013 On the reliability of mapping studies in software engineering
Claes Wohlin, Per Runeson, Paulo Anselmo da Mota Silveira Neto, Emelie Engström, Ivan do Carmo Machado, Eduardo Santana de Almeida
J. Syst. Softw.2
2013 A Second Replicated Quantitative Analysis of Fault Distributions in Complex Software Systems
abstract
Background: Software engineering is searching for general principles that apply across contexts, for example, to help guide software quality assurance. Fenton and Ohlsson presented such observations on fault distributions, which have been replicated once. Objectives: We aimed to replicate their study again to assess the robustness of the findings in a new environment, five years later. Method: We conducted a literal replication, collecting defect data from five consecutive releases of a large software system in the telecommunications domain, and conducted the same analysis as in the original study. Results: The replication confirms results on unevenly distributed faults over modules, and that fault proneness distributions persist over test phases. Size measures are not useful as predictors of fault proneness, while fault densities are of the same order of magnitude across releases and contexts. Conclusions: This replication confirms that the uneven distribution of defects motivates uneven distribution of quality assurance efforts, although predictors for such distribution of efforts are not sufficiently precise.
Tihana Galinac Grbac, Per Runeson, Darko Huljenic
IEEE Trans. Software Eng.2
2013 Trends in the Quality of Human-Centric Software Engineering Experiments-A Quasi-Experiment
abstract
Context: Several text books and papers published between 2000 and 2002 have attempted to introduce experimental design and statistical methods to software engineers undertaking empirical studies. Objective: This paper investigates whether there has been an increase in the quality of human-centric experimental and quasi-experimental journal papers over the time period 1993 to 2010. Method: Seventy experimental and quasi-experimental papers published in four general software engineering journals in the years 1992-2002 and 2006-2010 were each assessed for quality by three empirical software engineering researchers using two quality assessment methods (a questionnaire-based method and a subjective overall assessment). Regression analysis was used to assess the relationship between paper quality and the year of publication, publication date group (before 2003 and after 2005), source journal, average coauthor experience, citation of statistical text books and papers, and paper length. The results were validated both by removing papers for which the quality score appeared unreliable and using an alternative quality measure. Results: Paper quality was significantly associated with year, citing general statistical texts, and paper length (p <; 0.05). Paper length did not reach significance when quality was measured using an overall subjective assessment. Conclusions: The quality of experimental and quasi-experimental software engineering papers appears to have improved gradually since 1993.
Barbara A. Kitchenham, Dag I. K. Sjøberg, Tore Dybå, Pearl Brereton, David Budgen, Martin Höst, Per Runeson
IEEE Trans. Software Eng.7
2012 Evaluation of traceability recovery in context: A taxonomy for information retrieval tools
abstract
Background: Development of complex, software intensive systems generates large amounts of information. Several researchers have developed tools implementing information retrieval (IR) approaches to suggest traceability links among artifacts. Aim: We explore the consequences of the fact that a majority of the evaluations of such tools have been focused on benchmarking of mere tool output. Method: To illustrate this issue, we have adapted a framework of general IR evaluations to a context taxonomy specifically for IR-based traceability recovery. Furthermore, we evaluate a previously proposed experimental framework by conducting a study using two publicly available tools on two datasets originating from development of embedded software systems. Results: Our study shows that even though both datasets contain software artifacts from embedded development, the characteristics of the two datasets differ considerably, and consequently the traceability outcomes. Conclusions: To enable replications and secondary studies, we suggest that datasets should be thoroughly characterized in future studies on traceability recovery, especially when they can not be disclosed. Also, while we conclude that the experimental framework provides useful support, we argue that our proposed context taxonomy is a useful complement. Finally, we discuss how empirical evidence of the feasibility of IR-based traceability recovery can be strengthened in future research.
Markus Borg, Per Runeson, Lina Broden
EASE2
2012 It Takes Two to Tango - An Experience Report on Industry - Academia Collaboration
abstract
Industry - academia collaboration is critical for empirical research to exist. However, there are many obstacles in the collaboration process. This paper reports on the experiences gained by the author, in a 2-year collaboration project on software testing which involved on-site work by the researcher in the industry premises. Based on notes, minutes of meetings, and progress reports, the project history is outlined. The project is analyzed, using collaboration models as a frame of reference. We conclude that there must be a balance between company 'pull' and academia 'push' in the collaboration Management support is inevitably a key factor to success, while other factors like cross-cultural skills and interfaces towards key resources also contribute.
Per Runeson
ICST1
2012 Software Product Line Testing - A 3D Regression Testing Problem
abstract
In software product line engineering, testing for regression concerns not only versions, as in one-off product development, but also regression across variants. We propose a 3D process model, with the dimensions of level, version and variant, to help analyze, plan and manage software product line testing. We derive the model from empirical observations of regression testing practice and software product line testing theory and practice, and look forward to see the model evaluated in practitioner-oriented research.
Per Runeson, Emelie Engström
ICST1
2012 Three empirical studies on the agreement of reviewers about the quality of software engineering experiments
abstract
During systematic literature reviews it is necessary to assess the quality of empirical papers. Current guidelines suggest that two researchers should independently apply a quality checklist and any disagreements must be resolved. However, there is little empirical evidence concerning the effectiveness of these guidelines. This paper investigates the three techniques that can be used to improve the reliability (i.e. the consensus among reviewers) of quality assessments, specifically, the number of reviewers, the use of a set of evaluation criteria and consultation among reviewers. We undertook a series of studies to investigate these factors. Two studies involved four research papers and eight reviewers using a quality checklist with nine questions. The first study was based on individual assessments, the second study on joint assessments with a period of inter-rater discussion. A third more formal randomised block experiment involved 48 reviewers assessing two of the papers used previously in teams of one, two and three persons to assess the impact of discussion among teams of different size using the evaluations of the “teams” of one person as a control. For the first two studies, the inter-rater reliability was poor for individual assessments, but better for joint evaluations. However, the results of the third study contradicted the results of Study 2. Inter-rater reliability was poor for all groups but worse for teams of two or three than for individuals. When performing quality assessments for systematic literature reviews, we recommend using three independent reviewers and adopting the median assessment. A quality checklist seems useful but it is difficult to ensure that the checklist is both appropriate and understood by reviewers. Furthermore, future experiments should ensure participants are given more time to understand the quality checklist and to evaluate the research papers.
Barbara A. Kitchenham, Dag I. K. Sjøberg, Tore Dybå, Dietmar Pfahl, Pearl Brereton, David Budgen, Martin Höst, Per Runeson
Inf. Softw. Technol.8
2011 Case Studies Synthesis: Brief Experience and Challenges for the Future
abstract
Synthesis of case studies is different from synthesis of purely quantitative studies, for example, in that sampling and analysis in primary studies have been carried out differently, and that primary results are of a different nature. The objective of this research is to identify what challenges should be considered when choosing and using a method for synthesis of case studies. We collected experience from independent synthesis of two published case studies (on trust in outsourcing) by two teams, one team applied cross-case analysis, the other team applied thematic synthesis. The two teams reached both supporting and complimentary conclusions. Identified challenges relate to the goals and research questions of the cases to be synthesized, the number of case studies, temporal and spatial variations, and access to raw data.
Daniela S. Cruzes, Tore Dybå, Per Runeson, Martin Höst
ESEM3
2011 A Self-assessment Framework for Finding Improvement Objectives with ISO/IEC 29119 Test Standard
Jussi Kasurinen, Per Runeson, Leah Muthoni Riungu, Kari Smolander
EuroSPI2
2011 Improving Regression Testing Transparency and Efficiency with History-Based Prioritization - An Industrial Case Study
abstract
Background: History based regression testing was proposed as a basis for automating regression test selection, for the purpose of improving transparency and test efficiency, at the function test level in a large scale software development organization. Aim: The study aims at investigating the current manual regression testing process as well as adopting, implementing and evaluating the effect of the proposed method. Method: A case study was launched including: identification of important factors for prioritization and selection of test cases, implementation of the method, and a quantitative and qualitative evaluation. Results: 10 different factors, of which two are history-based, are identified as important for selection. Most of the information needed is available in the test management and error reporting systems while some is embedded in the process. Transparency is increased through a semi-automated method. Our quantitative evaluation indicates a possibility to improve efficiency, while the qualitative evaluation supports the general principles of history-based testing but suggests changes in implementation details.
Emelie Engström, Per Runeson, Andreas Ljung
ICST2
2011 Usage of Open Source in Commercial Software Product Development - Findings from a Focus Group Meeting
Martin Höst, Alma Orucevic-Alagic, Per Runeson
PROFES3
2011 A Factorial Experimental Evaluation of Automated Test Input Generation - - Java Platform Testing in Embedded Devices
Per Runeson, Per Heed, Alexander Westrup
PROFES1
2011 Software product line testing - A systematic mapping study
Emelie Engström, Per Runeson
Inf. Softw. Technol.2
2011 ICST 2009 Special Issue
abstract
This special issue contains extended versions of four papers from the second IEEE International Conference on Software Testing Verification and Validation (ICST 2009). These four papers were selected based on the reviews from members of the program committee and subsequently subjected to additional rounds of review and revision. We issued 10 invitations to this special issue. Three authors declined to submit, two were not selected after submission, one chose not to revise the paper after being reviewed, and four papers were eventually accepted. The first paper is Reducing Logic Test Set Size While Preserving Fault Detection, by Kaminski and Ammann. This paper introduces a new logic criterion, Minimal-MUMCUT, which has less overlap in terms of faults detected than previous logic criteria, and thus requires fewer tests. It is also provably stronger than the widely used MCDC. The second paper is Improving Penetration Testing through Static and Dynamic Analysis, by Halfond, Choudhary, and Orso. Penetration testing seeks vulnerabilities in software by simulating attacks. Halfond et al. use new software analysis techniques to improve penetration testing. The conference version won the best paper award at ICST 2009. The third paper is An Approach for Testing Pointcut Descriptiors in AspectJ, by Romain Delamare, Baudry, Ghosh, Gupta, and Le Traon. Aspect-oriented programming uses ‘crosscutting concerns’ to increase modularity in software, by encapsulating pieces of software that appear in separate units into their own objects. Delamare et al. have invented testing techniques to find faults in the descriptions of the aspects. The fourth paper is JDAMA: Java Database Application Mutation Analyzer, by Zhou and Frankl. Modern software uses databases much more than in the past, opening up a place for software faults to appear. Database queries are often long, complicated logical expressions, with lots of potential for faults. When the faults are subtle, most queries will return the correct information, and when they do fail, the tables that are returned often look correct. Zhou and Frankl extend a previous mutation-based approach to use analysis and instrumentation of bytecode. All four papers in the ICST 2009 special issue are partially drawn from each of the first authors' PhD Dissertations. The ICST organizers take this as a strong measure of success of the conference. Six years ago, if a student asked ‘where is the best place to publish a paper on testing software?’ the answer was far from clear. The field had several special purpose testing conferences, and many conferences included testing as one topic, but not the main topic. We also had conferences that were intentionally kept very small, even as field was growing. The conference that published the most testing papers was the International Symposium on Software Reliability Engineering, even though its main topic was not testing. It was clear that we had a real need; a need for a large-scale, high quality, broad conference that included all facets of testing, verification, and validation. Thus, a small group (initially Anneliese Andrews, Lionel Briand, and Jeff Offutt) started developing ideas for a new conference in the summer of 2005. This group grew quickly to 30 or 40 people who shared this vision of a new major conference in software testing. A steering committee was elected from that group in 2007 (Andrews, Baudry, Briand, Harman, Hierons, Le Traon, Mathur, Offutt, Williams) and a preliminary charter was approved in 2007, followed by a proposal to the IEEE Computer Society in 2007. The first ICST conference was held in Lillehammer, Norway in 2008, the second in Denver, USA in 2009, the third in Paris, France in 2010, and the fourth in Berlin, Germany in 2011. Because of the obvious overlap, STVR committed to offering a special issue to ICST on a yearly basis. The special issue for ICST 2008 appeared earlier this year, and the special issues for ICST 2010 and 2011 should appear next year. On a final note, we express our gratitude to the many hard working people who helped organize ICST 2009 and this special issue. The steering committee members provided valuable advice and we were assisted by industry chairs Wolfgang Grieskamp, Christian Zapf, and Robert Binder, and general chair Anneliese Andrews. The success of ICST 2009 and the high quality papers in this special issue is due to the efforts of these individuals as well as the Technical Program Committee for ICST 2009 and the additional reviewers for this special issue. We also appreciate the hard work from the ICST 2009 local arrangements chair, Susanne Sherba, the web chair, Orest Pilskalns, and our publicity chairs, Robert Feldt, Jens Krink, Shaoying Liu, Jose Maldonado, and Tao Xie. And last but certainly not least, we thank all the authors who spent valuable time in preparing papers for ICST 2009 and this special issue. 7 July 2011
A. Jefferson Offutt, Per Runeson
Softw. Test. Verification Reliab.2
2010 Can we evaluate the quality of software engineering experiments?
abstract
Context: The authors wanted to assess whether the quality of published human-centric software engineering experiments was improving. This required a reliable means of assessing the quality of such experiments. Aims: The aims of the study were to confirm the usability of a quality evaluation checklist, determine how many reviewers were needed per paper that reports an experiment, and specify an appropriate process for evaluating quality. Method: With eight reviewers and four papers describing human-centric software engineering experiments, we used a quality checklist with nine questions. We conducted the study in two parts: the first was based on individual assessments and the second on collaborative evaluations. Results: The inter-rater reliability was poor for individual assessments but much better for joint evaluations. Four reviewers working in two pairs with discussion were more reliable than eight reviewers with no discussion. The sum of the nine criteria was more reliable than individual questions or a simple overall assessment. Conclusions: If quality evaluation is critical, more than two reviewers are required and a round of discussion is necessary. We advise using quality criteria and basing the final assessment on the sum of the aggregated criteria. The restricted number of papers used and the relatively extensive expertise of the reviewers limit our results. In addition, the results of the second part of the study could have been affected by removing a time restriction on the review as well as the consultation process.
Barbara A. Kitchenham, Dag I. K. Sjøberg, Pearl Brereton, David Budgen, Tore Dybå, Martin Höst, Dietmar Pfahl, Per Runeson
ESEM8
2010 An Empirical Evaluation of Regression Testing Based on Fix-Cache Recommendations
abstract
Background: The fix-cache approach to regression test selection was proposed to identify the most fault-prone files and corresponding test cases through analysis of fixed defect reports. Aim: The study aims at evaluating the efficiency of this approach, compared to the previous regression test selection strategy in a major corporation, developing embedded systems. Method: We launched a post-hoc case study applying the fix-cache selection method during six iterations of development of a multi-million LOC product. The test case execution was monitored through the test management and defect reporting systems of the company. Results: From the observations, we conclude that the fix-cache method is more efficient in four iterations. The difference is statistically significant at alpha = 0.05. Conclusions: The new method is significantly more efficient in our case study. The study will be replicated in an environment with better control of the test execution.
Emelie Engström, Per Runeson, Greger Wikstrand
ICST2
2010 A Qualitative Survey of Regression Testing Practices
Emelie Engström, Per Runeson
PROFES2
2010 Challenges in Aligning Requirements Engineering and Verification in a Large-Scale Industrial Context
Giedre Sabaliauskaite, Annabella Loconsole, Emelie Engström, Michael Unterkalmsteiner, Björn Regnell, Per Runeson, Tony Gorschek, Robert Feldt
REFSQ6
2010 A systematic review on regression test selection techniques
Emelie Engström, Per Runeson, Mats Skoglund
Inf. Softw. Technol.2
2009 Reference-based search strategies in systematic reviews
Mats Skoglund, Per Runeson
EASE2
2009 Tutorial: Case Studies in Software Engineering
Per Runeson, Martin Höst
PROFES1
2009 Guidelines for conducting and reporting case study research in software engineering
abstract
Case study is a suitable research methodology for software engineering research since it studies contemporary phenomena in its natural context. However, the understanding of what constitutes a case study varies, and hence the quality of the resulting studies. This paper aims at providing an introduction to case study methodology and guidelines for researchers conducting case studies and readers studying reports of such studies. The content is based on the authors’ own experience from conducting and reading case studies. The terminology and guidelines are compiled from different methodology handbooks in other research domains, in particular social science and information systems, and adapted to the needs in software engineering. We present recommended practices for software engineering case studies as well as empirically derived and evaluated checklists for researchers and readers of case study research.
Per Runeson, Martin Höst
Empir. Softw. Eng.1
2008 Empirical evaluations of regression test selection techniques: a systematic review
abstract
Regression testing is the verification that previously functioning software remains after a change. In this paper we report on a systematic review of empirical evaluations of regression test selection techniques, published in major software engineering journals and conferences. Out of 2,923 papers analyzed in this systematic review, we identified 28 papers reporting on empirical comparative evaluations of regression test selection techniques. They report on 38 unique studies (23 experiments and 15 case studies), and in total 32 different techniques for regression test selection are evaluated. Our study concludes that no clear picture of the evaluated techniques can be provided based on existing empirical evidence, except for a small group of related techniques. Instead, we identified a need for more and better empirical studies were concepts are evaluated rather than small variations. It is also necessary to carefully consider the context in which studies are undertaken.
Emelie Engström, Mats Skoglund, Per Runeson
ESEM3
2007 Investigating Test Teams' Defect Detection in Function test
abstract
Art. 4343715
Carina Andersson 0001, Per Runeson
ESEM2
2007 Checklists for Software Engineering Case Study Research
abstract
Case study is an important research methodology for software engineering. We have identified the need for checklists supporting researchers and reviewers in conducting and reviewing case studies. We derived checklists for researchers and reviewers respectively, using systematic qualitative procedures. Based on nine sources on case studies, checklists are derived and validated, and hereby presented for further use and improvement.
Martin Höst, Per Runeson
ESEM2
2007 Detection of Duplicate Defect Reports Using Natural Language Processing
abstract
Defect reports are generated from various testing and development activities in software engineering. Sometimes two reports are submitted that describe the same problem, leading to duplicate reports. These reports are mostly written in structured natural language, and as such, it is hard to compare two reports for similarity with formal methods. In order to identify duplicates, we investigate using natural language processing (NLP) techniques to support the identification. A prototype tool is developed and evaluated in a case study analyzing defect reports at Sony Ericsson mobile communications. The evaluation shows that about 2/3 of the duplicates can possibly be found using the NLP techniques. Different variants of the techniques provide only minor result differences, indicating a robust technology. User testing shows that the overall attitude towards the technique is positive and that it has a growth potential.
Per Runeson, Magnus Alexandersson, Oskar Nyholm
ICSE1
2007 Improving Class Firewall Regression Test Selection by Removing the Class Firewall
abstract
One regression test selection technique proposed for object-oriented programs is the Class firewall regression test selection technique. The selection technique selects test cases for regression test, which test changed classes and classes depending on changed classes. However, in empirical studies of the application of the technique, we observed that another technique found the same defects, selected fewer tests and required a simpler, less costly, analysis. The technique, which we refer to as the Change-based regression test selection technique, is basically the Class firewall technique, but with the class firewall removed. In this paper we formulate a hypothesis stating that these empirical observations are not incidental, but an inherent property of the Class firewall technique. We prove that the hypothesis holds for Java in a stable testing environment, and conclude that the effectiveness of the Class firewall regression testing technique can be improved without sacrificing the defect detection capability of the technique, by removing the class firewall.
Mats Skoglund, Per Runeson
Int. J. Softw. Eng. Knowl. Eng.2
2007 A Replicated Quantitative Analysis of Fault Distributions in Complex Software Systems
abstract
To contribute to the body of empirical research on fault distributions during development of complex software systems, a replication of a study of Fenton and Ohlsson is conducted. The hypotheses from the original study are investigated using data taken from an environment that differs in terms of system size, project duration, and programming language. We have investigated four sets of hypotheses on data from three successive telecommunications projects: 1) the Pareto principle, that is, a small number of modules contain a majority of the faults (in the replication, the Pareto principle is confirmed), 2) fault persistence between test phases (a high fault incidence in function testing is shown to imply the same in system testing, as well as prerelease versus postrelease fault incidence), 3) the relation between number of faults and lines of code (the size relation from the original study could be neither confirmed nor disproved in the replication), and 4) fault density similarities across test phases and projects (in the replication study, fault densities are confirmed to be similar across projects). Through this replication study, we have contributed to what is known on fault distributions, which seem to be stable across environments.
Carina Andersson 0001, Per Runeson
IEEE Trans. Software Eng.2
2006 Simulation of Experiments for Data Collection - a Replicated Study
abstract
Simulations can be used as a means for extension of data collection from empirical studies. A simulation model is developed, based on the data from experiments, and new data is generated from the simulation model. This paper replicates an initial investigation by Münch and Armbrust, with the purpose of evaluating the generality of their approach. We replicate their study using data from two inspection experiments. We conclude that the replicated study corroborates the original one. The deviation between the detection rate of the underlying experiment and the simulation models was 2% for the original study and is 4% in the replicated study. Both figures are acceptable for using the approach further. Still the model is based on some adjustment variables that are not directly possible to interpret in terms of the original experiment, and hence the model is subject to improvement.
Per Runeson, Mathias Wiberg
EASE1
2006 Software Process Improvement - EuroSPI 2006 Conference
Richard Messnarz, Ita Richardson, Per Runeson
EuroSPI3
2006 Integrating agile software development into stage-gate managed product development
Daniel Karlström, Per Runeson
Empir. Softw. Eng.2
2005 A Framework for Design Tradeoffs
Anneliese Amschler Andrews, Ed Mancebo, Per Runeson, Robert B. France
Softw. Qual. J.3
2005 A minimal test practice framework for emerging software organizations
abstract
Testing takes a large share of software development efforts, and hence is of interest when seeking improvements. Several test process improvement frameworks exist, but they are extensive and much too large to be effective for smaller organizations. This paper presents a minimal test practice framework (MTPF) that allows the incremental introduction of appropriate practices at the appropriate time in rapidly expanding organizations. The process for introducing the practice framework tries to minimize resistance to change by maximizing the involvement of the entire organization in the improvement effort and ensuring that changes are made in small steps with a low threshold for each step. The practice framework created and its method of introduction have been evaluated at one company by applying the framework for a one-year period. Twelve local software development companies have also evaluated the framework in a survey. Copyright © 2005 John Wiley & Sons, Ltd.
Daniel Karlström, Per Runeson, Sara Nordén
Softw. Test. Verification Reliab.2
2004 A Case Study on Regression Test Suite Maintenance in System Evolution
abstract
When a system is maintained, its automated test suites must also be maintained to keep the tests up to date. Even though practice indicates that test suite maintenance can be very costly we have seen few studies considering the actual efforts for maintenance of test-ware. We conducted a case study on an evolving system with three updated versions, changed with three different change strategies. Test suites for automated unit and functional tests were used for regression testing the extended applications. With one change strategy more changes were made in the tests code than in the system that was tested, and with another strategy no changes were needed for the unit tests to work.
Mats Skoglund, Per Runeson
ICSM2
2004 Are Found Defects an Indicator of Software Correctness? An Investigation in a Controlled Case Study
abstract
In quality assurance programs, we want indicators of software quality, especially software correctness. The number of found defects during inspection and testing are often used as the basis for indicators of software correctness. However, there is a paradox in this approach, since the remaining defects is what impacts negatively on software correctness, not the found ones. In order to investigate the validity of using found defects or other product or process metrics as indicators of software correctness, a controlled case study is launched. 57 sets of 10 different programs from the PSP course are assessed using acceptance test suites for each program. In the analysis, the number of defects found during the acceptance test are compared to the number of defects found during development, code size, share of development time spent on testing etc. It is concluded from a correlation analysis that 1) fewer defects remain in larger programs 2) more defects remain when larger share of development effort is spent on testing, and 3) no correlation exist between found defects and correctness. We interpret these observations as 1) the smaller programs do not fulfill the expected requirements 2) that large share effort spent of testing indicates a "hacker" approach to software development, and 3) more research is needed to elaborate this issue.
Per Runeson, Måns Holmstedt Jönsson, Fredrik Scheja
ISSRE1
2004 Evaluation of Usage-Based Reading-Conclusions after Three Experiments
Thomas Thelin, Per Runeson, Claes Wohlin, Thomas Olsson 0001, Carina Andersson 0001
Empir. Softw. Eng.2
2004 Capture-recapture in software inspections after 10 years research--theory, evaluation and application
Håkan Petersson, Thomas Thelin, Per Runeson, Claes Wohlin
J. Syst. Softw.3
2004 Applying sampling to improve software inspections
Thomas Thelin, Håkan Petersson, Per Runeson, Claes Wohlin
J. Syst. Softw.3
2003 Detection or Isolation of Defects? An Experimental Comparison of Unit Testing and Code Inspection
abstract
Code inspections and white-box testing have both been used for unit testing. One is a static analysis technique, the other, a dynamic one, since it is based on executing test cases. Naturally, the question arises whether one is superior to the other, or, whether either technique is better suited to detect or isolate certain types of defects. We investigated this question with an experiment with a focus on detection of the defects (failures) and isolation of the underlying sources of the defects (faults). The results indicate that there exist significant differences for some of the effects of using code inspection versus testing. White-box testing is more effective, i.e. detects significantly more defects while inspection isolates the underlying source of a larger share of the defects detected. Testers spend significantly more time, hence the difference in efficiency is smaller, and is not statistically significant. The two techniques are also shown to detect and identify different defects, hence motivating the use of a combination of methods.
Per Runeson, Anneliese Amschler Andrews
ISSRE1
2003 Scaling Extreme Programming in a Market Driven Development Context
Daniel Karlström, Per Runeson
XP2
2003 Test processes in software product evolution - a qualitative survey on the state of practice
abstract
Abstract In order to understand the state of test process practices in the software industry, we have conducted a qualitative survey, covering software development departments at 11 companies in Sweden of different sizes and application domains. The companies develop products in an evolutionary manner, which means either new versions are released regularly, or new product variants under new names are released. The survey was conducted through workshop and interview sessions, loosely guided by a questionnaire scheme. The main conclusions of the survey are that the documented development process is emphasized by larger organizations as a key asset, while smaller organizations tend to lean more on experienced people. Further, product volution is performed primarily as new product variants for embedded systems, and as new versions for packaged software. The development is structured using incremental development or a daily build approach; increments are used among more process‐focused organizations, and daily build is more frequently utilized in less process‐focused organizations. Test automation is performed using scripts for products with focus on functionality, and recorded data for products with focus on non‐functional properties. Test automation is an issue which most organizations want to improve; handling the legacy parts of the product and related documentation presents a common problem in improvement efforts for product evolution. Copyright © 2003 John Wiley & Sons, Ltd.
Per Runeson, Carina Andersson 0001, Martin Höst
J. Softw. Maintenance Res. Pract.1
2003 Efficient Evaluation of Multifactor Dependent System Performance Using Fractional Factorial Design
abstract
Performance of computer-based systems may depend on many different factors, internal and external. In order to design a system to have the desired performance or to validate that the system has the required performance, the effect of the influencing factors must be known. Common methods give no or little guidance on how to vary the factors during prototyping or validation. Varying the factors in all possible combinations would be too expensive and too time-consuming. This paper introduces a systematic approach to the prototyping and the validation of a system's performance, by treating the prototyping or validation as an experiment, in which the fractional factorial design methodology is commonly used. To show that this is possible, a case study evaluating the influencing factors of the false and real target rate of a radar system is described. Our findings show that prototyping and validation of system performance become structured and effective when using the fractional factorial design. The methodology enables planning, performance, structured analysis, and gives guidance for appropriate test cases. The methodology yields not only main factors, but also interacting factors. The effort is minimized for finding the results, due to the methodology. The case study shows that after 112 test cases, of 1024 possible, the knowledge gained was enough to draw conclusions on the effects and interactions of 10 factors. This is a reduction with a factor 5-9 compared to alternative methods.
Tomas Berling, Per Runeson
IEEE Trans. Software Eng.2
2003 An Experimental Comparison of Usage-Based and Checklist-Based Reading
abstract
Software quality can be defined as the customers' perception of how a system works. Inspection is a method to monitor and control the quality throughout the development cycle. Reading techniques applied to inspections help reviewers to stay focused on the important parts of an artifact when inspecting. However, many reading techniques focus on finding as many faults as possible, regardless of their importance. Usage-based reading helps reviewers to focus on the most important parts of a software artifact from a user's point of view. We present an experiment, which compares usage-based and checklist-based reading. The results show that reviewers applying usage-based reading are more efficient and effective in detecting the most critical faults from a user's point of view than reviewers using checklist-based reading. Usage-based reading may be preferable for software organizations that utilize or start utilizing use cases in their software development.
Thomas Thelin, Per Runeson, Claes Wohlin
IEEE Trans. Software Eng.2
2002 Decision support for extreme programming introduction and practice selection
abstract
This paper presents an investigation concerning the introduction of Extreme Programming (XP) in software development organisations. More specifically the concept of using a decision support method known as the Analytical Hierarchy Process (AHP) is evaluated by a group of students and a group of developers and the outcome is compared to experiences from an XP case study. The results provide an indication that different practices are thought to be easier and more effective to implement in the two groups. A company considering implementing only a few practices can use this as help for deciding which practices to implement. Companies introducing all practices can use the results of this kind of method to see where more attention might be needed after or during the introduction of XP.
Daniel Karlström, Per Runeson
SEKE2
2002 Confidence intervals for capture-recapture estimations in software inspections
Thomas Thelin, Per Runeson
Inf. Softw. Technol.2
2001 Experiences from Teaching PSP for Freshmen
abstract
The Personal Software Process (PSP) is launched as a means for improving software development capabilities for the individual engineer. It is proposed that it should be used in software engineering curricula; some authors propose it to be used already during the first student year. The PSP course is successfully given for graduate students at Lund University since 1996. During the spring semester of 1999, it was given to undergraduate students at the software engineering program in their first year of university studies, directly after their first programming course. This paper reports results and experiences from the course given to these freshmen students. A quantitative analysis is conducted to compare the freshmen student data to data from graduate student courses, and a qualitative evaluation is conducted concerning the differences between the courses. It can be concluded that mostly the same improvement trends can be identified with freshmen students as with graduate students. However, the qualitative analysis shows that the freshmen students are more concerned with programming than with the software process issues. Based on the results, it is decided to move the PSP course to the second year in order to enable the programming skills to be better established before the PSP is launched.
Per Runeson
CSEE&T1
2001 Defect Content Estimation for Two Reviewers
abstract
Estimation of the defect content is important to enable quality control throughout the software development process. Capture-recapture methods and curve fitting methods have been suggested as tools to estimate the defect content after a review. The methods are highly reliant on the quality of the data. If the number of reviewers is fairly small, it becomes difficult or even impossible to get reliable estimates. This paper presents a comprehensive study of estimates based on two reviewers, using real data from reviews. Three experience-based defect content estimation methods are evaluated vs. methods that use data only from the current review. Some models are possible to distinguish from each other in terms of statistical significance. In order to gain an even better understanding, the best models are compared subjectively. It is concluded that the experience-based methods provide some good opportunities to estimate the defect content after a review.
Håkan Petersson, Claes Wohlin, Per Runeson, Martin Höst
ISSRE3
2001 A Classification Scheme for Studies on Fault-Prone Components
Per Runeson, Magnus C. Ohlsson, Claes Wohlin
PROFES1
2001 Usage-based readingan experiment to guide reviewers with use cases
Thomas Thelin, Per Runeson, Björn Regnell
Inf. Softw. Technol.2
2000 A New Software Engineering Program - Structure and Initial Experiences
abstract
Software plays an increasingly important role in new products of different kinds. Therefore, the need for engineers developing software is continuously increasing. However computer science education programmes are not enough to fulfil the industrial needs. Software engineering programmes are required with a holistic approach to the software life-cycle and its economics, as well as education in monitoring and managing the software process. At Lund University, Sweden, a new Bachelors software engineering programme was launched 1998. In this paper, the education programme's principles and structures are presented, as well as experiences from the first year. Software processes and methods play an important part in the education programme. It is concluded that the basic principles do function as expected, although the programme must be changed in the direction of requiring more programming skills before the introduction of systematic software processes.
Per Runeson
CSEE&T1
2000 Assessing the Sensitivity to Usage Profile Changes in Test Planning
abstract
Software reliability is an important characteristic for most systems, but due to its dynamic properties it is hard to assess until very late in the development. Nevertheless, the testing must be planned to meet the reliability requirements. In test planning, the notion of usage coverage may be used as an indicator of reliability as they are correlated with each other. When the testing is planned, the test cases to be run and the usage profile are derived. The usage profile is estimated using the available information of the expected usage. There are therefore some uncertainties in the estimated usage profile. The paper presents and evaluates a method for analysing the impact of uncertainties in the usage profile on the usage coverage. The method is applied during test planning to evaluate how sensitive the usage coverage is to usage profile uncertainties. Different usage profiles are simulated using the expected usage profile and an uncertainty expressed as a percentage value or a range within which the changed usage profile may vary. Test cases are derived from the expected usage profile and the resulting usage coverage is estimated for each simulated usage profile. Thus the impact in the usage coverage can be analysed. The presented analysis method is illustrated with an example system. One of the conclusions from this first study of the method is that the sparseness of the usage profile, which is defined by limitations in the usage pattern, reduces the impact of the usage profile uncertainties. Further, the number of test cases is an important factor in the sensitivity to uncertainties in the usage profile.
Anders Wesslén, Per Runeson, Björn Regnell
ISSRE2
2000 An Evaluation of Functional Size Methods and a Bespoke Estimation Method for Real-Time Systems
Per Runeson, Niklas Borgquist, Markus Landin, Wladyslaw Bolanowski
PROFES1
2000 Are the Perspectives Really Different? - Further Experimentation on Scenario-Based Reading of Requirements
Björn Regnell, Per Runeson, Thomas Thelin
Empir. Softw. Eng.2
2000 A survey of lead-time challenges in the development and evolution of distributed real-time systems
Lars Bratthall, Per Runeson, K. Adelswärd, W. Eriksson
Inf. Softw. Technol.2
2000 Towards integration of use case modelling and usage-based testing
Björn Regnell, Per Runeson, Claes Wohlin
J. Syst. Softw.2
2000 Robust estimations of fault content with capture-recapture and detection profile estimators
Thomas Thelin, Per Runeson
J. Syst. Softw.2
1999 Technical Requirements for the Implementation of an Experience Base
Mikael Broomé, Per Runeson
SEKE2
1999 Architecture Design Recovery of a Family of Embedded Software Systems
Lars Bratthall, Per Runeson
WICSA2
1998 Defect Content Estimations from Review Data
abstract
Reviews are essential for defect detection and they provide an opportunity to control the software development process. This paper focuses upon methods for estimating the defect content after a review and hence to provide support for process control. Two new estimation methods are introduced as the assumptions of the existing statistical methods are not fulfilled. The new methods are compared with a maximum-likelihood approach. Data from several reviews are used to evaluate the different methods. It is concluded that the new estimation methods provide new opportunities to estimate the defect content.
Claes Wohlin, Per Runeson
ICSE2
1998 Derivation of an integrated operational profile and use case model
abstract
Requirements engineering and software reliability engineering both involve model building related to the usage of the intended system; requirements models and test case models respectively are built. Use case modelling for requirements engineering and operational profile testing for software reliability engineering are techniques which are evolving into software engineering practice. Approaches towards integration of the use case model and the operational profile model are proposed. By integrating the derivation of the models, effort may be saved in both development and maintenance of software artifacts. Two integration approaches are presented, transformation and extension. It is concluded that the use case model structure can be transformed into an operational profile model adding the profile information. As a next step, the use case model can be extended to include the information necessary for the operational profile. Through both approaches, modelling and maintenance effort as well as risks for inconsistencies can be reduced. A positive spin-off effect is that quantitative information on usage frequencies is available in the requirements, enabling planning and prioritizing based on that information.
Per Runeson, Björn Regnell
ISSRE1
1998 Combining Scenario-based Requirements with Static Verification and Dynamic Testing
Björn Regnell, Per Runeson
REFSQ2
1998 An Experimental Evaluation of an Experience-Based Capture-Recapture Method in Software Code Inspections
Per Runeson, Claes Wohlin
Empir. Softw. Eng.1
1995 An Experimental Evaluation of Capture-Recapture in Software Inspections
abstract
Abstract The use of capture‐recapture to estimate the residual faults in a software artifact has evolved as a promising method. However, the assumptions needed to make the estimates are not completely fulfilled in software development, leading to an underestimation of the residual fault content. Therefore, a method employing a filtering technique with an experience factor to improve the estimate of the residual faults is proposed in this paper. An experimental study of the capture‐recapture method with this correction method has been conducted. It is concluded that the correction method improves the capture‐recapture estimate of the number of residual defects in the inspected document.
Claes Wohlin, Per Runeson, Johan Brantestam
Softw. Test. Verification Reliab.2
1994 Certification of Software Components
abstract
Reuse is becoming one of the key areas in dealing with the cost and quality of software systems. An important issue is the reliability of the components, hence making certification of software components a critical area. The objective of this article is to try to describe methods that can be used to certify and measure the ability of software components to fulfil the reliability requirements placed on them. A usage modelling technique is presented, which can be used to formulate usage models for components. This technique will make it possible not only to certify the components, but also to certify the system containing the components. The usage model describes the usage from a structural point of view, which is complemented with a profile describing the expected usage in figures. The failure statistics from the usage test form the input of a hypothesis certification model, which makes it possible to certify a specific reliability level with a given degree of confidence. The certification model is the basis for deciding whether the component can be accepted, either for storage as a reusable component or for reuse. It is concluded that the proposed method makes it possible to certify software components, both when developing for and with reuse.>
Claes Wohlin, Per Runeson
IEEE Trans. Software Eng.2
1992 A method proposal for early software reliability estimation
abstract
Presents a method proposal for estimation of software reliability before the implementation phase. The method is based upon a formal description technique and that it is possible to develop a tool for performing dynamic analysis, i.e. locating semantic faults in the design. The analysis is performed by applying a usage profile as input as well as doing a full analysis, i.e. locating all faults that the tool can find. The tool must provide failure data in terms of time since the last failure was detected. The mapping of the dynamic failures to the failures encountered during statistical usage testing and operation is discussed. The method can be applied either on the software specification or as a step in the development process by applying it on the design descriptions. The proposed method allows for software reliability estimations that can be used both as a quality indicator, and also for planning and controlling resources, development times etc. at an early stage in the development of software systems.>
Claes Wohlin, Per Runeson
ISSRE2