Miroslaw Staron

dblp:16/5647 · DBLP profile ↗
← Back
101ranked-venue papers
32as first author
26since 2021 · last 2026
0000-0002-9052-0864ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 95 · 29 first-author · 24 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 An Investigation of the AUTOSAR Adaptive Platform from an Industry Perspective
abstract
Context: The reliance on software as a distinguishing factor in the automotive industry is increasing. With a combined reliance on vendor-supplied software and cost-effective implementation, the AUTOSAR consortium was initialized to provide standardized platform specifications that enable re-use. Specifically, the AUTOSAR Adaptive Platform (AP) specification aims to provide a high-performance service-oriented architecture. Objective: The goal of this study is to investigate what pain-points emerge when developing AUTOSAR Adaptive applications and whether they originate from the platform specification, its vendor-implementation, or its local usage. Methods: We conduct a Design Science Research study, developing a minimal AP that serves as an experimental prototype for our investigation. Results: We find that a combination of specification-inherent, implementation-based, and local practices contributes to the emergence of pain-points. Conclusions: We conclude that there are AUTOSAR specification-inherent reasons for pain-points, resulting from architectural choices and re-use goals. The implication for development organizations is the need to mitigate these effects through tooling that better supports configuration file management and reduces developer training time to properly understand the adaptive application runtime life-cycle.
Bengt Haraldsson, Srijita Basu, Miroslaw Staron, Erika Mayer
ICSA3
2026 Evaluating Train-Test Data Leakage in Automotive Image Datasets
abstract
Reliable evaluation of machine learning (ML)-enabled perception systems for intelligent vehicles critically depends on the integrity of training and test datasets. A major risk arises when near-duplicate or visually similar images appear across subsets, leading to inflated performance estimates. This study systematically quantifies train-test similarity in six widely used automotive datasets - KITTI, ZOD, BDD100k, ONCE, Cirrus, and SODA10M - in their default splits. We employ perceptual hashing (pHash) and deep feature embeddings to measure image-level redundancy. Results show 4,513 pairs of KITTI and 562 of ZOD images were almost identical, corresponding to 25% and 5% of their test images, respectively. The other examined datasets contain only at most 3 pairs (almost 0%) of almost identical images in their existing train-test splits. These findings underscore the need for similarity analysis during dataset preparation, particularly for video-based collections with strong spatio-temporal dependencies. By exposing dataset-specific risks of data leakage in popular datasets, this study contributes practical insights for both dataset curators and ML practitioners. These insights are valuable when using public benchmark datasets in safety-critical domains such as autonomous driving (AD).
Md. Abu Ahammed Babu, Miroslaw Staron, András Bálint, Darko Durisic, Sushant Kumar Pandey
IV2
2026 From LLMS to Agents in Programming: The Impact of Providing an LLM with a Compiler
abstract
Large Language Models have demonstrated a remarkable capability in natural language and program generation and software development. However, the source code generated by the LLMs does not always meet quality requirements and may fail to compile. Therefore, many studies evolve into agents that can reason about the problem before generating the source code for the solution. The goal of this paper is to study the degree to which such agents benefit from access to software development tools, in our case, a gcc compiler. We conduct a computational experiment on the RosettaCode dataset, on 699 programming tasks in C. We evaluate how the integration with a compiler shifts the role of the language model from a passive generator to an active agent capable of iteratively developing runnable programs based on feedback from the compiler. We evaluated 16 language models with sizes ranging from small (135 million) to medium (3 billion) and large (70 billion). Our results show that access to a compiler improved the compilation success by 5.3 to 79.4 percentage units in compilation without affecting the semantics of the generated program. Syntax errors dropped by 75 %, and errors related to undefined references dropped by 87 % for the tasks where the agents outperformed the baselines. We also observed that in some cases, smaller models with a compiler outperform larger models with a compiler. We conclude that it is essential for LLMs to have access to software engineering tools to enhance their performance and reduce the need for large models in software engineering, such as reducing our energy footprint.
Viktor Kjellberg, Miroslaw Staron, Farnaz Fotrousi
SANER2
2026 Agentic Pipelines in Embedded Software Engineering: Emerging Practices and Challenges
abstract
A new transformation is underway in software engineering, driven by the rapid adoption of generative AI in development workflows. Similar to how version control systems once automated manual coordination, AI tools are now beginning to automate many aspects of programming. For embedded software engineering organizations, however, this marks their first experience integrating AI into safety-critical and resourceconstrained environments. The strict demands for determinism, reliability, and traceability pose unique challenges for adopting generative technologies. In this paper, we present findings from a qualitative study with ten senior experts from four companies who are evaluating generative AI-augmented development for embedded software. Through semi-structured focus group interviews and structured brainstorming sessions, we identified eleven emerging practices and fourteen challenges related to the orchestration, responsible governance, and sustainable adoption of generative AI tools. Our results show how embedded software engineering teams are rethinking workflows, roles, and toolchains to enable a sustainable transition toward agentic pipelines and generative AI-augmented development.
Simin Sun, Miroslaw Staron
SANER2
2026 How not to get your paper rejected - From the editors' notebook
Miroslaw Staron, Guilherme Horta Travassos, Barbara Russo, Sudipto Ghosh 0001
Inf. Softw. Technol.1
2026 The Impact of Class Noise-handling on the Effectiveness of Machine Learning-based Methods for Build Outcome and Code Change Request Predictions
abstract
Machine learning-based methods are increasingly used to optimize build processes and accelerate the integration of software code. These methods leverage large volumes of historical code changes to train models on predicting and preventing issues in the codebase that could delay code integrations and features delivery to end-users. The objective of this study is to examine the impact of handling class noise present in software code changes collected from Continuous Integration (CI) systems on the predictive performance of machine learning models for predicting the execution outcome of CI builds and negative code reviews. In this study, we conduct a series of computational experiments using data from 110 Java open-source projects, examining the effectiveness of two removal-based statistical techniques - Majority Filter (MF) and Consensus Filter (CF) - and two corrective techniques - Domain Knowledge-based (DB) and CleanLab. Our results show that removal-based techniques significantly improve model predictive performance in both build outcome and negative code review prediction tasks. For build outcome prediction, applying MF increased the F1-score from 82% to 97%, and MCC from 0.13 to 0.58. In negative code review predictions, MF improved the F1-score from 17% to 53%, and MCC from −0.03 to 0.57. The DB technique was effective primarily in the context of code review comments but less so for build outcome predictions. While CleanLab yielded more consistent predictions, its overall impact on model performance was more moderate compared to removal-based techniques. Additionally, our findings show that hyperparameter tuning, applied independently or in combination with CleanLab, can further improve model performance; however, these gains did not surpass those achieved by removal-based techniques alone. We conclude that applying removal-based techniques to the training data of code changes is necessary to improve the prediction of build outcomes and negative code review comments.
Khaled Walid Al-Sabbagh, Miroslaw Staron, Regina Hebig
ACM Trans. Softw. Eng. Methodol.2
2026 Literate Programming With LLMs? - A Study on Rosetta Code and CodeNet
abstract
Literate programming, a concept introduced by Knuth in 1984, emphasized the importance of combining human-readable documentation with machine-readable code as writing literate programs is a prerequisite for software quality. Our objective with this paper is to evaluate whether generative AI models, Large Language Models (LLM) like GPT-4, LLaMA or Falcon, are capable of literate programming because of their extensive use in software engineering. To truly achieve literate programming, LLMs must generate natural language descriptions and corresponding code with aligned semantics based on user prompts. In addition, their internal representation of programs should allow us to recognize both programming languages and their descriptions. To evaluate their capabilities, we conducted a study using the Rosetta Code and CodeNet repositories. We perform four computational experiments using the Rosetta Code repository, encompassing 1,228 tasks across 926 programming languages, and validate our findings on the larger CodeNet dataset, which includes 55 tasks and 52 languages. Our findings show that LLMs in the trillion-parameter class are capable of literate programming, while models in the million- and billion-parameter classes are better at recognizing programming languages than tasks. Based on these results, we conclude that modern LLMs inhibit a deeper ability to encode programming languages and the semantics of programming tasks, bringing us closer to realizing the full potential of literate programming.
Simin Sun, Miroslaw Staron
IEEE Trans. Software Eng.2
2025 Aspects of complexity in automotive software systems and their relation to maintainability effort. A case study
abstract
Context: Large embedded systems in vehicles tend to grow in size and complexity, which causes challenges when maintaining these systems. Objective: We explore how developers perceive the relation between maintainability effort and various sources of complexity. Methods: We conduct a case study at Scania AB, a heavy vehicle OEM. The units of analysis are two large software systems and their development teams/organizations. Results: Our results show that maintainability effort is driven by system internal complexity in the form of variant management and complex hardware control tasks. The maintainability is also influenced by emergent complexity caused by the system’s longevity and constant growth. Besides these system-internal complexities, maintainability effort is also influenced by external complexities, such as organizational coordination and business needs. During the study, developer trade-off strategies for minimizing maintainability effort emerged. Conclusions: Complexity is a good proxy of maintainability effort, and allows developers to create strategies for managing the maintainability effort. Adequate complexity metrics include both external aspects—e.g., coordination complexity—and internal ones—e.g., McCabe Cyclomatic Complexity.
Bengt Haraldsson, Miroslaw Staron
EASE2
2025 An Empirical study on LLM-based Log Retrieval for Software Engineering Metadata Management
abstract
Developing autonomous driving systems (ADSs) involves generating and storing extensive log data from test drives, which is essential for verification, research, and simulation. However, these high-frequency logs, recorded over varying durations, pose challenges for developers attempting to locate specific driving scenarios. This difficulty arises due to the wide range of signals representing various vehicle components and driving conditions, as well as unfamiliarity of some developers’ with the detailed meaning of these signals. Traditional SQL-based querying exacerbates this challenge by demanding both domain expertise and database knowledge, often yielding results that are difficult to verify for accuracy.
Simin Sun, Yuchuan Jin, Miroslaw Staron
EASE3
2025 Variant Management Impact on Architectural Maintainability in Embedded Systems - A Case Study
Bengt Haraldsson, Miroslaw Staron
ECSA2
2025 "Good" and "Bad" Failures in Industrial CI/CD-Balancing Cost and Quality Assurance
Simin Sun, David Friberg, Miroslaw Staron
SEAA (3)3
2025 D-LeDe: A Data Leakage Detection Method for Automotive Perception Systems
abstract
Data leakage is a very common problem that is often overlooked during splitting data into train and test sets before training any ML/DL model. The model performance gets artificially inflated with the presence of data leakage during the evaluation phase which often leads the model to erroneous prediction on real-time deployment. However, detecting the presence of such leakage is challenging, particularly in the object detection context of perception systems where the model needs to be supplied with image data for training. In this study, we conduct a computational experiment to develop a method for detecting data leakage. We then conducted an initial evaluation of the method as a first step on a public dataset, “Kitti”, which is a popular and widely accepted benchmark dataset in the automotive domain. The evaluation results show that our proposed D-LeDe method are able to successfully detect potential data leakage caused by image similarity. A further validation was also provided to justify the evaluation outcome by conducting pair-wise image similarity analysis using perceptual hash (pHash) distance.
Md. Abu Ahammed Babu, Sushant Kumar Pandey, Darko Durisic, Ashok Chaitanya Koppisetty, Miroslaw Staron
VEHITS5
2025 Design pattern recognition: a study of large language models
abstract
Abstract Context As Software Engineering (SE) practices evolve due to extensive increases in software size and complexity, the importance of tools to analyze and understand source code grows significantly. Objective This study aims to evaluate the abilities of Large Language Models (LLMs) in identifying DPs in source code, which can facilitate the development of better Design Pattern Recognition (DPR) tools. We compare the effectiveness of different LLMs in capturing semantic information relevant to the DPR task. Methods We studied Gang of Four (GoF) DPs from the P-MARt repository of curated Java projects. State-of-the-art language models, including Code2Vec, CodeBERT, CodeGPT, CodeT5, and RoBERTa, are used to generate embeddings from source code. These embeddings are then used for DPR via a k-nearest neighbors prediction. Precision, recall, and F1-score metrics are computed to evaluate performance. Results RoBERTa is the top performer, followed by CodeGPT and CodeBERT, which showed mean F1 Scores of 0.91, 0.79, and 0.77, respectively. The results show that LLMs without explicit pre-training can effectively store semantics and syntactic information, which can be used in building better DPR tools. Conclusion The performance of LLMs in DPR is comparable to existing state-of-the-art methods but with less effort in identifying pattern-specific rules and pre-training. Factors influencing prediction performance in Java files/programs are analyzed. These findings can advance software engineering practices and show the importance and abilities of LLMs for effective DPR in source code.
Sushant Kumar Pandey, Sivajeet Chand, Jennifer Horkoff, Miroslaw Staron, Miroslaw Ochodek, Darko Durisic
Empir. Softw. Eng.4
2024 Using Generative AI to Support Standardization Work - the Case of 3GPP
abstract
Standardization processes build upon consensus between partners, which depends on their ability to identify points of disagreement and resolving them. Large standardization organizations, like the 3GPP or ISO, rely on leaders of work packages who can correctly, and efficiently, identify disagreements, discuss them and reach a consensus. This task, however, is effort-, labor-intensive and costly. In this paper, we address the problem of identifying similarities, dissimilarities and discussion points using large language models. In a design science research study, we work with one of the organizations which leads several workgroups in the 3GPP standard. Our goal is to understand how well the language models can support the standardization process in becoming more cost-efficient, faster and more reliable. Our results show that generic models for text summarization correlate well with domain expert's and delegate's assessments (Pearson correlation between 0.66 and 0.98), but that there is a need for domain-specific models to provide better discussion materials for the standardization groups.
Miroslaw Staron, Jonathan Ström, Albin Karlsson, Wilhelm Meding
SEAA1
2023 TransDPR: Design Pattern Recognition Using Programming Language Models
abstract
Current Design Pattern Recognition (DPR) methods have limitations, such as the reliance on semantic information, limited recognition of novel or modified pattern versions, and other factors. We present an introductory DPR technique by using a Programming Language Model (PLM) called TransDPR, which utilizes a Facebook pre-trained model (TransCoder), which is a Cross-lingual programming Language Model (XLM) based on a transformer architecture. We leverage an n-dimensional vector representation of programs and apply logistic regression to learn design patterns (DPs). Our approach utilizes the GitHub repository to collect singleton and prototype DP programs written in$C$++ source code. Our results indicate that TransDPR achieves 90% accuracy and an F1-score of 0.88 on open-source projects. We evaluate the proposed model on two developed modules from Volvo Cars and invite the original developers to validate the prediction results.
Sushant Kumar Pandey, Miroslaw Staron, Jennifer Horkoff, Miroslaw Ochodek, Nicholas Mucci, Darko Durisic
ESEM2
2023 Defect Backlog Size Prediction for Open-Source Projects with the Autoregressive Moving Average and Exponential Smoothing Models
abstract
Context: predicting the number of defects in a defect backlog in a given time horizon can help allocate project resources and organize software development.Goal: to compare the accuracy of three defect backlog prediction methods in the context of large open-source (OSS) projects, i.e., ARIMA, Exponential Smoothing (ETS), and the state-of-the-art method developed at Ericsson AB (MS).Method: we perform a simulation study on a sample of 20 open-source projects to compare the prediction accuracy of the methods.Also, we use the Naïve prediction method as a baseline for sanity check.We use statistical inference tests and effect size coefficients to compare the prediction errors.Results: ARIMA, ETS, and MS were more accurate than the Naïve method.Also, the prediction errors were statistically lower for ETS than for MS (however, the effect size was negligible).Conclusions: ETS seems slightly more accurate than MS when predicting defect backlog size of OSS projects.
Paulina Aniola, Sushant Kumar Pandey, Miroslaw Staron, Miroslaw Ochodek
FedCSIS3
2023 Comparing Machine Learning Algorithms for Medical Time-Series Data
Alex Helmersson, Faton Hoti, Sebastian Levander, Aliasgar Shereef, Emil Svensson, Ali El-Merhi, Richard Vithal, Jaquette Liljencrantz, Linda Block, Helena Odenstedt Hergés, Miroslaw Staron
PROFES (1)11
2023 Design Patterns Understanding and Use in the Automotive Industry: An Interview Study
Sushant Kumar Pandey, Sivajeet Chand, Jennifer Horkoff, Miroslaw Staron
PROFES (1)4
2023 Editorial
Miroslaw Staron
Inf. Softw. Technol.1
2022 Comparing Input Prioritization Techniques for Testing Deep Learning Algorithms
abstract
Deep learning (DL) systems are becoming an essential part of software systems, so it is necessary to test them thoroughly. This is a challenging task since the test sets can grow over time as the new data is being acquired, and it becomes time-consuming. Input prioritization is necessary to reduce the testing time since prioritized test inputs are more likely to reveal the erroneous behavior of a DL system earlier during test execution. Input prioritization approaches have been rudimentary analyzed against each other, this study compares different input prioritization techniques regarding their effectiveness and efficiency. This work considers surprise adequacy, autoencoder-based, and similarity-based input prioritization approaches in the example of testing a DL image classification algorithms applied on MNIST, Fashion-MNIST, CIFAR-10, and STL-10 datasets. To measure effectiveness and efficiency, we use a modified APFD (Average Percentage of Fault Detected), and set up & execution time, respectively. We observe that the surprise adequacy is the most effective (0.785 to 0.914 APFD). The autoencoder-based and similarity-based techniques are less effective, with the performance from 0.532 to 0.744 APFD and 0.579 to 0.709 APFD, respectively. In contrast, the similarity-based and surprise adequacy-based approaches are the most and least efficient, respectively. The findings in this work demonstrate the trade-off between the considered input prioritization techniques to understanding their practical applicability for testing DL algorithms.
Vasilii Mosin, Miroslaw Staron, Darko Durisic, Francisco Gomes de Oliveira Neto, Sushant Kumar Pandey, Ashok Chaitanya Koppisetty
SEAA2
2022 Improving Software Regression Testing Using a Machine Learning-Based Method for Test Type Selection
Khaled Walid Al-Sabbagh, Miroslaw Staron, Regina Hebig
PROFES2
2022 Improving test case selection by handling class and attribute noise
abstract
Big data and machine learning models have been increasingly used to support software engineering processes and practices. One example is the use of machine learning models to improve test case selection in continuous integration. However, one of the challenges in building such models is the large volume of noise that comes in data, which impedes their predictive performance. In this paper, we address this issue by studying the effect of two types of noise, called class and attribute, on the predictive performance of a test selection model. For this purpose, we analyze the effect of class noise by using an approach that relies on domain knowledge for relabeling contradictory entries and removing duplicate ones. Thereafter, an existing approach from the literature is used to experimentally study the effect of attribute noise removal on learning. The analysis results show that the best learning is achieved when training a model on class-noise cleaned data only — irrespective of attribute noise. Specifically, the learning performance of the model reported 81% precision, 87% recall, and 84% f-score compared with 44% precision, 17% recall, and 25% f-score for a model built on uncleaned data. Finally, no causality relationship between attribute noise removal and the learning of a model for test case selection was drawn.
Khaled Walid Al-Sabbagh, Miroslaw Staron, Regina Hebig
J. Syst. Softw.2
2021 Understanding Metrics Team-Stakeholder Communication in Agile Metrics Service Delivery
abstract
In this paper, we explore challenges in communication between metrics teams and stakeholders in metrics service delivery. Drawing on interviews and interactive workshops with team members and stakeholders at two different Swedish agile software development organizations, we identify interrelated challenges such as aligning expectations, prioritizing demands, providing regular feedback, and maintaining continuous dialogue, which influence team-stakeholder interaction, relationships and performance. Our study shows the importance of understanding communicative hurdles and provides suggestions for their mitigation, therefore meriting further empirical research.
Nataliya Berbyuk Lindström, Dina Koutsikouri, Miroslaw Staron, Wilhelm Meding, Ola Soder
APSEC3
2021 A Method for Modeling Data Anomalies in Practice
abstract
As technology has allowed us to collect large amounts of industrial data, it has become critical to analyze and understand the data collected, in particular to find data anomalies. Anomaly analysis allows a company to detect, analyze and understand anomalous or unusual data patterns. This is an important activity to understand, for example, deviations in service which may indicate potential problems, or differing customer behavior which may reveal new business opportunities. Much previous work has focused on anomaly detection, in particular using machine learning. Such approaches allow clustering of data patterns by common attributes, and, although useful, clusters often do not correspond to the root causes of anomalies, meaning that more manual analysis is needed. In this paper we report on a design science study with two different teams, in a partner company which focuses on modeling and understanding the attributes and root causes of data anomalies. After iteration, for each team, we have created general and anomaly-specific UML class diagrams and goal models to capture anomaly details. We use our experiences to create an example taxonomy, classifying anomalies by their root causes, and to create a general method for modeling and understanding data anomalies. This work paves the way for a better understanding of anomalies and their root causes, leading towards creating a training set which may be used for machine learning approaches.
Jennifer Horkoff, Miroslaw Staron, Wilhelm Meding
SEAA2
2021 Software engineering and advanced applications conference 2019 - selected papers
Rafael Capilla, Miroslaw Staron
Inf. Softw. Technol.2
2021 MeTeaM - A method for characterizing mature software metrics teams
abstract
Metrics teams play an increasingly important role in handling data and information in modern software development organizations; they manage their companies’ measurement programs, collect and process data, and develop and distribute information products. Metrics teams can comprise several roles, and their set-up can differ between companies, as can the metrics maturity of host organizations. These differences impact the effectiveness and quality of a team’s measurement program. Our objective was to design and evaluate a model to describe the characteristics of a mature metrics team, which efficiently designs, develops, maintains, and evolves its organization’s measurement program. We conducted an action research study on four metrics teams of four distinct companies. We designed and evaluated a domain-specific model for assessing the maturity of metrics teams – MeTeaM – and also assessed the four metrics teams per se. Our results were two-fold: the creation of the metrics team maturity model MeTeaM and a template to assess metrics teams. Our evaluation showed that the model captures the characteristics of successful metrics teams and quantifies the maturity status of both the metrics teams and their host organizations. More mature metrics teams score higher in the MeTeaM model than less mature teams. The assessment provides less mature metrics teams with valuable insights on what factors to improve. Such insights can be shared with and acted upon successfully with their organizations.
Wilhelm Meding, Miroslaw Staron, Ola Soder
J. Syst. Softw.2
2020 Making Software Measurement Standards Understandable
abstract
Every discipline, e.g. medicine and engineering, has its own vocabulary to describe situations and tools. This dedicated language is important, because it allows for being specific, detailed and precise. On the other hand, this language, specific to each discipline, becomes a barrier for communication across disciplines. International software measurement standards are examples of such language. The standards are documents that provide definitions of terms used and describe processes specific to the discipline of software measurement. However, one major problem we have observed is that standards are difficult to read and to understand; even for the stakeholders that they are intended for. In this paper, we present our experience report from introducing standards to the work of a large software development organization at an infrastructure provider company. We describe one of the concepts that we found essential in making the measurement standards understandable, namely the notion of the indicator.
Wilhelm Meding, Miroslaw Staron
EASE2
2020 Improving Data Quality for Regression Test Selection by Reducing Annotation Noise
abstract
Big data and machine learning models have been increasingly used to support software engineering processes and practices. One example is the use of machine learning models to improve test case selection in continuous integration. However, one of the challenges in building such models is the identification and reduction of noise that often comes in large data. In this paper, we present a noise reduction approach that deals with the problem of contradictory training entries. We empirically evaluate the effectiveness of the approach in the context of selective regression testing. For this purpose, we use a curated training set as input to a tree-based machine learning ensemble and compare the classification precision, recall, and f-score against a non-curated set. Our study shows that using the noise reduction approach on the training instances gives better results in prediction with an improvement of 37% on precision, 70% on recall, and 59% on f-score.
Khaled Walid Al-Sabbagh, Miroslaw Staron, Regina Hebig, Wilhelm Meding
SEAA2
2020 Using Machine Learning to Identify Code Fragments for Manual Review
abstract
Code reviews are one of the first quality assurance tasks in continuous software integration and delivery. The goal of our work is to reduce the need for manual reviews by automatically identify which code fragments should be further reviewed manually. We conducted an action research study with two companies where we extracted code reviews and build machine learning classifiers (AdaBoost and Convolutional Neural Network– CNN). Our results show that the accuracy of recognizing code fragments that require manual review, measured with Matthews Correlation Coefficient, was 0.70 in the combination of our own feature extraction and CNN. We conclude that this way of combining automation with manual code reviews can improve the speed of reviews while providing organizations with the possibility to support knowledge transfer among the designers.
Miroslaw Staron, Miroslaw Ochodek, Wilhelm Meding, Ola Soder
SEAA1
2020 The Effect of Class Noise on Continuous Test Case Selection: A Controlled Experiment on Industrial Data
Khaled Walid Al-Sabbagh, Regina Hebig, Miroslaw Staron
PROFES3
2020 Recognizing lines of code violating company-specific coding guidelines using machine learning
abstract
Abstract Software developers in big and medium-size companies are working with millions of lines of code in their codebases. Assuring the quality of this code has shifted from simple defect management to proactive assurance of internal code quality. Although static code analysis and code reviews have been at the forefront of research and practice in this area, code reviews are still an effort-intensive and interpretation-prone activity. The aim of this research is to support code reviews by automatically recognizing company-specific code guidelines violations in large-scale, industrial source code. In our action research project, we constructed a machine-learning-based tool for code analysis where software developers and architects in big and medium-sized companies can use a few examples of source code lines violating code/design guidelines (up to 700 lines of code) to train decision-tree classifiers to find similar violations in their codebases (up to 3 million lines of code). Our action research project consisted of (i) understanding the challenges of two large software development companies, (ii) applying the machine-learning-based tool to detect violations of Sun’s and Google’s coding conventions in the code of three large open source projects implemented in Java, (iii) evaluating the tool on evolving industrial codebase, and (iv) finding the best learning strategies to reduce the cost of training the classifiers. We were able to achieve the average accuracy of over 99% and the average F-score of 0.80 for open source projects when using ca. 40K lines for training the tool. We obtained a similar average F-score of 0.78 for the industrial code but this time using only up to 700 lines of code as a training dataset. Finally, we observed the tool performed visibly better for the rules requiring to understand a single line of code or the context of a few lines (often allowing to reach the F-score of 0.90 or higher). Based on these results, we could observe that this approach can provide modern software development companies with the ability to use examples to teach an algorithm to recognize violations of code/design guidelines and thus increase the number of reviews conducted before the product release. This, in turn, leads to the increased quality of the final software.
Miroslaw Ochodek, Regina Hebig, Wilhelm Meding, Gert Frost, Miroslaw Staron
Empir. Softw. Eng.5
2020 PHANTOM: Curating GitHub for engineered software projects using time-series clustering
abstract
Abstract Context Within the field of Mining Software Repositories, there are numerous methods employed to filter datasets in order to avoid analysing low-quality projects. Unfortunately, the existing filtering methods have not kept up with the growth of existing data sources, such as GitHub, and researchers often rely on quick and dirty techniques to curate datasets. Objective The objective of this study is to develop a method capable of filtering large quantities of software projects in a resource-efficient way. Method This study follows the Design Science Research (DSR) methodology. The proposed method, PHANTOM, extracts five measures from Git logs. Each measure is transformed into a time-series, which is represented as a feature vector for clustering using the k-means algorithm. Results Using the ground truth from a previous study, PHANTOM was shown to be able to rediscover the ground truth on the training dataset, and was able to identify “engineered” projects with up to 0.87 Precision and 0.94 Recall on the validation dataset. PHANTOM downloaded and processed the metadata of 1,786,601 GitHub repositories in 21.5 days using a single personal computer, which is over 33% faster than the previous study which used a computer cluster of 200 nodes. The possibility of applying the method outside of the open-source community was investigated by curating 100 repositories owned by two companies. Conclusions It is possible to use an unsupervised approach to identify engineered projects. PHANTOM was shown to be competitive compared to the existing supervised approaches while reducing the hardware requirements by two orders of magnitude.
Peter Pickerill, Heiko Joshua Jungen, Miroslaw Ochodek, Michal Mackowiak, Miroslaw Staron
Empir. Softw. Eng.5
2020 Deep learning model for end-to-end approximation of COSMIC functional size based on use-case names
Miroslaw Ochodek, Sylwia Kopczynska, Miroslaw Staron
Inf. Softw. Technol.3
2019 Predicting Test Case Verdicts Using Textual Analysis of Committed Code Churns
Khaled Walid Al-Sabbagh, Miroslaw Staron, Regina Hebig, Wilhelm Meding
IWSM-Mensura2
2019 Evolution of Technical Debt: An Exploratory Study
Md. Abdullah Al Mamun 0001, Antonio Martini 0001, Miroslaw Staron, Christian Berger 0001, Jörgen Hansson
IWSM-Mensura3
2019 Information Needs for SAFe Teams and Release Train Management: A Design Science Research Study
Miroslaw Staron, Wilhelm Meding, Poupak Baniasad
IWSM-Mensura1
2019 Action Research in Software Engineering: Metrics' Research Perspective (Invited Talk)
Miroslaw Staron
SOFSEM1
2019 Simsax: A measure of project similarity based on symbolic approximation method and software defect inflow
Miroslaw Ochodek, Miroslaw Staron, Wilhelm Meding
Inf. Softw. Technol.2
2019 Assessing the impact of meta-model evolution: a measure and its automotive application
Darko Durisic, Miroslaw Staron, Matthias Tichy, Jörgen Hansson
Softw. Syst. Model.2
2018 Using Self-Healing to Increase Robustness of Handling In-Browser Third-Party Content
abstract
Monitoring of third-party content, such as ads, in web applications is one of the growing business areas in the web industry. In order to increase the impact of the ads and optimize the content on the web-page, companies measure which ads are displayed and how long they stay on the visible part of the screen. However, the challenge is that the third-party content can be of varying type, come from third-party servers or have active content. In this paper, we applied self-healing MAPE-K model in a monitor of the ads. Our results showed that the majority of faults could be repaired and that the resulting architecture is more maintainable than the one without self-healing; measured by architecture maintainability index. Therefore, we conclude that using self-healing can be applied to web-systems can increase both robustness and maintainability of these systems.
Sarah Nadi, Jimmy Hedstrom, Miroslaw Staron
SEAA3
2018 Measure early and decide fast: transforming quality management and measurement to continuous deployment
abstract
Continuous deployment has become software companies' inevitable response to the economic pressures of the market. At the same time, software quality is crucial in order to meet customers' expectations and hence succeed in the market. Therefore, current quality management processes require transformation in order to keep up with the fast pace of the market while at the same time meeting customers' expectations. In order to figure out how the current quality management process should be transformed to keep up with the fast pace of the market while ensuring both product quality and continuous deployment, we conducted a qualitative study at a large infrastructure provider company. During the interviews we conducted with the quality manager, developer and test architect, we used a metrics portfolio consisting of 59 candidate metrics that can be used in the transformed quality management process. Our findings show that, out of these candidate metrics, 9 metrics should be used in the internal quality measurement dashboard for quality check at the end of the software development life-cycle (SDLC) before the software is released to customer site, while 3 metrics should be used by quality manager to monitor earlier phases of SDLC and 5 metrics should also be delegated to earlier phases of SDLC but without the involvement of the quality manager. To summarize, our study support the claim that quality managers should not be only gatekeepers, but also proactive controllers of quality by monitoring earlier phases of the SDLC.
Gül Çalikli, Miroslaw Staron, Wilhelm Meding
ICSSP2
2018 Vetting Automatically Generated Trace Links: What Information is Useful to Human Analysts?
abstract
Automated traceability has been investigated for over a decade with promising results. However, a human analyst is needed to vet the generated trace links to ensure their quality. The process of vetting trace links is not trivial and while previous studies have analyzed the performance of the human analyst, they have not focused on the analyst's information needs. The aim of this study is to investigate what context information the human analyst needs. We used design science research, in which we conducted interviews with ten practitioners in the traceability area to understand the information needed by human analysts. We then compared the information collected from the interviews with existing literature. We created a prototype tool that presents this information to the human analyst. To further understand the role of context information, we conducted a controlled experiment with 33 participants. Our interviews reveal that human analysts need information from three different sources: 1) from the artifacts connected by the link, 2) from the traceability information model, and 3) from the tracing algorithm. The experiment results show that the content of the connected artifacts is more useful to the analyst than the contextual information of the artifacts.
Salome Maro, Jan-Philipp Steghöfer, Jane Huffman Hayes, Jane Cleland-Huang, Miroslaw Staron
RE5
2018 Special section on Visual Analytics in Software Engineering
Miroslaw Staron, Houari Sahraoui, Alexandru C. Telea
Inf. Softw. Technol.1
2018 Software traceability in the automotive domain: Challenges and solutions
Salome Maro, Jan-Philipp Steghöfer, Miroslaw Staron
J. Syst. Softw.3
2018 Industrial experiences from evolving measurement systems into self-healing systems for improved availability
abstract
Summary Automated measurement programs are an efficient way of collecting, processing, and visualizing measures in large software development companies. The number of measurements in these programs is usually large, which is caused by a diversity of the needs of the stakeholders. In this paper, we present the application of the self‐healing concepts to assure the availability of measurements to the stakeholders without the need for effort‐intensive and costly manual interventions of the operators. We study the measurement infrastructure at one of the development units of a large infrastructure provider. In this paper, we present how the Monitor, Analyze, Plane, and Execute with Knowledge model was instantiated in a simplistic manner to reduce the need for manual intervention in the operation of the measurement systems. Based on the experiences from the 2 cases studied in this paper, we show how an evolution toward self‐healing measurement systems is done both with a dedicated failure taxonomy and with an effective straightforward handling of the most common errors in the execution. The mechanisms studied and presented in this paper show that self‐healing provides significant improvements to the operation of the measurement program and reduces the need for daily oversight by an operator for the measurement systems.
Miroslaw Staron, Wilhelm Meding, Matthias Tichy, Jonas Bjurhede, Holger Giese, Ola Soder
Softw. Pract. Exp.1
2017 Predicting and Evaluating Software Model Growth in the Automotive Industry
abstract
The size of a software artifact influences the software quality and impacts the development process. In industry, when software size exceeds certain thresholds, memory errors accumulate and development tools might not be able to cope anymore, resulting in a lengthy program start up times, failing builds, or memory problems at unpredictable times. Thus, foreseeing critical growth in software modules meets a high demand in industrial practice. Predicting the time when the size grows to the level where maintenance is needed prevents unexpected efforts and helps to spot problematic artifacts before they become critical.Although the amount of prediction approaches in literature is vast, it is unclear how well they fit with prerequisites and expectations from practice. In this paper, we perform an industrial case study at an automotive manufacturer to explore applicability and usability of prediction approaches in practice. In a first step, we collect the most relevant prediction approaches from literature, including both, approaches using statistics and machine learning. Furthermore, we elicit expectations towards predictions from practitioners using a survey and stakeholder workshops. At the same time, we measure software size of 48 software artifacts by mining four years of revision history, resulting in 4,547 data points. In the last step, we assess the applicability of state-of-the-art prediction approaches using the collected data by systematically analyzing how well they fulfill the practitioners' expectations.Our main contribution is a comparison of commonly used prediction approaches in a real world industrial setting while considering stakeholder expectations. We show that the approaches provide significantly different results regarding prediction accuracy and that the statistical approaches fit our data best.
Jan Schroeder, Christian Berger 0001, Alessia Knauss, Harri Preenja, Mohammad Ali 0002, Miroslaw Staron, Thomas Herpel
ICSME6
2017 Co-Evolution of Meta-Modeling Syntax and Informal Semantics in Domain-Specific Modeling Environments - A Case Study of AUTOSAR
abstract
One domain-specific modeling environment is centered around a domain-specific meta-model which defines syntax (modeling elements, e.g., classes) for the domain models. However, in order for the system designers to be able to construct meaningful models, semantics of the domain-specific meta-model needs to be described as well. This semantics is often provided in a form of informal natural language specifications that contain a set of design requirements, each describing the intended use of one or more modeling elements. Intuitively, introduction of new concepts into the modeling environment is expected to require changes in both meta-modeling syntax and informal semantics in such a way that their co-evolution is highly correlated. In order to test this hypothesis, we analyzed the relation between added classes, attributes, and connectors, as meta-modeling syntax, and modified/added design requirements, as meta-modeling semantics, in a case study of the AUTOSAR meta-modeling environment. We found that new AUTOSAR concepts usually require both new modeling elements and new design requirements, but surprisingly adding more elements is not always followed by more requirements. This finding is also validated by the moderately strong correlation between the evolution of these two AUTOSAR meta-modeling artifacts (Spearman's rho 0,63 and Kendall's tau 0,49). For system designers, this means that both meta-modeling syntax and informal semantics is important to be considered in the analysis of domain-specific meta-model evolution, but it may not be enough for understanding the use of all modeling elements. For designers responsible for the maintenance of domain-specific meta-models, this means that more effort shall be put into describing the semantics of all introduced modeling elements.
Darko Durisic, Corrado Motta, Miroslaw Staron, Matthias Tichy
MoDELS3
2017 Measuring the Evolution of Meta-models - A Case Study of Modelica and UML Meta-models
abstract
The evolution of both general purpose and domain-specific meta-models and its impact on the existing models and modeling tools has been discussed extensively in the modeling research community. To assess the impact of domain-specific meta-model evolution on the modeling tools, a number of measures have been proposed by Durisic et al., NoC (Number of Changes) being the most prominent one. The proposed measures are evaluated on a case of AUTOSAR meta-model that specifies the language for designing automotive system architectures. In this paper, we assess the applicability of these measure and the underlying data-model for their calculation in a case study of Modelica and UML meta-models. Our preliminary results show that the proposed data-model and the measures can be applied to both analyzed meta-models as we were able to capture 68/77 changes on average per Modelica/UML release. However, only a subset of the data-model elements is applicable for analyzing the evolution of Modelica and also certain transformation of the data-model is required in case of UML. Despite these encouraging results, further studies are needed to assess the usefulness of the actual measures, e.g., NoC, in assessing the impact of Modelica/UML meta-model evolution on the modeling tools.
Maxime Jimenez, Darko Durisic, Miroslaw Staron
MODELSWARD3
2017 Proactive reviews of textual requirements
abstract
In large software development products the number of textual requirements can reach tens of thousands. When such a large number of requirements is delivered to software developers, there is a risk that vague or complex requirements remain undetected until late in the design process. In order to detect such requirements, companies conduct manual reviews of requirements. Manual reviews, however, take substantial amount of effort, and the efficiency is low. The goal of this paper is to present the application of a method for proactive requirements reviews. The method, that was developed and evaluated in a previous study, is now used in three companies. We show how the method evolved from an isolated scripted use to a fully integrated use in the three companies. The results showed that software engineers in the three companies use the method as a help in their job for continuous improvements of requirements.
Vard Antinyan, Miroslaw Staron
SANER2
2017 Evaluating code complexity triggers, use of complexity measures and the influence of code complexity on maintenance time
abstract
Code complexity has been studied intensively over the past decades because it is a quintessential characterizer of code’s internal quality. Previously, much emphasis has been put on creating code complexity measures and applying these measures in practical contexts. To date, most measures are created based on theoretical frameworks, which determine the expected properties that a code complexity measure should fulfil. Fulfilling the necessary properties, however, does not guarantee that the measure characterizes the code complexity that is experienced by software engineers. Subsequently, code complexity measures often turn out to provide rather superficial insights into code complexity. This paper supports the discipline of code complexity measurement by providing empirical insights into the code characteristics that trigger complexity, the use of code complexity measures in industry, and the influence of code complexity on maintenance time. Results of an online survey, conducted in seven companies and two universities with a total of 100 respondents, show that among several code characteristics, two substantially increase code complexity, which subsequently have a major influence on the maintenance time of code. Notably, existing code complexity measures are poorly used in industry.
Vard Antinyan, Miroslaw Staron, Anna Börjesson Sandberg
Empir. Softw. Eng.2
2017 2nd International Workshop on Automotive Systems and Software Architectures (WASA) - Introduction to special section
Yanjindulam Dajsuren, Harald Altinger, Miroslaw Staron
J. Syst. Archit.3
2017 Rendex: A method for automated reviews of textual requirements
Vard Antinyan, Miroslaw Staron
J. Syst. Softw.2
2017 Preface to the special issue on advances in software measurement
Miroslaw Staron, Wilhelm Meding, Alain Abran, Jan Bosch
Sci. Comput. Program.1
2016 Validating software measures using action research a method and industrial experiences
abstract
Validating software measures for using them in practice is a challenging task. Usually more than one complementary validation methods are applied for rigorously validating software measures: Theoretical methods help with defining the measures with expected properties and empirical methods help with evaluating the predictive power of measures. Despite the variety of these methods there still remain cases when the validation of measures is difficult. Particularly when the response variables of interest are not accurately measurable and the practical context cannot be reduced to an experimental setup the abovementioned methods are not effective. In this paper we present a complementary empirical method for validating measures. The method relies on action research principles and is meant to be used in combination with theoretical validation methods. The industrial experiences documented in this paper show that in many practical cases the method is effective.
Vard Antinyan, Miroslaw Staron, Anna Börjesson Sandberg, Jörgen Hansson
EASE2
2016 Unveiling anomalies and their impact on software quality in model-based automotive software revisions with software metrics and domain experts
abstract
The validation of simulation models (e.g., of electronic control units for vehicles) in industry is becoming increasingly challenging due to their growing complexity. To systematically assess the quality of such models, software metrics seem to be promising. In this paper we explore the use of software metrics and outlier analysis as a means to assess the quality of model-based software. More specifically, we investigate how results from regression analysis applied to measurement data received from size and complexity metrics can be mapped to software quality. Using the moving averages approach, models were fit to data received from over 65,000 software revisions for 71 simulation models that represent different electronic control units of real premium vehicles. Consecutive investigations using studentized deleted residuals and Cook’s Distance revealed outliers among the measurements. From these outliers we identified a subset, which provides meaningful information (anomalies) by comparing outlier scores with expert opinions. Eight engineers were interviewed separately for outlier impact on software quality. Findings were validated in consecutive workshops. The results show correlations between outliers and their impact on four of the considered quality characteristics. They also demonstrate the applicability of this approach in industry.
Jan Schroeder, Christian Berger 0001, Miroslaw Staron, Thomas Herpel, Alessia Knauss
ISSTA3
2016 Data veracity in intelligent transportation systems: The slippery road warning scenario
abstract
Intelligent transportation systems rely on the availability of high quality data in order to allow its multiple actors to make correct decisions in diverse traffic situations. Traditionally, high quality is associated with the correctness of the data, its timeliness or integrity. Going beyond data quality, this paper explores the notion of data veracity, which we approach from the perspective of the truthfulness of the data with respect to reality, or, in other words, its ability to be free from `lies'. Starting from the concrete case of the slippery road warning scenario (which comes from an industrial player), we define an initial taxonomy of data veracity (which is derived from the study of the literature) and use such taxonomy as a means to analyze the threats to data veracity in the above mentioned scenario. Additionally, this paper has the ambition to draw the attention of researchers and practitioners on the emerging challenges in the fields of data veracity and to define a research roadmap to tackle such challenges.
Miroslaw Staron, Riccardo Scandariato
Intelligent Vehicles Symposium1
2016 A Complexity Measure for Textual Requirements
abstract
Unequivocally understandable requirements are vital for software design process. However, in practice it is hard to achieve the desired level of understandability, because in large software products a substantial amount of requirements tend to have ambiguous or complex descriptions. Over time such requirements decelerate the development speed and increase the risk of late design modifications, therefore finding and improving them is an urgent task for software designers. Manual reviewing is one way of addressing the problem, but it is effort-intensive and critically slow for large products. Another way is using measurement, in which case one needs to design effective measures. In recent years there have been great endeavors in creating and validating measures for requirements understandability: most of the measures focused on ambiguous patterns. While ambiguity is one property that has major effect on understandability, there is also another important property, complexity, which also has major effect on understandability, but is relatively less investigated. In this paper we define a complexity measure for textual requirements through an action research project in a large software development organization. We also present its evaluation results in three large companies. The evaluation shows that there is a significant correlation between the measurement values and the manual assessment values of practitioners. We recommend this measure to be used with earlier created ambiguity measures as means for automated identification of complex specifications.
Vard Antinyan, Miroslaw Staron, Anna Börjesson Sandberg, Jörgen Hansson
IWSM-Mensura2
2016 A Key Performance Indicator Quality Model and Its Industrial Evaluation
abstract
Background: Modern software development companies increasingly rely on quantitative data in their decision-making for product releases, organizational performance assessment and monitoring of product quality. KPIs (Key Performance Indicators) are a critical element in the transformation of raw data (numbers) into decisions (indicators). The goal of the paper is to develop, document and evaluate a quality model for KPIs - addressing the research question of What characterizes a good KPI? In this paper we consider a KPI to be "good" when it is actionable and supports the organization in achieving its strategic goals. We use an action research collaborative project with an infrastructure provider company and an automotive OEM to develop and evaluate the model. We analyze a set of KPIs used at both companies and verify whether the organization's perception of these evaluated KPIs is aligned with the KPI's assessment according to our model. The results show that the model organizes good practices of KPI development and that it is easily used by the stakeholders to improve the quality of the KPIs or reduce the number of the KPIs. Using the KPI quality model provides the possibility to increase the effect of the KPIs in the organization and decreases the risk of wasting resources for collecting KPI data which cannot be used in practice.
Miroslaw Staron, Wilhelm Meding, Kent Niesel, Alain Abran
IWSM-Mensura1
2016 Addressing the Need for Strict Meta-modeling in Practice - A Case Study of AUTOSAR
abstract
Meta-modeling has been a topic of interest in the modeling community for many years, yielding substantialnumber of papers describing its theoretical concepts. Many of them are aiming to solve the problem of traditionalUML based domain-specific meta-modeling related to its non-compliance to the strict meta-modelingprinciple, such as the deep meta-modeling approach. In this paper, we show the practical use of meta-modelsin the automotive development process based on AUTOSAR and visualize places in the AUTOSAR metamodelwhich are broken according to the strict meta-modeling principle. We then explain how the AUTOSARmeta-modeling environment can be re-worked in order to comply to this principle by applying three individualapproaches, each one combined with the concept of Orthogonal Classification Architecture: UML extension,prototypical pattern and deep instantiation. Finally we discuss the applicability of these approaches in practiceand contrast the identified issues with the actual problems faced by the automotive meta-modeling practitioners.Our objective is to bridge the current gap between the theoretical and practical concerns in meta-modeling.
Darko Durisic, Miroslaw Staron, Matthias Tichy, Jörgen Hansson
MODELSWARD2
2016 Should We Adopt a New Version of a Standard? - A Method and Its Evaluation on AUTOSAR
Corrado Motta, Darko Durisic, Miroslaw Staron
PROFES3
2016 Guest editorial on special section: Automotive Software Architecture
Yanjindulam Dajsuren, Harald Altinger, Miroslaw Staron
Inf. Softw. Technol.3
2016 Analyzing defect inflow distribution and applying Bayesian inference method for software defect prediction in large software projects
Rakesh Rana, Miroslaw Staron, Christian Berger 0001, Jörgen Hansson, Martin Nilsson 0002, Wilhelm Meding
J. Syst. Softw.2
2016 MeSRAM - A method for assessing robustness of measurement programs in large software development organizations and its industrial evaluation
Miroslaw Staron, Wilhelm Meding
J. Syst. Softw.1
2015 ARCA - Automated Analysis of AUTOSAR Meta-model Changes
abstract
The software architecture of automotive software systems on the European market and wider is designed following the AUTOSAR standard. This requires continuous adoption of new AUTOSAR releases in the development projects in order to enable new innovative solutions in cars. Under these circumstances, the analysis of impact of the AUTOSAR meta-model changes on the modeling tools used in the development is crucial for avoiding delays and increased cost. However due to tens of new features combined with thousands of meta-model changes between consecutive releases of AUTOSAR, tool support is needed for such analysis. In this paper we present a systematic method and a tool - ARCA - for automated analysis of the AUTOSAR meta-model changes. The tool is able to identify relevant changes affecting modeling tools used by different roles in the development process and present the optimal set of new features to be adopted in the projects. The goal of the tool is to enable faster and cheaper software innovation cycles in cars.
Darko Durisic, Miroslaw Staron, Matthias Tichy
MiSE@ICSE2
2015 Measurement-as-a-Service - A New Way of Organizing Measurement Programs in Large Software Development Companies
Miroslaw Staron, Wilhelm Meding
IWSM/Mensura1
2015 Selecting the Right Visualization of Indicators and Measures - Dashboard Selection Model
Miroslaw Staron, Kent Niesel, Wilhelm Meding
IWSM/Mensura1
2015 Barriers and enablers for shortening software development lead-time in mechatronics organizations: a case study
abstract
The automotive industry adopts various approaches to reduce the production lead time in order to be competitive on the market. Due to the increasing amount of in-house software development, this industry gets new opportunities to decrease the software development lead-time. This can have a significant impact on decreasing time to market and fewer resources spent in projects. In this paper we present a study of software development areas where we perceived barriers for fast development and where we have identified enablers to overcome these barriers. We conducted a case study at one of the vehicle manufacturers in Sweden using structured interviews. Our results show that there are 21 barriers and 21 corresponding enablers spread over almost all phases of software development.
Mahshad M. Mahally, Miroslaw Staron, Jan Bosch
ESEC/SIGSOFT FSE2
2014 On the effect of using SysML requirement diagrams to comprehend requirements: results from two controlled experiments
abstract
We carried out a controlled experiment and an external replication to investigate whether the use of requirement diagrams of the System Modeling Language (SysML) helps in the comprehensibility of requirements. The original experiment was conducted at the University of Basilicata in Italy with Bachelor students, while its replication was executed at the University of Gothenburg in Sweden with Bachelor and Master students. A total of 87 participants took part in the experiment and its replication. The achieved results indicated that the comprehension of requirements is statistically significant when requirements specification documents include requirement diagrams without any impact on the time to accomplish comprehension tasks. On the basis of our results, we also present and discuss possible implications from the practitioner and researcher perspectives.
Giuseppe Scanniello, Miroslaw Staron, Håkan Burden, Rogardt Heldal
EASE2
2014 Defining Technical Risks in Software Development
abstract
Challenges of technical risk assessment is difficult to address, while its success can benefit software organizations appreciably. Classical definition of risk as a "combination of probability and impact of adverse event" appears not working with technical risk assessment. The main reason of this is the nature of adverse event's outcome which is rather continuous than discrete. The objective of this study was to scrutinize different aspects of technical risks and provide a definition, which will support effective risk assessment and management in software development organizations. In this study we defined the risk considering the nature of actual risks, emerged in software development. Afterwards, we summarized the software engineers' view on technical risks as results of three workshops with 15 engineers of four software development companies. The results show that technical risks could be viewed as a combination of uncertainty and magnitude of difference between actual and optimal design of product artifacts and processes. The presented definition is congruent with practitioners view on technical risk. It supports risk assessment in a quantitative manner and enables identification of potential product improvement areas.
Vard Antinyan, Miroslaw Staron, Wilhelm Meding, Anders Henriksson, Jörgen Hansson, Anna Börjesson Sandberg
IWSM/Mensura2
2014 Quantifying Long-Term Evolution of Industrial Meta-Models - A Case Study
abstract
Measurement in software engineering is an important activity for successful planning and management of projects under development. However knowing what to measure and how is crucial for the correct interpretation of the measurement results. In this paper, we assess the applicability of a number of software metrics for measuring a set of meta-model properties - size, length, complexity, coupling and cohesion. The goal is to identify which of these properties are mostly affected by the evolution of industrial meta-models and also which metrics should be used for their successful monitoring. In order to assess the applicability of the chosen set of metrics, we calculate them on a set of releases of the standardized meta-model used in the development of automotive software systems - the AUTOSAR meta-model - in a case study at Volvo Car Corporation. To identify the most applicable metrics, we used Principal Component Analysis (PCA). The results of these metrics shall be used by software designers in planning software development projects based on multiple AUTOSAR meta-model versions. We concluded that the evolution of the AUTOSAR meta-model is quite even with respect to all 5 properties and that the metrics based on fan-in complexity and package cohesion quantify the evolution most accurately.
Darko Durisic, Miroslaw Staron, Matthias Tichy, Jörgen Hansson
IWSM/Mensura2
2014 Identifying and Managing Complex Modules in Executable Software Design Models-Empirical Assessment of a Large Telecom Software Product
abstract
Using design models instead of executable code has shown itself to be an efficient way of increasing abstraction level of software development. However, applying established code-based software engineering methods to design models can be a challenge - due to different abstraction levels, the same metrics as for code are not applicable for the design models. One of practical challenges in using metrics at the model level is applying complexity-prediction formulas developed using code-based metrics to design models. The existing formulas do not apply as they do not take into consideration the behavior part of the models - e.g. State charts. In this paper we address this challenge by conducting a case study at one of the large telecom products at Ericsson with the goal to identify which metrics can predict complex, hard to understand and hard to maintain software modules based on their design models. We use both statistical methods like regression to build prediction formulas and qualitative interviews to codify expert designers' perception of which software modules are complex. The results of this case study show that such measures as the number of non-self-transitions, transition per states or state depth can be combined in order to identify software units that are perceived as complex by expert designers. Our conclusion is that these metrics can be used in other companies to predict complex modules, but the coefficients should be recalculated per product to increase the prediction accuracy.
Hengameh Rezaei, Filippa Ebersjo, Kristian Sandahl, Miroslaw Staron
IWSM/Mensura4
2014 Consequences of Mispredictions of Software Reliability: A Model and its Industrial Evaluation
abstract
Predicting reliability of software under development is an important part of estimations in software engineering projects. In many organizations as the goal is that software products are released with no known defects, the process of finding and removing defects correlates with the effort for software projects. Software development projects estimate the resources needed to design, develop, test and release software products, and the number of defects which have to be handled. In this paper we present a model for consequence analysis of inaccurate predictions of quality in software projects. The model is a result of multiple case studies and is evaluated at two companies. The model recognizes the most common mispredictions - e.g. Over- and under-prediction, early- and late-predictions - and the combination of theses. The results from the industrial evaluation show that the consequences can be grouped according to under- and over-predictions and that the late- and early-predictions have the same consequences. The results show also that mispredicting the shape of the reliability curve has a significant consequence with regard to assessment of release readiness and resource planning.
Miroslaw Staron, Rakesh Rana, Wilhelm Meding, Martin Nilsson 0002
IWSM/Mensura1
2014 Performance in software development - Special issue editorial
Miroslaw Staron, Jörgen Hansson, Jan Bosch
Inf. Softw. Technol.1
2014 Selecting software reliability growth models and improving their predictive accuracy using historical projects data
Rakesh Rana, Miroslaw Staron, Christian Berger 0001, Jörgen Hansson, Martin Nilsson 0002, Fredrik Törner, Wilhelm Meding, Christoffer Höglund
J. Syst. Softw.2
2013 Increasing Efficiency of ISO 26262 Verification and Validation by Combining Fault Injection and Mutation Testing with Model based Development
abstract
The rapid growth of software intensive active safety functions in modern cars resulted in adoption of new safety development standards like ISO 26262 by the automotive industry. Hazard analysis, safety assessment and adequate verification and validation methods for software and car electronics require effort but in the long run save lives. We argue that in the face of complex software development set-up with distributed functionality, Model-Based Development (MBD) and safety criticality of software embedded in modern cars, there is a need for evolving existing methods of MBD and complementing them with methods already used in the development of other systems (Fault Injection and Mutation Testing). Our position is that significant effectiveness and efficiency improvements can be made by applying fault injection techniques combined with mutation testing approach for verification and validation of automotive software at the model level. The improvements include such aspects as identification of safety related defects early in the development process thus providing enough time to remove the defects. The argument is based on our industrial case studies, the studies of ISO 26262 standard and academic experiments with new verification and validation methods applied to models.
Rakesh Rana, Miroslaw Staron, Christian Berger 0001, Jörgen Hansson, Martin Nilsson 0002, Fredrik Törner
ICSOFT2
2013 Evaluating long-term predictive power of standard reliability growth models on automotive systems
abstract
Software is today an integral part of providing improved functionality and innovative features in the automotive industry. Safety and reliability are important requirements for automotive software and software testing is still the main source of ensuring dependability of the software artifacts. Software Reliability Growth Models (SRGMs) have been long used to assess the reliability of software systems; they are also used for predicting the defect inflow in order to allocate maintenance resources. Although a number of models have been proposed and evaluated, much of the assessment of their predictive ability is studied for short term (e.g. last 10% of data). But in practice (in industry) the usefulness of SRGMs with respect to optimal resource allocation depends heavily on the long term predictive power of SRGMs i.e. much before the project is close to completion. The ability to reasonably predict the expected defect inflow provides important insight that can help project and quality managers to take necessary actions related to testing resource allocation on time to ensure high quality software at the release. In this paper we evaluate the long-term predictive power of commonly used SRGMs on four software projects from the automotive sector. The results indicate that Gompertz and Logistic model performs best among the tested models on all fit criterias as well as on predictive power, although these models are not reliable for long-term prediction with partial data.
Rakesh Rana, Miroslaw Staron, Christian Berger 0001, Jörgen Hansson, Martin Nilsson 0002, Fredrik Törner
ISSRE2
2013 Comparing between Maximum Likelihood Estimator and Non-linear Regression Estimation Procedures for NHPP Software Reliability Growth Modelling
abstract
Software Reliability Growth Models (SRGMs) have been used by engineers and managers for tracking and managing the reliability change of software to ensure required standard of quality is achieved before the software is released to the customer. SRGMs can be used during the project to help make testing resource allocation decisions and/ or it can be used after the testing phase to determine the latent faults prediction to assess the maturity of software artifact. A number of SRGMs have been proposed and to apply a given reliability model, defect inflow data is fitted to model equations. Two of the widely known and recommended techniques for parameter estimation are maximum likelihood and method of least squares. In this paper we compare between the two estimation procedures for their applicability in context of NHPP SRGMs. We also highlight a couple of practical considerations, reliability practitioners must be aware of when applying SRGMs.
Rakesh Rana, Miroslaw Staron, Christian Berger 0001, Jörgen Hansson, Martin Nilsson 0002, Fredrik Törner
IWSM/Mensura2
2013 Measuring and Visualizing Code Stability - A Case Study at Three Companies
abstract
Monitoring performance of software development organizations can be achieved from a number of perspectives - e.g. using such tools as Balanced Scorecards or corporate dashboards. In this paper we present results from a study on using code stability indicators as a tool for product stability and organizational performance, conducted at three different software development companies - Ericsson AB, Saab AB Electronic Defense Systems (Saab) and Volvo Group Trucks Technology (Volvo Group). The results show that visualizing the source code changes using heat maps and linking these visualizations to defect inflow profiles provide indicators of how stable the product under development is and whether quality assurance efforts should be directed to specific parts of the product. Observing the indicator and making decisions based on its visualization leads to shorter feedback loops between development and test, thus resulting in lower development costs, shorter lead time and increased quality. The industrial case study in the paper shows that the indicator and its visualization can show whether the modifications of software products are focused on parts of the code base or are spread widely throughout the product.
Miroslaw Staron, Jörgen Hansson, Robert Feldt, Anders Henriksson, Wilhelm Meding, Sven Nilsson, Christoffer Höglund
IWSM/Mensura1
2013 Why Do We Not Learn from Defects? - Towards Defect-Driven Software Process Improvement
abstract
In this paper, we put forth the thesis that state-of-the-art defect classification schemes – such as ODC and IEEE Std. 1044 – have failed to meet their target; limited industrial adoption is taken as part of the evidence combined with published studies on model driven software development. Notwithstanding, a number of publications show that defect reports can provide valuable information about common, important, or dangerous problems with software products. In this paper, we present the synthesis of two industrial case studies that illustrate that even expert judgement can be deceptive; demonstrating the need for more objective evidence to allow project stakeholder to make informed decisions, and that defect classification is one effective means to that end. Finally, we propose a roadmap that will contribute to improving the defect classification approach, which in consequence will lead to a wider industrial adoption.
Niklas Mellegård, Miroslaw Staron, Fredrik Törner
MODELSWARD2
2013 Evaluation of Standard Reliability Growth Models in the Context of Automotive Software Systems
Rakesh Rana, Miroslaw Staron, Niklas Mellegård, Christian Berger 0001, Jörgen Hansson, Martin Nilsson 0002, Fredrik Törner
PROFES2
2013 Measuring the impact of changes to the complexity and coupling properties of automotive software systems
Darko Durisic, Martin Nilsson 0002, Miroslaw Staron, Jörgen Hansson
J. Syst. Softw.3
2012 A Light-Weight Defect Classification Scheme for Embedded Automotive Software and Its Initial Evaluation
abstract
Objective: Defect classification is an essential part of software development process models as a means of early identification of patterns in defect inflow profiles. Such classification, however, may often be a tedious task requiring analysis work in addition to what is necessary to resolve the issue. To increase classification efficiency, adapted schemes are needed. In this paper a light-weight defect classification scheme adapted for minimal process footprint -- in terms of learning and classification effort -- is proposed and initially evaluated. Method: A case study was conducted at Volvo Car Corporation to adapt the IEEE Std. 1044 for automotive embedded software. An initial evaluation was conducted by applying the adapted scheme to defects from an existing software product with industry professionals as subjects. Results: The results showed that the classification scheme was quick to learn and understand -- required classification time stabilized around 5-10 minutes already after practicing on 3-5 defects. The results also showed that the patterns in the classified defects were interesting for the professionals, although in order to apply statistical methods more data was needed. Conclusions: We conclude that the adapted classification scheme captures what is currently tacit knowledge and has the potential of revealing patterns in the defects detected in different project phases. Furthermore, we were, in the initial evaluation, able to contribute with new information about the development process. As a result we are currently in the process of incorporating the classification scheme into the company's defect reporting system.
Niklas Mellegård, Miroslaw Staron, Fredrik Törner
ISSRE2
2012 Release Readiness Indicator for Mature Agile and Lean Software Development Projects
Miroslaw Staron, Wilhelm Meding, Klas Palm
XP1
2012 Critical role of measures in decision processes: Managerial and technical measures in the context of large software development organizations
Miroslaw Staron
Inf. Softw. Technol.1
2011 Monitoring Bottlenecks in Agile and Lean Software Development Projects - A Method and Its Industrial Use
Miroslaw Staron, Wilhelm Meding
PROFES1
2011 Developing measurement systems: an industrial case study
abstract
Abstract The process of measuring in software engineering has already been standardized in the ISO/IEC 15939 standard, where activities related to identifying, creating, and evaluating of measures are described. In the process of measuring software entities, however, an organization usually needs to create custom measurement systems, which are intended to collect, analyze, and present data for a specific purpose. In this paper, we present a proven industrial process for developing measurement systems including the artifacts and deliverables important for a successful deployment of measurement systems in industry. The process has been elicited during a case study at Ericsson and is used in the organization for over 3 years when the paper was written. The process is supported by a framework that simplifies the implementation of the measurement systems and shortens the time from the initial idea to a working measurement system by the factor of 5 compared with using a standard development process not tailored for measurement systems. Copyright © 2010 John Wiley & Sons, Ltd.
Miroslaw Staron, Wilhelm Meding, Göran Karlsson, Christer Nilsson
J. Softw. Maintenance Res. Pract.1
2010 Improving Efficiency of Change Impact Assessment Using Graphical Requirement Specifications: An Experiment
Niklas Mellegård, Miroslaw Staron
PROFES2
2010 A method for forecasting defect backlog in large streamline software development projects and its industrial evaluation
Miroslaw Staron, Wilhelm Meding, Bo Söderqvist
Inf. Softw. Technol.1
2009 Ensuring Reliability of Information Provided by Measurement Systems
Miroslaw Staron, Wilhelm Meding
IWSM/Mensura1
2009 Using Models to Develop Measurement Systems: A Method and Its Industrial Use
Miroslaw Staron, Wilhelm Meding
IWSM/Mensura1
2009 A framework for developing measurement systems and its industrial evaluation
Miroslaw Staron, Wilhelm Meding, Christer Nilsson
Inf. Softw. Technol.1
2008 Ontology Guided Evolution of Complex Embedded Systems Projects in the Direction of MDA
Lars Pareto, Miroslaw Staron, Peter S. Eriksson
MoDELS2
2008 Predicting weekly defect inflow in large software projects based on project planning and test status
Miroslaw Staron, Wilhelm Meding
Inf. Softw. Technol.1
2007 Using Students as Subjects in Experiments--A Quantitative Analysis of the Influence of Experimentation on Students' Learning Proces
abstract
The purpose of this study is to evaluate whether using students as subjects in software engineering experiments improves their learning process. In this paper we describe experiences with experimenting with students from several experiments. We show how these experiments were related to the courses and how we used them as an auxiliary tool for teaching. We perform a survey among the students who participated in the experiments and we analyze the final grades from the courses with and without experiments. The results show that the percentage of the high-passes increased after introducing experiments by 41%, 91% of students found the experiments to be useful in other courses, a vast majority of students perceived experiments as a good way of applying their knowledge in practice and as a good way of getting acquainted with empirical methods in software engineering (which was not the subject of the course). The implications of our research are that including carefully designed experiments into courses improve students' learning process and increase their motivation. The novelty of the paper is the perspective of students' learning perspective rather than researchers' benefits. The results help the teachers to take advantage of experiments to increase the students' learning.
Miroslaw Staron
CSEE&T1
2007 Predicting Short-Term Defect Inflow in Large Software Projects - An Initial Evaluation
Miroslaw Staron, Wilhelm Meding
EASE1
2007 Using Experiments in Software Engineering as an Auxiliary Tool for Teaching-A Qualitative Evaluation from the Perspective of Students' Learning Process
abstract
In this paper we discuss issues of using students as subjects from the perspective of their benefits in terms of learning. We conduct interviews with students who participated in the experiments and present their perceptions of the experiments in order to validate the claims posed in the existing literature. Finally we show quantitative data on the rate of obtaining the distinctive pass grades for courses with and without experiments. The results show that integrating experiments with courses could lead to improvements of the performance of students on courses.
Miroslaw Staron
ICSE1
2006 Adopting Model Driven Software Development in Industry - A Case Study at Two Companies
Miroslaw Staron
MoDELS1
2006 An Industrial Case Study on the Choice Between Language Customization Mechanisms
Miroslaw Staron, Claes Wohlin
PROFES1
2006 Empirical assessment of using stereotypes to improve comprehension of UML models: A set of experiments
Miroslaw Staron, Ludwik Kuzniarz, Claes Wohlin
J. Syst. Softw.1
2005 Evaluation of a Framework for Reverse Engineering Tool Construction
abstract
This paper describes an experiment of rapidly constructing reverse engineering tools from predefined components/tools. We perform two ways of construction. The first uses ad-hoc composition. The second exploits a framework design to support the customization of reverse engineering tools. We evaluate the quality of the assembled tools and their assembling time for each approach. For this, we describe in this paper briefly our framework, the metrics chosen for quality measurement based on the GQM approach, the experiment itself and our results. The paper gives empirical evidence that framework customization, here using the VizzAnalyzer framework, is superior to ad-hoc composition.
Thomas Panas, Miroslaw Staron
ICSM2
2002 Extracting Initial UML Domain Models from Daml+OIL Encoded Ontologies
Ludwik Kuzniarz, Miroslaw Staron
PROFES2