EDBT 2026 Demo / reviewers in the wild / expert
Nauman Bin Ali
dblp:37/10608
· DBLP profile ↗
37ranked-venue papers
10as first author
19since 2021 · last 2026
0000-0001-7266-5632ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 37 · 10 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Monitoring Data for Anomaly Detection in Cloud-Based Systems: A Systematic Mapping StudyabstractContext : Anomaly detection is crucial for maintaining cloud-based software systems, as it enables early identification and resolution of unexpected failures. Given rapid and significant advances in the anomaly detection domain and the complexity of its industrial implementation, an overview of techniques that utilize real-world operational data is needed. Aim : This study aims to complement existing research with an extensive catalog of the techniques and monitoring data used for detecting anomalies affecting the performance or reliability of cloud-based software systems that have been developed and/or evaluated in a real-world context. Method : We perform a systematic mapping study to examine the literature on anomaly detection in cloud-based systems, particularly focusing on the usage of real-world monitoring data, with the aim of identifying key data categories, tools, data preprocessing, and anomaly detection techniques. Results : Based on a review of 104 papers, we categorize monitoring data by structure, types, and origins and the tools used for data collection and processing. We offer a comprehensive overview of data preprocessing and anomaly detection techniques mapped to different data categories. Our findings highlight practical challenges and considerations in applying these techniques in real-world cloud environments. Conclusion : The findings help practitioners and researchers identify relevant data categories and select appropriate data preprocessing and anomaly detection techniques for their specific operational environments, which is important for improving the reliability and performance of cloud-based systems. Adha Hrusto, Nauman Bin Ali, Emelie Engström, Yuqing Wang 0002 |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2025 | Towards Lean Research Inception: Assessing Practical Relevance of Formulated Research Problemsabstract[Context] The lack of practical relevance in many Software Engineering (SE) research contributions is often rooted in oversimplified views of industrial practice, weak industry connections, and poorly defined research problems. Clear criteria for evaluating SE research problems can help align their value, feasibility, and applicability with industrial needs. [Goal] In this paper, we introduce the Lean Research Inception (LRI) framework, designed to support the formulation and assessment of practically relevant research problems in SE. We describe its initial evaluation strategy conducted in a workshop with a network of SE researchers experienced in industry-academia collaboration and report the evaluation of its three assessment criteria (valuable, feasible, and applicable) regarding their importance and completeness in assessing practical relevance. [Method] We applied LRI retroactively to a published research paper, engaging workshop participants in discussing and assessing the research problem by applying the proposed criteria using a semantic differential scale. Participants provided feedback on the criteria’s importance and completeness, drawn from their own experiences in industry-academia collaboration. [Results] The findings reveal an overall agreement on the importance of the three criteria – valuable (83.3%), feasible (76.2%), and applicable (73.8%) – for aligning research problems with industrial needs. Qualitative feedback suggested adjustments in terminology with a clearer distinction between feasible and applicable, and refinements for valuable by more clearly considering business value, ROI, and originality. [Conclusion] While LRI still constitutes ongoing research and requires further evaluation, our emerging results strengthen our confidence that the three criteria applied using the semantic differential scale can already help the community assess the practical relevance of SE research problems. Anrafel Fernandes Pereira, Marcos Kalinowski, Maria Teresa Baldassarre, Jürgen Börstler, Nauman Bin Ali, Daniel Méndez 0001 |
EASE | 5 |
| 2025 | Another Systematic Review? A Critical Analysis of Systematic Literature Reviews on Agile Effort and Cost EstimationabstractBackground: Systematic literature reviews (SLRs) have become prevalent in software engineering research. Several researchers may conduct SLRs on similar topics without a prospective register for SLR protocols. However, even ignoring these unavoidable duplications of effort in the simultaneous conduct of SLRs, the proliferation of overlapping and often repetitive SLRs indicates that researchers are not extensively checking for existing SLRs on a topic. Given how effort-intensive it is to design, conduct, and report an SLR, the situation is less than ideal for software engineering research. Aim: To understand how authors justify additional SLRs on a topic. Method: To illustrate the issue and develop suggestions for improvement to address this issue, we have intentionally picked a sufficiently narrow but well-researched topic, i.e., effort estimation in Agile software development. We identify common justification patterns through a qualitative content analysis of$\mathbf{1 8}$published SLRs. We further consider the citation data, publication years, publication venues, and the quality of the SLRs when interpreting the results. Results: The common justification patterns include authors claiming gaps in coverage, methodological limitations in prior studies, temporal obsolescence of previous SLRs, or rapid technological and methodological advancements necessitating updated syntheses. Conclusion: Our in-depth analysis of SLRs on a fairly narrow topic provides insights into SLRs in software engineering in general. By emphasizing the need for identifying existing SLRs and for justifying the undertaking of further SLRs, both in design and review guidelines and as a policy of conferences and journals, we can reduce the likelihood of duplication of effort and increase the rate of progress in the field. Henry Edison, Nauman Bin Ali |
ESEM | 2 |
| 2025 | What Do We Know About Software Analytics Research? A Critical Review of Secondary Studies
Muhammad Laiq, Nauman Bin Ali, Jürgen Börstler, Emelie Engström |
SEAA (2) | 2 |
| 2025 | A comparative analysis of ML techniques for bug report classificationabstractSeveral studies have evaluated various ML techniques and found promising results in classifying bug reports . However, these studies have used different evaluation designs, making it difficult to compare their results. Furthermore, they have focused primarily on accuracy and did not consider other potentially relevant factors such as generalizability , explainability, and maintenance cost. These two aspects make it difficult for practitioners and researchers to choose an appropriate ML technique for a given context. Therefore, we compare promising ML techniques against practitioners’ concerns using evaluation criteria that go beyond accuracy. Based on an existing framework for adopting ML techniques, we developed an evaluation framework for ML techniques for bug report classification. We used this framework to compare nine ML techniques on three datasets. The results enable a tradeoff analysis between various promising ML techniques. The results show that an ML technique with the highest predictive accuracy might not be the most suitable technique for some contexts. The overall approach presented in the paper supports making informed decisions when choosing ML techniques. It is not locked to the specific techniques, datasets, or factors we have selected here, and others could easily use and adapt it for additional techniques or concerns. Editor’s note: Open Science material was validated by the Journal of Systems and Software Open Science Board. Muhammad Laiq, Nauman Bin Ali, Jürgen Börstler, Emelie Engström |
J. Syst. Softw. | 2 |
| 2025 | A proposal and assessment of an improved heuristic for the Eager Test smell detectionabstractThe evidence for the prevalence of test smells at the unit testing level has relied on the accuracy of detection tools, which have seen intense research in the last two decades. The Eager Test smell, one of the most prevalent, is often identified using simplified detection rules that practitioners find inadequate. We aim to improve the rules for detecting the Eager Test smell. We reviewed the literature on test smells to analyze the definitions and detection rules of the Eager Test smell. We proposed a novel, unambiguous definition of the test smell and a heuristic to address the limitations of the existing rules. We evaluated our heuristic against existing detection rules by manually applying it to 300 unit test cases in Java. Our review identified 56 relevant studies. We found that inadequate interpretations of original definitions of the Eager Test smell led to imprecise detection rules, resulting in a high level of disagreement in detection outcomes. Also, our heuristic detected patterns of eager and non-eager tests that existing rules missed. Our heuristic captures the essence of the Eager Test smell more precisely; hence, it may address practitioners’ concerns regarding the adequacy of existing detection rules. • An overview of the definitions and detection rules of the Eager Test smell. • A novel, unambiguous definition of Eager Test based on the original definitions. • A novel heuristic with pseudocode for detecting eager tests. • The identification of common patterns of eager tests and non-eager tests. Huynh Khanh Vi Tran, Nauman Bin Ali, Michael Unterkalmsteiner, Jürgen Börstler |
J. Syst. Softw. | 2 |
| 2025 | Supporting the identification of prevalent quality issues in code changes by analyzing reviewers' feedbackabstractAbstract Context: Code reviewers provide valuable feedback during the code review. Identifying common issues described in the reviewers’ feedback can provide input for devising context-specific software development improvements. However, the use of reviewer feedback for this purpose is currently less explored. Objective: In this study, we assess how automation can derive more interpretable and informative themes in reviewers’ feedback and whether these themes help to identify recurring quality-related issues in code changes. Method: We conducted a participatory case study using the JabRef system to analyze reviewers’ feedback on merged and abandoned code changes. We used two promising topic modeling methods (GSDMM and BERTopic) to identify themes in 5,560 code review comments. The resulting themes were analyzed and named by a domain expert from JabRef. Results: The domain expert considered the identified themes from the two topic models to represent quality-related issues. Different quality issues are pointed out in code reviews for merged and abandoned code changes. While BERTopic provides higher objective coherence, the domain expert considered themes from short-text topic modeling more informative and easy to interpret than BERTopic-based topic modeling. Conclusions: The identified prevalent code quality issues aim to address the maintainability-focused issues. The analysis of code review comments can enhance the current practices for JabRef by improving the guidelines for new developers and focusing discussions in the developer forums. The topic model choice impacts the interpretability of the generated themes, and a higher coherence (based on objective measures) of generated topics did not lead to improved interpretability by a domain expert. Umar Iftikhar, Jürgen Börstler, Nauman Bin Ali, Oliver Kopp |
Softw. Qual. J. | 3 |
| 2025 | Quality attributes of test cases and test suites - importance & challenges from practitioners' perspectivesabstractAbstract The quality of the test suites and the constituent test cases significantly impacts confidence in software testing. While research has identified several quality attributes of test cases and test suites, there is a need for a better understanding of their relative importance in practice. We investigate practitioners’ perceptions regarding the relative importance of quality attributes of test cases and test suites and the challenges that they face in ensuring the perceived important quality attributes. To capture the practitioners’ perceptions, we conducted an industrial survey using a questionnaire based on the quality attributes identified in an extensive literature review. We used a sampling strategy that leverages LinkedIn to draw a large and heterogeneous sample of professionals with experience in software testing. We collected 354 responses from practitioners with a wide range of experience (from less than one year to 42 years of experience). We found that the majority of practitioners rated Fault Detection, Usability, Maintainability, Reliability, and Coverage to be the most important quality attributes. Resource Efficiency, Reusability, and Simplicity received the most divergent opinions, which, according to our analysis, depend on the software-testing contexts. Also, we identified common challenges that apply to the important attributes, namely inadequate definition, lack of useful metrics, lack of an established review process, and lack of external support. The findings point out where practitioners actually need further support with respect to achieving high-quality test cases and test suites under different software testing contexts. Hence, the findings can serve as a guideline for academic researchers when looking for research directions on the topic. Furthermore, the findings can be used to encourage companies to provide more support to practitioners to achieve high-quality test cases and test suites. Huynh Khanh Vi Tran, Nauman Bin Ali, Michael Unterkalmsteiner, Jürgen Börstler, Panagiota Chatzipetrou |
Softw. Qual. J. | 2 |
| 2024 | Industrial adoption of machine learning techniques for early identification of invalid bug reportsabstractAbstract Despite the accuracy of machine learning (ML) techniques in predicting invalid bug reports, as shown in earlier research, and the importance of early identification of invalid bug reports in software maintenance, the adoption of ML techniques for this task in industrial practice is yet to be investigated. In this study, we used a technology transfer model to guide the adoption of an ML technique at a company for the early identification of invalid bug reports. In the process, we also identify necessary conditions for adopting such techniques in practice. We followed a case study research approach with various design and analysis iterations for technology transfer activities. We collected data from bug repositories, through focus groups, a questionnaire, and a presentation and feedback session with an expert. As expected, we found that an ML technique can identify invalid bug reports with acceptable accuracy at an early stage. However, the technique’s accuracy drops over time in its operational use due to changes in the product, the used technologies, or the development organization. Such changes may require retraining the ML model. During validation, practitioners highlighted the need to understand the ML technique’s predictions to trust the predictions. We found that a visual (using a state-of-the-art ML interpretation framework) and descriptive explanation of the prediction increases the trustability of the technique compared to just presenting the results of the validity predictions. We conclude that trustability, integration with the existing toolchain, and maintaining the techniques’ accuracy over time are critical for increasing the likelihood of adoption. Muhammad Laiq, Nauman Bin Ali, Jürgen Börstler, Emelie Engström |
Empir. Softw. Eng. | 2 |
| 2024 | Acceptance behavior theories and models in software engineering - A mapping studyabstractThe adoption or acceptance of new technologies or ways of working in software development activities is a recurrent topic in the software engineering literature. The topic has, therefore, been empirically investigated extensively. It is, however, unclear which theoretical frames of reference are used in this research to explain acceptance behaviors. In this study, we explore how major theories and models of acceptance behavior have been used in the software engineering literature to empirically investigate acceptance behavior. We conduct a systematic mapping study of empirical studies using acceptance behavior theories in software engineering. We identified 47 primary studies covering 56 theory uses. The theories were categorized into six groups. Technology acceptance models (TAM and its extensions) were used in 29 of the 47 primary studies, innovation theories in 10, and the theories of planned behavior/ reasoned action (TPB/TRA) in six. All other theories were used in at most two of the primary studies. The usage and operationalization of the theories were, in many cases, inconsistent with the underlying theories. Furthermore, we identified 77 constructs used by these studies of which many lack clear definitions. Our results show that software engineering researchers are aware of some of the leading theories and models of acceptance behavior, which indicates an attempt to have more theoretical foundations. However, we identified issues related to theory usage that make it difficult to aggregate and synthesize results across studies. We propose mitigation actions that encourage the consistent use of theories and emphasize the measurement of key constructs. Jürgen Börstler, Nauman Bin Ali, Kai Petersen, Emelie Engström |
Inf. Softw. Technol. | 2 |
| 2024 | A tertiary study on links between source code metrics and external quality attributesabstractSeveral secondary studies have investigated the relationship between internal quality attributes, source code metrics and external quality attributes. Sometimes they have contradictory results. We synthesize evidence of the link between internal quality attributes, source code metrics and external quality attributes along with the efficacy of the prediction models used. We conducted a tertiary review to identify, evaluate and synthesize secondary studies. We used several characteristics of secondary studies as indicators for the strength of evidence and considered them when synthesizing the results. From 711 secondary studies, we identified 15 secondary studies that have investigated the link between source code and external quality. Our results show : (1) primarily, the focus has been on object-oriented systems, (2) maintainability and reliability are most often linked to internal quality attributes and source code metrics, with only one secondary study reporting evidence for security, (3) only a small set of complexity, coupling, and size-related source code metrics report a consistent positive link with maintainability and reliability, and (4) group method of data handling (GMDH) based prediction models have performed better than other prediction models for maintainability prediction. Based on our results, lines of code, coupling, complexity and the cohesion metrics from Chidamber & Kemerer (CK) metrics are good indicators of maintainability with consistent evidence from high and moderate-quality secondary studies. Similarly, four CK metrics related to coupling, complexity and cohesion are good indicators of reliability, while inheritance and certain cohesion metrics show no consistent evidence of links to maintainability and reliability. Further empirical studies are needed to explore the link between internal quality attributes, source code metrics and other external quality attributes, including functionality, portability, and usability. The results will help researchers and practitioners understand the body of knowledge on the subject and identify future research directions. Umar Iftikhar, Nauman Bin Ali, Jürgen Börstler, Muhammad Usman 0002 |
Inf. Softw. Technol. | 2 |
| 2024 | Experiences from conducting rapid reviews in collaboration with practitioners - Two industrial casesabstractEvidence-based software engineering (EBSE) aims to improve research utilization in practice. It relies on systematic methods to identify, appraise, and synthesize existing research findings to answer questions of interest for practice. However, the lack of practitioners’ involvement in these studies’ design, execution, and reporting indicates a lack of appreciation for the need for knowledge exchange between researchers and practitioners. The resultant systematic literature studies often lack relevance for practice. This paper explores the use of Rapid Reviews (RRs), in fostering knowledge exchange between academia and industry. Through the lens of two case studies, we delve into the practical application and experience of conducting RRs. We analyzed the conduct of two rapid reviews by two different groups of researchers and practitioners. We collected data through interviews, and the documents produced during the review (like review protocols, search results, and presentations). The interviews were analyzed using thematic analysis. We report how the two groups of researchers and practitioners performed the rapid reviews. We observed some benefits, like promoting dialogue and paving the way for future collaborations. We also found that practitioners entrusted the researchers to develop and follow a rigorous approach and were more interested in the applicability of the findings in their context. The problems investigated in these two cases were relevant but not the most immediate ones. Therefore, rapidness was not a priority for the practitioners. The study illustrates that rapid reviews can support researcher-practitioner communication and industry-academia collaboration. Furthermore, the recommendations based on the experiences from the two cases complement the detailed guidelines researchers and practitioners may follow to increase interaction and knowledge exchange. Sergio Rico, Nauman Bin Ali, Emelie Engström, Martin Höst |
Inf. Softw. Technol. | 2 |
| 2023 | On potential improvements in the analysis of the evolution of themes in code review commentsabstractContext: The modern code review process is considered an essential quality assurance step in software development. The code review comments generated can provide insights regarding source code quality and development practices. However, the large number of code review comments makes it challenging to identify interesting patterns manually. In a recent study, Wen et al. used traditional topic modeling to analyze the evolution of code review comments. Their approach could identify interesting patterns that may lead to improved development practices.Objective: In this study, we investigate potential improvements to Wen et al.’s state-of-the-art approach to analyze the evolution of code review comments.Method: We used 209,166 code review comments from three open-source systems to explore and empirically analyze alternative design and implementation choices and demonstrate their impact.Results: We identified the following potential improvements to the current state-of-the-art as described by Wen et al.: 1) utilize a topic modeling method that is optimized for short texts, 2) a refined approach for identifying a suitable number of topics, and 3) a more elaborate approach for analyzing topic evolution. Our results indicate that the proposed changes have quantitatively different results than the current approach. The qualitative interpretation of the topics generated with our changes indicates their usefulness.Conclusions: Our results indicate the potential usefulness of changes to state-of-the-art approaches to analyzing the evolution of code review comments, with practical implications for researchers and practitioners. However, further research is required to compare the effectiveness of both approaches. Umar Iftikhar, Jürgen Börstler, Nauman Bin Ali |
SEAA | 3 |
| 2023 | Software selection in large-scale software engineering: A model and criteria based on interactive rapid reviewsabstractContext: Software selection in large-scale software development continues to be ad hoc and ill-structured. Previous proposals for software component selection tend to be technology-specific and/or do not consider business or ecosystem concerns. Objective: Our main aim is to develop an industrially relevant technology-agnostic method that can support practitioners in making informed decisions when selecting software components for use in tools or in products based on a holistic perspective of the overall environment. Method: We used method engineering to iteratively develop a software selection method for Ericsson AB based on a combination of published research and practitioner insights. We used interactive rapid reviews to systematically identify and analyse scientific literature and to support close cooperation and co-design with practitioners from Ericsson. The model has been validated through a focus group and by practical use at the case company. Results: The model consists of a high-level selection process and a wide range of criteria for assessing and for evaluating software to include in business products and tools. Conclusions: We have developed an industrially relevant model for component selection through active engagement from a company. Co-designing the model based on previous knowledge demonstrates a viable approach to industry-academia collaboration and provides a practical solution that can support practitioners in making informed decisions based on a holistic analysis of business, organisation and technical factors. Elizabeth Bjarnason, Patrik Åberg, Nauman Bin Ali |
Empir. Softw. Eng. | 3 |
| 2023 | Double-counting in software engineering tertiary studies - An overlooked threat to validityabstractDouble-counting in a literature review occurs when the same data, population, or evidence is erroneously counted multiple times during synthesis. Detecting and mitigating the threat of double-counting is particularly challenging in tertiary studies. Although this topic has received much attention in the health sciences, it seems to have been overlooked in software engineering. We describe issues with double-counting in tertiary studies, investigate the prevalence of the issue in software engineering, and propose ways to identify and address the issue. We analyze 47 tertiary studies in software engineering to investigate in which ways they address double-counting and whether double-counting might be a threat to validity in them. In 19 of the 47 tertiary studies, double-counting might bias their results. Of those 19 tertiary studies, only 5 consider double-counting a threat to their validity, and 7 suggest strategies to address the issue. Overall, only 9 of the 47 tertiary studies, acknowledge double-counting as a potential general threat to validity for tertiary studies. Double-counting is an overlooked issue in tertiary studies in software engineering, and existing design and evaluation guidelines do not address it sufficiently. Therefore, we propose recommendations that may help to identify and mitigate double-counting in tertiary studies. Jürgen Börstler, Nauman Bin Ali, Kai Petersen |
Inf. Softw. Technol. | 2 |
| 2023 | A data-driven approach for understanding invalid bug reports: An industrial case studyabstractBug reports created during software development and maintenance do not always describe deviations from a system’s valid behavior. Such invalid bug reports may consume significant resources and adversely affect the prioritization and resolution of valid bug reports. There is a need to identify preventive actions to reduce the inflow of invalid bug reports. Existing research has shown that manually analyzing invalid bug report descriptions provides cues regarding preventive actions. However, such a manual approach is not cost-effective due to the time required to analyze a sufficiently large number of bug reports needed to identify useful patterns. Furthermore, the analysis needs to be repeated as the underlying causes of invalid bug reports change over time. In this study, we propose and evaluate the use of Latent Dirichlet Allocation (LDA), a topic modeling approach, to support practitioners in suggesting preventive actions to avoid the creation of similar invalid bug reports in the future. In an industrial case study, we first manually analyzed descriptions of invalid bug reports to identify common patterns in their descriptions. We further investigated to what extent LDA can support this manual process. We used expert-based validation to evaluate the relevance of identified common patterns and their usefulness in suggesting preventive measures. We found that invalid bug reports have common patterns that are perceived as relevant, and they can be used to devise preventive measures. Furthermore, the identification of common patterns can be supported with automation. Using LDA, practitioners can effectively identify representative groups of bug reports (i.e., relevant common patterns) from a large number of bug reports and analyze them further to devise preventive measures. Muhammad Laiq, Nauman Bin Ali, Jürgen Börstler, Emelie Engström |
Inf. Softw. Technol. | 2 |
| 2023 | Investigating acceptance behavior in software engineering - Theoretical perspectivesabstractSoftware engineering research aims to establish software development practice on a scientific basis. However, the evidence of the efficacy of technology is insufficient to ensure its uptake in industry. In the absence of a theoretical frame of reference, we mainly rely on best practices and expert judgment from industry-academia collaboration and software process improvement research to improve the acceptance of the proposed technology. To identify acceptance models and theories and discuss their applicability in the research of acceptance behavior related to software development. We analyzed literature reviews within an interdisciplinary team to identify models and theories relevant to software engineering research. We further discuss acceptance behavior from the human information processing perspective of automatic and affect-driven processes (“fast” system 1 thinking) and rational and rule-governed processes (“slow” system 2 thinking). We identified 30 potentially relevant models and theories. Several of them have been used in researching acceptance behavior in contexts related to software development, but few have been validated in such contexts. They use constructs that capture aspects of (automatic) system 1 and (rational) system 2 oriented processes. However, their operationalizations focus on system 2 oriented processes indicating a rational view of behavior, thus overlooking important psychological processes underpinning behavior. Software engineering research may use acceptance behavior models and theories more extensively to understand and predict practice adoption in the industry. Such theoretical foundations will help improve the impact of software engineering research. However, more consideration should be given to their validation, overlap, construct operationalization, and employed data collection mechanisms when using these models and theories. Jürgen Börstler, Nauman Bin Ali, Martin Svensson, Kai Petersen |
J. Syst. Softw. | 2 |
| 2022 | Early Identification of Invalid Bug Reports in Industrial Settings - A Case Study
Muhammad Laiq, Nauman Bin Ali, Jürgen Börstler, Emelie Engström |
PROFES | 2 |
| 2021 | Assessing test artifact quality - A tertiary studyabstractContext: Modern software development increasingly relies on software testing for an ever more frequent delivery of high quality software. This puts high demands on the quality of the central artifacts in software testing, test suites and test cases. Objective: We aim to develop a comprehensive model for capturing the dimensions of test case/suite quality, which are relevant for a variety of perspectives. Methods: We have carried out a systematic literature review to identify and analyze existing secondary studies on quality aspects of software testing artifacts. Results: We identified 49 relevant secondary studies. Of these 49 studies, less than half did some form of quality appraisal of the included primary studies and only 3 took into account the quality of the primary study when synthesizing the results. We present an aggregation of the context dimensions and factors that can be used to characterize the environment in which the test case/suite quality is investigated. We also provide a comprehensive model of test case/suite quality with definitions for the quality attributes and measurements based on findings in the literature and ISO/IEC 25010:2011. Conclusion: The test artifact quality model presented in the paper can be used to support test artifact quality assessment and improvement initiatives in practice. Furthermore, the model can also be used as a framework for documenting context characteristics to make research results more accessible for research and practice. Huynh Khanh Vi Tran, Michael Unterkalmsteiner, Jürgen Börstler, Nauman Bin Ali |
Inf. Softw. Technol. | 4 |
| 2020 | The Impact of a Proposal for Innovation Measurement in the Software IndustryabstractBackground: Measuring an organization's capability to innovate and assessing its innovation output and performance is a challenging task. Previously, a comprehensive model and a suite of measurements to support this task were proposed. Aims: In the current paper, seven years since the publication of the paper titled Towards innovation measurement in the software industry, we have reflected on the impact of the work. Method: We have mainly relied on quantitative and qualitative analysis of the citations of the paper using an established classification schema. Results: We found that the article has had a significant scientific impact (indicated by the number of citations), i.e., (1) cited in literature from both software engineering and other fields, (2) cited in grey literature and peer-reviewed literature, and (3) substantial citations in literature not published in the English language. However, we consider a majority of the citations in the peer-reviewed literature (75 out of 116) as neutral, i.e., they have not used the innovation measurement paper in any substantial way. All in all, 38 out of 116 have used, modified or based their work on the definitions, measurements or the model proposed in the article. This analysis revealed a significant weakness of the citing work, i.e., among the citing papers, we found only two explicit comparisons to the innovation measurement proposal, and we found no papers that identify weaknesses of said proposal. Conclusions: This work highlights the need for being cautious of relying solely on the number of citations for understanding impact, and the need for further improving and supporting the peer-review process to identify unwarranted citations in papers. Nauman Bin Ali, Henry Edison, Richard Torkar |
ESEM | 1 |
| 2020 | Revisiting the Impact of Concept Drift on Just-in-Time Quality AssuranceabstractThe performance of software defect prediction(SDP) models is known to be dependent on the datasets used for training the models. Evolving data in a dynamic software development environment such as significant refactoring and organizational changes introduces new concept to the prediction model, thus making improved classification performance difficult. In this study, we investigate and assess the existence and impact of concept drift on SDP performances. We empirically asses the prediction performance of five models by conducting cross-version experiments using fifty-five releases of five open-source projects. Prediction performance fluctuated as the training datasets changed over time. Our results indicate that the quality and the reliability of defect prediction models fluctuate over time and that this instability should be considered by software quality teams when using historical datasets. The performance of a static predictor constructed with data from historical versions may degrade over time due to the challenges posed by concept drift. Kwabena Ebo Bennin, Nauman Bin Ali, Jürgen Börstler, Xiao Yu 0008 |
QRS | 2 |
| 2019 | Test-Case Quality - Understanding Practitioners' Perspectives
Huynh Khanh Vi Tran, Nauman Bin Ali, Jürgen Börstler, Michael Unterkalmsteiner |
PROFES | 2 |
| 2019 | On the search for industry-relevant regression testing researchabstractRegression testing is a means to assure that a change in the software, or its execution environment, does not introduce new defects. It involves the expensive undertaking of rerunning test cases. Several techniques have been proposed to reduce the number of test cases to execute in regression testing, however, there is no research on how to assess industrial relevance and applicability of such techniques. We conducted a systematic literature review with the following two goals: firstly, to enable researchers to design and present regression testing research with a focus on industrial relevance and applicability and secondly, to facilitate the industrial adoption of such research by addressing the attributes of concern from the practitioners’ perspective. Using a reference-based search approach, we identified 1068 papers on regression testing. We then reduced the scope to only include papers with explicit discussions about relevance and applicability (i.e. mainly studies involving industrial stakeholders). Uniquely in this literature review, practitioners were consulted at several steps to increase the likelihood of achieving our aim of identifying factors important for relevance and applicability. We have summarised the results of these consultations and an analysis of the literature in three taxonomies, which capture aspects of industrial-relevance regarding the regression testing techniques. Based on these taxonomies, we mapped 38 papers reporting the evaluation of 26 regression testing techniques in industrial settings. Nauman Bin Ali, Emelie Engström, Masoumeh Taromirad, Mohammad Reza Mousavi 0001, Nasir Mehmood Minhas, Daniel Helgesson, Sebastian Kunze, Mahsa Varshosaz |
Empir. Softw. Eng. | 1 |
| 2019 | A critical appraisal tool for systematic literature reviews in software engineeringabstractContext: Methodological research on systematic literature reviews (SLRs) in Software Engineering (SE) has so far focused on developing and evaluating guidelines for conducting systematic reviews. However, the support for quality assessment of completed SLRs has not received the same level of attention. Objective: To raise awareness of the need for a critical appraisal tool (CAT) for assessing the quality of SLRs in SE. To initiate a community-based effort towards the development of such a tool. Method: We reviewed the literature on the quality assessment of SLRs to identify the frequently used CATs in SE and other fields. Results: We identified that the CATs currently used is SE were borrowed from medicine, but have not kept pace with substantial advancements in the field of medicine. Conclusion: In this paper, we have argued the need for a CAT for quality appraisal of SLRs in SE. We have also identified a tool that has the potential for application in SE. Furthermore, we have presented our approach for adapting this state-of-the-art CAT for assessing SLRs in SE. Nauman Bin Ali, Muhammad Usman 0002 |
Inf. Softw. Technol. | 1 |
| 2019 | Special issue on evaluation and assessment in software engineering
Emilia Mendes, Nauman Bin Ali, Steve Counsell, Maria Teresa Baldassarre |
J. Syst. Softw. | 2 |
| 2019 | An evaluation of effort estimation supported by change impact analysis in agile software developmentabstractAbstract In agile software development, functionality is added to the system in an incremental and iterative manner. Practitioners often rely on expert judgment to estimate the effort in this context. However, the impact of a change on the existing system can provide objective information to practitioners to arrive at an informed estimate. In this regard, we have developed a hybrid method, that utilizes change impact analysis information for improving effort estimation. We also developed an estimation model based on gradient boosted trees (GBT). In this study, we evaluate the performance and usefulness of our hybrid method with tool support and the GBT model in a live iteration at Insiders Technologies GmbH, a German software company. Additionally, the solution was also assessed for perceived usefulness and understandability in a study with graduate and post‐graduate students. The results from the industrial evaluation show that the proposed method produces more accurate estimates than only expert‐based or only model‐based estimates. Furthermore, both students and practitioners perceived the usefulness and understandability of the method positively. Binish Tanveer, Anna Maria Vollmer, Nauman Bin Ali |
J. Softw. Evol. Process. | 4 |
| 2018 | Reliability of search in systematic reviews: Towards a quality assessment framework for the automated-search strategy
Nauman Bin Ali, Muhammad Usman 0002 |
Inf. Softw. Technol. | 1 |
| 2018 | Towards a benefits dependency network for DevOps based on a systematic literature reviewabstractAbstract DevOps as a new way of thinking for software development and operations has received much attention in the industry, while it has not been thoroughly investigated in academia yet. The objective of this study is to characterize DevOps by exploring its central components in terms of principles, practices and their relations to the principles, challenges of DevOps adoption, and benefits reported in the peer‐reviewed literature. As a key objective, we also aim to realize the relations between DevOps practices and benefits in a systematic manner. A systematic literature review was conducted. Also, we used the concept of benefits dependency network to synthesize the findings, in particular, to specify dependencies between DevOps practices and link the practices to benefits. We found that in many cases, DevOps characteristics, ie, principles, practices, benefits, and challenges, were not sufficiently defined in detail in the peer‐reviewed literature. In addition, only a few empirical studies are available, which can be attributed to the nascency of DevOps research. Also, an initial version of the DevOps benefits dependency network has been derived. The definition of DevOps principles and practices should be emphasized given the novelty of the concept. Further empirical studies are needed to improve the benefits dependency network presented in this study. Ramtin Jabbari, Nauman Bin Ali, Kai Petersen, Binish Tanveer |
J. Softw. Evol. Process. | 2 |
| 2017 | SERP-test: a taxonomy for supporting industry-academia communication
Emelie Engström, Kai Petersen, Nauman Bin Ali, Elizabeth Bjarnason |
Softw. Qual. J. | 3 |
| 2016 | Is effectiveness sufficient to choose an intervention?: Considering resource use in empirical software engineeringabstractContext: Software Engineering (SE) research with a scientific foundation aims to influence SE practice to enable and sustain efficient delivery of high quality software. Goal: To improve the impact of SE research, one objective is to facilitate practitioners in choosing empirically vetted interventions. Method: Literature from evidence-based medicine, economic evaluations in SE and software economics is reviewed. Results: In empirical SE research, the emphasis has been on substantiating the claims about the benefits of proposed interventions. However, to support informed decision making by practitioners regarding technology adoption, we must present a business case for these interventions, which should comprise not just effectiveness, but also the evidence of cost-effectiveness. Conclusions: This paper highlights the need to investigate and report the resources required to adopt an intervention. It also provides some guidelines and examples to improve support for practitioners in decisions regarding technology adoption. Nauman Bin Ali |
ESEM | 1 |
| 2016 | FLOW-assisted value stream mapping in the early phases of large-scale software development
Nauman Bin Ali, Kai Petersen, Kurt Schneider |
J. Syst. Softw. | 1 |
| 2015 | Evaluation of simulation-assisted value stream mapping for software product development: Two industrial cases
Nauman Bin Ali, Kai Petersen, Breno B. N. de França |
Inf. Softw. Technol. | 1 |
| 2014 | Evaluating strategies for study selection in systematic literature studiesabstractContext: The study selection process is critical to improve the reliability of secondary studies. Goal: To evaluate the selection strategies commonly employed in secondary studies in software engineering. Method: Building on these strategies, a study selection process was formulated and evaluated in a systematic review. Results: The selection process used a more inclusive strategy than the one typically used in secondary studies, which led to additional relevant articles. Conclusions: The results indicates that a good-enough sample could be obtained by following a less inclusive but more efficient strategy, if the articles identified as relevant for the study are a representative sample of the population, and there is a homogeneity of results and quality of the articles. Nauman Bin Ali, Kai Petersen |
ESEM | 1 |
| 2014 | A systematic literature review on the industrial use of software process simulation
Nauman Bin Ali, Kai Petersen, Claes Wohlin |
J. Syst. Softw. | 1 |
| 2013 | Towards innovation measurement in the software industry
Henry Edison, Nauman Bin Ali, Richard Torkar |
J. Syst. Softw. | 2 |
| 2012 | Testing highly complex system of systems: an industrial case studyabstractContext: Systems of systems (SoS) are highly complex and are integrated on multiple levels (unit, component, system, system of systems). Many of the characteristics of SoS (such as operational and managerial independence, integration of system into system of systems, SoS comprised of complex systems) make their development and testing challenging. Nauman Bin Ali, Kai Petersen, Mika Mäntylä |
ESEM | 1 |
| 2011 | Identifying Strategies for Study Selection in Systematic Reviews and MapsabstractStudy selection in systematic reviews is prone to bias and there exist no commonly defined strategies of how to reduce the bias and resolve disagreement between researchers. This study aims at identifying strategies for bias reduction and disagreement resolution. A review of existing systematic reviews is conducted for study selection strategy identification. In total 13 different strategies have been identified. Kai Petersen, Nauman Bin Ali |
ESEM | 2 |