Wasif Afzal

dblp:59/5422 · DBLP profile ↗
← Back
49ranked-venue papers
10as first author
15since 2021 · last 2025
0000-0003-0611-2655ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 40 · 8 first-author · 11 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 1 since 2021Systems, architecture and hardware · 4 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Security and privacy · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 PyLC+: A Scalable Python Framework for Automated Translation and Testing of Industrial PLC Programs
abstract
As industrial PLC programs become more complex, automated testing and verification methods are needed to ensure their reliability and correctness. This paper presents PyLC+, a modular framework that translates PLC programs into Python, allowing for automated AI-driven test generation. PyLC+ builds upon our previous work, addressing limitations by adopting a class-based modular architecture that improves the tool’s scalability, maintainability, and extensibility. This structural refinement eliminates reliance on nested functions, facilitating the translation of large-scale, real-world PLC programs while maintaining precise use of cyclic execution. Furthermore, PyLC+ introduces automated handling of stateful FBs, ensuring compliance with IEC 61131-3 execution semantics.Additionally, the tool proposes integrating LLM-driven test generation with search-based test generation to improve the efficiency and effectiveness of testing PLC software. We tested PyLC+ in a large-scale company developing train control systems, demonstrating its efficiency and effectiveness in handling complex industrial PLC programs.
Mikael Ebrahimi Salari, Eduard Paul Enoiu, Alessio Bucaioni, Wasif Afzal, Cristina Cerschi Seceleanu
COMPSAC4
2025 Gaps in Software Testing Education: A Survey of Academic Courses in Sweden
abstract
A cross-sectional, questionnaire-based survey of software testing courses offered at Swedish universities was undertaken in the final quarter of 2023. With a return rate of 44%, the survey delved into the contents of these software testing courses to gain an understanding of how the courses differ in terms of depth and breadth of content. Information was also sought about administrative and course planning activities related to the courses. Some key findings are that there is in-depth coverage of unit testing in all the courses, with none of the courses offering in-depth testing of other test levels such as acceptance testing. Also notable is the difference in test types. As an example, functional testing is taught in-depth in all the courses, while accessibility testing is not taught at all in half of the courses. It is suggested that a greater range of software testing topics is needed in future education if more stakeholders, such as business analysts and software developers, not just software testers, are to have a quality-centred approach to software development.
Ayodele A. Barrett, Eduard Paul Enoiu, Wasif Afzal
CSEE&T3
2025 Passive Testing of Vehicular Embedded Systems: An Industrial Case Study with T-EARS and Napkin Studio
abstract
Abstract Passive testing is an approach to verify system behavior by observing logs from normal operation, without actively injecting test stimuli. This paper presents an industrial case study of applying passive testing in the domain of vehicular embedded systems, utilizing two specialized tools: Timed Easy Approach to Requirements Syntax (T-EARS) for specifying temporal requirements, and Napkin Studio for evaluating these requirements against real system execution logs. We collaborated with Volvo Construction Equipment (VCE) to translate a set of natural language requirements into structured T-EARS specifications. Then we used Napkin Studio to test these requirements against recorded machine log data passively. We evaluate the feasibility of this approach, the extent to which it can detect requirement violations or injected faults, and the perceptions of industry stakeholders regarding the adoption of such passive tests in their verification process. The results show that a majority of functional requirements can be expressed as Guarded Assertions (GAs) and validated on logs, uncovering specific issues. Stakeholders found the method promising for improving test coverage and efficiency, although integration challenges (e.g., log signal inconsistencies and tool usability issues) were noted. Overall, this work provides empirical evidence that passive testing with T-EARS and Napkin Studio can complement traditional hardware-in-the-loop testing, offering a scalable and non-intrusive verification approach in developing vehicular systems.
Aleksandra Nicaj, Daniel Flemström, Eduard Paul Enoiu, Wasif Afzal
ICTSS4
2025 Comparative analysis of text mining and clustering techniques for assessing functional dependency between manual test cases
abstract
Abstract Text mining techniques, particularly those leveraging machine learning for natural language processing, have gained significant attention for qualitative data analysis in software testing. However, their complexity and lack of transparency can pose challenges, especially in safety-critical domains where simpler, interpretable solutions are often preferred unless accuracy is heavily compromised. This study investigates the trade-offs between complexity, effort, accuracy, and utility in text mining and clustering techniques, focusing on their application for detecting functional dependencies among manual integration test cases in safety-critical systems. Using empirical data from an industrial testing project at ALSTOM Sweden, we evaluate various string distance methods, NCD compressors, and machine learning approaches. The results highlight the impact of preprocessing techniques, such as tokenization, and intrinsic factors, such as text length, on algorithm performance. Findings demonstrate how text mining and clustering can be optimized for safety-critical contexts, offering actionable insights for researchers and practitioners aiming to balance simplicity and effectiveness in their testing workflows.
Sahar Tahvili, Leo Hatvani, Michael Felderer, Francisco Gomes de Oliveira Neto, Wasif Afzal, Robert Feldt
Softw. Qual. J.5
2024 Can serious gaming tactics bolster spear-phishing and phishing resilience? : Securing the human hacking in Information Security
abstract
In the digital age, there is a notable increase in fraudulent activities perpetrated by social engineers who exploit individuals’ limited knowledge of digital devices. These actors strategically manipulate human psychology, targeting IT devices to gain unauthorized access to sensitive data. Our study is centered around two distinct objectives to be accomplished through the utilization of a serious game: (i) The primary objective entails delivering training and educational content to participants with a focus on phishing attacks; (ii) The secondary objective aims to heighten participants’ awareness regarding the perils associated with divulging excessive information online. To address these objectives, we have employed the following techniques and methods: (i) A comprehensive literature review was conducted to establish foundational knowledge in areas such as social engineering, game design, learning principles, human interaction, and game-based learning; (ii) We meticulously aligned the game design with the philosophical concept of social engineering attacks; (iii) We devised and crafted an advanced hybrid version of the game, incorporating the use of QR codes to generate game card data; (iv) We conducted an empirical evaluation encompassing surveys, observations, discussions, and URL assessments to assess the effectiveness of the proposed hybrid game version. Quantitative data and qualitative observations suggest the “PhishDefend Quest” game successfully improved players’ comprehension of phishing threats and how to detect them through an interactive learning experience. The results highlight the potential of serious games to educate people about social engineering risks. Through the evaluation, we can readily arrive at the following conclusions: (i) Game-based learning proves to be a viable approach for educating participants about phishing awareness and the associated risks tied to the unnecessary disclosure of sensitive information online; (ii) Furthermore, game-based learning serves as an effective means of disseminating awareness among participants and players concerning prevalent phishing attacks.
Affan Yasin, Rubia Fatima, Wasif Afzal, Shahid Raza
Inf. Softw. Technol.4
2023 Automating Test Generation of Industrial Control Software Through a PLC-to-Python Translation Framework and Pynguin
abstract
Numerous industrial sectors employ Programmable Logic Controllers (PLC) software to control safety-critical systems. These systems necessitate extensive testing and stringent coverage measurements, which can be facilitated by automated test-generation techniques. Existing such techniques have not been applied to PLC programs, and therefore do not directly support the latter regarding automated test-case generation. To address this deficit, in this work, we introduce PyLC, a tool designed to automate the conversion of PLC programs to Python code, assisted by an existing test generator called Pynguin. Our framework is capable of handling PLC programs written in the Function Block Diagram language. To demonstrate its capabilities, we employ PyLC to transform safety-critical programs from industry and illustrate how our approach can facilitate the manual and automatic creation of tests. Our study highlights the efficacy of leveraging Python as an intermediary language to bridge the gap between PLC development tools, Python-based unit testing, and automated test generation.
Mikael Ebrahimi Salari, Eduard Paul Enoiu, Cristina Cerschi Seceleanu, Wasif Afzal, Filip Sebek
APSEC4
2023 DevOps for Cyber-Physical Systems: Objectives, Results and Lessons Learned from the Adeptness H2020 Project
abstract
While most large web-based software systems (e.g., Amazon, Google) release a new software version every almost a minute, in the context of Cyber-Physical Systems (CPSs), this is still far. However, the software of CPSs needs to evolve while these are in operation to fix bugs, add new functionalities, carry out refactoring activities and deal with unforeseen situations that were discovered while the CPS was operating. In the last three years, the Adeptness project has been developing in a solution to help speedup the software release of CPSs that are in operation while guaranteeing their reliability. In this paper, we summarize the objectives, results and lessons learned from this H2020 project.
Aitor Arrieta, Goiuria Sagardui Mendieta, Aitor Agirre, Wasif Afzal, Shaukat Ali 0001
DSD4
2023 An Empirical Evaluation of System-Level Test Effectiveness for Safety-Critical Software
Muhammad Nouman Zafar, Wasif Afzal, Eduard Paul Enoiu
ENASE2
2023 Retrieving arXiv, SocArXiv, and SSRN metadata for initial review screening
Rubia Fatima, Affan Yasin, Lin Liu 0001, Jianmin Wang 0001, Wasif Afzal
Inf. Softw. Technol.5
2022 Evaluating System-Level Test Generation for Industrial Software: A Comparison between Manual, Combinatorial and Model-Based Testing
abstract
Adequate testing of safety-critical systems is vital to ensure correct functional and non-functional operations. Previous research has shown that testing such systems requires a lot of effort, thus automated testing techniques have found a certain degree of success. However, automated testing has not replaced the need for manual testing, rather a common industrial practice exhibits a balance between automated and manual testing. In this respect, comparing manual testing with automated testing techniques continues to be an interesting topic to investigate. The need for this investigation is most apparent at system-level testing of industrial systems, where there is a lack of results on how different testing techniques perform concerning both structural and system-level metrics such as Modified Condition/Decision Coverage (MC/DC) and requirement coverage. In addition to the coverage, the cost of these techniques will also determine their efficiency and thus practical viability. In this paper, we have developed cost models for efficiency measurement and performed an experimental evaluation of manual testing, model-based testing (MBT), and combinatorial testing (CT) in terms of MC/DC and requirement coverage. The evaluation is done in an industrial context of a safety-critical system that controls several functions on-board the passenger trains. We have reported the dominant conditions of MC/DC affected by each technique while generating MC/DC adequate test suites. Moreover, we investigated differences and overlaps of test cases generated by each of the three techniques. The results showed that all test suites achieved 100% requirement coverage except the test suite generated by the pairwise testing strategy. However, MBT-generated test suites were more MC/DC adequate and provided a higher number of both similar and unique test cases. Moreover, unique test cases generated by MBT had an observable effect on MC/DC, which will complement manual testing to increase MC/DC coverage. The least dominant MC/DC condition fulfilled by the generated test cases by all three techniques is the 'independent effect of a condition on the outcomes of a decision'. Lastly, the evaluation also showed CT as the most efficient testing technique amongst the three in terms of time required for test generation and execution, but with an added cost parameter of manual identification of expected outcomes.
Muhammad Nouman Zafar, Wasif Afzal, Eduard Paul Enoiu
AST2
2022 SmartDelta: Automated Quality Assurance and Optimization in Incremental Industrial Software Systems Development
abstract
A common phenomenon in software development is that as a system is being built and incremented with new features, certain quality aspects of the system begin to deteriorate. Therefore, it is important to be able to accurately analyze and determine the quality implications of each change and increment to a system. To address this topic, the multinational SmartDelta project develops automated solutions for quality assessment of product deltas in a continuous engineering environment. The project will provide smart analytics from development artifacts and system executions, offering insights into quality degradation or improvements across different product versions, and providing recommendations for next builds.
Mehrdad Saadatmand, Eduard Paul Enoiu, Holger Schlingloff, Michael Felderer, Wasif Afzal
DSD5
2022 Resilient Conflict-free Replicated Data Types without Atomic Broadcast
abstract
In a distributed system, applications can perform both reads and updates without costly synchronous network round-trips by using Conflict-free Replicated Data Types (CRDTs). Most CRDTs are based on some variant of atomic broadcast, as that enables them to support causal dependencies between updates of multiple objects. However, the overhead of this atomic broadcast is unnecessary in systems handling only independent CRDT objects. We identified a set of use cases for tracking resource usage where there is a need for a replication mechanism with less complexity and network usage as compared to using atomic broadcast. In this paper, we present the design of such a replication protocol that efficiently leverages the commutativity of CRDTs. The proposed protocol CReDiT (CRDT enhanced with intelligence) uses up to four communication steps per update, but these steps can be batched as needed. It uses network resources only when updates need to be communicated. Furthermore, it is less sensit ive to server failures than current state-of-the-art solutions as other nodes can use new values already after the first communication step, instead of after two or more.
Daniel Brahneborg, Wasif Afzal, Saad Mubeen
ICSOFT2
2022 Machine learning techniques for software vulnerability prediction: a comparative study
Gul Jabeen, Sabit Rahim, Wasif Afzal, Dawar Khan, Aftab Ahmed Khan, Tehmina Bibi
Appl. Intell.3
2022 Software test results exploration and visualization with continuous integration and nightly testing
abstract
Abstract Software testing is key for quality assurance of embedded systems. However, with increased development pace, the amount of test results data risks growing to a level where exploration and visualization of the results are unmanageable. This paper covers a tool, Tim, implemented at a company developing embedded systems, where software development occurs in parallel branches and nightly testing is partitioned over software branches, test systems and test cases. Tim aims to replace a previous solution with problems of scalability, requirements and technological flora. Tim was implemented with a reference group over several months. For validation, data were collected both from reference group meetings and logs from the usage of the tool. Data were analyzed quantitatively and qualitatively. The main contributions from the study include the implementation of eight views for test results exploration and visualization, the identification of four solutions patterns for these views (filtering, aggregation, previews and comparisons), as well as six challenges frequently discussed at reference group meetings (expectations, anomalies, navigation, integrations, hardware details and plots). Results are put in perspective with related work and future work is proposed, e.g., enhanced anomaly detection and integrations with more systems such as risk management, source code and requirements repositories.
Per Erik Strandberg, Wasif Afzal, Daniel Sundmark
Int. J. Softw. Tools Technol. Transf.2
2021 Industrial Scale Passive Testing with T-EARS
abstract
Passive testing continuously observes the system or system execution logs without any interference or instrumentation to test diverse combinations of functions, resulting in a more thorough evaluation over time. However, reaching a working solution to passive testing is not without challenges. While there have been some efforts to extract information from system requirements to create passive test cases, to our knowledge, no such efforts are mature enough to be applied in a real, industrial safety-critical context. Our passive testing approach uses the Timed Easy Approach to Requirements Syntax (T-EARS) specification language and its accompanying tool-chain. This study reports challenges and solutions to introducing system-level passive testing for a vehicular safety-critical system through industrial data analysis, including 116 safety-related requirements. Our results show that passive testing using the T-EARS language and its tool-chain can be used for system-level testing in an industrial setting for 64% of the studied requirements. We identified several sources of false positive results and show how to tune test cases to reduce such false positives systematically. Finally, we show the requirement coverage achieved by a manual test session and that passive testing using T-EARS can find a set of injected faults that are considered hard to find with other test techniques.
Daniel Flemström, Henrik Jonsson, Eduard Paul Enoiu, Wasif Afzal
ICST4
2020 Superlinear and Bandwidth Friendly Geo-replication for Store-and-forward Systems
Daniel Brahneborg, Wasif Afzal, Adnan Causevic, Mats Björkman
ICSOFT2
2020 Intermittently failing tests in the embedded systems domain
abstract
Software testing is sometimes plagued with intermittently failing tests and finding the root causes of such failing tests is often difficult. This problem has been widely studied at the unit testing level for open source software, but there has been far less investigation at the system test level, particularly the testing of industrial embedded systems. This paper describes our investigation of the root causes of intermittently failing tests in the embedded systems domain, with the goal of better understanding, explaining and categorizing the underlying faults. The subject of our investigation is a currently-running industrial embedded system, along with the system level testing that was performed. We devised and used a novel metric for classifying test cases as intermittent. From more than a half million test verdicts, we identified intermittently and consistently failing tests, and identified their root causes using multiple sources. We found that about 1-3% of all test cases were intermittently failing. From analysis of the case study results and related work, we identified nine factors associated with test case intermittence. We found that a fix for a consistently failing test typically removed a larger number of failures detected by other tests than a fix for an intermittent test. We also found that more effort was usually needed to identify fixes for intermittent tests than for consistent tests. An overlap between root causes leading to intermittent and consistent tests was identified. Many root causes of intermittence are the same in industrial embedded systems and open source software. However, when comparing unit testing to system level testing, especially for embedded systems, we observed that the test environment itself is often the cause of intermittence.
Per Erik Strandberg, Thomas J. Ostrand, Elaine J. Weyuker, Wasif Afzal, Daniel Sundmark
ISSTA4
2020 Towards a Taxonomy for Eliciting Design-Operation Continuum Requirements of Cyber-Physical Systems
abstract
Software systems that are embedded in autonomous Cyber-Physical Systems (CPSs) usually have a large life-cycle, both during its development and in maintenance. This software evolves during its life-cycle in order to incorporate new requirements, bug fixes, and to deal with hardware obsolescence. The current process for developing and maintaining this software is very fragmented, which makes developing new software versions and deploying them in the CPSs extremely expensive. In other domains, such as web engineering, the phases of development and operation are tightly connected, making it possible to easily perform software updates of the system, and to obtain operational data that can be analyzed by engineers at development time. However, in spite of the rise of new communication technologies (e.g., 5G) providing an opportunity to acquire Design-Operation Continuum Engineering methods in the context of CPSs, there are still many complex issues that need to be addressed, such as the ones related with hardware-software co-design. Therefore, the process of Design-Operation Continuum Engineering for CPSs requires substantial changes with respect to the current fragmented software development process. In this paper, we build a taxonomy for Design-Operation Continuum Engineering of CPSs based on case studies from two different industrial domains involving CPSs (elevation and railway). This taxonomy is later used to elicit requirements from these two case studies in order to present a blueprint on adopting Design-Operation Continuum Engineering in any organization developing CPSs.
Jon Ayerdi, Aitor Gartziandia, Aitor Arrieta, Wasif Afzal, Eduard Paul Enoiu, Aitor Agirre, Goiuria Sagardui Mendieta, Maite Arratibel, Ola Sellin
RE4
2020 A novel methodology to classify test cases using natural language processing and imbalanced learning
abstract
Detecting the dependency between integration test cases plays a vital role in the area of software test optimization. Classifying test cases into two main classes – dependent and independent – can be employed for several test optimization purposes such as parallel test execution, test automation, test case selection and prioritization, and test suite reduction. This task can be seen as an imbalanced classification problem due to the test cases’ distribution. Often the number of dependent and independent test cases is uneven, which is related to the testing level, testing environment and complexity of the system under test. In this study, we propose a novel methodology that consists of two main steps. Firstly, by using natural language processing we analyze the test cases’ specifications and turn them into a numeric vector. Secondly, by using the obtained data vectors, we classify each test case into a dependent or an independent class. We carry out a supervised learning approach using different methods for handling imbalanced datasets. The feasibility and possible generalization of the proposed methodology is evaluated in two industrial projects at Bombardier Transportation, Sweden, which indicates promising results.
Sahar Tahvili, Leo Hatvani, Enislay Ramentol, Rita Pimentel, Wasif Afzal, Francisco Herrera
Eng. Appl. Artif. Intell.5
2020 An evaluation of Monte Carlo-based hyper-heuristic for interaction testing of industrial embedded software applications
abstract
Abstract Hyper-heuristic is a new methodology for the adaptive hybridization of meta-heuristic algorithms to derive a general algorithm for solving optimization problems. This work focuses on the selection type of hyper-heuristic, called the exponential Monte Carlo with counter (EMCQ). Current implementations rely on the memory-less selection that can be counterproductive as the selected search operator may not (historically) be the best performing operator for the current search instance. Addressing this issue, we propose to integrate the memory into EMCQ for combinatorial t-wise test suite generation using reinforcement learning based on the Q-learning mechanism, called Q-EMCQ. The limited application of combinatorial test generation on industrial programs can impact the use of such techniques as Q-EMCQ. Thus, there is a need to evaluate this kind of approach against relevant industrial software, with a purpose to show the degree of interaction required to cover the code as well as finding faults. We applied Q-EMCQ on 37 real-world industrial programs written in Function Block Diagram (FBD) language, which is used for developing a train control management system at Bombardier Transportation Sweden AB. The results show that Q-EMCQ is an efficient technique for test case generation. Addition- ally, unlike the t-wise test suite generation, which deals with the minimization problem, we have also subjected Q-EMCQ to a maximization problem involving the general module clustering to demonstrate the effectiveness of our approach. The results show the Q-EMCQ is also capable of outperforming the original EMCQ as well as several recent meta/hyper-heuristic including modified choice function, Tabu high-level hyper-heuristic, teaching learning-based optimization, sine cosine algorithm, and symbiotic optimization search in clustering quality within comparable execution time.
Bestoun S. Ahmed, Eduard Paul Enoiu, Wasif Afzal, Kamal Zuhairi Zamli
Soft Comput.3
2019 On the Use of Hackathons to Enhance Collaboration in Large Collaborative Projects : - A Preliminary Case Study of the MegaM@Rt2 EU Project -
abstract
In this paper, we present the MegaM@Rt2 ECSEL project and discuss in details our approach for fostering collaboration in this project. We choose to use an internal hackathon approach that focuses on technical collaboration between case study owners and tool/method providers. The novelty of the approach is that we organize the technical workshop at our regular project progress meetings as a challenge-based contest involving all partners in the project. Case study partners submit their challenges related to the project goals and their use cases in advance. These challenges are concise enough to be experimented within approximately 4 hours. Teams are then formed to address those challenges. The teams include tool/method providers, case study owners and researchers/developers from other consortium members. On the hackathon day, partners work together to come with results addressing the challenges that are both interesting to encourage collaboration and convincing to continue further deeper investigations. Obtained results demonstrate that the hackathon approach stimulated knowledge exchanges among project partners and triggered new collaborations, notably between tool providers and use case owners.
Andrey Sadovykh, Dragos Truscan, Pierluigi Pierini, Gunnar Widforss, Adnan Ashraf, Hugo Bruneliere, Pavel Smrz, Alessandra Bagnato, Wasif Afzal, Alexandra Espinosa Hortelano
DATE9
2019 Improving software requirements reasoning by novices: a story-based approach
abstract
Requirements elicitation is one of the essential steps towards software design and construction. Business analysts and stakeholders often face challenges in gathering or conveying key software requirements. There are many methods and tools designed by researchers and practitioners but with the persistent development of new technologies, there is a need to make requirements gathering and design‐rationale process more efficient and adaptable. Storytelling is an emerging concept and researchers are witnessing its effectiveness in education, community building, information system, and requirement elicitation. Objectives of this study are to devise a method for requirements elicitation and improving design‐rationales using story‐based techniques and evaluate the effectiveness of the proposed activity. To answer the research objectives, the authors have conducted open‐ended interviews to get feedback on the proposed method; the authors have case requirement from a running project to map how this method can be useful; and performed empirical evaluation of the proposed card‐based activity. The estimated regression model, in our study, has shown that participants' perception about the simplicity/easiness and the joy of playing the game has an eventual positive effect on requirements elicitation through enhancing user's desire to play the game, which in turn increases the collaborative learning outcomes of the game.
Rubia Fatima, Affan Yasin, Lin Liu 0001, Jianmin Wang 0001, Wasif Afzal, Atif Yasin
IET Softw.5
2019 Sharing information online rationally: An observation of user privacy concerns and awareness using serious game
Rubia Fatima, Affan Yasin, Lin Liu 0001, Jianmin Wang 0001, Wasif Afzal, Awaid Yasin
J. Inf. Secur. Appl.5
2019 An improved software reliability prediction model by using high precision error iterative analysis method
abstract
Summary Software reliability deals with the probability that software will not cause the failure of a system in a specified time interval. Software reliability growth models (SRGMs) are used to predict future behaviour from known characteristics of software, like historical failures. With the increasing demand to deliver quality software, more accurate SRGMs are required to estimate the software release time and cost of the testing effort. Software failure predictions at early phases also provide an opportunity for investing in proper quality assurance and upfront resource planning. Up till now, many parametric software reliability growth models (PSRGMs) have been proposed. However, several limitations of them mean that their predictive capacities differ from one dataset to others. In this paper, to enhance the prediction accuracy of existing PSRGMs, a high precision error iterative analysis method (HPEIAM) has been proposed based on the residual errors. In HPEIAM, residual errors from the estimated results of SRGMs are considered as another source of data that can combine the residual error modification with artificial neural network sign estimator. The repeated computation of residual errors by SRGMs improves and corrects the prediction accuracy up to the expected level. The performance of HPEIAM is tested with several PSRGMs using two sets of real software failure data based on three performance criteria. Moreover, we have compared the estimated failures predicted by HPEIAM with genetic algorithm (GA)‐based prediction improvement. The results demonstrate that HPEIAM gives an improvement in goodness‐of‐fit and predictive performance for every PSRGM in initial few iterations.
Gul Jabeen, Ping Luo 0004, Wasif Afzal
Softw. Test. Verification Reliab.3
2018 Decision making and visualizations based on test results
abstract
Background: Testing is one of the main methods for quality assurance in the development of embedded software, as well as in software engineering in general. Consequently, test results (and how they are reported and visualized) may substantially influence business decisions in software-intensive organizations. Aims: This case study examines the role of test results from automated nightly software testing and the visualizations for decision making they enable at an embedded systems company in Sweden. In particular, we want to identify the use of the visualizations for supporting decisions from three aspects: in daily work, at feature branch merge, and at release time. Method: We conducted an embedded case study with multiple units of analysis by conducting interviews, questionnaires, using archival data and participant observations. Results: Several visualizations and reports built on top of the test results database are utilized in supporting daily work, merging a feature branch to the master and at release time. Some important visualizations are: lists of failing test cases, easy access to log files, and heatmap trend plots. The industrial practitioners perceived the visualizations and reporting as valuable, however they also mentioned several areas of improvement such as better ways of visualizing test coverage in a functional area as well as better navigation between different views. Conclusions: We conclude that visualizations of test results are a vital decision making tool for a variety of roles and tasks in embedded software development, however the visualizations need to be continuously improved to keep their value for its stakeholders.
Per Erik Strandberg, Wasif Afzal, Daniel Sundmark
ESEM2
2018 Automated test mapping and coverage for network topologies
abstract
Communication devices such as routers and switches play a critical role in the reliable functioning of embedded system networks. Dozens of such devices may be part of an embedded system network, and they need to be tested in conjunction with various computational elements on actual hardware, in many different configurations that are representative of actual operating networks. An individual physical network topology can be used as the basis for a test system that can execute many test cases, by identifying the part of the physical network topology that corresponds to the configuration required by each individual test case. Given a set of available test systems and a large number of test cases, the problem is to determine for each test case, which of the test systems are suitable for executing the test case, and to provide the mapping that associates the test case elements (the logical network topology) with the appropriate elements of the test system (the physical network topology).
Per Erik Strandberg, Thomas J. Ostrand, Elaine J. Weyuker, Daniel Sundmark, Wasif Afzal
ISSTA5
2018 From Natural Language Requirements to Passive Test Cases Using Guarded Assertions
abstract
In large-scale embedded system development, requirements are often expressed in natural language. Translating these requirements to executable test cases, while keeping the test cases and requirements aligned, is a challenging task. While such a transformation typically requires extensive domain knowledge, we show that a systematic process in combination with passive testing would facilitate the translation as well as linking the requirements to tests. Passive testing approaches observe the behavior of the system and test their correctness without interfering with the normal behavior. We use a specific approach to passive testing: guarded assertions (G/A). This paper presents a method for transforming system requirements expressed in natural language into G/As. We further present a proof of concept evaluation, performed at Bombardier Transportation Sweden AB, in which we show how the process would be used, together with practical advice of the reasoning behind the translation steps.
Daniel Flemström, Eduard Paul Enoiu, Wasif Afzal, Daniel Sundmark, Thomas Gustafsson, Avenir Kobetski
QRS3
2018 Model-based product line engineering in an industrial automotive context: an exploratory case study
abstract
Product Line Engineering is an approach to reuse assets of complex systems by taking advantage of commonalities between product families. Reuse within complex systems usually means reuse of artifacts from different engineering domains such as mechanical, electronics and software engineering. Model-based systems engineering is becoming a standard for systems engineering and collaboration within different domains. This paper presents an exploratory case study on initial efforts of adopting Product Line Engineering practices within the model-based systems engineering process at Volvo Construction Equipment (Volvo CE), Sweden. We have used SysML to create overloaded models of the engine systems at Volvo CE. The variability within the engine systems was captured by using the Orthogonal Variability Modeling language. The case study has shown us that overloaded SysML models tend to become complex even on small scale systems, which in turn makes scalability of the approach a major challenge. For successful reuse and to, possibly, tackle scalability, it is necessary to have a database of reusable assets from which product variants can be derived.
Damir Bilic, Daniel Sundmark, Wasif Afzal, Peter Wallin, Adnan Causevic, Christoffer Amlinger
SPLC (2)3
2018 Round-Trip Time Anomaly Detection
abstract
Mobile text messages (SMS) are sometimes used for authentication, which requires short and reliable delivery times. The observed round-trip times when sending an SMS message provide valuable information on the quality of the connection. In this industry paper, we propose a method for detecting round-trip time anomalies, where the exact distribution is unknown, the variance is several orders of magnitude, and there are lots of shorter spikes that should be ignored. In particular, we show that using an adaption of Double Seasonal Exponential Smoothing to reduce the content dependent variations, followed by the Remedian to find short-term and long-term medians, successfully identifies larger groups of outliers. As training data for our method we use log files from a live SMS gateway. In order to verify the effectiveness of our approach, we utilize simulated data. Our contributions are a description on how to isolate content dependent variations, and the sequence of steps to find significant anomalies in big data.
Daniel Brahneborg, Wasif Afzal, Adnan Causevic, Daniel Sundmark, Mats Björkman
ICPE2
2018 ESPRET: A tool for execution time estimation of manual test cases
Sahar Tahvili, Wasif Afzal, Mehrdad Saadatmand, Markus Bohlin, Sharvathul Hasan Ameerjan
J. Syst. Softw.2
2018 Similarity-based prioritization of test case automation
abstract
The importance of efficient software testing procedures is driven by an ever increasing system complexity as well as global competition. In the particular case of manual test cases at the system integration level, where thousands of test cases may be executed before release, time must be well spent in order to test the system as completely and as efficiently as possible. Automating a subset of the manual test cases, i.e, translating the manual instructions to automatically executable code, is one way of decreasing the test effort. It is further common that test cases exhibit similarities, which can be exploited through reuse when automating a test suite. In this paper, we investigate the potential for reducing test effort by ordering the test cases before such automation, given that we can reuse already automated parts of test cases. In our analysis, we investigate several approaches for prioritization in a case study at a large Swedish vehicular manufacturer. The study analyzes the effects with respect to test effort, on four projects with a total of 3919 integration test cases constituting 35,180 test steps, written in natural language. The results show that for the four projects considered, the difference in expected manual effort between the best and the worst order found is on average 12 percentage points. The results also show that our proposed prioritization method is nearly as good as more resource demanding meta-heuristic approaches at a fraction of the computational time. Based on our results, we conclude that the order of automation is important when the set of test cases contain similar steps (instructions) that cannot be removed, but are possible to reuse. More precisely, the order is important with respect to how quickly the manual test execution effort decreases for a set of test cases that are being automated.
Daniel Flemström, Pasqualina Potena, Daniel Sundmark, Wasif Afzal, Markus Bohlin
Softw. Qual. J.4
2017 The MegaM@Rt2 ECSEL Project: MegaModelling at Runtime - Scalable Model-Based Framework for Continuous Development and Runtime Validation of Complex Systems
abstract
A major challenge for the European electronic industry is to enhance productivity while reducing costs and ensuring quality in development, integration and maintenance. Model-Driven Engineering (MDE) principles and techniques have already shown promising capabilities but still need to scale to support real-world scenarios implied by the full deployment and use of complex electronic components and systems. Moreover, maintaining efficient traceability, integration and communication between two fundamental system life-time phases (design time and runtime) is another challenge facing scalability of MDE. This paper presents an overview of the ECSEL project entitled "MegaModelling at runtime -- Scalable model-based framework for continuous development and runtime validation of complex systems" (MegaM@Rt2), whose aim is to address the above mentioned challenges facing MDE. Driven by both large and small industrial enterprises, with the support of research partners and technology providers, MegaM@Rt2 aims to deliver a framework of tools and methods for: 1) system engineering/design & continuous development, 2) related runtime analysis and 3) global model & traceability management, respectively. The diverse industrial use cases (covering domains such as aeronautics, railway, construction and telecommunications) will integrate and apply such a framework that shall demonstrate the validation of the MegaM@Rt2 solution.
Wasif Afzal, Hugo Bruneliere, Davide Di Ruscio, Andrey Sadovykh, Silvia Mazzini, Eric Cariou, Dragos Truscan, Jordi Cabot, Daniel Field, Luigi Pomante, Pavel Smrz
DSD1
2017 Towards Execution Time Prediction for Manual Test Cases from Test Specification
abstract
Knowing the execution time of test cases is important to perform test scheduling, prioritization and progress monitoring. This work in progress paper presents a novel approach for predicting the execution time of test cases based on test specifications and available historical data on previously executed test cases. Our approach works by extracting timing information (measured and maximum execution time)for various steps in manual test cases. This information is then used to estimate the maximum time for test steps that have not previously been executed, but for which textual specifications exist. As part of our approach, natural language parsing of the specifications is performed to identify word combinations to check whether existing timing information on various test activities is already available or not. Finally, linear regression is used to predict the actual execution time for test cases. A proof-of-concept use case at Bombardier Transportation serves to evaluate the proposed approach.
Sahar Tahvili, Mehrdad Saadatmand, Markus Bohlin, Wasif Afzal, Sharvathul Hasan Ameerjan
SEAA4
2017 Handling constraints in combinatorial interaction testing in the presence of multi objective particle swarm and multithreading
Bestoun S. Ahmed, Luca Maria Gambardella, Wasif Afzal, Kamal Zuhairi Zamli
Inf. Softw. Technol.3
2017 10 Years of research on debugging concurrent and multicore software: a systematic mapping study
Sara Abbaspour Asadollah, Daniel Sundmark, Sigrid Eldh, Hans A. Hansson, Wasif Afzal
Softw. Qual. J.5
2016 Capturing cost avoidance through reuse: systematic literature review and industrial evaluation
abstract
Background: Cost avoidance through reuse shows the benefits gained by the software organisations when reusing an artefact. Cost avoidance captures benefits that are not captured by cost savings e.g. spending that would have increased in the absence of the cost avoidance activity. This type of benefit can be combined with quality aspects of the product e.g. costs avoided because of defect prevention. Cost avoidance is a key driver for software reuse. Objectives: The main objectives of this study are: (1) To assess the status of capturing cost avoidance through reuse in the academia; (2) Based on the first objective, propose improvements in capturing of reuse cost avoidance, integrate these into an instrument, and evaluate the instrument in the software industry. Method: The study starts with a systematic literature review (SLR) on capturing of cost avoidance through reuse. Later, a solution is proposed and evaluated in the industry to address the shortcomings identified during the systematic literature review. Results: The results of a systematic literature review describe three previous studies on reuse cost avoidance and show that no solution, to capture reuse cost avoidance, was validated in industry. Afterwards, an instrument and a data collection form are proposed that can be used to capture the cost avoided by reusing any type of reuse artefact. The instrument and data collection form (describing guidelines) were demonstrated to a focus group, as part of static evaluation. Based on the feedback, the instrument was updated and evaluated in industry at 6 development sites, in 3 different countries, covering 24 projects in total. Conclusion: The proposed solution performed well in industrial evaluation. With this solution, practitioners were able to do calculations for reuse costs avoidance and use the results as decision support for identifying potential artefacts to reuse.
Mohsin Irshad, Richard Torkar, Kai Petersen, Wasif Afzal
EASE4
2016 Exploring Test Overlap in System Integration: An Industrial Case Study
abstract
Tougher safety regulations, global competition and ever increasing complexity of embedded software puts extensive pressure on the effectiveness of the software testing procedures. Previous studies have found that there exist overlaps (i.e., multiple instances of highly similar test cases) and even redundancies in the software testing process. Such overlap has been found between versions, variants and integration levels, but primarily at unit test level. Given large embedded systems involving many subsystems, does overlap exist within the system integration testing as well? In this paper, we present an industrial case study, aiming to a) evaluate if there exist test overlaps within the given context, b) if so, investigate how these overlaps are distributed, and c) find ways of reducing test effort by investigating how the knowledge of overlaps and their distribution may be used for finding candidate test cases for automation, maintenance or even removal. We have studied manual test cases, written in natural language, at a large vehicular manufacturer in Sweden. In particular, we have collected and analyzed test cases from the system integration testing levels of four different projects of a vehicle control management system. Using a similarity function, we evaluate if any overlaps between test cases exist, and where. We found that overlaps do exist within the system integration level, particularly in the form of partial test step sequences. However, very few test cases overlapped in their entirety. Some candidates for test step automation and update propagation were identified, but none for easy removal.
Daniel Flemström, Wasif Afzal, Daniel Sundmark
SEAA2
2016 Experience Report: Automated System Level Regression Test Prioritization Using Multiple Factors
abstract
We propose a new method of determining an effective ordering of regression test cases, and describe its implementation as an automated tool called SuiteBuilder developed by Westermo Research and Development AB. The tool generates an efficient order to run the cases in an existing test suite by using expected or observed test duration and combining priorities of multiple factors associated with test cases, including previous fault detection success, interval since last executed, and modifications to the code tested. The method and tool were developed to address problems in the traditional process of regression testing, such as lack of time to run a complete regression suite, failure to detect bugs in time, and tests that are repeatedly omitted. The tool has been integrated into the existing nightly test framework for Westermo software that runs on large-scale data communication systems. In experimental evaluation of the tool, we found significant improvement in regression testing results. The re-ordered test suites finish within the available time, the majority of fault-detecting test cases are located in the first third of the suite, no important test case is omitted, and the necessity for manual work on the suites is greatly reduced.
Per Erik Strandberg, Daniel Sundmark, Wasif Afzal, Thomas J. Ostrand, Elaine J. Weyuker
ISSRE3
2016 Cost-Benefit Analysis of Using Dependency Knowledge at Integration Testing
Sahar Tahvili, Markus Bohlin, Mehrdad Saadatmand, Stig Larsson 0002, Wasif Afzal, Daniel Sundmark
PROFES5
2016 Software test process improvement approaches: A systematic literature review and an industrial case study
Wasif Afzal, Snehal Alone, Kerstin Glocksien, Richard Torkar
J. Syst. Softw.1
2015 An experiment on the effectiveness and efficiency of exploratory testing
Wasif Afzal, Ahmad Nauman Ghazi, Juha Itkonen, Richard Torkar, Anneliese Amschler Andrews, Khurram Bhatti
Empir. Softw. Eng.1
2014 Search-Based Testing for Embedded Telecom Software with Complex Input Structures
Kivanc Doganay, Sigrid Eldh, Wasif Afzal, Markus Bohlin
ICTSS3
2014 Prediction of faults-slip-through in large software projects: an empirical evaluation
Wasif Afzal, Richard Torkar, Robert Feldt, Tony Gorschek
Softw. Qual. J.1
2012 Resampling Methods in Software Quality Classification
abstract
In the presence of a number of algorithms for classification and prediction in software engineering, there is a need to have a systematic way of assessing their performances. The performance assessment is typically done by some form of partitioning or resampling of the original data to alleviate biased estimation. For predictive and classification studies in software engineering, there is a lack of a definitive advice on the most appropriate resampling method to use. This is seen as one of the contributing factors for not being able to draw general conclusions on what modeling technique or set of predictor variables are the most appropriate. Furthermore, the use of a variety of resampling methods make it impossible to perform any formal meta-analysis of the primary study results. Therefore, it is desirable to examine the influence of various resampling methods and to quantify possible differences. Objective and method: This study empirically compares five common resampling methods (hold-out validation, repeated random sub-sampling, 10-fold cross-validation, leave-one-out cross-validation and non-parametric bootstrapping) using 8 publicly available data sets with genetic programming (GP) and multiple linear regression (MLR) as software quality classification approaches. Location of (PF, PD) pairs in the ROC (receiver operating characteristics) space and area under an ROC curve (AUC) are used as accuracy indicators. Results: The results show that in terms of the location of (PF, PD) pairs in the ROC space, bootstrapping results are in the preferred region for 3 of the 8 data sets for GP and for 4 of the 8 data sets for MLR. Based on the AUC measure, there are no significant differences between the different resampling methods using GP and MLR. Conclusion: There can be certain data set properties responsible for insignificant differences between the resampling methods based on AUC. These include imbalanced data sets, insignificant predictor variables and high-dimensional data sets. With the current selection of data sets and classification techniques, bootstrapping is a preferred method based on the location of (PF, PD) pair data in the ROC space. Hold-out validation is not a good choice for comparatively smaller data sets, where leave-one-out cross-validation (LOOCV) performs better. For comparatively larger data sets, 10-fold cross-validation performs better than LOOCV.
Wasif Afzal, Richard Torkar, Robert Feldt
Int. J. Softw. Eng. Knowl. Eng.1
2011 On the application of genetic programming for software engineering predictive modeling: A systematic review
Wasif Afzal, Richard Torkar
Expert Syst. Appl.1
2010 Using Faults-Slip-Through Metric as a Predictor of Fault-Proneness
abstract
Background: The majority of software faults are present in small number of modules, therefore accurate prediction of fault-prone modules helps improve software quality by focusing testing efforts on a subset of modules. Aims: This paper evaluates the use of the faults-slip-through (FST) metric as a potential predictor of fault-prone modules. Rather than predicting the fault-prone modules for the complete test phase, the prediction is done at the specific test levels of integration and system test. Method: We applied eight classification techniques, to the task of identifying fault prone modules, representing a variety of approaches, including a standard statistical technique for classification (logistic regression), tree-structured classifiers (C4.5 and random forests), a Bayesian technique (Naïve Bayes), machine-learning techniques (support vector machines and back-propagation artificial neural networks) and search-based techniques (genetic programming and artificial immune recognition systems) on FST data collected from two large industrial projects from the telecommunication domain. Results: Using area under the receiver operating characteristic (ROC) curve and the location of (PF, PD) pairs in the ROC space, the faults slip-through metric showed impressive results with the majority of the techniques for predicting fault-prone modules at both integration and system test levels. There were, however, no statistically significant differences between the performance of different techniques based on AUC, even though certain techniques were more consistent in the classification performance at the two test levels. Conclusions: We can conclude that the faults-slip-through metric is a potentially strong predictor of fault-proneness at integration and system test levels. The faults-slip-through measurements interact in ways that is conveniently accounted for by majority of the data mining techniques.
Wasif Afzal
APSEC1
2009 A systematic review of search-based testing for non-functional system properties
Wasif Afzal, Richard Torkar, Robert Feldt
Inf. Softw. Technol.1
2008 A Comparative Evaluation of Using Genetic Programming for Predicting Fault Count Data
abstract
There have been a number of software reliability growth models (SRGMs) proposed in literature. Due to several reasons, such as violation of models' assumptions and complexity of models, the practitioners face difficulties in knowing which models to apply in practice. This paper presents a comparative evaluation of traditional models and use of genetic programming (GP) for modeling software reliability growth based on weekly fault count data of three different industrial projects. The motivation of using a GP approach is its ability to evolve a model based entirely on prior data without the need of making underlying assumptions. The results show the strengths of using GP for predicting fault count data.
Wasif Afzal, Richard Torkar
ICSEA1
2008 A Systematic Mapping Study on Non-Functional Search-based Software Testing
Wasif Afzal, Richard Torkar, Robert Feldt
SEKE1