Antonia Bertolino

dblp:03/3919 · DBLP profile ↗
← Back
115ranked-venue papers
58as first author
26since 2021 · last 2026
0000-0001-8749-1356ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 97 · 51 first-author · 25 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 5 first-authorDatabases, data management, data science and information retrieval · 4 · 2 first-authorSecurity and privacy · 3 · 1 first-authorSystems, architecture and hardware · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2
YearPublicationVenuePosition
2026 A Framework for Similarity-based and Resource-aware Orchestration of End-to-End Test Cases
Cristian Augusto, Antonia Bertolino, Guglielmo De Angelis, Claudio de la Riva, Francesca Lonetti, Jesús Morán
AST2
2026 Recent developments in software engineering for systems-of-systems and software ecosystems
abstract
The increasing scale, distribution, and interconnection of software-intensive systems continue to change the landscape of software engineering research and practice, bringing the diffusion of two similar paradigms: Systems-of-Systems (SoS) and Software Ecosystems (SECO). SoS are characterized by the integration of operationally and managerially independent systems that join together to accomplish a common mission. SoS are central to domains such as transportation, healthcare, defense, smart cities, and industrial automation where heterogeneous systems must cooperate, exchange information, and adapt to evolving missions. On the other hand, SECO describe environments in which a platform and its surrounding network of developers, partners, and organizations co-create software offerings. In SECO, technical artifacts interact with economic and social processes. Modern digital platforms, mobile operating systems, cloud services, and enterprise platforms leverage the capabilities of their ecosystems. This editorial brings together perspectives that explore the shared challenges and complementary insights of SoS and SECO research, aiming to foster a richer understanding of complex software-intensive systems and highlight new opportunities for collaboration across communities. From the long-running, successful series of the International Workshop on Software Engineering for Systems-of-Systems and Software Ecosystems (SESoS), co-located with the IEEE/ACM International Conference on Software Engineering (ICSE), we present this special issue of the Journal of Systems and Software on the topics of SESoS 2024 in Lisbon, Portugal. From a total of 18 submissions, 7 articles were accepted in this special issue. The articles in this collection address fundamental questions of SoS and SECO, including the evolution of functional relations and experimentation practices, the key factors affecting developer experience, also related to women’s inclusion, as well as the adoption of GenAI-driven approaches for vulnerability fixing. These articles offer updates on current advances of SoS and SECO engineering to researchers and practitioners, highlighting opportunities for future research.
Francesca Lonetti, Antonia Bertolino, Pablo Oliveira Antonino, Doo-Hwan Bae
J. Syst. Softw.2
2026 Recent developments in software engineering for systems-of-systems and software ecosystems
Francesca Lonetti, Antonia Bertolino, Pablo Oliveira Antonino, Doo-Hwan Bae
J. Syst. Softw.2
2026 Using Metamorphic Relations in Redundancy-based Fault/Intrusion Tolerance
abstract
Redundancy is widely used as a method for fault and intrusion tolerance. However, if the redundant components lack sufficient diversity, potentially dangerous common mode failures may go undetected. To address this issue, the design diversity approach has been proposed in the literature for decades. In this article, we take an innovative approach to this problem by introducing a broader notion of diversity, which leverages Metamorphic Relations (MRs), i.e., necessary properties that must hold among diverse inputs and diverse outputs. We define two generic categories of MRs that establish data diversity and functional diversity. Furthermore, we elaborate on two corresponding logical architectures, paying particular attention to the necessary conditions for the adjudicator component. Finally, we present an initial evaluation of the proposed architectures, which points out the advantages with respect to their counterparts based on the traditional design diversity method, and discuss future research directions for this novel conceptual approach to redundancy-based fault/intrusion tolerance.
Felicita Di Giandomenico, Giulio Masetti, Francesca Lonetti, Antonia Bertolino
ACM Trans. Softw. Eng. Methodol.4
2025 An Adaptive Testing Approach Based on Field Data
abstract
The growing need to test systems post-release has led to extending testing activities into production environments, where uncertainty and dynamic conditions pose significant challenges. Field testing approaches, especially Self-Adaptive Testing in the Field (SATF), face hurdles like managing unpredictability, minimizing system overhead, and reducing human intervention, among others. Despite its importance, SATF remains underexplored in the literature. This work introduces AdapTA (Adaptive Testing Approach), a novel SATF strategy tailored for testing Body Sensor Networks (BSNs). BSNs are networks of wearable or implantable sensors designed to monitor physiological and environmental data. AdapTA employs an ex-vivo approach, using real-world data collected from the field to simulate patient behavior in in-house experiments. Field data are used to derive Discrete-Time Markov Chain (DTMC) models, which simulate patient profiles and generate test input data for the BSN. The BSN’s outputs are compared against a proposed oracle to evaluate test outcomes. AdapTA’s adaptive logic continuously monitors the system under test and the simulated patient, triggering adaptations as needed. Results demonstrate that AdapTA achieves greater effectiveness compared to a non-adaptive version of the proposed approach across three adaptation scenarios, emphasizing the value of its adaptive logic.
Samira Silva, Ricardo Caldas, Patrizio Pelliccione, Antonia Bertolino
AST4
2025 RETORCH*: A Cost and Resource aware Model for E2E Testing in the Cloud
abstract
Moving testing to the Cloud overcomes time/resource constraints by leveraging an unlimited and elastic infrastructure, especially for testing levels like End-to-End (E2E) that require a high number of resources and/or execution time. However, it introduces new challenges to those already faced on-premises, like selecting the most suitable Cloud infrastructure and billing scheme. We propose the RETORCH* test execution model that estimates and compares the monetary cost of executing an E2E test suite with different Cloud alternatives, billing schemes, and test configurations. RETORCH* goes beyond the mere cost billed, and selects the solution that best aligns with the test team strategy using the data of on-premises prior executions and the tester's experience. This cost is broken down into the cost incurred to execute the test suite (testing cost) and possible unused infrastructure (overprovisioning cost). Based on these distinct costs, the test team can compare different Cloud and test configurations. RETORCH* has been evaluated using a real-world application's E2E test suite. We analyze how the different decisions taken when the suite is migrated to the Cloud impact the cost, highlighting how RETORCH* can help the tester during Cloud and test configuration to make a more informed decision.
Cristian Augusto, Jesús Morán, Antonia Bertolino, Claudio de la Riva, Javier Tuya
J. Syst. Softw.3
2025 Different approaches for testing body sensor network applications
abstract
Body Sensor Networks (BSNs) offer a cost-effective way to monitor patients’ health and detect potential risks. Despite the growing interest attracted by BSNs, there is a lack of testing approaches for them. Testing a Body Sensor Network (BSN) is challenging due to its evolving nature, the complexity of sensor scenarios and their fusion, the potential necessity of third-party testing for certification, and the need to prioritize critical failures given limited resources. This paper addresses these challenges by proposing three BSN testing approaches: PASTA, ValComb, and TransCov. These approaches share common characteristics, which are described through a general framework called GATE4BSN. PASTA simulates patients with sensors and models sensor trends using a Discrete Time Markov Chain (DTMC). ValComb explores various health conditions by considering all sensor risk level combinations, while TransCov ensures full coverage of DTMC transitions. We empirically evaluate these approaches, comparing them with a baseline approach in terms of failure detection. The results demonstrate that PASTA, ValComb, and TransCov uncover previously undetected failures in an open-source BSN and outperform the baseline approach. Statistical analysis reveals that PASTA is the most effective, while ValComb is 76 times faster than PASTA and nearly as effective.
Samira Silva, Ricardo Caldas, Patrizio Pelliccione, Antonia Bertolino
J. Syst. Softw.4
2025 Advances in Software Engineering Research for Systems-of-Systems and Software Ecosystems
abstract
ABSTRACT For more than a decade, software engineering for systems‐of‐systems (SoS) and software ecosystems (SECO) has been largely investigated in order to cope with complexity in software‐intensive systems. SoS research addresses several aspects related to software system architecture comprising a set of constituent systems that relate to each other to perform missions. As such, SoS have key characteristics such as operational and managerial independence, distribution, emergent behavior, and evolutionary development. Full interoperability and dynamic architecture become critical challenges in this context. On the hand, SECO research refers to modeling and analysis of a socio‐technical network of actors and artifacts formed on top of common technological platforms, in which business factors directly influence software maintenance and evolution. Software sustainability and diversity as well as quality attributes that affect the SECO platform health represent challenges in the field. From the long‐running, successful series of the International Workshop on Software Engineering for systems‐of‐systems and Software Ecosystems (SESoS), co‐located with the IEEE/ACM International Conference on Software Engineering (ICSE), we present this special issue on the topics in the Journal of Software: Evolution and Process from SESoS 2023 in Melbourne, Australia. Four articles were accepted and published in this special issue, covering a longitudinal analysis of SoS research, as well as strategic patterns, services, and trust in SECO. These articles provide researchers and practitioners with advances in the state of the art and point out opportunities for further research.
Rodrigo Pereira dos Santos, Antonia Bertolino, Pablo Oliveira Antonino, Doo-Hwan Bae
J. Softw. Evol. Process.2
2024 Software System Testing Assisted by Large Language Models: An Exploratory Study
Cristian Augusto, Jesús Morán, Antonia Bertolino, Claudio de la Riva, Javier Tuya
ICTSS3
2024 Flakiness goes live: Insights from an In Vivo testing simulation study
abstract
Test flakiness is a topmost concern in software test automation. While conducting pre-deployment testing, those tests that are flagged as flaky are put aside for being either repaired or discarded. We hypothesise that some flaky tests could provide useful insights if run in the field, i.e., they could help identify failures that manifest themselves sporadically during In House testing, but are later experienced in operation. We present the first simulation study to investigate the behaviour of flaky tests when moved to the field. The work compares the behaviour of known flaky tests from an open-source library when executed in the development environment vs. when executed in a simulation of the field. Our experimentation over 52 test methods labelled as flaky provides a first confirmation that moving from the development environment to the field, the behaviour of tests changes. In particular, the failure frequency of intermittently failing tests can increase, and we could also identify few cases of field failures that would have been hardly detected during In House testing due to the numerous combinations of inputs and states. In most cases, such flakiness was rooted in the design of the test method itself, however we could also identify an actual bug. The results of our study suggest that the identification of an intermittently failing behaviour could be a valuable hint for a test engineer, and hence flaky tests should not be dismissed right away.
Morena Barboni, Antonia Bertolino, Guglielmo De Angelis
Inf. Softw. Technol.2
2024 A framework for the design of fault-tolerant systems-of-systems
Francisco Henrique Ferreira, Elisa Yumi Nakagawa, Antonia Bertolino, Francesca Lonetti, Vânia de Oliveira Neves, Rodrigo Pereira dos Santos
J. Syst. Softw.3
2024 Self-Adaptive Testing in the Field
abstract
We are increasingly surrounded by systems connecting us with the digital world and facilitating our life by supporting our work, leisure, activities at home, health, and so on. These systems are pressed by two forces. On the one side, they operate in environments that are increasingly challenging due to uncertainty and uncontrollability. On the other side, they need to evolve, often in a continuous fashion, to meet changing needs, to offer new functionalities, or also to fix emerging failures. To make the picture even more complex, these systems rarely work in isolation and often need to collaborate with other systems, as well as humans. All such facets call for moving their validation during operation, as offered by approaches called testing in the field. In this article, we observe that even the field-based testing approaches should change over time to follow and adapt to the changes and evolution of collaborating systems or environments or users’ behaviors. We provide a taxonomy of this new category of testing that we call self-adaptive testing in the field (SATF) , together with a reference architecture for SATF approaches. To achieve this objective, we surveyed the literature and collected feedback and contributions from experts in the domain via a questionnaire and interviews.
Samira Silva, Patrizio Pelliccione, Antonia Bertolino
ACM Trans. Auton. Adapt. Syst.3
2024 Automatic Debugging of Design Faults in MapReduce Applications
abstract
Among the current technologies to analyse large data, the MapReduce processing model stands out in Big Data. MapReduce is implemented in frameworks such as Hadoop, Spark or Flink that are able to manage the program executions according to the resources available at runtime. The developer should design the program in order to support all possible non-deterministic executions. However, the program may fail due to a design fault. Debugging these kinds of faults is difficult because the data are executed non-deterministically in parallel and the fault is not caused directly by the code, but by its design. This paper presents a framework called MRDebug which includes two debugging techniques focused on the MapReduce design faults. A spectrum-based fault localization technique locates the root cause of these faults analysing several executions of the test case, and a Delta Debugging technique isolates the data relevant to trigger the failure. An empirical evaluation with 13 programs shows that MRDebug is effective in debugging the faults, especially when the localization is done with the reduced data. In summary, MRDebug automatically provides valuable information to understandMapReducedesign faults as it helps locate their root cause and obtains a minimal data that triggers the failure.
Jesús Morán, Antonia Bertolino, Claudio de la Riva, Javier Tuya
IEEE Trans. Software Eng.2
2023 Cross-coverage testing of functionally equivalent programs
abstract
Cross-coverage of a program P refers to the test coverage measured over a different program Q that is functionally equivalent to P. The novel concept of cross-coverage can find useful applications in the test of redundant software. We apply here cross-coverage for test suite augmentation and show that additional test cases generated from the coverage of an equivalent program, referred to as cross tests, can increase the coverage of a program in more effective way than a random baseline. We also observe that -contrary to traditional coverage testing-cross coverage could help finding (artificially created) missing functionality faults.
Antonia Bertolino, Guglielmo De Angelis, Felicita Di Giandomenico, Francesca Lonetti
AST1
2023 Orchestration Strategies for Regression Test Suites
abstract
Regression testing is widely studied in the literature, although most research on the topic is concerned with improving specific sub-challenges of a wider goal. Test suite orchestration proposes a more comprehensive view of the challenge of regression testing, by merging and combining different techniques with a variety of objectives, including prioritizing, selecting, reducing and amplifying tests, detecting flaky tests and potentially more. This paper presents the key approaches and techniques that form test suite orchestration, along with common evaluation metrics, and discusses how they can be used together to ultimately provide an efficient and effective regression testing strategy. To illustrate the benefits of orchestration, we provide some examples of existing papers that take steps towards this goal, even if the specific terminology is not yet used. Orchestrated strategies utilizing existing regression testing techniques provide a pathway to practicality and real-world usage of the academic literature.
Renan Greca, Breno Miranda, Antonia Bertolino
AST3
2023 Model-based security testing in IoT systems: A Rapid Review
abstract
Security testing is a challenging and effort-demanding task in IoT scenarios. The heterogeneous devices expose different vulnerabilities that can influence the methods and cost of security testing. Model-based security testing techniques support the systematic generation of test cases for the assessment of security requirements by leveraging the specifications of the IoT system model and of the attack templates. This paper aims to review the adoption of model-based security testing in the context of IoT, and then provides the first systematic and up-to-date comprehensive classification and analysis of research studies in this topic. We conducted a systematic literature review analysing 803 publications and finally selecting 17 primary studies, which satisfied our inclusion criteria and were classified according to a set of relevant analysis dimensions. We report the state-of-the-art about the used formalisms, the test techniques, the objectives, the target applications and domains; we also identify the targeted security attacks, and discuss the challenges, gaps and future research directions. Our review represents the first attempt to systematically analyze and classify existing studies on model-based security testing for IoT. According to the results, model-based security testing has been applied in core IoT domains. Models complexity and the need of modeling evolving scenarios that include heterogeneous open software and hardware components remain the most important shortcomings. Our study shows that model-based security testing of IoT applications is a promising research direction. The principal future research directions deal with: extending the existing modeling formalisms in order to capture all peculiarities and constraints of complex and large scale IoT networks; the definition of context-aware and dynamic evolution modelling approaches of IoT entities; and the combination of model-based testing techniques with other security test strategies such as penetration testing or learning techniques for model inference.
Francesca Lonetti, Antonia Bertolino, Felicita Di Giandomenico
Inf. Softw. Technol.2
2023 Introduction to the special issue on test automation: Trends, benefits, and costs
Antonia Bertolino, Guglielmo De Angelis, Maurizio Leotta, Filippo Ricca
J. Syst. Softw.1
2023 DevOpRET: Continuous reliability testing in DevOps
abstract
Abstract To enter the production stage, in DevOps practices candidate software releases have to pass quality gates, where they are assessed to meet established target values for key indicators of interest. We believe software reliability should be an important such indicator, as it greatly contributes to the end‐user satisfaction. We propose DevOpRET , an approach for reliability testing as part of the acceptance testing stage in DevOps. DevOpRET relies on operational‐profile–based testing, a common reliability assessment technique. DevOpRET leverages usage and failure data monitored in operations to continuously refine its estimate. We evaluate accuracy and efficiency of DevOpRET through controlled experiments with a real‐world open source platform and with a microservice architectures benchmark. The results show that DevOpRET provides accurate and efficient estimates of the true reliability over subsequent DevOps cycles.
Antonia Bertolino, Guglielmo De Angelis, Antonio Guerriero, Breno Miranda, Roberto Pietrantuono, Stefano Russo 0001
J. Softw. Evol. Process.1
2023 Introduction to the special issue on automation of software test and test code quality
abstract
We are pleased to present the papers selected for inclusion in the special issue devoted to the 1st International Conference on Automation of Software Test, which was held virtually in colocation with the 42nd IEEE/ACM International Conference on Software Engineering (ICSE 2020). Software testing is an integral and important part of the software engineering (SE) discipline. Over the past decades, a significant amount of SE research has focused on automation of software test (AST), including the automation of test case generation, test case selection and prioritization, test execution, test verdict analysis, and debugging. AST practice has also moved forward significantly, and in recent years, many test tools, frameworks, and methodologies have been developed and have significantly enhanced the quality assurance and the productivity in SE practices. Despite the significant achievements, AST remains challenging. To achieve total automation of software testing, a huge amount of code for the test cases and the supporting test infrastructure is needed, which yields itself in turn to large maintenance costs. However, if on the one side it is now generally accepted that disciplined procedures and quality standards should be applied in production code development, on the other side, comparable levels of rigor and quality are not demanded for the code written for testing that production code. Indeed, several recent empirical studies point to how test code is affected by many problems, bugs, unjustified assumptions, hidden dependencies from other tests or environment, flakiness, and performance issues. Because of such problems, test effectiveness is impacted and several false alarms are raised in regression testing that increase the test costs. Recently, both researchers and practitioners have proposed solutions toward this problem by identifying test code smells or test code quality issues and providing techniques to automatically detect and repair test code bugs and flakiness. In consideration of this active research thread, the 1st ACM/IEEE International Conference on Automation of Software Test (AST 2020) featured “Who Tests The Tests?” as the conference theme. AST 2020 was run for the first time in the format of a colocated conference to ICSE, after a successful series of 14 workshops under the same name. This special issue offers a venue for researchers and practitioners to share the advances in software test automation, especially on test code quality. In particular, we invited the authors of the best papers of AST 2020 to submit extended versions of their conference publications. Moreover, we also encouraged the authors and participants of the AST series and the broader SE community to submit novel original papers on the themes of software test automation and test code quality, including approaches for the following: understanding the dimensions and characteristics of problems with unreliable, low-quality test code; identifying and preventing test code smells, test code bugs, and flaky tests; impact of test code quality in scaling up test automation of very large, complex system; metrics for test code quality and robustness; automated repair of test code bugs and flakiness; and employing Artificial Intelligence and Machine Learning methods to help test automation. Each of the 10 submitted papers underwent a rigorous review process by independent referees to ensure that the paper had a sound, novel, and original contribution to the field of software test automation. Finally, five papers were accepted for inclusion in the special issue; three of these are extended versions of the best papers of AST 2020, and two are novel external submissions. In particular, the paper titled “Quantum Software Testing—State of the Art,” by Antonio García de la Barrera, Ignacio García-Rodríguez de Guzmán, Macario Polo, and Mario Piattini, provides a systematic mapping study assessing the state of the art in testing of quantum computing applications, indeed an emerging new paradigm that promises exponential speed up in solving highly demanding computational problems. In “An Empirical Study on How Sapienz Achieves Coverage and Crash Detection”, Iván Arcuschin, Juan Pablo Galeotti, and Diego Garbervetsky report the results from an empirical study aiming at better understanding how the main features of Sapienz, a powerful tool for automated testing of Android applications using evolutionary algorithms, impact its effectiveness. Many program analysis and testing tools need to incorporate string solver mechanisms. After observing that adequate tools and benchmarks for comparing existing solvers were lacking, in “ZaligVinder: A Generic Test Framework for String Solvers”, Mitja Kulczynski, Florin Manea, Dirk Nowotka, and Danny Bøgsted Poulsen propose an extensible framework gathering several string solver benchmarks that can be used for analysis and debugging purposes. “ExVivoMicroTest: Ex-Vivo Testing of Microservices” by Luca Gazzola, Maayan Goldstein, Leonardo Mariani, Marco Mobilio, Itai Segall, Alessandro Tundo, and Luca Ussi focuses on ex vivo testing of microservices during deployment. The authors claim that regression testing of such services may not be adequate prior to deployment due to a lack of knowledge regarding new use scenarios. The authors then propose a technique, named ExVivoMicroTest, that analyzes the behavior of deployed microservices and generates test cases for testing future updates. In “Fight Silent Horror Unit Test Methods by Consulting a TestWizard”, Maura Cerioli, Giovanni Lagorio, Maurizio Leotta, and Filippo Ricca focus on a practical problem, that is, the existence of incorrect tests. While a totally automated solution to identify invalid tests seems impractical, the idea of TestWizard is to assess individual tests' quality from the point of view of their coherence to specifications. We would like to thank the Editors-in-Chief of the Journal of Software: Evolution and Process for giving us the opportunity to publish this special issue. We appreciate all reviewers for their efforts in providing thoughtful and constructive comments to improve the quality of the publications. We also thank the authors of all submissions and publications of this special issue.
Antonia Bertolino, Shin Hong, Aditya P. Mathur
J. Softw. Evol. Process.1
2023 In vivo test and rollback of Java applications as they are
abstract
Summary Modern software systems accommodate complex configurations and execution conditions that depend on the environment where the software is run. While in house testing can exercise only a fraction of such execution contexts, in vivo testing can take advantage of the execution state observed in the field to conduct further testing activities. In this paper, we present the Groucho approach to in vivo testing. Groucho can suspend the execution, run some in vivo tests, rollback the side effects introduced by such tests, and eventually resume normal execution. The approach can be transparently applied to the original application, even if only available as compiled code, and it is fully automated. Our empirical studies of the performance overhead introduced by Groucho under various configurations showed that this may be kept to a negligible level by activating in vivo testing with low probability. Our empirical studies about the effectiveness of the approach confirm previous findings on the existence of faults that are unlikely exposed in house and become easy to expose in the field. Moreover, we include the first study to quantify the coverage increase gained when in vivo testing is added to complement in house testing.
Antonia Bertolino, Guglielmo De Angelis, Breno Miranda, Paolo Tonella
Softw. Test. Verification Reliab.1
2022 Testing non-testable programs using association rules
abstract
We propose a novel scalable approach for testing non-testable programs denoted as ARMED testing. The approach leverages efficient Association Rules Mining algorithms to determine relevant implication relations among features and actions observed while the system is in operation. These relations are used as the specification of positive and negative tests, allowing for identifying plausible or suspicious behaviors: for those cases when oracles are inherently unknownable, such as in social testing, ARMED testing introduces the novel concept of testing for plausibility. To illustrate the approach we walk-through an application example.
Antonia Bertolino, Emilio Cruciani, Breno Miranda, Roberto Verdecchia
AST1
2022 Comparing and Combining File-based Selection and Similarity-based Prioritization towards Regression Test Orchestration
abstract
Test case selection (TCS) and test case prioritization (TCP) techniques can reduce time to detect the first test failure. Although these techniques have been extensively studied in combination and isolation, they have not been compared one against the other. In this paper, we perform an empirical study directly comparing TCS and TCP approaches, represented by the tools Ekstazi and FAST, respectively. Furthermore, we develop the first combination, named Fastazi, of file-based TCS and similarity-based TCP and evaluate its benefit and cost against each individual technique. We performed our experiments using 12 Java-based open-source projects. Our results show that, in the median case, the combined approach detects the first failure nearly two times faster than either Ekstazi alone (with random test ordering) or FAST alone (without TCS). Statistical analysis shows that the effectiveness of Fastazi is higher than that of Ekstazi, which in turn is higher than that of FAST. On the other hand, FAST adds the least overhead to testing time, while the difference between the additional time needed by Ekstazi and Fastazi is negligible. Fastazi can also improve failure detection in scenarios where the time available for testing is restricted.
Renan Greca, Breno Miranda, Milos Gligoric 0001, Antonia Bertolino
AST4
2022 Self-adaptive Testing in the Field: Are We There Yet?
abstract
Testing in the field is gaining momentum, as a means to detect those failures that escape in-house testing by continuing the testing even while a system is operating in production. Among several approaches that are proposed, this paper focuses on the important notion of self-adaptivity of testing in the field, as such techniques need to adapt in many ways their strategy to the context and the emerging behaviors of the system under test. In this work, we investigate the topic by conducting a scoping review of the literature on self-adaptive testing in the field. We rely on a taxonomy organized in some categories that include the object to adapt, the adaptation trigger, the temporal characteristics, the realization issues, the interaction concerns, the type of field-based approach, and the impact/cost. Our study sheds light on self-adaptive testing in the field by identifying related key concepts and key characteristics and extracting some knowledge gaps to better guide future research.
Samira Silva, Antonia Bertolino, Patrizio Pelliccione
SEAMS2
2022 A Delphi study to recognize and assess systems of systems vulnerabilities
abstract
System of Systems (SoS) is an emerging paradigm by which independent systems collaborate by sharing resources and processes to achieve objectives that they could not achieve on their own. In this context, a number of emergent behaviors may arise that can undermine the security of the constituent systems. We apply the Delphi method with the aims to improve our understanding of SoS security and related problems, and to investigate their possible causes and remedies. Experts on SoS expressed their opinions and reached consensus in a series of rounds by following a structured questionnaire. The results show that the experts found more consensus in disagreement than in agreement about some SoS characteristics, and on how SoS vulnerabilities could be identified and prevented. From this study we learn that more work is needed to reach a shared understanding of SoS vulnerabilities, and we leverage expert feedback to outline some future research directions.
Miguel Angel Olivero, Antonia Bertolino, Francisco José Domínguez Mayo, Ilaria Matteucci, María José Escalona Cuaresma
Inf. Softw. Technol.2
2022 Designing and testing systems of systems: From variability models to test cases passing through desirability assessment
abstract
Abstract In the early stages of a system of systems (SoS) conception, several constituent systems could be available that provide similar functionalities. An SoS design methodology should provide adequate means to model variability in order to support the opportunistic selection of the most desirable SoS configuration. We propose the VANTESS approach that (i) supports SoS modeling taking into account the variation points implied by the considered constituent systems; (ii) includes a heuristics to weight benefits and costs of potential architectural choices (called as SoS variants) for the selection of the constituent systems; and finally (iii) also helps test planning for the selected SoS variant by deriving a simulation model on which test objectives and scenarios can be devised. We illustrate an application example of VANTESS to the “educational” SoS and discuss its pros and cons within a focus group.
Francesca Lonetti, Vânia de Oliveira Neves, Antonia Bertolino
J. Softw. Evol. Process.3
2021 Adaptive Test Case Allocation, Selection and Generation Using Coverage Spectrum and Operational Profile
abstract
We present an adaptive software testing strategy for test case allocation, selection and generation, based on the combined use of operational profile and coverage spectrum, aimed at achieving high delivered reliability of the program under test. Operational profile-based testing is a black-box technique considered well suited when reliability is a major concern, as it selects the test cases having the largest impact on failure probability in operation. Coverage spectrum is a characterization of a program's behavior in terms of the code entities (e.g., branches, statements, functions) that are covered as the program executes. The proposed strategy - named covrel+ - complements operational profile information with white-box coverage measures, so as to adaptively select/generate the most effective test cases for improving reliability as testing proceeds. We assess covrel+ through experiments with subjects commonly used in software testing research, comparing results with traditional operational testing. The results show that exploiting operational and coverage data in an integrated adaptive way allows generally to outperform operational testing at achieving a given reliability target, or at detecting faults under the same testing budget, and that covrel+ has greater ability than operational testing in detecting hard-to-detect faults.
Antonia Bertolino, Breno Miranda, Roberto Pietrantuono, Stefano Russo 0001
IEEE Trans. Software Eng.1
2020 Learning-to-rank vs ranking-to-learn: strategies for regression testing in continuous integration
abstract
In Continuous Integration (CI), regression testing is constrained by the time between commits. This demands for careful selection and/or prioritization of test cases within test suites too large to be run entirely. To this aim, some Machine Learning (ML) techniques have been proposed, as an alternative to deterministic approaches. Two broad strategies for ML-based prioritization are learning-to-rank and what we call ranking-to-learn (i.e., reinforcement learning). Various ML algorithms can be applied in each strategy. In this paper we introduce ten of such algorithms for adoption in CI practices, and perform a comprehensive study comparing them against each other using subjects from the Apache Commons project. We analyze the influence of several features of the code under test and of the test process. The results allow to draw criteria to support testers in selecting and tuning the technique that best fits their context.
Antonia Bertolino, Antonio Guerriero, Breno Miranda, Roberto Pietrantuono, Stefano Russo 0001
ICSE1
2020 Run Java Applications and Test Them In-Vivo Meantime
abstract
The outcome of test case execution depends on the state of the object under test. While testers can carefully choose meaningful and representative object states for test execution, it is unaffordable to cover the combinatorial space of possible object states exhaustively. An appealing option is to delegate part of the testing activities to the runtime and to execute test cases in the field whenever a new or uncommon state is observed. We have designed and developed Groucho, a framework for in-vivo testing of Java applications. Among the challenges that we faced, the most important ones are isolation of the test session from the user session and minimal performance overhead. Experimental results show that if the activation probability is kept reasonably small (e.g., $10 ^{- {4}}$), the impact of the framework is imperceptible(i.e., either statistically insignificant or with a negligible effect size).
Antonia Bertolino, Guglielmo De Angelis, Breno Miranda, Paolo Tonella
ICST1
2020 What is the Vocabulary of Flaky Tests?
abstract
Flaky tests are tests whose outcomes are non-deterministic. Despite the recent research activity on this topic, no effort has been made on understanding the vocabulary of flaky tests. This work proposes to automatically classify tests as flaky or not based on their vocabulary. Static classification of flaky tests is important, for example, to detect the introduction of flaky tests and to search for flaky tests after they are introduced in regression test suites.
Gustavo Pinto 0001, Breno Miranda, Supun Dissanayake, Marcelo d'Amorim, Christoph Treude, Antonia Bertolino
MSR6
2020 JTeC: A Large Collection of Java Test Classes for Test Code Analysis and Processing
abstract
The recent push towards test automation and test-driven development continues to scale up the dimensions of test code that needs to be maintained, analysed, and processed side-by-side with production code. As a consequence, on the one side regression testing techniques, e.g., for test suite prioritization or test case selection, capable to handle such large-scale test suites become indispensable; on the other side, as test code exposes own characteristics, specific techniques for its analysis and refactoring are actively sought. We present JTeC, a large-scale dataset of test cases that researchers can use for benchmarking the above techniques or any other type of tool expressly targeting test code. JTeC collects more than 2.5M test classes belonging to 31K+ GitHub projects and summing up to more than 430 Million SLOCs of ready-to-use real-world test code.
Federico Coro, Roberto Verdecchia, Emilio Cruciani, Breno Miranda, Antonia Bertolino
MSR5
2020 Quality-of-Experience driven configuration of WebRTC services through automated testing
abstract
Quality of Experience (QoE) refers to the end users level of satisfaction with a real-time service, in particular in relation to its audio and video quality. Advances in WebRTC technology have favored the spread of multimedia services through use of any browser. Provision of adequate QoE in such services is of paramount importance. The assessment of QoE is costly and can be done only late in the service lifecycle. In this work we propose a simple approach for QoE-driven non-functional testing of WebRTC services that relies on the ElasTest open-source platform for end-to-end testing of large complex systems. We describe the ElasTest platform, the proposed approach and an experimental study. In this study, we compared qualitatively and quantitatively the effort required in the ElasTest supported scenario with respect to a "traditional" solution, showing great savings in terms of effort and time.
Antonia Bertolino, Antonello Calabrò, Guglielmo De Angelis, Francisco Gortázar, Francesca Lonetti, Michel Maes-Bermejo, Guiomar Tunon de Hita
QRS1
2020 Cloud testing automation: industrial needs and ElasTest response
abstract
While great emphasis is given in the current literature about the potential of leveraging the cloud for testing purposes, the authors have scarce factual evidence from real‐world industrial contexts about the motivations, drawbacks and benefits related to the adoption of automated cloud testing technology. In this study, the authors present an empirical study undertaken within the ongoing European Project ElasTest, which has developed an open source platform for end‐to‐end testing of large distributed systems. This study aims at validating the ElasTest solution, and consists of the assessment of four demonstrators belonging to different application domains, namely e‐commerce, 5G networking, WebRTC and Internet of Things. For each demonstrator, they collected differing requirements, and achieved varying results, both positive and negative, showing that cloud testing needs careful assessment before adoption.
Antonia Bertolino, Antonello Calabrò, Eda Marchetti, Anton Cervantes Sala, Guiomar Tunon de Hita, Ilie-Daniel Gheorghe-Pop, Varun Gowtham
IET Softw.1
2020 Digital persona portrayal: Identifying pluridentity vulnerabilities in digital life
Miguel Angel Olivero, Antonia Bertolino, Francisco José Domínguez Mayo, María José Escalona Cuaresma, Ilaria Matteucci
J. Inf. Secur. Appl.2
2020 FlakyLoc: Flakiness Localization for Reliable Test Suites in Web Applications
abstract
Web application testing is a great challenge due to the management of complex asynchronous communications, the concurrency between the clients-servers, and the heterogeneity of resources employed. It is difficult to ensure that a test case is re-running in the same conditions because it can be executed in undesirable ways according to several environmental factors that are not easy to fine-grain control such as network bottlenecks, memory issues or screen resolution. These environmental factors can cause flakiness, which occurs when the same test case sometimes obtains one test outcome and other times another outcome in the same application due to the execution of environmental factors. The tester usually stops relying on flaky test cases because their outcome varies during the re-executions. To fix and reduce the flakiness it is very important to locate and understand which environmental factors cause the flakiness. This paper is focused on the localization of the root cause of flakiness in web applications based on the characterization of the different environmental factors that are not controlled during testing. The root cause of flakiness is located by means of spectrum-based localization techniques that analyse the test execution under different combinations of the environmental factors that can trigger the flakiness. This technique is evaluated with an educational web platform called FullTeaching. As a result, our technique was able to locate automatically the root cause of flakiness and provide enough information to both understand it and fix it.
Jesús Morán, Cristian Augusto, Antonia Bertolino, Claudio de la Riva, Javier Tuya
J. Web Eng.3
2020 RETORCH: an approach for resource-aware orchestration of end-to-end test cases
Cristian Augusto, Jesús Morán, Antonia Bertolino, Claudio de la Riva, Javier Tuya
Softw. Qual. J.3
2020 Advances in test automation for software with special focus on artificial intelligence and machine learning
J. Jenny Li 0001, Andreas Ulrich, Xiaoying Bai, Antonia Bertolino
Softw. Qual. J.4
2020 Testing Relative to Usage Scope: Revisiting Software Coverage Criteria
abstract
Coverage criteria provide a useful and widely used means to guide software testing; however, indiscriminately pursuing full coverage may not always be convenient or meaningful, as not all entities are of interest in any usage context. We aim at introducing a more meaningful notion of coverage that takes into account how the software is going to be used. Entities that are not going to be exercised by the user should not contribute to the coverage ratio. We revisit the definition of coverage measures, introducing a notion of relative coverage. According to this notion, we provide a definition and a theoretical framework of relative coverage, within which we discuss implications on testing theory and practice. Through the evaluation of three different instances of relative coverage, we could observe that relative coverage measures provide a more effective strategy than traditional ones: we could reach higher coverage measures, and test cases selected by relative coverage could achieve higher reliability. We hint at several other useful implications of relative coverage notion on different aspects of software testing.
Breno Miranda, Antonia Bertolino
ACM Trans. Softw. Eng. Methodol.2
2019 Scalable approaches for test suite reduction
abstract
Test suite reduction approaches aim at decreasing software regression testing costs by selecting a representative subset from large-size test suites. Most existing techniques are too expensive for handling modern massive systems and moreover depend on artifacts, such as code coverage metrics or specification models, that are not commonly available at large scale. We present a family of novel very efficient approaches for similarity-based test suite reduction that apply algorithms borrowed from the big data domain together with smart heuristics for finding an evenly spread subset of test cases. The approaches are very general since they only use as input the test cases themselves (test source code or command line input). We evaluate four approaches in a version that selects a fixed budget B of test cases, and also in an adequate version that does the reduction guaranteeing some fixed coverage. The results show that the approaches yield a fault detection loss comparable to state-of-the-art techniques, while providing huge gains in terms of efficiency. When applied to a suite of more than 500K real world test cases, the most efficient of the four approaches could select B test cases (for varying B values) in less than 10 seconds.
Emilio Cruciani, Breno Miranda, Roberto Verdecchia, Antonia Bertolino
ICSE4
2019 Debugging Flaky Tests on Web Applications
abstract
International Conference on Web Information Systems and Technologies, WEBIST (15th. 2019. Vienna, Austria)
Jesús Morán, Cristian Augusto, Antonia Bertolino, Claudio de la Riva, Javier Tuya
WEBIST3
2018 FAST approaches to scalable similarity-based test case prioritization
abstract
Many test case prioritization criteria have been proposed for speeding up fault detection. Among them, similarity-based approaches give priority to the test cases that are the most dissimilar from those already selected. However, the proposed criteria do not scale up to handle the many thousands or even some millions test suite sizes of modern industrial systems and simple heuristics are used instead. We introduce the FAST family of test case prioritization techniques that radically changes this landscape by borrowing algorithms commonly exploited in the big data domain to find similar items. FAST techniques provide scalable similarity-based test case prioritization in both white-box and black-box fashion. The results from experimentation on real world C and Java subjects show that the fastest members of the family outperform other black-box approaches in efficiency with no significant impact on effectiveness, and also outperform white-box approaches, including greedy ones, if preparation time is not counted. A simulation study of scalability shows that one FAST technique can prioritize a million test cases in less than 20 minutes.
Breno Miranda, Emilio Cruciani, Roberto Verdecchia, Antonia Bertolino
ICSE4
2018 CARS: Context Aware Reputation Systems to Evaluate Vehicles' Behaviour
abstract
The introduction of new generation ICT systems into vehicles makes them highly connected with the external World. As drawback, vehicle becomes potentially vulnerable to security attacks. Here, we consider a scenario in which Vehicular Networks and a Urban Network work together to realize a defence mechanism based on Reputation Systems. In this way, we are able to identify and isolate possible malicious vehicles acting that could send messages with the aim of reducing the availability of the network. We propose Context Aware Reputation Systems, CARS, able to identify insider attackers and isolate them taking into account contextual conditions derived from sensors spread along the entire urban network. Then, we experimentally evaluate CARS on a real data-set of mobility traces of taxis in Rome to compare the proposed systems with existing ones that do not consider contextual conditions. The preliminary results obtained are promising and show the feasibility and potentiality of CARS.
Gianpiero Costantino, Fabio Martinelli, Ilaria Matteucci, Antonia Bertolino, Antonello Calabrò, Eda Marchetti
PDP4
2018 Coverage Testing and Reliability Improvement: A Marriage of Convenience?
Antonia Bertolino
WEBIST1
2018 A categorization scheme for software engineering conference papers and its application
Antonia Bertolino, Antonello Calabrò, Francesca Lonetti, Eda Marchetti, Breno Miranda
J. Syst. Softw.1
2018 A tour of secure software engineering solutions for connected vehicles
Antonia Bertolino, Antonello Calabrò, Felicita Di Giandomenico, Giuseppe Lami, Francesca Lonetti, Eda Marchetti, Fabio Martinelli, Ilaria Matteucci, Paolo Mori
Softw. Qual. J.1
2018 An assessment of operational coverage as both an adequacy and a selection criterion for operational profile based testing
abstract
While the relation between code coverage measures and fault detection is actively studied, only few works have investigated the correlation between measures of coverage and of reliability. In this work, we introduce a novel approach to measuring code coverage, called the operational coverage, that takes into account how much the program’s entities are exercised so to reflect the profile of usage into the measure of coverage. Operational coverage is proposed as (i) an adequacy criterion, i.e., to assess the thoroughness of a black box test suite derived from the operational profile, and as (ii) a selection criterion, i.e., to select test cases for operational profile-based testing. Our empirical evaluation showed that operational coverage is better correlated than traditional coverage with the probability that the next test case derived according to the user’s profile will not fail. This result suggests that our approach could provide a good stopping rule for operational profile-based testing. With respect to test case selection, our investigations revealed that operational coverage outperformed the traditional one in terms of test suite size and fault detection capability when we look at the average results.
Breno Miranda, Antonia Bertolino
Softw. Qual. J.2
2018 Automatic Testing of Design Faults in MapReduce Applications
abstract
New processing models are being adopted in Big Data engineering to overcome the limitations of traditional technology. Among them, MapReduce stands out by allowing for the processing of large volumes of data over a distributed infrastructure that can change during runtime. The developer only designs the functionality of the program and its execution is managed by a distributed system. As a consequence, a program can behave differently at each execution because it is automatically adapted to the resources available at each moment. Therefore, when the program has a design fault, this could be revealed in some executions and masked in others. However, during testing, these faults are usually masked because the test infrastructure is stable, and they are only revealed in production because the environment is more aggressive with infrastructure failures, among other reasons. This paper proposes new testing techniques that aimed to detect these design faults by simulating different infrastructure configurations. The testing techniques generate a representative set of infrastructure configurations that as whole are more likely to reveal failures using random testing, and partition testing together with combinatorial testing. The techniques are automated by using a test execution engine called MRTest that is able to detect these faults using only the test input data, regardless of the expected output. Our empirical evaluation shows that MRTest can automatically detect these design faults within a reasonable time.
Jesús Morán, Antonia Bertolino, Claudio de la Riva, Javier Tuya
IEEE Trans. Reliab.2
2017 Adaptive coverage and operational profile-based testing for reliability improvement
abstract
We introduce covrel, an adaptive software testing approach based on the combined use of operational profile and coverage spectrum, with the ultimate goal of improving the delivered reliability of the program under test. Operational profile-based testing is a black-box technique that selects test cases having the largest impact on failure probability in operation, as such, it is considered well suited when reliability is a major concern. Program spectrum is a characterization of a program's behavior in terms of the code entities (e.g., branches, statements, functions) that are covered as the program executes. The driving idea of covrel is to complement operational profile information with white-box coverage measures based on count spectra, so as to dynamically select the most effective test cases for reliability improvement. In particular, we bias operational profile-based test selection towards those entities covered less frequently. We assess the approach by experiments with 18 versions from 4 subjects commonly used in software testing research, comparing results with traditional operational and coverage testing. Results show that exploiting operational and coverage data in a combined adaptive way actually pays in terms of reliability improvement, with covrel overcoming conventional operational testing in more than 80% of the cases.
Antonia Bertolino, Breno Miranda, Roberto Pietrantuono, Stefano Russo 0001
ICSE1
2017 Towards Ex Vivo Testing of MapReduce Applications
abstract
Big Data programs are those that process large data exceeding the capabilities of traditional technologies. Among newly proposed processing models, MapReduce stands out as it allows the analysis of schema-less data in large distributed environments with frequent infrastructure failures. Functional faults in MapReduce are hard to detect in a testing/preproduction environment due to its distributed characteristics. We propose an automatic test framework implementing a novel testing approach called Ex Vivo. The framework employs data from production but executes the tests in a laboratory to avoid side-effects on the application. Faults are detected automatically without human intervention by checking if the same data would generate different outputs with different infrastructure configurations. The framework (MrExist) is validated with a real-world program. MrExist can identify a fault in a few seconds, then the program can be stopped, not only avoiding an incorrect output, but also saving money, time and energy of production resources.
Jesús Morán, Antonia Bertolino, Claudio de la Riva, Javier Tuya
QRS2
2017 Towards Automated Deployment of Self-adaptive Applications on Hybrid Clouds (Short Paper)
Lom-Messan Hillah, Rodrigo Elia Assad, Antonia Bertolino, Márcio Eduardo Delamaro, Fabio De Rosa, Vinicius Cardoso Garcia, Francesca Lonetti, Ariele-Paolo Maesano, Libero Maesano, Eda Marchetti, Breno Miranda, Auri M. R. Vincenzi, Juliano Iyoda
SEFM3
2017 Scope-aided test prioritization, selection and minimization for software reuse
Breno Miranda, Antonia Bertolino
J. Syst. Softw.2
2016 Learning Path Specification for Workplace Learning based on Business Process Management
abstract
In modern society, workers are continuously challenged to acquire new skills and competencies while at work. Novel approaches and tools to support effective and efficient workplace learning in collaborative and engaging ways are needed. On the other hand, Business Process Management (BPM) is more and more employed to support and manage the complex processes carried out within organizations. We propose to use BPM also to drive workplace learning, with the advantage of aligning real tasks to training tasks. We introduce a specification of learning path that maps BPM tasks and activities into sequences of learning tasks that can be customized to learners competence. The learning path specification can be used to both drive learning sessions, and to inform a monitor that can assess learner's progress. We describe a platform that is under development, and provide a simple motivational example to illustrate the approach. The goal is to combine work and learning in natural and effective way.
Venkatapathy Subramanian, Antonia Bertolino
CSEDU (1)2
2016 A Tool-Supported Methodology for Validation and Refinement of Early-Stage Domain Models
abstract
Model-driven engineering (MDE) promotes automated model transformations along the entire development process. Guaranteeing the quality of early models is essential for a successful application of MDE techniques and related tool-supported model refinements. Do these models properly reflect the requirements elicited from the owners of the problem domain? Ultimately, this question needs to be asked to the domain experts. The problem is that a gap exists between the respective backgrounds of modeling experts and domain experts. MDE developers cannot show a model to the domain experts and simply ask them whether it is correct with respect to the requirements they had in mind. To facilitate their interaction and make such validation more systematic, we propose a methodology and a tool that derive a set of customizable questionnaires expressed in natural language from each model to be validated. Unexpected answers by domain experts help to identify those portions of the models requiring deeper attention. We illustrate the methodology and the current status of the developed tool MOTHIA, which can handle UML Use Case, Class, and Activity diagrams. We assess MOTHIA effectiveness in reducing the gap between domain and modeling experts, and in detecting modeling faults on the European Project CHOReOS.
Marco Autili, Antonia Bertolino, Guglielmo De Angelis, Davide Di Ruscio, Alessio Di Sandro
IEEE Trans. Software Eng.2
2015 Similarity testing for access control
Antonia Bertolino, Said Daoudagh, Donia El Kateb, Christopher Henard, Yves Le Traon, Francesca Lonetti, Eda Marchetti, Tejeddine Mouelhi, Mike Papadakis
Inf. Softw. Technol.1
2014 Extending UML Testing Profile Towards Non-functional Test Modeling
abstract
The research community has broadly recognized the importance of the validation of non-functional properties including performance and dependability requirements. However, the results of a systematic survey we carried out evidenced the lack of a standard notation for designing non-functional test cases. For some time, the greatest attention of Model-Based Testing (MBT) research has focused on functional aspects. The only exception is represented by the UML Testing Profile (UML-TP) that is a lightweight extension of UML to support the design of testing artifacts, but it only provides limited support for non-functional testing. In this paper we provide a first attempt to extend UML-TP for improving the design of non-functional tests. The proposed extension deals with some important concepts of non-functional testing such as the workload and the global verdicts. As a proof of concept we show how the extended UML-TP can be used for modeling non-functional test cases of an application example.
Federico Toledo Rodríguez, Francesca Lonetti, Antonia Bertolino, Macario Polo, Beatriz Pérez Lamancha
MODELSWARD3
2014 A Requirements-Led Approach for Specifying QoS-Aware Service Choreographies: An Experience Report
Neil A. M. Maiden, James Lockerbie, Konstantinos Zachos, Antonia Bertolino, Guglielmo De Angelis, Francesca Lonetti
REFSQ4
2014 Testing of PolPA-based usage control systems
Antonia Bertolino, Said Daoudagh, Francesca Lonetti, Eda Marchetti, Fabio Martinelli, Paolo Mori
Softw. Qual. J.1
2014 Editorial for the special issue of STVR on the 5th IEEE International Conference on Software Testing, Verification, and Validation (ICST 2012)
abstract
The 5th IEEE International Conference on Software Testing, Verification, and Validation (ICST 2012) was held on 17–21 April 2012, in Montreal, Québec, Canada. ICST has established itself as the premier forum for presentation of leading edge results on outstanding topics in all areas related to software quality. The conference brings together researchers and practitioners who study the theory, techniques, technologies, and applications of software testing, verification, and validation. ICST 2012 originally attracted 177 submissions, among which 41 research papers and seven industry papers were selected for inclusion in the proceedings. All papers were refereed by at least three members of the ICST 2012 Program Committee. Of the 48 papers accepted, five papers were selected by the Program Co-Chairs, Antonia Bertolino and Yvan Labiche, for consideration for this special issue of STVR. These papers were extended from their conference version by the authors and subjected to the rigorous STVR reviewing process, thus undergoing additional rounds of reviews and revisions. Three papers successfully completed the review process and are contained in this special issue. The first paper, “Automatic Testing of GUI-Based Applications” by Leonardo Mariani, Mauro Pezzè, Oliviero Riganelli, and Mauro Santoro, describes a technique to automatically create a test suite, or augment an existing one, to exercise a GUI-based application with black box, system level test cases. The test cases are synthesized incrementally by employing reinforcement learning. Results indicate improvements over other state of the art techniques in terms of amount of behaviour being exercised by the augmented test suite, and therefore in potential of fault detection. The second paper, “Automatic Test Case Evolution” by Mehdi Mirzaaghaei, Fabrizio Pastore, and Mauro Pezzè, presents a framework to repair test cases that have become obsolete due to evolution and also use information in an existing test suite to generate new test cases. The authors refer to this process as the evolution of a test suite. Experimental evaluation on five different case study systems provides very encouraging results, with a high effectiveness at repairing broken test cases. The third paper, “An Efficient Regression Testing Approach for PHP Web Applications: A Controlled Experiment” by Hyunsook Do and Md. Hossain, introduces a regression testing technique for PHP web applications based on impact analysis and program slicing. A controlled experiment on five non-trivial web applications taken from SourceForge shows a promising reduction of regression testing costs. The special issue editors would like to express their gratitude to the many people who contributed to the successful organization of ICST 2012. Although it is not possible to list them all, a mention goes to the General Chair, Giulio Antoniol, the Industrial Track Chairs, Thomas Ostrand, Saurabh Sinha, and Peter Zimmerer, and the ICST Steering Committee members. The high quality of the papers in this special issue is certainly due to all ICST authors, who dedicated effort and energy in sharing their results with the community, and made up a great pool from which invitations to submit to the journal could be issued; it is due as well to the ICST 2012 Program Committee and the STVR additional reviewers, who provided extensive competent feedback. Last but not least, the STVR chief editors, Robert Hierons and Jeff Offutt, provided expert guidance and important advice throughout the process. Enjoy!
Antonia Bertolino, Yvan Labiche
Softw. Test. Verification Reliab.1
2013 A Toolchain for Designing and Testing XACML Policies
abstract
In modern pervasive application domains, such as Service Oriented Architectures (SOAs) and Peer-to-Peer (P2P) systems, security aspects are critical. Justified confidence in the security mechanisms that are implemented for assuring proper data access is a key point. In the last years XACML has become the de facto standard for specifying policies for access control decisions in many application domains. Briefly, an XACML policy defines the constraints and conditions that a subject needs to comply with for accessing a resource and doing an action in a given environment. Due to the complexity of the language, XACML policy specification is a difficult and error prone process that requires specific knowledge and a high effort to be properly managed.
Antonia Bertolino, Marianne Busch, Said Daoudagh, Nora Koch, Francesca Lonetti, Eda Marchetti
ICST1
2013 A Generative Approach for the Adaptive Monitoring of SLA in Service Choreographies
Antonia Bertolino, Antonello Calabrò, Guglielmo De Angelis
ICWE1
2013 Adequate monitoring of service compositions
abstract
Monitoring is essential to validate the runtime behaviour of dynamic distributed systems. However, monitors can inform of relevant events as they occur, but by their very nature they will not report about all those events that are not happening. In service-oriented applications it would be desirable to have means to assess the thoroughness of the interactions among the services that are being monitored. In case some events or message sequences or interaction patterns have not been observed for a while, in fact, one could timely check whether this happens because something is going wrong. In this paper, we introduce the novel notion of monitoring adequacy, which is generic and can be defined on different entities. We then define two adequacy criteria for service compositions and implement a proof-of-concept adequate monitoring framework. We validate the approach on two case studies, the Travel Reservation System and the Future Market choreographies.
Antonia Bertolino, Eda Marchetti, Andrea Morichetta 0001
ESEC/SIGSOFT FSE1
2013 Guest Editorial for Special Section from Component-based Software Engineering (CBSE) 2011
Antonia Bertolino, Kendra M. L. Cooper
Inf. Softw. Technol.1
2013 Special section on automation of software test
Antonia Bertolino, Howard Foster, J. Jenny Li 0001, Hong Zhu 0002
J. Syst. Softw.1
2012 Automatic XACML Requests Generation for Policy Testing
abstract
Access control policies are usually specified by the XACML language. However, policy definition could be an error prone process, because of the many constraints and rules that have to be specified. In order to increase the confidence on defined XACML policies, an accurate testing activity could be a valid solution. The typical policy testing is performed by deriving specific test cases, i.e. XACML requests, that are executed by means of a PDP implementation, so to evidence possible security lacks or problems. Thus the fault detection effectiveness of derived test suite is a fundamental property. To evaluate the performance of the applied test strategy and consequently of the test suite, a commonly adopted methodology is using mutation testing. In this paper, we propose two different methodologies for deriving XACML requests, that are defined independently from the policy under test. The proposals exploit the values of the XACML policy for better customizing the generated requests and providing a more effective test suite. The proposed methodologies have been compared in terms of their fault detection effectiveness by the application of mutation testing on a set of real policies.
Antonia Bertolino, Said Daoudagh, Francesca Lonetti, Eda Marchetti
ICST1
2012 Validation and Verification Policies for Governance of Service Choreographies
Guglielmo De Angelis, Antonia Bertolino, Andrea Polini
WEBIST2
2012 Quality Requirements for Service Choreographies
Cesare Bartolini, Antonia Bertolino, Andrea Ciancone, Guglielmo De Angelis, Raffaela Mirandola
WEBIST2
2012 The X-CREATE Framework - A Comparison of XACML Policy Testing Strategies
Antonia Bertolino, Said Daoudagh, Francesca Lonetti, Eda Marchetti
WEBIST1
2011 Teaching software testing: Experiences, lessons learned and the path forward
abstract
According to a study commissioned by the National Institute of Standards and Technology in 2002, software bugs cost the U.S. economy an estimated $59.5 billion annually, or about 0.6 percent of the nation's gross domestic product (GDP). The same study also found that more than one-third of these costs, or an estimated $22.2 billion, could be eliminated by an improved testing infrastructure. These numbers would be significantly higher if the study were conducted today.
W. Eric Wong, Antonia Bertolino, Vidroha Debroy, Aditya P. Mathur, A. Jefferson Offutt, Mladen A. Vouk
CSEE&T2
2011 Sixth international workshop on automation of software test: (AST 2011)
abstract
The Sixth International Workshop on Automation of Software Test (AST 2011) is associated with the 33rd International Conference on Software Engineering (ICSE 2011). This edition of AST was focused on the special theme of Software Design and the Automation of Software Test and authors were encouraged to submit work in this area. The workshop covers two days with presentations of regular research papers, industrial case studies and experience reports. The workshop also aims to have extensive discussions on collaborative solutions in the form of charette sessions. This paper summarizes the organization of the workshop, the special theme, as well as the sessions.
Howard Foster, Antonia Bertolino, J. Jenny Li 0001
ICSE2
2011 Towards Ensuring Eternal Connectability
Antonia Bertolino
ICSOFT (1)1
2011 Automated Refinement of Dependability Analysis through Monitoring in Dynamically Connected Systems
abstract
Model-based analysis is a well-established method to assess the dependability of a system before deployment. It is well known that, in highly dynamic contexts, the accuracy of the analysis results can be limited because unpredictable phenomena may affect the system during its operation. In such contexts, the analysis typically needs to be refined with data obtained from real system executions. In this paper we tackle the issue of refining model-based dependability analysis in automated systems through monitoring. Specifically, we report on our preliminary results on the development of a system that exploits the synergic use of an automated approach for model-based dependability analysis and a flexible monitoring architecture.
Antonia Bertolino, Antonello Calabrò, Felicita Di Giandomenico, Marco Martinucci, Paolo Masci 0001
ISADS1
2011 (role)CAST: A Framework for On-line Service Testing
Antonia Bertolino, Guglielmo De Angelis, Andrea Polini
WEBIST1
2011 Bringing white-box testing to Service Oriented Architectures through a Service Oriented Approach
Cesare Bartolini, Antonia Bertolino, Sebastian G. Elbaum, Eda Marchetti
J. Syst. Softw.2
2011 Is my model right? Let me ask the expert
Antonia Bertolino, Guglielmo De Angelis, Alessio Di Sandro, Antonino Sabetta
J. Syst. Softw.1
2010 On-the-Fly Interoperability through Automated Mediator Synthesis and Monitoring
Antonia Bertolino, Paola Inverardi, Valérie Issarny, Antonino Sabetta, Romina Spalazzese
ISoLA (2)1
2009 CONNECT Challenges: Towards Emergent Connectors for Eternal Networked Systems
abstract
The CONNECT European project that started in February 2009 aims at dropping the interoperability barrier faced by todaypsilas distributed systems. It does so by adopting a revolutionary approach to the seamless networking of digital systems, that is, synthesizing on the fly the connectors via which networked systems communicate. CONNECT then investigates formal foundations for connectors together with associated automated support for learning, reasoning about and adapting the interaction behavior of networked systems.
Valérie Issarny, Bernhard Steffen, Bengt Jonsson 0001, Gordon S. Blair, Paul Grace, Marta Z. Kwiatkowska, Radu Calinescu, Paola Inverardi, Massimo Tivoli, Antonia Bertolino, Antonino Sabetta
ICECCS10
2009 WS-TAXI: A WSDL-based Testing Tool for Web Services
abstract
Web services (WSs) are the W3C-endorsed realization of the Service-Oriented Architecture (SOA). Since they are supposed to be implementation-neutral, WSs are typically tested black-box at their interface. Such an interface is generally specified in an XML-based notation called the WS Description Language (WSDL). Conceptually, these WSDL documents are eligible for fully automated WS test generation using syntax-based testing approaches. Towards such goal, we introduce the WS-TAXI framework, in which we combine the coverage of WS operations with data-driven test generation. In this paper we present an early-stage implementation of WS-TAXI, obtained by the integration of two existing softwares: soapUI, a popular tool for WS testing, and TAXI, an application we have previously developed for the automated derivation of XML instances from a XML schema. WS-TAXI delivers a complete suite of test messages ready for execution. Test generation is driven by basic coverage criteria and by the application of some heuristics. The application of WS-TAXI to a real case study gave encouraging results.
Cesare Bartolini, Antonia Bertolino, Eda Marchetti, Andrea Polini
ICST2
2009 Whitening SOA testing
abstract
Service Oriented Architectures (SOAs) are becoming increasingly popular and powerful. Fueling that growth is the availability of independent web services that can be cost-effectively composed with other services to provide richer functionality. The reasons that make these systems easier to build, however, also make them more challenging to test. Independent web services usually provide just an interface, enough to invoke them and develop some general (black-box) tests, but insufficient for a tester to develop an adequate understanding of the integration quality between the application and independent web services. To address this lack we propose a "whitening" approach to make web services more transparent through the addition of an intermediate coverage service. The approach, named Service Oriented Coverage Testing (SOCT), provides a tester with feedback about how a whitened service, called a Testable Service, is exercised. In this paper we introduce the SOCT approach, implement an instance of it, and perform a preliminary study to show its feasibility and potential value. SOCT enables SOA white-box testing, while maintaining SOA flexibility, dynamism and loose coupling.
Cesare Bartolini, Antonia Bertolino, Sebastian G. Elbaum, Eda Marchetti
ESEC/SIGSOFT FSE2
2009 Automatic synthesis of behavior protocols for composable web-services
abstract
Web-services are broadly considered as an effective means to achieve interoperability between heterogeneous parties of a business process and offer an open platform for developing new composite web-services out of existing ones. In the literature many approaches have been proposed with the aim to automatically compose web-services. All of them assume that, along with the web-service signature, some information is provided about how clients interacting with the web-service should behave when invoking it.
Antonia Bertolino, Paola Inverardi, Patrizio Pelliccione, Massimo Tivoli
ESEC/SIGSOFT FSE1
2008 Towards Automated WSDL-Based Testing of Web Services
Cesare Bartolini, Antonia Bertolino, Eda Marchetti, Andrea Polini
ICSOC2
2008 A Framework for Analyzing and Testing the Performance of Software Services
Antonia Bertolino, Guglielmo De Angelis, Antinisca Di Marco, Paola Inverardi, Antonino Sabetta, Massimo Tivoli
ISoLA1
2008 VCR: Virtual Capture and Replay for Performance Testing
abstract
This paper proposes a novel approach to performance testing, called virtual capture and replay (VCR), that couples capture and replay techniques with the checkpointing capabilities provided by the latest virtualization technologies. VCR enables software performance testers to automatically take a snapshot of a running system when certain critical conditions are verified, and to later replay the scenario that led to those conditions. Several in-depth analyses can be separately carried out in the laboratory just by rewinding the captured scenario and replaying it using different probes and analysis tools.
Antonia Bertolino, Guglielmo De Angelis, Antonino Sabetta
ASE1
2008 Software Testing Forever: Old and New Processes and Techniques for Validating Today's Applications
Antonia Bertolino
PROFES1
2007 A QoS Test-Bed Generator for Web Services
Antonia Bertolino, Guglielmo De Angelis, Andrea Polini
ICWE1
2007 Welcome to the WISE track
abstract
This year ESCE/FSE launches the new Widened Software Engineering (WISE) track with an explicit aim to widen international participation, especially from countries which are usually under-represented in the conference audience.
Antonia Bertolino, Henry Muccini
ESEC/SIGSOFT FSE1
2007 XModel-Based Testing of XSLT Applications
Antonia Bertolino, Jinghua Gao, Eda Marchetti, Andrea Polini
WEBIST (2)1
2007 Testing software components for integration: a survey of issues and techniques
abstract
Abstract Component‐based development has emerged as a system engineering approach that promises rapid software development with fewer resources. Yet, improved reuse and reduced cost benefits from software components can only be achieved in practice if the components provide reliable services, thereby rendering component analysis and testing a key activity. This paper discusses various issues that can arise in component testing by the component user at the stage of its integration within the target system. The crucial problem is the lack of information for analysis and testing of externally developed components. Several testing techniques for component integration have recently been proposed. These techniques are surveyed here and classified according to a proposed set of relevant attributes. The paper thus provides a comprehensive overview which can be useful as introductory reading for newcomers in this research field, as well as to stimulate further investigation. Copyright © 2006 John Wiley & Sons, Ltd.
Muhammad Jaffar-Ur Rehman, Fakhra Jabeen, Antonia Bertolino, Andrea Polini
Softw. Test. Verification Reliab.3
2006 Modeling and Early Performance Estimation for Network Processor Applications
Antonia Bertolino, Alvise Bonivento, Guglielmo De Angelis, Alberto L. Sangiovanni-Vincentelli
MoDELS1
2006 XML Every-Flavor Testing
Antonia Bertolino, Jinghua Gao, Eda Marchetti
WEBIST (1)1
2004 Using Software Architecture for Code Testing
abstract
Our research deals with the use of software architecture (SA) as a reference model for testing the conformance of an implemented system with respect to its architectural specification. We exploit the specification of SA dynamics to identify useful schemes of interactions between system components and to select test classes corresponding to relevant architectural behaviors. The SA dynamics is modeled by labeled transition systems (LTSs). The approach consists of deriving suitable LTS abstractions called ALTSs. ALTSs offer specific views of SA dynamics by concentrating on relevant features and abstracting away from uninteresting ones. Intuitively, deriving an adequate set of test classes entails deriving a set of paths that appropriately cover the ALTS. Next, a relation between these abstract SA tests and more concrete, executable tests needs to be established so that the architectural tests derived can be refined into code-level tests. We use the TRMCS case study to illustrate our hands-on experience. We discuss the insights gained and highlight some issues, problems, and solutions of general interest in architecture-based testing.
Henry Muccini, Antonia Bertolino, Paola Inverardi
IEEE Trans. Software Eng.2
2003 A Framework for Component Deployment Testing
abstract
Component-based development is the emerging paradigm in software production, though several challenges still slow down its full taking up. In particular, the "component trust problem" refers to how adequate guarantees and documentation about a component's behaviour can be transferred from the component developer to its potential users. The capability to test a component when deployed within the target application environment can help establish the compliance of a candidate component to the customer's expectations and certainly contributes to "increase trust". To this purpose, we propose the CDT framework for Component Deployment Testing. CDT provides the customer with both a technique to early specify a deployment test suite and an environment for running and reusing the specified tests on any component implementation. The framework can also be used to deliver the component developer's test suite and to later re-execute it. The central feature of CDT is the complete decoupling between the specification of the tests and the component implementation.
Antonia Bertolino, Andrea Polini
ICSE1
2003 Use case-based testing of product lines
abstract
This paper presents PLUTO, a simple and intuitive methodology to manage the testing process of product lines, described as Product Lines Use Cases (PLUCs). PLUCs are an extension of the well-known Cockburn's Use Cases, a notation based on natural language descriptions of requirements. The proposed test methodology is based on the Category Partition method, and can be used to derive a generic Test Specification for the product line, and a set of relevant test scenarios for a customer specific application.
Antonia Bertolino, Stefania Gnesi
ESEC / SIGSOFT FSE1
2003 Using Spanning Sets for Coverage Testing
abstract
A test coverage criterion defines a set E/sub r/ of entities of the program flowgraph and requires that every entity in this set is covered under some test Case. Coverage criteria are also used to measure the adequacy of the executed test cases. In this paper, we introduce the notion of spanning sets of entities for coverage testing. A spanning set is a minimum subset of E/sub r/, such that a test suite covering the entities in this subset is guaranteed to cover every entity in E/sub r/. When the coverage of an entity always guarantees the coverage of another entity, the former is said to subsume the latter. Based on the subsumption relation between entities, we provide a generic algorithm to find spanning sets for control flow and data flow-based test coverage criteria. We suggest several useful applications of spanning sets: They help reduce and estimate the number of test cases needed to satisfy coverage criteria. We also empirically investigate how the use of spanning sets affects the fault detection effectiveness.
Martina Marré, Antonia Bertolino
IEEE Trans. Software Eng.2
2002 ISSTA 2002 panel: is ISSTA research relevant to industrial users?
Antonia Bertolino
ISSTA1
2002 Preventing untestedness in data-flow based testing
abstract
Abstract A large number of path‐oriented testing criteria have been proposed in the last twenty years. Surprisingly, almost all of them suffer from a serious weakness, which is called the untestedness syndrome: even though a criterion is satisfied, some statements of the program under test may remain ‘untested’, i.e., the observed test output does not depend on them. A new data‐flow based testing criterion is introduced which does not suffer from untestedness, called the All Program Function (APF) criterion. Intuitively, it requires that each possible computation to every output statement in a program be covered by some test; but for lots of programs APF would require an infinite number of tests. A second, applicable criterion is thus introduced, derived from APF and called the Basic Program Function (BPF) criterion. BPF leaves no statement untested and yields finite test suites. Some examples show the application of BPF and investigate the failure‐detection capability of the proposed criterion. Copyright © 2001 John Wiley & Sons, Ltd.
István Forgács, Antonia Bertolino
Softw. Test. Verification Reliab.2
2002 Guest Editors' Introduction: 2000 International Symposium on Software Testing and Analysis
Mary Jean Harrold, Antonia Bertolino
IEEE Trans. Software Eng.2
2001 An Explorative Journey from Architectural Tests Definition downto Code Tests Execution
Antonia Bertolino, Paola Inverardi, Henry Muccini
ICSE1
2000 Deriving test plans from architectural descriptions
abstract
The paper presents an approach to derive test plans for the conformance testing of a system implementation with respect to the formal description of its Software Architecture (SA). The SA describes a system in terms of its components and connections, therefore the derived test plans address the integration testing phase. We base our approach on a Labelled Transition System (LTS) modeling the SA dynamics, and on suitable abstractions of it, the Abstract Labelled Transition Systems (ALTSs). ALTSs oer specic views of the SA dynamics by concentrating on relevant features and abstracting away from uninteresting ones. ALTS is a tool we provide the software architect with allow him/her to focus on relevant behavioral patterns and more easily identify those ones that are more meaningful for validation purposes. Intuitively deriving an adequate set of functional test classes means deriving a set of paths appropriately covering the ALTS. In the paper we describe our approach in the scope of a...
Antonia Bertolino, Flavio Corradini, Paola Inverardi, Henry Muccini
ICSE1
2000 An overview of the ICSE 2000 workshop program
abstract
Past ICSE attendees will recognize—with pleasure, we hope—workshops that have been successful in previous years. Indeed, we have tried to balance the program between workshops based on novel and promising ideas, with those strongly continuing the work started in previous ICSEs. In two cases, the program also includes workshops that already have some tradition, but are associated with ICSE for the first time: the ISAW workshop (4th edition) and the DSV-IS workshop (7th edition).
Antonia Bertolino, Gail C. Murphy
ICSE1
1999 Towards Statistical Control of an Industrial Test Process
Gaetano Lombardi, Emilia Peciola, Raffaela Mirandola, Antonia Bertolino, Eda Marchetti
SAFECOMP4
1998 Assessing the Risk due to Software Faults: Estimates of Failure Rate versus Evidence of Perfection
abstract
In the debate over the assessment of software reliability (or safety), as applied to critical software, two extreme positions can be discerned: the ‘statistical’ position, which requires that the claims of reliability be supported by statistical inference from realistic testing or operation, and the ‘perfectionist’ position, which requires convincing indications that the software is free from defects. These two positions naturally lead to requiring different kinds of supporting evidence, and actually to stating the dependability requirements in different ways, not allowing any direct comparison. There is often confusion about the relationship between statements about software failure rates and about software correctness, and about which evidence can support either kind of statement. This note clarifies the meaning of the two kinds of statement and how they relate to the probability of failure-free operation, and discusses their practical merits, especially for high required reliability or safety. © 1998 John Wiley & Sons, Ltd.
Antonia Bertolino, Lorenzo Strigini
Softw. Test. Verification Reliab.1
1997 An approach to integration testing based on architectural descriptions
abstract
Software architectures can play a role in improving the testing process of complex systems. In particular descriptions of the software architecture can be useful to drive integration testing, since they supply information about how the software is structured in parts and how those parts (are expected to) interact. We propose to use formal architectural descriptions to model the "interesting" behaviour of the system. This model is at a right level of abstraction to be used as a formal base on which integration test strategies can be devised. Starting from a formal description of the software architecture (given in the CHAM formalism), we first derive a graph of all the possible behaviours of the system in terms of the interactions between its components. This graph contains altogether the information we need for the planning of integration testing. On this comprehensive model, we then identify a suitable set of reduced graphs, each highlighting specific architectural properties of the system. These reduced graphs can be used for the generation of integration tests according to a coverage strategy, analogously to what happens with the control and data flow graphs in unit testing.
Antonia Bertolino, Paola Inverardi, Henry Muccini, Andrea Rosetti
ICECCS1
1997 A case study in branch testing automation
Antonia Bertolino, Raffaela Mirandola, Emilia Peciola
J. Syst. Softw.1
1996 Reducing and Estimating the Cost of Test Coverage Criteria
Martina Marré, Antonia Bertolino
ICSE2
1996 Unconstrained Duals and Their Use in Achieving All-Uses Coverage
abstract
Testing takes a considerable amount of the time and resources spent on producing software. It would therefore be useful to have ways 1) to reduce the cost of testing and 2) to estimate this cost. In particular, the number of tests to be executed is an important and useful attribute of the entity "testing effort". All-uses coverage is a data flow testing strategy widely researched in recent years. In this paper we present spanning sets of duas for the all-uses coverage criterion. A spanning set of duas is a minimum set of duas (definition-use associations) such that a set of test paths covering them covers every dua in the program. We give a method to find a spanning set of duas using the relation of subsumption between duas. Intuitively, there exists a natural ordering between the duas in a program: some duas are covered more easily than others, since coverage of the former is automatically guaranteed whenever the latter are covered. Those duas that are the most difficult to be covered according to this ordering are called unconstrained. A spanning set of duas is composed of unconstrained duas. Our results are useful for reducing the cost of testing, since the generation of test paths can be targeted to cover the smaller spanning set of duas, rather than all those in a program. On the other hand, assuming that a different path is taken to cover each dua in a spanning set, the cardinality of spanning sets can be used to estimate the cost of testing. Other interesting uses of spanning sets of duas are also discussed.
Martina Marré, Antonia Bertolino
ISSTA2
1996 Acceptance Criteria for Critical Software Based on Testability Estimates and Test Results
Antonia Bertolino, Lorenzo Strigini
SAFECOMP1
1996 How Many Paths are Needed for Branch Testing?
Antonia Bertolino, Martina Marré
J. Syst. Softw.1
1996 Software Assessment: Reliability, Safety, Testability, by Michael A. Friedman and Jeffrey M. Voas, Wiley, 1995 (Book Review)
Antonia Bertolino
Softw. Test. Verification Reliab.1
1996 On the Use of Testability Measures for Dependability Assessment
abstract
Program "testability" is informally, the probability that a program will fail under test if it contains at least one fault. When a dependability assessment has to be derived from the observation of a series of failure free test executions (a common need for software subject to "ultra high reliability" requirements), measures of testability can-in theory-be used to draw inferences on program correctness. We rigorously investigate the concept of testability and its use in dependability assessment, criticizing, and improving on, previously published results. We give a general descriptive model of program execution and testing, on which the different measures of interest can be defined. We propose a more precise definition of program testability than that given by other authors, and discuss how to increase testing effectiveness without impairing program reliability in operation. We then study the mathematics of using testability to estimate, from test results: the probability of program correctness and the probability of failures. To derive the probability of program correctness, we use a Bayesian inference procedure and argue that this is more useful than deriving a classical "confidence level". We also show that a high testability is not an unconditionally desirable property for a program. In particular, for programs complex enough that they are unlikely to be completely fault free, increasing testability may produce a program which will be less trustworthy, even after successful testing.
Antonia Bertolino, Lorenzo Strigini
IEEE Trans. Software Eng.1
1995 Using Testability Measures for Dependability Assessment
abstract
Program "testability"is the probability that a fault in a program, if present, will cause the program to fail.Measures of testability can be used to draw inferences on program correctness from the observation of a series of failure-free test executions, a common need for software with "ultra-high reliability" requirements.For a program that has passed a certain number of tests without failing, a high value of testability implies a high probability that the program is correct.We give a general descriptive model of program execution and testing, and propose a more precise definition of program testability than that given by other authors.We then study the use of testability in: i) providing, through testing, confidence in the absence of faults and ii) bounding the probability of failures, from the results of operational testing.We derive the probability of absence of faults through a Bayesian inference procedure, criticise previously proposed derivations of this probability, and study the relationship between the testability of a program and its failure probability in operation.We derive the conditions under which a high testability improves one's expectations about program reliability.Last, we discuss the potential of these methods in practical applications.
Antonia Bertolino, Lorenzo Strigini
ICSE1
1994 A Meaningful Bound for Branch Testing (Abstract)
abstract
Branch coverage is often used to evaluate testing thoroughness. For this reason, it is useful to set a lower bound on the number of test paths needed to achieve branch coverage when estimating how much effort will be needed to test a given program.
Antonia Bertolino, Martina Marré
ISSTA1
1994 Guest editor's corner achieving quality in software
Antonia Bertolino
J. Syst. Softw.1
1994 Automatic Generation of Path Covers Based on the Control Flow Analysis of Computer Programs
abstract
Branch testing a program involves generating a set of paths that will cover every arc in the program flowgraph, called a path cover, and finding a set of program inputs that will execute every path in the path cover. This paper presents a generalized algorithm that finds a path cover for a given program flowgraph. The analysis is conducted on a reduced flowgraph, called a ddgraph, and uses graph theoretic principles differently than previous approaches. In particular, the relations of dominance and implication which form two trees of the arcs of the ddgraph are exploited. These relations make it possible to identify a subset of ddgraph arcs, called unconstrained arcs, having the property that a set of paths exercising all the unconstrained arcs also cover all the arcs in the ddgraph. In fact, the algorithm has been designed to cover all the unconstrained arcs of a given ddgraph: the paths are derived one at a time, each path covering at least one as yet uncovered unconstrained arc. The greatest merits of the algorithm are its simplicity and its flexibility. It consists in just visiting recursively in combination the dominator and the implied trees, and is flexible in the sense that it can derive a path cover to satisfy different requirements, according to the strategy adopted for the selection of the unconstrained arc to be covered at each recursive iteration. This feature of the algorithm can be employed to address the problem of infeasible paths, by adopting the most suitable selection strategy for the problem at hand. Embedding of the algorithm into a software analysis and testing tool is recommended.>
Antonia Bertolino, Martina Marré
IEEE Trans. Software Eng.1
1993 Unconstrained edges and their application to branch analysis and testing of programs
Antonia Bertolino
J. Syst. Softw.1
1991 An overview of automated software testing
Antonia Bertolino
J. Syst. Softw.1
1988 An Approach to Efficient Distributed Transactions
Paolo Ancilotti, Antonia Bertolino, Mario Fusani
Distributed Comput.2