EDBT 2026 Demo / reviewers in the wild / expert
Breno Miranda
dblp:119/0420 · also Breno Alexandro Ferreira de Miranda
· DBLP profile ↗
32ranked-venue papers
8as first author
17since 2021 · last 2026
0000-0001-9608-9393ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 32 · 8 first-author · 17 since 2021Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Toward a non-invasive robotic-supported approach for data loss detection in android applicationsabstractContext: Maintaining data consistency in mobile applications is essential, as activity restarts—such as those triggered by screen rotations—can lead to user data loss. Existing automated approaches, including DLD and R-DLD, rely on semi-invasive testing methods that require system-level access through debugging tools like ADB or app instrumentation. These dependencies limit their applicability to specific devices and reduce realism in user interaction. Objective: This work aims to overcome the limitations of semi-invasive data loss detection by introducing a fully non-invasive, hardware-based framework for Android applications. The goal is to preserve the authenticity of user interaction while maintaining detection accuracy comparable to existing approaches. Method: We present a new version of R-DLD, which leverages an industrial robotic arm to reproduce realistic user interactions such as touch gestures, scrolling, and device rotations without system-level access. The oracle was completely redesigned, using an external camera and computer vision techniques that apply similarity and dissimilarity metrics to identify data loss after double rotation events. Results: Experimental evaluation on thirty benchmark Android applications showed that the new R-DLD detected 100% of known data loss cases, matching R-DLD’s performance while operating entirely non-invasively. The results also demonstrate the robustness of the visual oracle against environmental lighting variations. Conclusion: The new R-DLD represents a significant advancement in realistic, non-invasive testing for mobile applications. By eliminating the need for ADB or instrumentation, it broadens the applicability of automated data loss detection and establishes a foundation for future research in vision-based, robotics-supported testing. Davi Freitas, Breno Miranda, Juliano Iyoda |
Inf. Softw. Technol. | 2 |
| 2025 | A Tool-Assisted Training Approach for Empowering Localization and Internationalization Testing ProficiencyabstractSoftware testing is an important area in the Software Engineering domain, yet it faces significant gaps within Computer Science education. In that context, there is an even less addressed topic: internationalization (i18n) and localization (l10n) tests, which are essential for ensuring software quality in global markets, supporting multiple languages. A key challenge for both large and small companies is the onboarding of new testers without adequate skills, often resulting in insufficient real-world preparation. This paper proposes a tool-assisted training approach designed to enhance the proficiency of l10n and i18n testers through practical, hands-on exercises using real-world failure examples. The effectiveness of our approach was evaluated from two complementary perspectives: i) To assess the effectiveness of the training approach, a case study was conducted within a software industry training setting, comparing novice testers trained with our tool against those receiving conventional training. Results indicated a 150% improvement in fault identification and resolution for testers using the tool, underscoring its effectiveness in enhancing testing skills and overall software quality. ii) To evaluate the perceived usability of the tool developed to support the training activities, a System Usability Scale (SUS) questionnaire was given to five technical leaders responsible for training novice testers. The tool achieved a score of 94.5, positioning its usability between “excellent” and “best imaginable”. This paper highlights the tool's design, deployment context, and its potential for adoption by other practitioners, aiming to address the current gaps in i18n and l10n tester training, being a promising approach to be used in Software Engineering education. Maria Couto, Breno Miranda, Kiev Gama |
ICST | 2 |
| 2025 | Microservices testing: A systematic literature review
Francisco Ponce 0001, Roberto Verdecchia, Breno Miranda, Jacopo Soldani |
Inf. Softw. Technol. | 3 |
| 2023 | Orchestration Strategies for Regression Test SuitesabstractRegression testing is widely studied in the literature, although most research on the topic is concerned with improving specific sub-challenges of a wider goal. Test suite orchestration proposes a more comprehensive view of the challenge of regression testing, by merging and combining different techniques with a variety of objectives, including prioritizing, selecting, reducing and amplifying tests, detecting flaky tests and potentially more. This paper presents the key approaches and techniques that form test suite orchestration, along with common evaluation metrics, and discusses how they can be used together to ultimately provide an efficient and effective regression testing strategy. To illustrate the benefits of orchestration, we provide some examples of existing papers that take steps towards this goal, even if the specific terminology is not yet used. Orchestrated strategies utilizing existing regression testing techniques provide a pathway to practicality and real-world usage of the academic literature. Renan Greca, Breno Miranda, Antonia Bertolino |
AST | 2 |
| 2023 | DevOpRET: Continuous reliability testing in DevOpsabstractAbstract To enter the production stage, in DevOps practices candidate software releases have to pass quality gates, where they are assessed to meet established target values for key indicators of interest. We believe software reliability should be an important such indicator, as it greatly contributes to the end‐user satisfaction. We propose DevOpRET , an approach for reliability testing as part of the acceptance testing stage in DevOps. DevOpRET relies on operational‐profile–based testing, a common reliability assessment technique. DevOpRET leverages usage and failure data monitored in operations to continuously refine its estimate. We evaluate accuracy and efficiency of DevOpRET through controlled experiments with a real‐world open source platform and with a microservice architectures benchmark. The results show that DevOpRET provides accurate and efficient estimates of the true reliability over subsequent DevOps cycles. Antonia Bertolino, Guglielmo De Angelis, Antonio Guerriero, Breno Miranda, Roberto Pietrantuono, Stefano Russo 0001 |
J. Softw. Evol. Process. | 4 |
| 2023 | In vivo test and rollback of Java applications as they areabstractSummary Modern software systems accommodate complex configurations and execution conditions that depend on the environment where the software is run. While in house testing can exercise only a fraction of such execution contexts, in vivo testing can take advantage of the execution state observed in the field to conduct further testing activities. In this paper, we present the Groucho approach to in vivo testing. Groucho can suspend the execution, run some in vivo tests, rollback the side effects introduced by such tests, and eventually resume normal execution. The approach can be transparently applied to the original application, even if only available as compiled code, and it is fully automated. Our empirical studies of the performance overhead introduced by Groucho under various configurations showed that this may be kept to a negligible level by activating in vivo testing with low probability. Our empirical studies about the effectiveness of the approach confirm previous findings on the existence of faults that are unlikely exposed in house and become easy to expose in the field. Moreover, we include the first study to quantify the coverage increase gained when in vivo testing is added to complement in house testing. Antonia Bertolino, Guglielmo De Angelis, Breno Miranda, Paolo Tonella |
Softw. Test. Verification Reliab. | 3 |
| 2023 | Test Flakiness Across Programming LanguagesabstractRegression Testing (RT) is a quality-assurance practice commonly adopted in the software industry to check if functionality remains intact after code changes. Test flakiness is a serious problem for RT. A test is said to be flaky when it non-deterministically passes or fails on a fixed environment. Prior work studied test flakiness primarily on Java programs. It is unclear, however, how problematic is test flakiness for software written in other programming languages. This paper reports on a study focusing on three central aspects of test flakiness: concentration, similarity, and cost. Considering concentration, our results show that, for any given programming language that we studied (C, Go, Java, JS, and Python), most issues could be explained by a small fraction of root causes (5/13 root causes cover 78.07% of the issues) and could be fixed by a relatively small fraction of fix strategies (10/23 fix strategies cover 85.20% of the issues). Considering similarity, although there were commonalities in root causes and fixes across languages (e.g., concurrency and async wait are common causes of flakiness in most languages), we also found important differences (e.g., flakiness due to improper release of resources are more common in C), suggesting that there is opportunity to fine tuning analysis tools. Considering cost, we found that issues related to flaky tests are resolved either very early once they are posted ($< $10 days), suggesting relevance, or very late ($>$100 days), suggesting irrelevance. Keila Barbosa Costa, Ronivaldo Ferreira, Gustavo Pinto 0001, Marcelo d'Amorim, Breno Miranda |
IEEE Trans. Software Eng. | 5 |
| 2022 | Testing non-testable programs using association rulesabstractWe propose a novel scalable approach for testing non-testable programs denoted as ARMED testing. The approach leverages efficient Association Rules Mining algorithms to determine relevant implication relations among features and actions observed while the system is in operation. These relations are used as the specification of positive and negative tests, allowing for identifying plausible or suspicious behaviors: for those cases when oracles are inherently unknownable, such as in social testing, ARMED testing introduces the novel concept of testing for plausibility. To illustrate the approach we walk-through an application example. Antonia Bertolino, Emilio Cruciani, Breno Miranda, Roberto Verdecchia |
AST | 3 |
| 2022 | Comparing and Combining File-based Selection and Similarity-based Prioritization towards Regression Test OrchestrationabstractTest case selection (TCS) and test case prioritization (TCP) techniques can reduce time to detect the first test failure. Although these techniques have been extensively studied in combination and isolation, they have not been compared one against the other. In this paper, we perform an empirical study directly comparing TCS and TCP approaches, represented by the tools Ekstazi and FAST, respectively. Furthermore, we develop the first combination, named Fastazi, of file-based TCS and similarity-based TCP and evaluate its benefit and cost against each individual technique. We performed our experiments using 12 Java-based open-source projects. Our results show that, in the median case, the combined approach detects the first failure nearly two times faster than either Ekstazi alone (with random test ordering) or FAST alone (without TCS). Statistical analysis shows that the effectiveness of Fastazi is higher than that of Ekstazi, which in turn is higher than that of FAST. On the other hand, FAST adds the least overhead to testing time, while the difference between the additional time needed by Ekstazi and Fastazi is negligible. Fastazi can also improve failure detection in scenarios where the time available for testing is restricted. Renan Greca, Breno Miranda, Milos Gligoric 0001, Antonia Bertolino |
AST | 2 |
| 2022 | A Systematic Mapping Study on Robotic Testing of Mobile DevicesabstractContext: Test automation is often seen as a possible solution to overcome the challenges of testing mobile devices. However, most of the automation techniques adopted for mobile testing are intrusive and, sometimes, unrealistic. One possible solution for coping with intrusive and unrealistic testing is the use of robots. Despite the growing interest in the intersection between robotics and software testing, the motivations, the usefulness, and the return of investment of adopting robots for supporting testing activities are not clear. Objective: We aim at surveying the literature on the use of robotics for supporting mobile testing with a focus on the motivations, the types of tests that are automated, and the reported effectiveness/efficiency. Method: We conduct a systematic mapping study on robotic testing of mobile devices (hereafter, referred as robotic mobile testing). We searched primary studies published since 2000 by querying five digital libraries, and by performing backward and forward snowballing cycles. Results: We started with a set of 1353 papers and after applying our study protocol, we selected a final set of 20 primary studies. We provide both a quantitative analysis, and a qualitative evaluation of the motivations, types of tests automated and the effectiveness/efficiency reported by the selected studies. Conclusions: Based on the selected studies, allowing more realistic interactions is among the main motivations for adopting robotic mobile testing. The tests automated with the support of robots are usually system-level tests targeting stress, interface, and performance testing. More empirical evidence is needed for supporting the claimed benefits. Most of the surveyed work do not compare the effectiveness and efficiency of the proposed robotics-based approach against traditional automation techniques. We discuss the implications of our findings for researchers and practitioners, and outline a research agenda. Lucas D. Maciel, Alice Oliveira, Riei Rodrigues, Williams Santiago, Andresa Silva, Gustavo Carvalho, Breno Miranda |
SEAA | 7 |
| 2022 | Investigating the Adoption of History-based Prioritization in the Context of Manual Testing in a Real Industrial SettingabstractMany test case prioritization techniques have been proposed with the ultimate goal of speeding up fault detection. History-based prioritization, in particular, has been shown to be an effective strategy. Most of the empirical studies conducted on this topic, however, have focused on the context of automated testing. Investigating the effectiveness of history-based prioritization in the context of manual testing is important because, despite the popularity of automated approaches, manual testing is still largely adopted in industry. In this work we propose two history-based prioritization heuristics and evaluate them in the context of manual testing in a real industrial setting. For our evaluation we collected historical test execution information for 23 products, spanning over seven years of historical information, accounting for a total of 2,352 unique test cases and 3,993,863 test results. The results of our experiments showed that the effectiveness of the proposed approach is not far from a theoretical optimal prioritization, and that they are significantly better than alternative orderings of the test suite, including the order suggested by the test management tool and the execution order followed by the testers during the real execution of the test suites evaluated as part of our study. Vinícius Siqueira, Breno Miranda |
SEAA | 2 |
| 2022 | Guest editors' introduction to the special issue "Automatic Software Testing from the Trenches"abstractSoftware testing is an integral and important part of the software engineering discipline, and its automation has been actively pursued in both academia and industry to reduce its high costs. In the past decades, a considerable research effort has been devoted to automatic test case generation, automatic test selection, and automatic test oracles. The practice of software test automation has also moved forward significantly, and in recent years, a large number of software test tools have been developed and released to the market. However, progress in automatic software testing (AST) research is still required. The increasing complexity, pervasiveness and inter-connection of software systems, the ever-shrinking development cycles and time-to-market, and the scarcity of tools that can support all testing tasks within one environment, have posed new challenges and stricter constraints. Thus, despite significant achievements both in theory and practice, AST remains a challenging research area, and there is an urgent requirement to improve test automation to scale up productivity and quality in software development. Many times, however, industry needs differ from the research agenda, as companies need to prioritize reducing cost and time-to-market. Moreover, practitioners may have a hard time choosing a particular testing method or technology, since the real challenges that influence the decision are usually hidden.\nThis special issue includes revised and extended versions of the best papers presented at the 2nd ACM/IEEE International Conference on Automation of Software Test (AST 2021), held in conjunction with the 43rd International Conference on Software Engineering (ICSE 2021), as well as new original submissions on the theme of “Automatic Software Testing from the Trenches.” This issue initially received a total of 13 submissions.\nBoth the extended papers from AST 2021 as well as the new original submissions underwent a rigorous review process, and ultimately, 8 submissions were accepted for inclusion in this special issue. Breno Miranda, Javier Tuya, Alejandra Garrido 0001 |
J. Softw. Evol. Process. | 1 |
| 2022 | RVprio: A tool for prioritizing runtime verification violationsabstractSummary Runtime verification (RV) helps to find software bugs by monitoring formally specified properties during testing. A key problem in using RV during testing is how to reduce the manual inspection effort for checking whether property violations are true bugs. To date, there was no automated approach for determining the likelihood that property violations were true bugs to reduce tedious and time‐consuming manual inspection. We present RVprio, the first automated approach for prioritizing RV violations in order of likelihood of being true bugs. RVprio uses machine learning classifiers to prioritize violations. For training, we used a labelled dataset of 1170 violations from 110 projects. On that dataset, (1) RVprio reached 90% of the effectiveness of a theoretically optimal prioritizer that ranks all true bugs at the top of the ranked list, and (2) 88.1% of true bugs were in the top 25% of RVprio‐ranked violations; 32.7% of true bugs were in the top 10%. RVprio was also effective when we applied it to new unlabelled violations, from which we found previously unknown bugs—54 bugs in 8 open‐source projects. Our dataset is publicly available online. Lucas Cabral 0002, Breno Miranda, Igor Lima, Marcelo d'Amorim |
Softw. Test. Verification Reliab. | 2 |
| 2021 | Demystifying the Challenges of Formally Specifying API Properties for Runtime VerificationabstractRuntime Verification (RV) is a technique to monitor formally-specified properties of the software during its execution. RV has shown to be very effective for bug finding. Unfortunately, RV typically relies on formal specification languages and learning those languages be costly for developers. This paper reports on a study to assess the challenges to specify API properties for the purpose of RV. To that end, we wrote SIESTA, a minimalist specification language, extending Java with two features (the ability to catch calls to specified methods and the ability to access the event history of a given object), and asked inexperienced developers (students) to write specifications in that language for certain parts of the Java API. Among our findings, we observed that 40% of the specifications written by the students matched the ground truth perfectly. The main messages of this work are that 1) it is feasible to use a simple imperative language for specifying properties without significant loss of generality; and that 2) developers are capable of writing specifications in the (programming) language they feel comfortable. Leopoldo Teixeira, Breno Miranda, Henrique Rebêlo, Marcelo d'Amorim |
ICST | 2 |
| 2021 | Shaker: a Tool for Detecting More Flaky Tests FasterabstractA test case that intermittently passes or fails when performed under the same version of source code and test code is said to be flaky. The presence of flaky tests wastes testing time and effort. The most popular approach in industry to detect flakiness is ReRun. The idea behind ReRun is very simple: failing test cases are re-executed many times looking for inconsistencies in the output. Despite its simplicity, the ReRun strategy is very expensive both in terms of time and in terms of computational resources. This is particularly true for contexts where thousands of test cases are performed on a daily basis. Reducing the rerunning overhead is, thus, of utmost importance. This paper presents SHAKER, an open-source tool for detecting flakiness in time-constrained tests by adding noise in the execution environment. The main idea behind SHAKER is to add stressing tasks that compete with the test execution for the use of resources (CPU or memory). SHAKER is available as a GitHub Actions workflow that can be seamlessly integrated with any GitHub project. Alternatively, SHAKER can also be used via its provided Command Line Interface. In our evaluation, SHAKER was able to discover more flaky tests than ReRun and in a faster way (less re-executions); besides, our approach revealed tens of new flaky tests that went undetected by ReRun even after 50 re-executions. Thanks to its flexibility and ease of use, we believe that SHAKER can be useful for both practitioners and researchers.Demo video: https://youtu.be/7-aiQwOb4rAShaker website: https://star-rg.github.io/shaker Marcello Cordeiro, Denini Silva, Leopoldo Teixeira, Breno Miranda, Marcelo d'Amorim |
ASE | 4 |
| 2021 | Exposing bugs in JavaScript engines through test transplantation and differential testing
Igor Lima, Jefferson Silva, Breno Miranda, Gustavo Pinto 0001, Marcelo d'Amorim |
Softw. Qual. J. | 3 |
| 2021 | Adaptive Test Case Allocation, Selection and Generation Using Coverage Spectrum and Operational ProfileabstractWe present an adaptive software testing strategy for test case allocation, selection and generation, based on the combined use of operational profile and coverage spectrum, aimed at achieving high delivered reliability of the program under test. Operational profile-based testing is a black-box technique considered well suited when reliability is a major concern, as it selects the test cases having the largest impact on failure probability in operation. Coverage spectrum is a characterization of a program's behavior in terms of the code entities (e.g., branches, statements, functions) that are covered as the program executes. The proposed strategy - named covrel+ - complements operational profile information with white-box coverage measures, so as to adaptively select/generate the most effective test cases for improving reliability as testing proceeds. We assess covrel+ through experiments with subjects commonly used in software testing research, comparing results with traditional operational testing. The results show that exploiting operational and coverage data in an integrated adaptive way allows generally to outperform operational testing at achieving a given reliability target, or at detecting faults under the same testing budget, and that covrel+ has greater ability than operational testing in detecting hard-to-detect faults. Antonia Bertolino, Breno Miranda, Roberto Pietrantuono, Stefano Russo 0001 |
IEEE Trans. Software Eng. | 2 |
| 2020 | Learning-to-rank vs ranking-to-learn: strategies for regression testing in continuous integrationabstractIn Continuous Integration (CI), regression testing is constrained by the time between commits. This demands for careful selection and/or prioritization of test cases within test suites too large to be run entirely. To this aim, some Machine Learning (ML) techniques have been proposed, as an alternative to deterministic approaches. Two broad strategies for ML-based prioritization are learning-to-rank and what we call ranking-to-learn (i.e., reinforcement learning). Various ML algorithms can be applied in each strategy. In this paper we introduce ten of such algorithms for adoption in CI practices, and perform a comprehensive study comparing them against each other using subjects from the Apache Commons project. We analyze the influence of several features of the code under test and of the test process. The results allow to draw criteria to support testers in selecting and tuning the technique that best fits their context. Antonia Bertolino, Antonio Guerriero, Breno Miranda, Roberto Pietrantuono, Stefano Russo 0001 |
ICSE | 3 |
| 2020 | Run Java Applications and Test Them In-Vivo MeantimeabstractThe outcome of test case execution depends on the state of the object under test. While testers can carefully choose meaningful and representative object states for test execution, it is unaffordable to cover the combinatorial space of possible object states exhaustively. An appealing option is to delegate part of the testing activities to the runtime and to execute test cases in the field whenever a new or uncommon state is observed. We have designed and developed Groucho, a framework for in-vivo testing of Java applications. Among the challenges that we faced, the most important ones are isolation of the test session from the user session and minimal performance overhead. Experimental results show that if the activation probability is kept reasonably small (e.g., $10 ^{- {4}}$), the impact of the framework is imperceptible(i.e., either statistically insignificant or with a negligible effect size). Antonia Bertolino, Guglielmo De Angelis, Breno Miranda, Paolo Tonella |
ICST | 3 |
| 2020 | Prioritizing Runtime Verification ViolationsabstractRuntime Verification (RV) can help find software bugs by monitoring formally specified properties during testing. A key problem when using RV during testing is how to reduce the manual inspection effort for checking whether property violations are true bugs. To date, there was no automated approach for determining the likelihood that property violations were true bugs to reduce tedious and time-consuming manual inspection.We present RVPRIO, the first automated approach for prioritizing RV violations in order of likelihood of being true bugs. RVPRIO uses machine learning classifiers to prioritize violations. For training, we used a labeled dataset of 1,170 violations from 110 projects. On that dataset, (1) RVPRIO reached 90% of the effectiveness of a theoretically optimal prioritizer that ranks all true bugs at the top of the ranked list, and (2) 88.1% of true bugs were in the top 25% of RVPRIO-ranked violations; 32.7% of true bugs were in the top 10%. RVPRIO was also effective when we applied it to new unlabeled violations, from which we found previously unknown bugs-29 bugs in 7 projects and two bugs in two properties. Our dataset is publicly available online. Breno Miranda, Igor Lima, Owolabi Legunsen, Marcelo d'Amorim |
ICST | 1 |
| 2020 | What is the Vocabulary of Flaky Tests?abstractFlaky tests are tests whose outcomes are non-deterministic. Despite the recent research activity on this topic, no effort has been made on understanding the vocabulary of flaky tests. This work proposes to automatically classify tests as flaky or not based on their vocabulary. Static classification of flaky tests is important, for example, to detect the introduction of flaky tests and to search for flaky tests after they are introduced in regression test suites. Gustavo Pinto 0001, Breno Miranda, Supun Dissanayake, Marcelo d'Amorim, Christoph Treude, Antonia Bertolino |
MSR | 2 |
| 2020 | JTeC: A Large Collection of Java Test Classes for Test Code Analysis and ProcessingabstractThe recent push towards test automation and test-driven development continues to scale up the dimensions of test code that needs to be maintained, analysed, and processed side-by-side with production code. As a consequence, on the one side regression testing techniques, e.g., for test suite prioritization or test case selection, capable to handle such large-scale test suites become indispensable; on the other side, as test code exposes own characteristics, specific techniques for its analysis and refactoring are actively sought. We present JTeC, a large-scale dataset of test cases that researchers can use for benchmarking the above techniques or any other type of tool expressly targeting test code. JTeC collects more than 2.5M test classes belonging to 31K+ GitHub projects and summing up to more than 430 Million SLOCs of ready-to-use real-world test code. Federico Coro, Roberto Verdecchia, Emilio Cruciani, Breno Miranda, Antonia Bertolino |
MSR | 4 |
| 2020 | Testing Relative to Usage Scope: Revisiting Software Coverage CriteriaabstractCoverage criteria provide a useful and widely used means to guide software testing; however, indiscriminately pursuing full coverage may not always be convenient or meaningful, as not all entities are of interest in any usage context. We aim at introducing a more meaningful notion of coverage that takes into account how the software is going to be used. Entities that are not going to be exercised by the user should not contribute to the coverage ratio. We revisit the definition of coverage measures, introducing a notion of relative coverage. According to this notion, we provide a definition and a theoretical framework of relative coverage, within which we discuss implications on testing theory and practice. Through the evaluation of three different instances of relative coverage, we could observe that relative coverage measures provide a more effective strategy than traditional ones: we could reach higher coverage measures, and test cases selected by relative coverage could achieve higher reliability. We hint at several other useful implications of relative coverage notion on different aspects of software testing. Breno Miranda, Antonia Bertolino |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2019 | Scalable approaches for test suite reductionabstractTest suite reduction approaches aim at decreasing software regression testing costs by selecting a representative subset from large-size test suites. Most existing techniques are too expensive for handling modern massive systems and moreover depend on artifacts, such as code coverage metrics or specification models, that are not commonly available at large scale. We present a family of novel very efficient approaches for similarity-based test suite reduction that apply algorithms borrowed from the big data domain together with smart heuristics for finding an evenly spread subset of test cases. The approaches are very general since they only use as input the test cases themselves (test source code or command line input). We evaluate four approaches in a version that selects a fixed budget B of test cases, and also in an adequate version that does the reduction guaranteeing some fixed coverage. The results show that the approaches yield a fault detection loss comparable to state-of-the-art techniques, while providing huge gains in terms of efficiency. When applied to a suite of more than 500K real world test cases, the most efficient of the four approaches could select B test cases (for varying B values) in less than 10 seconds. Emilio Cruciani, Breno Miranda, Roberto Verdecchia, Antonia Bertolino |
ICSE | 2 |
| 2018 | FAST approaches to scalable similarity-based test case prioritizationabstractMany test case prioritization criteria have been proposed for speeding up fault detection. Among them, similarity-based approaches give priority to the test cases that are the most dissimilar from those already selected. However, the proposed criteria do not scale up to handle the many thousands or even some millions test suite sizes of modern industrial systems and simple heuristics are used instead. We introduce the FAST family of test case prioritization techniques that radically changes this landscape by borrowing algorithms commonly exploited in the big data domain to find similar items. FAST techniques provide scalable similarity-based test case prioritization in both white-box and black-box fashion. The results from experimentation on real world C and Java subjects show that the fastest members of the family outperform other black-box approaches in efficiency with no significant impact on effectiveness, and also outperform white-box approaches, including greedy ones, if preparation time is not counted. A simulation study of scalability shows that one FAST technique can prioritize a million test cases in less than 20 minutes. Breno Miranda, Emilio Cruciani, Roberto Verdecchia, Antonia Bertolino |
ICSE | 1 |
| 2018 | A categorization scheme for software engineering conference papers and its application
Antonia Bertolino, Antonello Calabrò, Francesca Lonetti, Eda Marchetti, Breno Miranda |
J. Syst. Softw. | 5 |
| 2018 | An assessment of operational coverage as both an adequacy and a selection criterion for operational profile based testingabstractWhile the relation between code coverage measures and fault detection is actively studied, only few works have investigated the correlation between measures of coverage and of reliability. In this work, we introduce a novel approach to measuring code coverage, called the operational coverage, that takes into account how much the program’s entities are exercised so to reflect the profile of usage into the measure of coverage. Operational coverage is proposed as (i) an adequacy criterion, i.e., to assess the thoroughness of a black box test suite derived from the operational profile, and as (ii) a selection criterion, i.e., to select test cases for operational profile-based testing. Our empirical evaluation showed that operational coverage is better correlated than traditional coverage with the probability that the next test case derived according to the user’s profile will not fail. This result suggests that our approach could provide a good stopping rule for operational profile-based testing. With respect to test case selection, our investigations revealed that operational coverage outperformed the traditional one in terms of test suite size and fault detection capability when we look at the average results. Breno Miranda, Antonia Bertolino |
Softw. Qual. J. | 1 |
| 2017 | Adaptive coverage and operational profile-based testing for reliability improvementabstractWe introduce covrel, an adaptive software testing approach based on the combined use of operational profile and coverage spectrum, with the ultimate goal of improving the delivered reliability of the program under test. Operational profile-based testing is a black-box technique that selects test cases having the largest impact on failure probability in operation, as such, it is considered well suited when reliability is a major concern. Program spectrum is a characterization of a program's behavior in terms of the code entities (e.g., branches, statements, functions) that are covered as the program executes. The driving idea of covrel is to complement operational profile information with white-box coverage measures based on count spectra, so as to dynamically select the most effective test cases for reliability improvement. In particular, we bias operational profile-based test selection towards those entities covered less frequently. We assess the approach by experiments with 18 versions from 4 subjects commonly used in software testing research, comparing results with traditional operational and coverage testing. Results show that exploiting operational and coverage data in a combined adaptive way actually pays in terms of reliability improvement, with covrel overcoming conventional operational testing in more than 80% of the cases. Antonia Bertolino, Breno Miranda, Roberto Pietrantuono, Stefano Russo 0001 |
ICSE | 2 |
| 2017 | Towards Automated Deployment of Self-adaptive Applications on Hybrid Clouds (Short Paper)
Lom-Messan Hillah, Rodrigo Elia Assad, Antonia Bertolino, Márcio Eduardo Delamaro, Fabio De Rosa, Vinicius Cardoso Garcia, Francesca Lonetti, Ariele-Paolo Maesano, Libero Maesano, Eda Marchetti, Breno Miranda, Auri M. R. Vincenzi, Juliano Iyoda |
SEFM | 11 |
| 2017 | Scope-aided test prioritization, selection and minimization for software reuse
Breno Miranda, Antonia Bertolino |
J. Syst. Softw. | 1 |
| 2014 | A proposal for revisiting coverage testing metricsabstractTest coverage information can be very useful for guiding testers in enhancing their test suites to exercise possible uncovered entities and in deciding when to stop testing. Since the concept of test criterion was born, several contributions have been made by both academia and industry in the definition and adaptation of adequacy criteria aiming at ensuring the discovery of more failures. Numerous contributions have also been done in the development of coverage tools. However, for complex applications that are reused in different contexts and for emerging paradigms (e.g., component-based development, service-oriented architecture, and cloud computing), traditional coverage metrics may no longer provide meaningful information to help testers on these tasks. Inspired by the idea of relative coverage this research focuses on the introduction of meaningful coverage metrics to cope with the challenges imposed by the current programming paradigms as well as on the definition of a theoretical framework for the development of relative coverage metrics. Breno Miranda |
ASE | 1 |
| 2012 | Recommender systems for manual testing: deciding how to assign tests in a test teamabstractBACKGROUND: Software testing can be an arduous and expensive activity. A typical activity to maximise testing productivity is to allocate test cases according to the testers' profile. However, optimising the allocation of manual test cases is not a trivial task: in big companies, test managers are responsible for allocating hundreds of test cases among several testers. OBJECTIVE: In this paper we propose and evaluate 2 assignment algorithms for test case allocation and 3 tester profiles based on recommender systems. Each assignment algorithm can be combined with 3 tester profiles, which results in six possible allocation systems. METHOD: We run a controlled experiment that uses 100 test suites, each one with at least 50 test cases, from a real industrial setting in order to compare our allocation systems to the manager's allocation in terms of precision, recall and unassignment (percentage of test cases the algorithm could not allocate). RESULTS: In our experiment, the statistical analysis shows that one of the systems outperforms the others with respect to the precision and recall metrics. For unassignment, three of our six allocation systems achieved zero (best value) for the unassignment rate. CONCLUSION: The results of our experiment suggest that, in similar environments, test managers can use our allocation systems to reduce the amount of time spent in the test case allocation task. In the real industrial setting in which our work was developed, managers spend from 16 to 30 working days a year on test case allocation. Our algorithms can help them do it faster and better. Breno Miranda, Eduardo Aranha, Juliano Iyoda |
ESEM | 1 |