EDBT 2026 Demo / reviewers in the wild / expert
Fitsum Meshesha Kifetew
dblp:118/6464
· DBLP profile ↗
35ranked-venue papers
12as first author
15since 2021 · last 2026
0000-0003-1860-8666ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 34 · 12 first-author · 15 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AugmenTest: A tool for improving test quality through automatic assertion generation
Shaker Khandaker, Fitsum Meshesha Kifetew, Davide Prandi, Angelo Susi |
Sci. Comput. Program. | 2 |
| 2026 | Programming Smart PlaytestingabstractUntil recently the game industry heavily relied on manual playtesting to test the games it produces. Even if the benefits of introducing automated testing are acknowledged, it is rarely done in practice. Some of the main hurdles include the lack of automated testing tools that can target computer games as well as the complexity of automated game plays which are much more difficult to program than typical simple test sequences. This article presents an agent-based testing framework called aplib that comes with a Domain Specific Language (DSL) that allows complex playtests to be programmed more abstractly. A so-called goal structure is used to abstractly formulate a playtest scenario in terms of main goals and their decomposition into subgoals. Scenarios that are not too complicated can be formulated using static goal structures. More complex scenarios may need a test agent that can dynamically adapt its play according to the situation that evolves during the play. To handle such cases, aplib allows dynamic goals to be expressed as well. Invariants and pre-/post-conditions are used to assert the properties that a play is expected to satisfy. They include differential properties that allow constraints on the current state to be related to that of past states. Three case studies are included in the article. The first one aims to evaluate the performance of playtests programmed with aplib . The second shows that the approach can also be combined with other automated testing approaches, in this case reinforcement learning. The third shows the applicability of such playtests in a 3D setup and for non-functional testing. I. S. W. B. Prasetya, Mehdi Dastani, Rui Prada, Tanja E. J. Vos, Frank Dignum, Fitsum Meshesha Kifetew, Guido Mintjes, Samira Shirzadehhajimahmood, Saba Gholizadeh Ansari |
ACM Trans. Softw. Eng. Methodol. | 6 |
| 2025 | AugmenTest: Enhancing Tests with LLM-Driven OraclesabstractAutomated test generation is crucial for ensuring the reliability and robustness of software applications while at the same time reducing the effort needed. While significant progress has been made in test generation research, generating valid test oracles still remains an open problem. To address this challenge, we present AugmenTest, an approach leveraging Large Language Models (LLMs) to infer correct test oracles based on available documentation of the software under test. Unlike most existing methods that rely on code, AugmenTest utilizes the semantic capabilities of LLMs to infer the intended behavior of a method from documentation and developer comments, without looking at the code. AugmenTest includes four variants: Simple Prompt, Extended Prompt, RAG with a generic prompt (without the context of class or method under test), and RAG with Simple Prompt, each offering different levels of contextual information to the LLMs. To evaluate our work, we selected 142 Java classes and generated multiple mutants for each. We then generated tests from these mutants, focusing only on tests that passed on the mutant but failed on the original class, to ensure that the tests effectively captured bugs. This resulted in 203 unique tests with distinct bugs, which were then used to evaluate AugmenTest. Results show that in the most conservative scenario, AugmenTest's Extended Prompt consistently outperformed the Simple Prompt, achieving a success rate of 30% for generating correct assertions. In comparison, the state-of-the-art TOGA approach achieved 8.2%. Contrary to our expectations, the RAG-based approaches did not lead to improvements, with performance of 18.2% success rate for the most conservative scenario. Our study demonstrates the potential of LLMs in improving the reliability of automated test generation tools, while also highlighting areas for future enhancement. Shaker Khandaker, Fitsum Meshesha Kifetew, Davide Prandi, Angelo Susi |
ICST | 2 |
| 2025 | On the Energy Consumption of Test GenerationabstractResearch in the area of automated test generation has seen remarkable progress in recent years, resulting in several approaches and tools for effective and efficient generation of test cases. In particular, the EvoSuite tool has been at the forefront of this progress embodying various algorithms for automated test generation of Java programs. EvoSuite has been used to generate test cases for a wide variety of programs as well. While there are a number of empirical studies that report results on the effectiveness, in terms of code coverage and other related metrics, of the various test generation strategies and algorithms implemented in EvoSuite, there are no studies, to the best of our knowledge, on the energy consumption associated to the automated test generation. In this paper, we set out to investigate this aspect by measuring the energy consumed by EvoSuite when generating tests. We also measure the energy consumed in the execution of the test cases generated, comparing them with those manually written by developers. The results show that the different test generation algorithms consumed different amounts of energy, in particular on classes with high cyclomatic complexity. Furthermore, we also observe that manual tests tend to consume more energy as compared to automatically generated tests, without necessarily achieving higher code coverage. Our results also give insight into the methods that consume significantly higher levels of energy, indicating potential points of improvement both for EvoSuite as well as the different programs under test. Fitsum Meshesha Kifetew, Davide Prandi, Angelo Susi |
ICST | 1 |
| 2025 | Evolv-1 at the ICST 2025 Tool Competition - UAV Testing TrackabstractEvolv-1 is a test case generation tool designed using Evolutionary Algorithms (EAs) to optimize UAV testing scenarios. This short paper presents Evolv-1's implementation as part of the ICST 2025 UAV Testing Tool Competition. Pietro Lechthaler, Davide Prandi, Fitsum Meshesha Kifetew, Angelo Susi |
ICST | 3 |
| 2024 | Model-Based Testing of Railway Interlocking Systems
Alessandro Cimatti, Shaker Khandaker, Fitsum Meshesha Kifetew, Lorenzo Leone, Davide Prandi, Giuseppe Scaglione, Angelo Susi, Orazio Turboli |
ISoLA (5) | 3 |
| 2024 | PX-MBT: A framework for model-based player experience testingabstractAs video games become more complex and widespread, player experience (PX) testing becomes crucial in the game industry. Attracting and retaining players are key elements to guarantee the success of a game in the highly competitive market. Although a number of techniques have been introduced to measure the emotional aspect of the experience, automated testing of player experience still needs to be explored. This paper presents PX-MBT, a framework for automated player experience testing with emotion pattern verification. PX-MBT (1) utilizes a model-based testing approach for test suite generation, (2) employs a computational model of emotions developed based on a psychological theory of emotions to model players' emotions during game-plays with an intelligent agent, and (3) verifies emotion patterns given by game designers on executed test suites to identify PX-issues. We explain PX-MBT architecture and provide an example along with its result in emotion pattern verification, which asserts the evolution of emotions over time, and heat-maps to showcase the spatial distribution of emotions on the game map. Saba Gholizadeh Ansari, I. S. W. B. Prasetya, Mehdi Dastani, Gabriele Keller, Davide Prandi, Fitsum Meshesha Kifetew, Frank Dignum |
Sci. Comput. Program. | 6 |
| 2023 | Model-based Player Experience Testing with Emotion Pattern VerificationabstractAbstract Player eXperience (PX) testing has attracted attention in the game industry as video games become more complex and widespread. Understanding players’ desires and their experience are key elements to guarantee the success of a game in the highly competitive market. Although a number of techniques have been introduced to measure the emotional aspect of the experience, automated testing of player experience still needs to be explored. This paper presents a framework for automated player experience testing by formulating emotion patterns’ requirements and utilizing a computational model of players’ emotions developed based on a psychological theory of emotions along with a model-based testing approach for test suite generation. We evaluate the strength of our framework by performing mutation test. The paper also evaluates the performance of a search-based generated test suite and LTL model checking-based test suite in revealing various variations of temporal and spatial emotion patterns. Results show the contribution of both algorithms in generating complementary test cases for revealing various emotions in different locations of a game level. Saba Gholizadeh Ansari, I. S. W. B. Prasetya, Davide Prandi, Fitsum Meshesha Kifetew, Mehdi Dastani, Frank Dignum, Gabriele Keller |
FASE | 4 |
| 2023 | EvoMBT: Evolutionary model based testing
Raihana Ferdous, Chia-kang Hung, Fitsum Meshesha Kifetew, Davide Prandi, Angelo Susi |
Sci. Comput. Program. | 3 |
| 2023 | JUGE: An infrastructure for benchmarking Java unit test generatorsabstractSummary Researchers and practitioners have designed and implemented various automated test case generators to support effective software testing. Such generators exist for various languages (e.g., Java, C#, or Python) and various platforms (e.g., desktop, web, or mobile applications). The generators exhibit varying effectiveness and efficiency, depending on the testing goals they aim to satisfy (e.g., unit‐testing of libraries versus system‐testing of entire applications) and the underlying techniques they implement. In this context, practitioners need to be able to compare different generators to identify the most suited one for their requirements, while researchers seek to identify future research directions. This can be achieved by systematically executing large‐scale evaluations of different generators. However, executing such empirical evaluations is not trivial and requires substantial effort to select appropriate benchmarks, setup the evaluation infrastructure, and collect and analyse the results. In this Software Note, we present ourJUnit Generation Benchmarking Infrastructure(JUGE) supporting generators (search‐based, random‐based, symbolic execution, etc.) seeking to automate the production of unit tests for various purposes (validation, regression testing, fault localization, etc.). The primary goal is to reduce the overall benchmarking effort, ease the comparison of several generators, and enhance the knowledge transfer between academia and industry by standardizing the evaluation and comparison process. Since 2013, several editions of a unit testing tool competition, co‐located with the Search‐Based Software Testing Workshop, have taken place whereJUGEwas used and evolved. As a result, an increasing amount of tools (over 10) from academia and industry have been evaluated onJUGE, matured over the years, and allowed the identification of future research directions. Based on the experience gained from the competitions, we discuss the expected impact ofJUGEin improving the knowledge transfer on tools and approaches for test generation between academia and industry. Indeed, theJUGEinfrastructure demonstrated an implementation design that is flexible enough to enable the integration of additional unit test generation tools, which is practical for developers and allows researchers to experiment with new and advanced unit testing tools and approaches. Xavier Devroey, Alessio Gambi, Juan P. Galeotti, René Just, Fitsum Meshesha Kifetew, Annibale Panichella, Sebastiano Panichella |
Softw. Test. Verification Reliab. | 5 |
| 2022 | Message from the Program Co-ChairsabstractWelcome to the proceedings of the 15th International Conference on Software Testing, Verification and Validation (ICST 2022). The conference aims to provide a common forum for researchers, scientists, engineers, and practitioners throughout the world to present their latest research findings, ideas, developments, and applications in the area of Software Testing, Verification, and Validation. Fitsum Meshesha Kifetew, Annibale Panichella |
ICST | 1 |
| 2022 | Towards Agent-Based Testing of 3D Games using Reinforcement LearningabstractComputer game is a billion-dollar industry and is booming. Testing games has been recognized as a difficult task, which mainly relies on manual playing and scripting based testing. With the advances in technologies, computer games have become increasingly more interactive and complex, thus play-testing using human participants alone has become unfeasible. In recent days, play-testing of games via autonomous agents has shown great promise by accelerating and simplifying this process. Reinforcement Learning solutions have the potential of complementing current scripted and automated solutions by learning directly from playing the game without the need of human intervention. This paper presented an approach based on reinforcement learning for automated testing of 3D games. We make use of the notion of curiosity as a motivating factor to encourage an RL agent to explore its environment. The results from our exploratory study are promising and we have preliminary evidence that reinforcement learning can be adopted for automated testing of 3D games. Raihana Ferdous, Fitsum Meshesha Kifetew, Davide Prandi, Angelo Susi |
ASE | 2 |
| 2021 | Combining risk and variability modelling for requirements analysis in SAS engineeringabstractResearch on self-adaptive systems (SASs) has proliferated in the last fifteen years. Approaches resting on models at run-time have been proposed (e.g., to model system variants), as well as methods that aim at giving requirements a key role in driving the adaptation process (e.g., to choose the most appropriate system variant). More recent research focuses on automating model-based decisions, such as requirements revision, by exploiting data generated at execution time.Uncertainty is considered a first-class citizen in SAS engineering. A well recognised technique for dealing with uncertainty is risk management. Several risk management methods exist, as well as visual modelling languages that aim at supporting risk analysis.Our objective is to investigate how complementing requirements modelling with risk modelling could support automating risk-driven requirements analysis. While risk could be identified and modelled at design-time using domain knowledge and data generated by previous system executions, their estimation will be done at run-time, and guide the selection of system behaviour that minimises the risk of the system not being compliant with requirements.In this paper, we introduce our research objective that concerns the definition of an engineering framework, called Risk4SAS, that enables risk-driven requirements analysis in SASs life-cycle and describe first steps towards its realisation, including a meta-model, which captures the dependency between risk and the characteristics of a SAS’s variants. We conclude by presenting our research road-map. Denisse Muñante Arzapalo, Anna Perini, Fitsum Meshesha Kifetew, Angelo Susi |
RE | 3 |
| 2021 | Search-Based Automated Play Testing of Computer Games: A Model-Based Approach
Raihana Ferdous, Fitsum Meshesha Kifetew, Davide Prandi, I. S. W. B. Prasetya, Samira Shirzadehhajimahmood, Angelo Susi |
SSBSE | 2 |
| 2021 | Automating user-feedback driven requirements prioritization
Fitsum Meshesha Kifetew, Anna Perini, Angelo Susi, Alberto Siena, Denisse Muñante Arzapalo, Itzel Morales-Ramirez |
Inf. Softw. Technol. | 1 |
| 2020 | A Framework for In-Vivo Testing of Mobile ApplicationsabstractThe ecosystem in which mobile applications run is highly heterogeneous and configurable. All layers upon which mobile apps are built offer wide possibilities of variations, from the device and the hardware, to the operating system and middleware, up to the user preferences and settings. Testing all possible configurations exhaustively, before releasing the app, is unaffordable. As a consequence, the app may exhibit different, including faulty, behaviours when executed in the field, under specific configurations.In this paper, we describe a framework that can be instantiated to support in-vivo testing of a mobile app. The framework monitors the configuration in the field and triggers in-vivo testing when an untested configuration is recognized. Experimental results show that the overhead introduced by monitoring is unnoticeable to negligible (i.e., 0-6%) depending on the device being used (high- vs. low-end). In-vivo test execution required on average 3s: if performed upon screen lock activation, it introduces just a slight delay before locking the device. Mariano Ceccato, Davide Corradini, Luca Gazzola, Fitsum Meshesha Kifetew, Leonardo Mariani, Matteo Orrù, Paolo Tonella |
ICST | 4 |
| 2020 | Agent-based Testing of Extended Reality SystemsabstractTesting for quality assurance (QA) is a crucial step in the development of Extended Reality (XR) systems that typically follow iterative design and development cycles. Bringing automation to these testing procedures will increase the productivity of XR developers. However, given the complexity of the XR environments and the User Experience (UX) demands, achieving this is highly challenging. We propose to address this issue through the creation of autonomous cognitive test agents that will have the ability to cope with the complexity of the interaction space by intelligently explore the most prominent interactions given a test goal and support the assessment of affective properties of the UX by playing the role of users. Rui Prada, I. S. W. B. Prasetya, Fitsum Meshesha Kifetew, Frank Dignum, Tanja E. J. Vos, Jason Lander, Jean-Yves Donnart, Alexandre Kazmierowski, Joseph Davidson, Pedro M. Fernandes |
ICST | 3 |
| 2019 | Speech-acts based analysis for requirements discovery from online discussions
Itzel Morales-Ramirez, Fitsum Meshesha Kifetew, Anna Perini |
Inf. Syst. | 2 |
| 2018 | Incremental Control Dependency Frontier Exploration for Many-Criteria Test Case GenerationabstractSeveral criteria have been proposed over the years for measuring test suite adequacy. Each criterion can be converted into a specific objective function to optimize with search-based techniques in an attempt to generate test suites achieving the highest possible coverage for that criterion. Recent work has tried to optimize for multiple-criteria at once by constructing a single objective function obtained as a weighted sum of the objective functions of the respective criteria. However, this solution suffers the problem of sum scalarization, i.e., differences along the various dimensions being optimized get lost when such dimensions are projected into a single value. Recent advances in SBST formulated coverage as a many-objective optimization problem rather than applying sum scalarization. Starting from this formulation, in this work, we apply many-objective test generation that handles multiple adequacy criteria simultaneously. To scale the approach to the big number of objectives to be optimized at the same time, we adopt an incremental strategy, where only coverage targets in the control dependency frontier are considered until the frontier is expanded by covering a previously uncovered target. Annibale Panichella, Fitsum Meshesha Kifetew, Paolo Tonella |
SSBSE | 2 |
| 2018 | A large scale empirical comparison of state-of-the-art search-based test case generators
Annibale Panichella, Fitsum Meshesha Kifetew, Paolo Tonella |
Inf. Softw. Technol. | 2 |
| 2018 | Automated Test Case Generation as a Many-Objective Optimisation Problem with Dynamic Selection of the TargetsabstractThe test case generation is intrinsically a multi-objective problem, since the goal is covering multiple test targets (e.g., branches). Existing search-based approaches either consider one target at a time or aggregate all targets into a single fitness function (whole-suite approach). Multi and many-objective optimisation algorithms (MOAs) have never been applied to this problem, because existing algorithms do not scale to the number of coverage objectives that are typically found in real-world software. In addition, the final goal for MOAs is to find alternative trade-off solutions in the objective space, while in test generation the interesting solutions are only those test cases covering one or more uncovered targets. In this paper, we present Dynamic Many-Objective Sorting Algorithm (DynaMOSA), a novel many-objective solver specifically designed to address the test case generation problem in the context of coverage testing. DynaMOSA extends our previous many-objective technique Many-Objective Sorting Algorithm (MOSA) with dynamic selection of the coverage targets based on the control dependency hierarchy. Such extension makes the approach more effective and efficient in case of limited search budget. We carried out an empirical study on 346 Java classes using three coverage criteria (i.e., statement, branch, and strong mutation coverage) to assess the performance of DynaMOSA with respect to the whole-suite approach (WS), its archive-based variant (WSA) and MOSA. The results show that DynaMOSA outperforms WSA in 28 percent of the classes for branch coverage (+8 percent more coverage on average) and in 27 percent of the classes for mutation coverage (+11 percent more killed mutants on average). It outperforms WS in 51 percent of the classes for statement coverage, leading to +11 percent more coverage on average. Moreover, DynaMOSA outperforms its predecessor MOSA for all the three coverage criteria in 19 percent of the classes with +8 percent more code coverage on average. Annibale Panichella, Fitsum Meshesha Kifetew, Paolo Tonella |
IEEE Trans. Software Eng. | 2 |
| 2017 | Analysis of Online Discussions in Support of Requirements Discovery
Itzel Morales-Ramirez, Fitsum Meshesha Kifetew, Anna Perini |
CAiSE | 2 |
| 2017 | Tool-Supported Collaborative Requirements PrioritisationabstractAutomated decision-making techniques are useful to support engineers when performing requirements engineering tasks. However, to be effectively used in practice they need to be integrated into the organisational context, in which stakeholder engagement becomes a critical adoption factor. In this paper, we propose a tool-supported collaborative requirements prioritisation process, called GRP, which exploits gamification elements to engage distributed stakeholders to contribute to the overall decision-making process. Analytic Hierarchy Process is used as key component of the game engine, and enables an iterative prioritisation process. The GRP process has been evaluated through an exploratory case study, which has been conducted at a small software company, providing us with preliminary evidence about the effectiveness of the proposed solution. The main findings and lessons learned from the case study are presented. Paolo Busetta, Fitsum Meshesha Kifetew, Denisse Muñante Arzapalo, Anna Perini, Alberto Siena, Angelo Susi |
COMPSAC (1) | 2 |
| 2017 | DMGame: A Gamified Collaborative Requirements Prioritisation ToolabstractAutomated decision-making techniques have been proposed to support engineers in selecting and prioritising requirements. However, to be effectively used in practice they need to be integrated into the organisational context, and their users, namely the members of the development team, and more generally the project's stakeholders, need to be engaged in the resulting tool-supported decision-making process. In this demo paper, we present a tool-supported collaborative requirements prioritisation process, which exploits game elements to engage distributed stakeholders to contribute to the overall decision-making process. AHP and Genetic Algorithms are used as key component of the game engine, which enables an iterative prioritisation process. The tool is part of the tool-suite developed in the SUPERSEDE project which aims at supporting a flexible feedback-anddata-driven software evolution approach. Fitsum Meshesha Kifetew, Denisse Muñante Arzapalo, Anna Perini, Angelo Susi, Alberto Siena, Paolo Busetta |
RE | 1 |
| 2017 | Gamifying Collaborative Prioritization: Does Pointsification Work?abstractGamification has been applied in software engineering contexts, and more recently in requirements engineering with the purpose of improving the motivation and engagement of people performing specific engineering tasks. But often an objective evaluation that the resulting gamified tasks successfully meet the intended goal is missing. On the other hand, current practices in designing gamified processes seem to rest on a try, test and learn approach, rather than on first principles design methods. Thus empirical evaluation should play an even more important role.We combined gamification and automated reasoning techniques to support collaborative requirements prioritization in software evolution. A first prototype has been evaluated in the context of three industrial use cases. To further investigate the impact of specific game elements, namely point-based elements, we performed a quasi-experiment comparing two versions of the tool, with and without pointsification. We present the results from these two empirical evaluations, and discuss lessons learned. Fitsum Meshesha Kifetew, Denisse Muñante Arzapalo, Anna Perini, Angelo Susi, Alberto Siena, Paolo Busetta, Danilo Valerio |
RE | 1 |
| 2017 | Exploiting User Feedback in Tool-Supported Multi-criteria Requirements PrioritizationabstractAs different types of user feedback are becoming available, from a variety of sources and in large amount, several analysis techniques have been developed with the purpose of extracting information that can be useful for requirements engineering purposes. For instance, automated extraction and prioritization of feature requests have been recently investigated for the specific case of app development, where the key prioritization criterion is value for the user. For other types of software applications and services, software evolution relies on multi-criteria requirements prioritization, which may take into account different stakeholders' perspectives, thus leading to a complex decision-making problem. Different automated reasoning techniques have been proposed to support multi-criteria requirements prioritization, aimed at reducing human effort and improving the quality of the resulting ranking of the candidate requirements.The goal of our research is to understand how we can exploit user feedback in tool-supported multi-criteria requirements prioritization processes. Towards this objective, we discuss the properties of user feedback which are relevant for requirements prioritization, formulate a multi-criteria requirements prioritization problem, and outline a possible solution that integrates state of the art automated reasoning techniques which we extend to cope with information derived from user feedback. Itzel Morales-Ramirez, Denisse Muñante Arzapalo, Fitsum Meshesha Kifetew, Anna Perini, Angelo Susi, Alberto Siena |
RE | 3 |
| 2017 | Grammar Based Genetic Programming for Software Configuration Problem
Fitsum Meshesha Kifetew, Denisse Muñante Arzapalo, Jesús Gorroñogoitia, Alberto Siena, Angelo Susi, Anna Perini |
SSBSE | 1 |
| 2017 | LIPS vs MOSA: A Replicated Empirical Study on Automated Test Case Generation
Annibale Panichella, Fitsum Meshesha Kifetew, Paolo Tonella |
SSBSE | 2 |
| 2017 | Generating valid grammar-based test inputs by means of genetic programming and annotated grammars
Fitsum Meshesha Kifetew, Roberto Tiella, Paolo Tonella |
Empir. Softw. Eng. | 1 |
| 2015 | Reformulating Branch Coverage as a Many-Objective Optimization ProblemabstractTest data generation has been extensively investigated as a search problem, where the search goal is to maximize the number of covered program elements (e.g., branches). Recently, the whole suite approach, which combines the fitness functions of single branches into an aggregate, test suite-level fitness, has been demonstrated to be superior to the traditional single-branch at a time approach. In this paper, we propose to consider branch coverage directly as a many-objective optimization problem, instead of aggregating multiple objectives into a single value, as in the whole suite approach. Since programs may have hundreds of branches (objectives), traditional many-objective algorithms that are designed for numerical optimization problems with less than 15 objectives are not applicable. Hence, we introduce a novel highly scalable many-objective genetic algorithm, called MOSA (Many-Objective Sorting Algorithm), suitably defined for the many- objective branch coverage problem. Results achieved on 64 Java classes indicate that the proposed many-objective algorithm is significantly more effective and more efficient than the whole suite approach. In particular, effectiveness (coverage) was significantly improved in 66% of the subjects and efficiency (search budget consumed) was improved in 62% of the subjects on which effectiveness remains the same. Annibale Panichella, Fitsum Meshesha Kifetew, Paolo Tonella |
ICST | 2 |
| 2014 | Reproducing Field Failures for Programs with Complex Grammar-Based InputabstractTo isolate and fix failures that occur in the field, after deployment, developers must be able to reproduce and investigate such failures in-house. In practice, however, bug reports rarely provide enough information to recreate field failures, thus making in-house debugging an arduous task. This task becomes even more challenging for programs whose input must adhere to a formal specification, such as a grammar. To help developers address this issue, we propose an approach for automatically generating inputs that recreate field failures in-house. Given a faulty program and a field failure for this program, our approach exploits the potential of grammar-guided genetic programming to iteratively find legal inputs that can trigger the observed failure using a limited amount of runtime data collected in the field. When applied to 11 failures of 5 real-world programs, our approach was able to reproduce all but one of the failures while imposing a limited amount of overhead. Fitsum Meshesha Kifetew, Wei Jin 0001, Roberto Tiella, Alessandro Orso, Paolo Tonella |
ICST | 1 |
| 2014 | Combining Stochastic Grammars and Genetic Programming for Coverage Testing at the System Level
Fitsum Meshesha Kifetew, Roberto Tiella, Paolo Tonella |
SSBSE | 1 |
| 2013 | Orthogonal exploration of the search space in evolutionary test case generationabstractThe effectiveness of evolutionary test case generation based on Genetic Algorithms (GAs) can be seriously impacted by genetic drift, a phenomenon that inhibits the ability of such algorithms to effectively diversify the search and look for alternative potential solutions. In such cases, the search becomes dominated by a small set of similar individuals that lead GAs to converge to a sub-optimal solution and to stagnate, without reaching the desired objective. This problem is particularly common for hard-to-cover program branches, associated with an extremely large solution space. In this paper, we propose an approach to solve this problem by integrating a mechanism for orthogonal exploration of the search space into standard GA. The diversity in the population is enriched by adding individuals in orthogonal directions, hence providing a more effective exploration of the solution space. To the best of our knowledge, no prior work has addressed explicitly the issue of evolution direction based diversification in the context of evolutionary testing. Results achieved on 17 Java classes indicate that the proposed enhancements make GA much more effective and efficient in automating the testing process. In particular, effectiveness (coverage) was significantly improved in 47% of the subjects and efficiency (search budget consumed) was improved in 85% of the subjects on which effectiveness remains the same. Fitsum Meshesha Kifetew, Annibale Panichella, Andrea De Lucia, Rocco Oliveto, Paolo Tonella |
ISSTA | 1 |
| 2013 | SBFR: A search based approach for reproducing failures of programs with grammar based inputabstractReproducing field failures in-house, a step developers must perform when assigned a bug report, is an arduous task. In most cases, developers must be able to reproduce a reported failure using only a stack trace and/or some informal description of the failure. The problem becomes even harder for the large class of programs whose input is highly structured and strictly specified by a grammar. To address this problem, we present SBFR, a search-based failure-reproduction technique for programs with structured input. SBFR formulates failure reproduction as a search problem. Starting from a reported failure and a limited amount of dynamic information about the failure, SBFR exploits the potential of genetic programming to iteratively find legal inputs that can trigger the failure. Fitsum Meshesha Kifetew, Wei Jin 0001, Roberto Tiella, Alessandro Orso, Paolo Tonella |
ASE | 1 |
| 2012 | A Search-Based Framework for Failure Reproduction
Fitsum Meshesha Kifetew |
SSBSE | 1 |