VLDB 2026 Research / reviewers in the wild / expert
Alessio Gambi
dblp:50/5924
· DBLP profile ↗
32ranked-venue papers
14as first author
18since 2021 · last 2027
0000-0002-0132-6497ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 27 · 13 first-author · 15 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Does road diversity really matter in testing automated driving systems?abstractAbstract Context The use of automated driving systems (ADSs) in the real world requires rigorous testing to ensure safety. To increase trust, ADSs should be tested on a large set of diverse road scenarios. Literature suggests that if a vehicle is driven along a set of geometrically diverse roads—measured using various diversity measures (DMs)—it will react in a wide range of behaviours, thereby increasing the chances of observing failures, or strengthening the confidence in its safety, if no failures are observed. However, this assumption has never been tested before, nor have road DMs been assessed for their properties. Objective Our goal was to perform an exploratory study on 53 currently used and new, potentially promising road DMs. Specifically, our research questions looked into the road DMs themselves, to analyse their properties (e.g. monotonicity , computation efficiency ), and to test correlation between DMs. Furthermore, we investigated the use of road DMs to determine whether the assumption that diverse test suites of roads expose diverse driving behaviour holds. Method Our empirical analysis relies on a state-of-the-art, open-source ADS testing infrastructure and uses a data set containing over 97,000 individual road geometries and matching simulation data that were collected using two driving agents. By considering test suites of various sizes and measuring their roads’ geometric diversity, we studied road DM properties, the correlation between road DMs, and the correlation between road DMs and the observed behaviour. Results Our findings reveal a strong correlation between road diversity and behavioural diversity, confirming that geometrically diverse test suites systematically exercise diverse driving behaviours. We identified and aggregations as most effective, with achieving the strongest correlation of 0.95 while requiring minimal computation time. The analysed measures maintain robust correlation with behavioural diversity across test suites containing roads of varying lengths, eliminating the need for length normalisation. Conclusions These results empirically validate the fundamental assumption underlying diversity-driven ADS testing: road geometry diversity serves as a reliable proxy for behavioural diversity. For practitioners, we recommend or as optimal choices, whilst -based measures should be avoided entirely. The near-identical correlation patterns observed across architecturally different driving agents indicate that our findings generalise beyond specific ADS implementations, providing a solid foundation for diversity-driven test generation and selection. Stefan Klikovits, Vincenzo Riccio, Ezequiel Castellano, Ahmet Cetinkaya, Alessio Gambi, Paolo Arcaini |
Empir. Softw. Eng. | 5 |
| 2026 | Efficient Exploration of Autonomous Driving System Safety Boundaries
Alves Marinov, Paolo Arcaini, Antony Bartlett, Alessio Gambi, Fuyuki Ishikawa, Annibale Panichella |
IV | 4 |
| 2025 | Taming Uncertainty in Critical Scenario Generation for Testing Automated Driving SystemsabstractScenario-based testing in simulation has become a cornerstone of industrial practice for systematically assessing autonomous driving systems across diverse and relevant situations. Generating critical scenarios is central to this methodology, yet it remains challenging due to the inherent uncertainties resulting from scenario parameterization. While parameterization is essential for modeling unpredictable factors, like weather, an excess of parameters hampers testing effectiveness. To address these challenges, this paper introduces a methodology that guides testers in selecting scenario parameters and managing the associated uncertainties. Our approach integrates specification-driven and optimization-based test generation with sensitivity analysis, enabling testers to assess the impact of scenario parameters on scenario criticality. We implemented our approach using well-established industry technologies and evaluated it in a highway case study on three reference search-based scenario generation methods with varying degrees of exploitativeness. Results from our evaluation suggest that reducing the parameter-induced uncertainty can improve the ability of some testing methods to identify critical scenarios while maintaining the diversity of input parameter values. Selma Grosse, Adam Molin, Dejan Nickovic, Alessio Gambi, Cristinel Mateis |
ICST | 4 |
| 2025 | Scenario-Based Testing with BeamNG.tech (Hands-On Training)abstractAutonomous Driving Systems (ADS) are safety-critical Cyber-Physical Systems that require thorough validation. Currently, scenario-based testing in simulations is the cornerstone of ADS validation, complementing expensive and dangerous natural field operational testing. Scenario-based testing in simulation can systematically assess ADS in diverse, relevant, and critical driving situations. However, it requires sophisticated tools and rich content to let developers quickly generate the intended testing scenarios. This hands-on training illustrates how to use the BeamNG.tech framework for effective scenario-based testing, focusing on manual scenario generation, which is often neglected in research despite its central role in practice. Learning material and additional descriptions are available at: https://github.com/BeamNG/scenario-based-testing-tutorial Chrysanthi Papamichail, David Stark, Alessio Gambi |
ICST | 3 |
| 2025 | Generation of Critical Interactive Scenarios for Trajectory PlanningabstractAutonomous Vehicles (AVs) must be thoroughly tested to meet high safety requirements. Scenario-based testing using simulation is a common approach for validating them. Usually, scenarios test only one ego vehicle against road users with pre-defined behaviors called Non-Playable Characters (NPC). Such scenarios ensure reproducibility but are not always relevant and realistic, as they do not capture interactions between (e.g., non-cooperative) AVs. Consequently, they are unsuitable for testing safety-critical emerging behaviors like those happening in the real world. To tackle this problem, we propose TIAV, an approach for generating interactive critical scenarios that allows developers to study how AVs influence each other. Experiments on the reference CommonRoad simulation framework show that TIAV can identify scenarios leading to collisions and disengagements and trigger significantly more failures than a random baseline. Thanks to its ability to expose unsafe AV interactions, TIAV allows developers to validate AVs' functional correctness and check the effects of AVs' simultaneous deployment. TIAV is available as open-source software: https://github.comJparcaini/TIAV Alessio Gambi, Paolo Arcaini, Dejan Nickovic |
IV | 1 |
| 2025 | Preface for the special issue on SBFT'23: Search-Based and Fuzz Testing - Tools
Alessio Gambi, Sebastiano Panichella |
Sci. Comput. Program. | 1 |
| 2024 | An Empirical Study on How Large Language Models Impact Software Testing LearningabstractSoftware testing is a challenging topic in software engineering education and requires creative approaches to engage learners. For example, the Code Defenders game has students compete over a Java class under test by writing effective tests and mutants. While such gamified approaches deal with problems of motivation and engagement, students may nevertheless require help to put testing concepts into practice. The recent widespread diffusion of Generative AI and Large Language Models raises the question of whether and how these disruptive technologies could address this problem, for example, by providing explanations of unclear topics and guidance for writing tests. However, such technologies might also be misused or produce inaccurate answers, which would negatively impact learning. To shed more light on this situation, we conducted the first empirical study investigating how students learn and practice new software testing concepts in the context of the Code Defenders testing game, supported by a smart assistant based on a widely known, commercial Large Language Model. Our study shows that students had unrealistic expectations about the smart assistant, “blindly” trusting any output it generated, and often trying to use it to obtain solutions for testing exercises directly. Consequently, students who resorted to the smart assistant more often were less effective and efficient than those who did not. For instance, they wrote 8.6% fewer tests, and their tests were not useful in 78.0% of the cases. We conclude that giving unrestricted and unguided access to Large Language Models might generally impair learning. Thus, we believe our study helps to raise awareness about the implications of using Generative AI and Large Language Models in Computer Science Education and provides guidance towards developing better and smarter learning tools. Simone Mezzaro, Alessio Gambi, Gordon Fraser 0001 |
EASE | 2 |
| 2024 | Empirical Evaluation of Frequency Based Statistical Models for Estimating Killable MutantsabstractBackground. Mutation analysis is the premier technique for evaluating test suite quality estimating residual software defects. However, the reliability of mutation analysis is hampered by equivalent mutants which are undetectable by test cases. Reliably detecting and eliminating killable mutants is difficult as it is highly program and location dependent. Statistical estimation of killable mutants seems to be a promising approach to tackle this problem. Aims. Frequency-based species estimation methods have been proposed as a solution for several related problems in software testing. This paper investigates whether such frequency-based estimation methods can accurately estimate the number of killable mutants. Method. We conducted a large-scale empirical study on the ability of twelve widely known frequency-based estimators to predict the number of killable mutants in ten mature software projects. Result. Our investigation finds limited or no evidence that any of the statistical estimators are able to consistently predict the number of killable mutants in projects evaluated. Conclusion. We found that the investigated estimators lack sufficient predictive power and cannot produce reliable and useful estimates of killable mutants. Konstantin Kuznetsov 0001, Alessio Gambi, Saikrishna Dhiddi, Julia Hess, Rahul Gopinath |
ESEM | 2 |
| 2024 | The Flexcrash Platform for Testing Autonomous Vehicles in Mixed-Traffic ScenariosabstractAutonomous vehicles (AV) leverage Artificial Intelligence to reduce accidents and improve fuel efficiency while sharing the roads with human drivers. Current AV prototypes have not yet reached these goals, highlighting the need for better development and testing methodologies. AV testing practices extensively rely on simulations, but existing AV tools focus on testing single AV instances or do not consider human drivers. Thus, they might generate many irrelevant mixed-traffic test scenarios. The Flexcrash platform addresses these issues by allowing the generation and simulation of mixed-traffic scenarios, thus enabling testers to identify realistic critical scenarios, traffic experts to create new datasets, and regulators to extend consumer testing benchmarks. Alessio Gambi, Shreya Mathews, Benedikt Steininger, Mykhailo Poienko, David Bobek |
ISSTA | 1 |
| 2023 | Artisan: An Action-Based Test Carving Tool for Android AppsabstractAndroid app developers should take advantage of end-to-end tests to validate user flows and unit tests to perform focused debugging activities. While there is considerable tool support for creating end-to-end tests in the form of GUI tests, Android app developers lack automated support for creating unit tests, in particular, unit tests that run on a Java Virtual Machine (JVM). These tests are complex to create, as there is a need to suitably handle the coupling between the app under test (AUT) and the Android framework while creating the tests.To help Android app developers create unit tests that run on the JVM, we present a tool called ARTISAN. The tool uses the idea of test carving to create unit tests from GUI tests. Specifically, ARTISAN traces method invocations in the AUT during GUI testing and then creates focused unit tests that replicate the behavior of executed methods. We used 152 end-to-end tests from five apps to perform an empirical evaluation of ARTISAN and the tool was able to generate unit tests that cover a significant portion of the code exercised by the GUI tests (i.e., 45% of the initial GUI-test-based statement coverage on average). ARTISAN and its related experimental data are publicly available at https://github.com/se-umn/artisan. We provide a video demo of the tool at https://youtu.be/pRhVAhquPGk. Alessio Gambi, Mattia Fazzini |
ICSME | 1 |
| 2023 | STRETCH: Generating Challenging Scenarios for Testing Collision Avoidance SystemsabstractCollision avoidance systems are fundamental for autonomous driving and need to be tested thoroughly to check whether they safely handle critical scenarios. Testing collision avoidance systems is generally done by means of scenario-based testing using simulators and comes with the main challenge of generating situations that are realistic but avoidable. In other words, driving scenarios must stress the collision avoidance functionalities while being representative. Existing crash databases and accident reports describe observed accidents and enable to (re)create realistic collisions in simulations; however, as those data sources focus on the impact, their data do not generally lead to avoidable collision scenarios. To address this issue, we propose STRETCH, which generates realistic, critical, and avoidable collision scenarios by extending focused collision descriptions using a multi-objective optimization algorithm. Thanks to STRETCH, developers and testers can automatically generate challenging test cases based on realistic crash scenarios. Franz Scheuer, Alessio Gambi, Paolo Arcaini |
IV | 2 |
| 2023 | TEASER: Simulation-Based CAN Bus Regression Testing for Self-Driving Cars SoftwareabstractSafety-critical systems such as self-driving cars (SDCs) must be rigorously tested. Especially electronic control units (ECUs) of SDCs should be tested with realistic input data. In this context, a communication protocol called Controller Area Network (CAN) is typically used to transfer sensor data to the SDC control units. A challenge for SDC maintainers and testers is the need to manually define the CAN inputs that realistically represent the state of the SDC in the real world. To address this challenge, we developed TEASER; a tool that generates realistic CAN signals for SDCs obtained from sensors from state-of-the-art car simulators. We evaluated TEASER based on its integration capability into a DevOps pipeline of aicas GmbH, a company in the automotive sector. Concretely, we integrated TEASER in a Continous Integration (CI) pipeline configured with Jenkins. The pipeline executes the test cases in simulation environments and sends the sensor data over the CAN bus to a physical CAN device, the test subject. Our evaluation shows the ability of TEASER to generate and execute CI test cases that expose simulation-based faults (using regression strategies); the tool produces CAN inputs that realistically represent the state of the SDC in the real world. This result is critically important for increasing the automation and effectiveness of simulation-based CAN bus regression testing for SDCs. Christian Birchler, Cyrill Rohrbach, Hyeongkyun Kim, Alessio Gambi, Tianhai Liu, Jens Horneber, Timo Kehrer, Sebastiano Panichella |
ASE | 4 |
| 2023 | Machine learning-based test selection for simulation-based testing of self-driving cars softwareabstractAbstract Simulation platforms facilitate the development of emerging Cyber-Physical Systems (CPS) like self-driving cars (SDC) because they are more efficient and less dangerous than field operational test cases. Despite this, thoroughly testing SDCs in simulated environments remains challenging because SDCs must be tested in a sheer amount of long-running test cases. Past results on software testing optimization have shown that not all the test cases contribute equally to establishing confidence in test subjects’ quality and reliability, and the execution of “safe and uninformative” test cases can be skipped to reduce testing effort. However, this problem is only partially addressed in the context of SDC simulation platforms. In this paper, we investigate test selection strategies to increase the cost-effectiveness of simulation-based testing in the context of SDCs. We propose an approach called SDC-Scissor (SDC coS t-effeC tI ve teS t S electOR) that leverages Machine Learning (ML) strategies to identify and skip test cases that are unlikely to detect faults in SDCs before executing them. Our evaluation shows that SDC-Scissor outperforms the baselines. With the Logistic model, we achieve an accuracy of 70%, a precision of 65%, and a recall of 80% in selecting tests leading to a fault and improved testing cost-effectiveness. Specifically, SDC-Scissor avoided the execution of 50% of unnecessary tests as well as outperformed two baseline strategies. Complementary to existing work, we also integrated SDC-Scissor into the context of an industrial organization in the automotive domain to demonstrate how it can be used in industrial settings. Christian Birchler, Sajad Khatiri, Bill Bosshard, Alessio Gambi, Sebastiano Panichella |
Empir. Softw. Eng. | 4 |
| 2023 | Cost-effective simulation-based test selection in self-driving cars softwareabstractSimulation environments are essential for the continuous development of complex cyber-physical systems such as self-driving cars (SDCs). Previous results on simulation-based testing for SDCs have shown that many automatically generated tests do not strongly contribute to the identification of SDC faults, hence do not contribute towards increasing the quality of SDCs. Because running such “uninformative” tests generally leads to a waste of computational resources and a drastic increase in the testing cost of SDCs, testers should avoid them. However, identifying “uninformative” tests before running them remains an open challenge. Hence, this paper proposes SDC-Scissor, a framework that leverages Machine Learning (ML) to identify SDC tests that are unlikely to detect faults in the SDC software under test, thus enabling testers to skip their execution and drastically increase the cost-effectiveness of simulation-based testing of SDCs software. Our evaluation concerning the usage of six ML models on two large datasets characterized by 22'652 tests showed that SDC-Scissor achieved a classification F1-score up to 96%. Moreover, our results show that SDC-Scissor outperformed a randomized baseline in identifying more failing tests per time unit. Webpage & Video: https://github.com/ChristianBirchler/sdc-scissor Christian Birchler, Nicolas Erni, Sajad Khatiri, Alessio Gambi, Sebastiano Panichella |
Sci. Comput. Program. | 4 |
| 2023 | JUGE: An infrastructure for benchmarking Java unit test generatorsabstractSummary Researchers and practitioners have designed and implemented various automated test case generators to support effective software testing. Such generators exist for various languages (e.g., Java, C#, or Python) and various platforms (e.g., desktop, web, or mobile applications). The generators exhibit varying effectiveness and efficiency, depending on the testing goals they aim to satisfy (e.g., unit‐testing of libraries versus system‐testing of entire applications) and the underlying techniques they implement. In this context, practitioners need to be able to compare different generators to identify the most suited one for their requirements, while researchers seek to identify future research directions. This can be achieved by systematically executing large‐scale evaluations of different generators. However, executing such empirical evaluations is not trivial and requires substantial effort to select appropriate benchmarks, setup the evaluation infrastructure, and collect and analyse the results. In this Software Note, we present ourJUnit Generation Benchmarking Infrastructure(JUGE) supporting generators (search‐based, random‐based, symbolic execution, etc.) seeking to automate the production of unit tests for various purposes (validation, regression testing, fault localization, etc.). The primary goal is to reduce the overall benchmarking effort, ease the comparison of several generators, and enhance the knowledge transfer between academia and industry by standardizing the evaluation and comparison process. Since 2013, several editions of a unit testing tool competition, co‐located with the Search‐Based Software Testing Workshop, have taken place whereJUGEwas used and evolved. As a result, an increasing amount of tools (over 10) from academia and industry have been evaluated onJUGE, matured over the years, and allowed the identification of future research directions. Based on the experience gained from the competitions, we discuss the expected impact ofJUGEin improving the knowledge transfer on tools and approaches for test generation between academia and industry. Indeed, theJUGEinfrastructure demonstrated an implementation design that is flexible enough to enable the integration of additional unit test generation tools, which is practical for developers and allows researchers to experiment with new and advanced unit testing tools and approaches. Xavier Devroey, Alessio Gambi, Juan P. Galeotti, René Just, Fitsum Meshesha Kifetew, Annibale Panichella, Sebastiano Panichella |
Softw. Test. Verification Reliab. | 2 |
| 2023 | Efficient and Effective Feature Space Exploration for Testing Deep Learning SystemsabstractAssessing the quality of Deep Learning (DL) systems is crucial, as they are increasingly adopted in safety-critical domains. Researchers have proposed several input generation techniques for DL systems. While such techniques can expose failures, they do not explain which features of the test inputs influenced the system’s (mis-) behaviour. DeepHyperion was the first test generator to overcome this limitation by exploring the DL systems’ feature space at large. In this article, we propose DeepHyperion-CS , a test generator for DL systems that enhances DeepHyperion by promoting the inputs that contributed more to feature space exploration during the previous search iterations. We performed an empirical study involving two different test subjects (i.e., a digit classifier and a lane-keeping system for self-driving cars). Our results proved that the contribution-based guidance implemented within DeepHyperion-CS outperforms state-of-the-art tools and significantly improves the efficiency and the effectiveness of DeepHyperion . DeepHyperion-CS exposed significantly more misbehaviours for five out of six feature combinations and was up to 65% more efficient than DeepHyperion in finding misbehaviour-inducing inputs and exploring the feature space. DeepHyperion-CS was useful for expanding the datasets used to train the DL systems, populating up to 200% more feature map cells than the original training set. Tahereh Zohdinasab, Vincenzo Riccio, Alessio Gambi, Paolo Tonella |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2022 | Cost-effective Simulation-based Test Selection in Self-driving Cars Software with SDC-ScissorabstractSimulation platforms facilitate the continuous development of complex systems such as self-driving cars (SDCs). However, previous results on testing SDCs using simulations have shown that most of the automatically generated tests do not strongly contribute to establishing confidence in the quality and reliability of the SDC. Therefore, those tests can be characterized as “uninformative”, and running them generally means wasting precious computational resources. We address this issue with SDC-Scissor, a framework that leverages Machine Learning to identify simulation-based tests that are unlikely to detect faults in the SDC software under test and skip them before their execution. Consequently, by filtering out those tests, SDC-Scissor reduces the number of long-running simulations to execute and drastically increases the cost-effectiveness of simulation-based testing of SDCs software. Our evaluation concerning two large datasets and around 12'000 tests showed that SDC-Scissor achieved a higher classification F1-score (between 47% and 90%) than a randomized baseline in identifying tests that lead to a fault and reduced the time spent running uninformative tests (speedup between 107% and 170%). Webpage & Video: https://github.com/ChristianBirchler/sdc-scissor Christian Birchler, Nicolas Erni, Sajad Khatiri, Alessio Gambi, Sebastiano Panichella |
SANER | 4 |
| 2021 | DeepHyperion: exploring the feature space of deep learning-based systems through illumination searchabstractDeep Learning (DL) has been successfully applied to a wide range of application domains, including safety-critical ones. Several DL testing approaches have been recently proposed in the literature but none of them aims to assess how different interpretable features of the generated inputs affect the system's behaviour. Tahereh Zohdinasab, Vincenzo Riccio, Alessio Gambi, Paolo Tonella |
ISSTA | 3 |
| 2019 | Automatically testing self-driving cars with search-based procedural content generationabstractSelf-driving cars rely on software which needs to be thoroughly tested. Testing self-driving car software in real traffic is not only expensive but also dangerous, and has already caused fatalities. Virtual tests, in which self-driving car software is tested in computer simulations, offer a more efficient and safer alternative compared to naturalistic field operational tests. However, creating suitable test scenarios is laborious and difficult. In this paper we combine procedural content generation, a technique commonly employed in modern video games, and search-based testing, a testing technique proven to be effective in many domains, in order to automatically create challenging virtual scenarios for testing self-driving car soft- ware. Our AsFault prototype implements this approach to generate virtual roads for testing lane keeping, one of the defining features of autonomous driving. Evaluation on two different self-driving car software systems demonstrates that AsFault can generate effective virtual road networks that succeed in revealing software failures, which manifest as cars departing their lane. Compared to random testing AsFault was not only more efficient, but also caused up to twice as many lane departures. Alessio Gambi, Marc Müller, Gordon Fraser 0001 |
ISSTA | 1 |
| 2019 | Gamifying a Software Testing Course with Code DefendersabstractSoftware testing is an essential skill for software developers, but it is challenging to get students engaged in this activity. The Code Defenders game addresses this problem by letting students compete over code under test by either introducing faults ("attacking") or by writing tests ("defending") to reveal these faults. In this paper, we describe how we integrated Code Defenders as a semester-long activity of an undergraduate and graduate level university course on software testing. We complemented the regular course sessions with weekly Code Defenders sessions, addressing challenges such as selecting suitable code to test, managing games, and assessing performance. Our experience and our data show that the integration of Code Defenders was well-received by students and led them to practice testing thoroughly. Positive learning effects are evident as student performance improved steadily throughout the semester. Gordon Fraser 0001, Alessio Gambi, Marvin Kreis, José Miguel Rojas |
SIGCSE | 2 |
| 2019 | Generating effective test cases for self-driving cars from police reportsabstractAutonomous driving carries the promise to drastically reduce the number of car accidents; however, recently reported fatal crashes involving self-driving cars show that such an important goal is not yet achieved. This calls for better testing of the software controlling self-driving cars, which is difficult because it requires producing challenging driving scenarios. To better test self-driving car soft- ware, we propose to specifically test car crash scenarios, which are critical par excellence. Since real car crashes are difficult to test in field operation, we recreate them as physically accurate simulations in an environment that can be used for testing self-driving car software. To cope with the scarcity of sensory data collected during real car crashes which does not enable a full reproduction, we extract the information to recreate real car crashes from the police reports which document them. Our extensive evaluation, consisting of a user study involving 34 participants and a quantitative analysis of the quality of the generated tests, shows that we can generate accurate simulations of car crashes in a matter of minutes. Compared to tests which implement non critical driving scenarios, our tests effectively stressed the test subject in different ways and exposed several shortcomings in its implementation. Alessio Gambi, Tri Huynh, Gordon Fraser 0001 |
ESEC/SIGSOFT FSE | 1 |
| 2018 | Practical Test Dependency DetectionabstractRegression tests should consistently produce the same outcome when executed against the same version of the system under test. Recent studies, however, show a different picture: in many cases simply changing the order in which tests execute is enough to produce different test outcomes. These studies also identify the presence of dependencies between tests as one likely cause of this behavior. Test dependencies affect the quality of tests and of the correlated development activities, like regression test selection, prioritization, and parallelization, which assume that tests are independent. Therefore, developers must promptly identify and resolve problematic test dependencies. This paper presents PRADET, a novel approach for detecting problematic dependencies that is both effective and efficient. PRADET uses a systematic, data-driven process to detect problematic test dependencies significantly faster and more precisely than prior work. PRADET scales to analyze large projects with thousands of tests that existing tools cannot analyze in reasonable amount of time, and found 27 previously unknown dependencies. Alessio Gambi, Jonathan Bell 0001, Andreas Zeller |
ICST | 1 |
| 2017 | O!Snap: Cost-Efficient Testing in the CloudabstractPorting a testing environment to a cloud infrastructure is not straightforward. This paper presents O!Snap, an approach to generate test plans to cost-efficiently execute tests in the cloud. O!Snap automatically maximizes reuse of existing virtual machines, and interleaves the creation of updated test images with the execution of tests to minimize overall test execution time and/or cost. In an evaluation involving 2,600+ packages and 24,900+ test jobs of the Debian continuous integration environment, O!Snap reduces test setup time by up to 88% and test execution time by up to 43.3% without additional costs. Alessio Gambi, Alessandra Gorla, Andreas Zeller |
ICST | 1 |
| 2017 | CUT: automatic unit testing in the cloudabstractUnit tests can be significantly sped up by running them in parallel over distributed execution environments, such as the cloud. However, manually setting up such environments and configuring the testing frameworks to effectively use them is cumbersome and requires specialized expertise that developers might lack. Alessio Gambi, Sebastian Kappler, Johannes Lampel, Andreas Zeller |
ISSTA | 1 |
| 2016 | Kriging-Based Self-Adaptive Cloud ControllersabstractCloud technology is rapidly substituting classic computing solutions, and challenges the community with new problems. In this paper we focus on controllers for cloud application elasticity, and propose a novel solution for self-adaptive cloud controllers based on Kriging models. Cloud controllers are application specific schedulers that allocate resources to applications running in the cloud, aiming to meet the quality of service requirements while optimizing the execution costs. General-purpose cloud resource schedulers provide sub-optimal solutions to the problem with respect to application-specific solutions that we call cloud controllers. In this paper we discuss a general way to design self-adaptive cloud controllers based on Kriging models. We present Kriging models, and show how they can be used for building efficient controllers thanks to their unique characteristics. We report experimental data that confirm the suitability of Kriging models to support efficient cloud control and open the way to the development of a new generation of cloud controllers. Alessio Gambi, Mauro Pezzè, Giovanni Toffetti Carughi |
IEEE Trans. Serv. Comput. | 1 |
| 2015 | Poster: Improving Cloud-Based Continuous Integration EnvironmentsabstractWe propose a novel technique for improving the efficiency of cloud-based continuous integration development environments. Our technique identifies repetitive, expensive and time-consuming setup activities that are required to run integration and system tests in the cloud, and consolidates them into preconfigured testing virtual machines such that the overall costs of test execution are minimized. We create such testing machines by reconfiguring and opportunistically snapshotting the virtual machines already registered in the cloud. Alessio Gambi, Rostyslav Zabolotnyi, Schahram Dustdar |
ICSE (2) | 1 |
| 2014 | CoMoT - A Platform-as-a-Service for Elasticity in the CloudabstractPlatform-as-a-Service (PaaS) should support the design, deployment, execution, test and monitoring of native elastic systems constructed from elastic service units based on multi-dimensional elasticity requirements. In this paper, we discuss fundamental building blocks for enabling multi-dimensional elasticity programming of software-defined elastic systems. We describe CoMoT, a novel PaaS for elasticity in the cloud that is developed based on these fundamental building blocks. Hong Linh Truong 0001, Schahram Dustdar, Georgiana Copil, Alessio Gambi, Waldemar Hummer, Duc-Hung Le, Daniel Moldovan |
IC2E | 4 |
| 2013 | Automated testing of cloud-based elastic systems with AUToCLESabstractCloud-based elastic computing systems dynamically change their resources allocation to provide consistent quality of service and minimal usage of resources in the face of workload fluctuations. As elastic systems are increasingly adopted to implement business critical functions in a cost-efficient way, their reliability is becoming a key concern for developers. Without proper testing, cloud-based systems might fail to provide the required functionalities with the expected service level and costs. Using system testing techniques, developers can expose problems that escaped the previous quality assurance activities and have a last chance to fix bugs before releasing the system in production. System testing of cloud-based systems accounts for a series of complex and time demanding activities, from the deployment and configuration of the elastic system, to the execution of synthetic clients, and the collection and persistence of execution data. Furthermore, clouds enable parallel executions of the same elastic system that can reduce the overall test execution time. However, manually managing the concurrent testing of multiple system instances might quickly overwhelm developers' capabilities, and automatic support for test generation, system test execution, and management of execution data is needed. In this demo we showcase AUToCLES, our tool for automatic testing of cloud-based elastic systems. Given specifications of the test suite and the system under test, AUToCLES implements testing as a service (TaaS): It automatically instantiates the SUT, configures the testing scaffoldings, and automatically executes test suites. If required, AUToCLES can generate new test inputs. Designers can inspect executions both during and after the tests. Alessio Gambi, Waldemar Hummer, Schahram Dustdar |
ASE | 1 |
| 2013 | Iterative test suites refinement for elastic computing systemsabstractElastic computing systems can dynamically scale to continuously and cost-effectively provide their required Quality of Service in face of time-varying workloads, and they are usually implemented in the cloud. Despite their wide-spread adoption by industry, a formal definition of elasticity and suitable procedures for its assessment and verification are still missing. Both academia and industry are trying to adapt established testing procedures for functional and non-functional properties, with limited effectiveness with respect to elasticity. In this paper we propose a new methodology to automatically generate test-suites for testing the elastic properties of systems. Elasticity, plasticity, and oscillations are first formalized through a convenient behavioral abstraction of the elastic system and then used to drive an iterative test suite refinement process. The outcomes of our approach are a test suite tailored to the violation of elasticity properties and a human-readable abstraction of the system behavior to further support diagnosis and fix. Alessio Gambi, Antonio Filieri, Schahram Dustdar |
ESEC/SIGSOFT FSE | 1 |
| 2012 | Modeling Cloud performance with KrigingabstractCloud infrastructures allow service providers to implement elastic applications. These can be scaled at runtime to dynamically adjust their resources allocation to maintain consistent quality of service in response to changing working conditions, like flash crowds or periodic peaks. Providers need models to predict the system performances of different resource allocations to fully exploit dynamic application scaling. Traditional performance models such as linear models and queueing networks might be simplistic for real Cloud applications; moreover, they are not robust to change. We propose a performance modeling approach that is practical for highly variable elastic applications in the Cloud and automatically adapts to changing working conditions. We show the effectiveness of the proposed approach for the synthesis of a self-adaptive controller. Alessio Gambi, Giovanni Toffetti Carughi |
ICSE | 1 |
| 2010 | Engineering Autonomic Controllers for Virtualized Web ApplicationsabstractModern Web applications are often hosted in a virtualized cloud computing infrastructure, and can dynamically scale in response to unpredictable changes in the workload to guarantee a given service level agreement. In this paper we propose to use Kriging surrogate models to approximate the performance profile of virtualized, multi-tier Web applications. The model is first built through a set of automated and controlled experiments at staging time, and can be later updated and refined by monitoring the Web application deployed in production. We claim that surrogate modeling makes a very good candidate for a model-driven approach to the engineering of an autonomic controller. Our experimental evaluation shows that the model predictions are faithful to the observed system's performance, they improve with an increasing amount of samples and they can be computed quickly. We also provide evidence that the model can be effectively used to synthetize an aggregated objective function, a critical component of the autonomic controller. The approach is evaluated in the context of a RESTful Web service composition case study deployed on the RESERVOIR cloud. Giovanni Toffetti Carughi, Alessio Gambi, Mauro Pezzè, Cesare Pautasso |
ICWE | 2 |
| 2007 | Negotiation of Service Level Agreements: An Architecture and a Search-Based Approach
Elisabetta Di Nitto, Massimiliano Di Penta, Alessio Gambi, Gianluca Ripa, Maria Luisa Villani |
ICSOC | 3 |