EDBT 2026 Demo / reviewers in the wild / expert
Matteo Biagiola
dblp:205/1273
· DBLP profile ↗
24ranked-venue papers
9as first author
19since 2021 · last 2026
0000-0002-7825-3409ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 24 · 9 first-author · 19 since 2021Artificial intelligence and machine learning · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Neural Embeddings for Web TestingabstractWeb test automation techniques often rely on crawlers to infer models of web applications for automated test generation. However, current crawlers rely on state equivalence algorithms that struggle to distinguish near-duplicate pages, often leading to redundant test cases and incomplete coverage of application functionality. In this paper, we present a model-based test generation approach that employs transformer-based Siamese neural networks (SNNs) to infer web application models more accurately. By learning similarity-based representations, SNNs capture structural and textual relationships among web pages, improving near-duplicate detection during crawling and enhancing the quality of inferred models, and thus, the effectiveness of generated test suites. Our evaluation across nine web apps shows that SNNs outperform state-of-the-art techniques in near-duplicate detection, resulting in superior web app models with an average F-1 score improvement of 56%. These enhanced models enable the generation of more effective test suites that achieve higher code coverage, with improvements ranging from 6% to 21% and averaging at 12%. Kasun Kanaththage, Luigi L. L. Starace, Matteo Biagiola, Paolo Tonella, Andrea Stocco 0001 |
ICST | 3 |
| 2026 | Towards Automated Page Object Generation for Web Testing using Large Language Models
Betupsilonl Karagöz, Filippo Ricca, Matteo Biagiola, Andrea Stocco 0001 |
ICST | 3 |
| 2026 | DeepNaqqal: Human-Aligned Automated Validation of Test Inputs for Deep LearningabstractTest input generators (TIGs) are widely used to assess the robustness of Deep Learning (DL) image classifiers, yet they often produce invalid inputs that fall outside the semantic domain of the task, misleading quality assessment. While several automated validators have been proposed, there is a critical mismatch between automated and human validation criteria and, thus, automated validators are merely a proxy of domain validity, as perceived by human testers. We introduce DeepNaqqal, a supervised test input validator that learns validity directly from human-annotated labels using transfer learning on deep vision models. Our empirical study on automated validation of misclassification-inducing inputs compares DeepNaqqal against six state-of-the-art validators across three image classification tasks and multiple TIG families, using independent human assessment as ground truth. Our results show that DeepNaqqal consistently achieves the highest agreement with human judgments, while generalizing to unseen TIGs and remaining effective with substantially reduced labeled data. Maryam, Matteo Biagiola, Paolo Tonella, Vincenzo Riccio |
ICST | 2 |
| 2026 | XMutant: XAI-based fuzzing for deep learning systemsabstractSemantic-based test generators are widely used to produce failure-inducing inputs for Deep Learning (DL) systems. They typically generate challenging test inputs by applying random perturbations to input semantic concepts until a failure is found or a timeout is reached. However, such randomness may hinder them from efficiently achieving their goal. This paper proposes XMutant, a technique that leverages explainable artificial intelligence (XAI) techniques to generate challenging test inputs. XMutant uses the local explanation of the input to inform the fuzz testing process and effectively guide it toward failures of the DL system under test. We evaluated different configurations of XMutant in triggering failures for different DL systems both for model-level (sentiment analysis, digit recognition) and system-level testing (advanced driving assistance). Our studies showed that XMutant enables more effective and efficient test generation by focusing on the most impactful parts of the input. XMutant generates up to $$125\%$$ more failure-inducing inputs compared to an existing baseline, up to 7 $$\times$$ faster. We also assessed the validity of these inputs, maintaining a validation rate above $$89\%$$ , according to automated and human validators. Xingcheng Chen, Matteo Biagiola, Vincenzo Riccio, Marcelo d'Amorim, Andrea Stocco 0001 |
Empir. Softw. Eng. | 2 |
| 2026 | Simulator ensembles for trustworthy autonomous driving systems testingabstractAbstract Scenario-based testing with driving simulators is extensively used to identify failing conditions of automated driving assistance systems (ADAS) and reduce the amount of in-field road testing. However, existing studies have shown that repeated test execution in the same as well as in distinct simulators can yield different outcomes, which can be attributed to sources of flakiness or different implementations of the physics, among other factors. In this paper, we present , a novel approach to multi-simulation ADAS testing based on a search-based testing approach that leverages an ensemble of simulators to identify failure-inducing, simulator-agnostic test scenarios. During the search, each scenario is evaluated jointly on multiple simulators. Scenarios that produce consistent results across simulators are prioritized for further exploration, while those that fail on only a subset of simulators are given less priority, as they may reflect simulator-specific issues rather than generalizable failures. Our empirical study, which involves testing three lane-keeping ADAS with increasing complexity on different pairs of three widely used simulators, demonstrates that outperforms single-simulator testing by achieving, on average, a higher rate of simulator-agnostic failures by 66%. Compared to a state-of-the-art multi-simulator approach that combines the outcome of independent test generation campaigns obtained in different simulators, identifies on average up to $$3.4\times$$ more simulator-agnostic failing tests and higher failure rates. To avoid the costly execution of test inputs on which simulators disagree, we propose an enhancement of that leverages surrogate models to predict simulator disagreements and bypass test executions. Our results show that utilizing a surrogate model during the search does not only retain the average number of valid failures but also improves its efficiency in finding the first valid failure. These findings indicate that combining an ensemble of simulators during the search is a promising approach for the automated cross-replication in ADAS testing. Lev Sorokin, Matteo Biagiola, Andrea Stocco 0001 |
Empir. Softw. Eng. | 2 |
| 2026 | Metamorphic Testing for Infrastructure-as-Code EnginesabstractInfrastructure-as-Code (IaC) engines, such as Terraform, OpenTofu, and Pulumi, automate the provisioning and management of cloud resources. They parse IaC specifications and orchestrate the required actions, making them the backbone of modern clouds, and critical to the reliability of both the underlying infrastructure and the software that depends on it. Despite this importance, this class of systems has received little attention: prior work largely targets the correctness of IaC programs rather than the IaC engines themselves. Existing test suites rely on manually written oracles and struggle to expose faults that manifest across multiple executions, leaving a significant reliability gap. We present EMI a C, a metamorphic testing framework for IaC engines. EMI a C defines metamorphic relations as graph-based transformations of IaC programs and checks invariants across executions of the original and transformed programs. A central novelty is our use of e-graphs in software testing, as both a test-input generator and an equivalence oracle. E-graphs compactly represent program equivalences, enabling the systematic generation of large spaces of equivalent IaC programs. To ground these relations, we analyze 43,593 real-world Terraform programs and show that IaC dependency graphs are typically small and sparse, making e-graphs a natural fit. Evaluating EMI a C on Pulumi, Terraform, and OpenTofu, we show that it complements existing test suites by exercising engine-critical code paths and covering 98 previously untested statements in Terraform and 1,313 in Pulumi. EMI a C also uncovers previously unknown issues in all three test suites, improving their adequacy. Three test cases have been merged into Terraform’s main branch, and Pulumi has merged a specification fix. David Spielmann, George Zakhour, Dominik Arnold, Matteo Biagiola, Roland Meier, Guido Salvaneschi |
Proc. ACM Program. Lang. | 4 |
| 2026 | GIFTbench: Generative image fuzz testing benchmarkabstractGIFTbench is a modular framework for testing Deep Learning image classifiers that combines Generative AI with genetic algorithms. Its architecture integrates pretrained generative models with a user-friendly Gradio interface, enabling automated, reproducible, and interpretable robustness testing. Supporting VAE, GAN, and Diffusion models, GIFTbench generates test inputs by perturbing latent representations to expose misbehaviors of the classifier under test. By automating test input generation and reducing the need for manual coding, GIFTbench accelerates experimentation and facilitates comparative evaluation of both classifiers and generative models. Designed for researchers and practitioners, it enables reproducible assessment of image classifiers, while supporting studies on classifier vulnerabilities, mutation strategies, and the role of generative models in robustness testing. Maryam, Matteo Biagiola, Andrea Stocco 0001, Vincenzo Riccio |
Sci. Comput. Program. | 2 |
| 2025 | $\mu \text{PRL}$: A Mutation Testing Pipeline for Deep Reinforcement Learning Based on Real FaultsabstractReinforcement Learning (RL) is increasingly adopted to train agents that can deal with complex sequential tasks, such as driving an autonomous vehicle or controlling a humanoid robot. Correspondingly, novel approaches are needed to ensure that RL agents have been tested adequately before going to production. Among them, mutation testing is quite promising, especially under the assumption that the injected faults (mutations) mimic the real ones. In this paper, we first describe a taxonomy of real RL faults obtained by repository mining. Then, we present the mutation operators derived from such real faults and implemented in the tool$\mu \text{PRL}$. Finally, we discuss the experimental results, showing that$\mu \text{PRL}$is effective at discriminating strong from weak test generators, hence providing useful feedback to developers about the adequacy of the generated test scenarios. Deepak-George Thomas, Matteo Biagiola, Nargiz Humbatova, Mohammad Wardat, Gunel Jahangirova, Hridesh Rajan, Paolo Tonella |
ICSE | 2 |
| 2025 | Improving the Readability of Automatically Generated Tests Using Large Language ModelsabstractSearch-based test generators are effective at producing unit tests with high coverage. However, such automatically generated tests have no meaningful test and variable names, making them hard to understand and interpret by developers. On the other hand, large language models (LLMs) can generate highly readable test cases, but they are not able to match the effectiveness of search-based generators, in terms of achieved code coverage. In this paper, we propose to combine the effectiveness of search-based generators with the readability of LLM generated tests. Our approach focuses on improving test and variable names produced by search-based tools, while keeping their semantics (i.e., their coverage) unchanged. Our evaluation on nine industrial and open source LLMs show that our readability improvement transformations are overall semantically-preserving and stable across multiple repetitions. Moreover, a human study with ten professional developers, show that our LLM-improved tests are as readable as developer-written tests, regardless of the LLM employed. Matteo Biagiola, Gianluca Ghislotti, Paolo Tonella |
ICST | 1 |
| 2025 | Benchmarking Generative AI Models for Deep Learning Test Input GenerationabstractTest Input Generators (TIGs) are crucial to assess the ability of Deep Learning (DL) image classifiers to provide correct predictions for inputs beyond their training and test sets. Recent advancements in Generative AI(GenAI) models have made them a powerful tool for creating and manipulating synthetic images, although these advancements also imply increased complexity and resource demands for training. In this work, we benchmark and combine different GenAI models with TIGs, assessing their effectiveness, efficiency, and quality of the generated test images, in terms of domain validity and label preservation. We conduct an empirical study involving three different GenAI architectures (VAEs, GANs, Diffusion Models), five classification tasks of increasing complexity, and 364 human evaluations. Our results show that simpler architectures, such as VAEs, are sufficient for less complex datasets like MNIST. However, when dealing with feature-rich datasets, such as ImageNet, more sophisticated architectures like Diffusion Models achieve superior performance by generating a higher number of valid, misclassification-inducing inputs. Maryam, Matteo Biagiola, Andrea Stocco 0001, Vincenzo Riccio |
ICST | 2 |
| 2025 | Reinforcement learning for online testing of autonomous driving systems: a replication and extension studyabstractIn a recent study, Reinforcement Learning (RL) used in combination with many-objective search, has been shown to outperform alternative techniques (random search and many-objective search) for online testing of Deep Neural Network-enabled systems. The empirical evaluation of these techniques was conducted on a state-of-the-art Autonomous Driving System (ADS). This work is a replication and extension of that empirical study. Our replication shows that RL does not outperform pure random test generation in a comparison conducted under the same settings of the original study, but with no confounding factor coming from the way collisions are measured. Our extension aims at eliminating some of the possible reasons for the poor performance of RL observed in our replication: (1) the presence of reward components providing contrasting feedback to the RL agent; (2) the usage of an RL algorithm (Q-learning) which requires discretization of an intrinsically continuous state space. Results show that our new RL agent is able to converge to an effective policy that outperforms random search. Results also highlight other possible improvements, which open to further investigations on how to best leverage RL for online ADS testing. Luca Giamattei, Matteo Biagiola, Roberto Pietrantuono, Stefano Russo 0001, Paolo Tonella |
Empir. Softw. Eng. | 2 |
| 2025 | STILE: A tool for optimizing E2E web test scripts parallelizationabstractWeb applications quality is commonly assessed by executing End-to-End (E2E) test scripts interacting with those systems as a human tester would. To avoid setting up the web application state for each test script, testers usually create test scripts that may depend on others previously executed. However, the presence of dependencies prevents parallelization, a fundamental technique for speedup the execution of large test suites. In this paper, we present Stile , a tool for parallelizing the execution of E2E web test scripts that generates and executes a set of test schedules satisfying two important constraints: (1) every schedule respects existing test dependencies, and (2) all test scripts in the test suite are executed at least once. Moreover, Stile optimizes the execution by running only once the test scripts that are shared among the schedules. We empirically evaluated Stile on eight E2E test suites by comparing the execution time of Stile both with the sequential execution and with the parallel execution based on Selenium Grid. Our results show that Stile can reduce the execution time up to 80% w.r.t. the sequential execution and up to 50% w.r.t. Grid. Moreover, Stile provides a reduction in the CPUs usage (i.e., overall CPU-time) up to 75%. Dario Olianas, Maurizio Leotta, Filippo Ricca, Matteo Biagiola, Paolo Tonella |
J. Syst. Softw. | 4 |
| 2025 | Parallelization in System-Level Testing: Novel Approaches to Manage Test Suite DependenciesabstractSystem-level testing is fundamental to ensure the reliability of software systems. However, the execution time for system tests can be quite long, sometimes prohibitively long, especially in a regimen of continuous integration and deployment. One way to speed things up is to run the tests in parallel, provided that the execution schedule respects any dependency between tests. We present two novel approaches to detect dependencies in system-level tests, namely PFAST and MEM-FAST, which are highly parallelizable and optimistically run test schedules to exclude many dependencies when there are no failures. We evaluated our approaches both asymptotically and practically, on six Web applications and their system-level test suites, as well as on MySQL system-level tests. Our results show that, in general, PFAST is significantly faster than the state-of-the-art PRADET dependency detection algorithm, while producing parallelizable schedules that achieve a significant reduction in the overall test suite execution time. Pasquale Polverino, Fabio Di Lauro, Matteo Biagiola, Paolo Tonella, Antonio Carzaniga |
IEEE Trans. Software Eng. | 3 |
| 2024 | Adversarial Testing with Reinforcement Learning: A Case Study on Autonomous DrivingabstractTesting autonomous driving systems (ADSs) is essential to ensure their safety. Existing testing techniques manipulate the objects of the driving environment in order to trigger a misbehavior of the ADS under test. Reinforcement learning (RL) approaches have been applied to effectively modify the dynamic objects of the environment (e.g., pedestrians and other vehicles), also known as Non-Playable Characters (NPCs). However, existing RL approaches implement centralized controllers of the environment, resulting in possibly unrealistic and even invalid behaviors of the NPCs. In this paper, we propose to model NPCs as independent and fully autonomous agents to challenge the ADS under test (i.e., the ego ADS). In the first step of our approach, we train an adversarial ADS by designing a reward function as a linear combination of two components: (1) a component that encourages it to drive well in the given driving scenario, and (2) an adversarial component, that smoothly guides the agent towards a collision with the ego ADS. In our second step, we resume training of the ego ADS, to increase its robustness towards the behaviors of the adversarial ADS. Our experiments on a highway driving scenario show that the adversarial ADS is significantly more effective at generating collisions of the ego ADS than a random baseline. Moreover, adversarial retraining induces safe behaviors of the ego ADS, preventing the adversarial ADS from colliding. Andréa Doreste, Matteo Biagiola, Paolo Tonella |
ICST | 2 |
| 2024 | Two is better than one: digital siblings to improve autonomous driving testingabstractAbstract Simulation-based testing represents an important step to ensure the reliability of autonomous driving software. In practice, when companies rely on third-party general-purpose simulators, either for in-house or outsourced testing, the generalizability of testing results to real autonomous vehicles is at stake. In this paper, we enhance simulation-based testing by introducing the notion ofdigital siblings—a multi-simulator approach that tests a given autonomous vehicle on multiple general-purpose simulators built with different technologies, that operate collectively as an ensemble in the testing process. We exemplify our approach on a case study focused on testing the lane-keeping component of an autonomous vehicle. We use two open-source simulators as digital siblings, and we empirically compare such a multi-simulator approach against a digital twin of a physical scaled autonomous vehicle on a large set of test cases. Our approach requires generating and running test cases for each individual simulator, in the form of sequences of road points. Then, test cases are migrated between simulators, using feature maps to characterize the exercised driving conditions. Finally, the joint predicted failure probability is computed, and a failure is reported only in cases of agreement among the siblings. Our empirical evaluation shows that the ensemble failure predictor by the digital siblings is superior to each individual simulator at predicting the failures of the digital twin. We discuss the findings of our case study and detail how our approach can help researchers interested in automated testing of autonomous driving software. Matteo Biagiola, Andrea Stocco 0001, Vincenzo Riccio, Paolo Tonella |
Empir. Softw. Eng. | 1 |
| 2024 | Testing of Deep Reinforcement Learning Agents with Surrogate ModelsabstractDeep Reinforcement Learning (DRL) has received a lot of attention from the research community in recent years. As the technology moves away from game playing to practical contexts, such as autonomous vehicles and robotics, it is crucial to evaluate the quality of DRL agents. In this article, we propose a search-based approach to test such agents. Our approach, implemented in a tool called Indago , trains a classifier on failure and non-failure environment (i.e., pass) configurations resulting from the DRL training process. The classifier is used at testing time as a surrogate model for the DRL agent execution in the environment, predicting the extent to which a given environment configuration induces a failure of the DRL agent under test. The failure prediction acts as a fitness function, guiding the generation towards failure environment configurations, while saving computation time by deferring the execution of the DRL agent in the environment to those configurations that are more likely to expose failures. Experimental results show that our search-based approach finds 50% more failures of the DRL agent than state-of-the-art techniques. Moreover, such failures are, on average, 78% more diverse; similarly, the behaviors of the DRL agent induced by failure configurations are 74% more diverse. Matteo Biagiola, Paolo Tonella |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2024 | Boundary State Generation for Testing and Improvement of Autonomous Driving SystemsabstractRecent advances in Deep Neural Networks (DNNs) and sensor technologies are enabling autonomous driving systems (ADSs) with an ever-increasing level of autonomy. However, assessing their dependability remains a critical concern. State-of-the-art ADS testing approaches modify the controllable attributes of a simulated driving environment until the ADS misbehaves. In such approaches, environment instances in which the ADS is successful are discarded, despite the possibility that they could contain hidden driving conditions in which the ADS may misbehave. In this paper, we presentGenBo(GENerator of BOundary state pairs), a novel test generator for ADS testing.GenBomutates the driving conditions of the ego vehicle (position, velocity and orientation), collected in a failure-free environment instance, and efficiently generates challenging driving conditions at the behavior boundary (i.e., where the model starts to misbehave) in the same environment instance. We use such boundary conditions to augment the initial training dataset and retrain the DNN model under test. Our evaluation results show that the retrained model has, on average, up to 3$\times$higher success rate on a separate set of evaluation tracks with respect to the original DNN model. Matteo Biagiola, Paolo Tonella |
IEEE Trans. Software Eng. | 1 |
| 2022 | Testing the Plasticity of Reinforcement Learning-based SystemsabstractThe dataset available for pre-release training of a machine-learning based system is often not representative of all possible execution contexts that the system will encounter in the field. Reinforcement Learning (RL) is a prominent approach among those that support continual learning, i.e., learning continually in the field, in the post-release phase. No study has so far investigated any method to test the plasticity of RL-based systems, i.e., their capability to adapt to an execution context that may deviate from the training one. We propose an approach to test the plasticity of RL-based systems. The output of our approach is a quantification of the adaptation and anti-regression capabilities of the system, obtained by computing the adaptation frontier of the system in a changed environment. We visualize such frontier as an adaptation/anti-regression heatmap in two dimensions, or as a clustered projection when more than two dimensions are involved. In this way, we provide developers with information on the amount of changes that can be accommodated by the continual learning component of the system, which is key to decide if online, in-the-field learning can be safely enabled or not. Matteo Biagiola, Paolo Tonella |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2021 | STILE: a Tool for Parallel Execution of E2E Web Test ScriptsabstractAutomated end-to-end (E2E) Web testing relying on frameworks such as Selenium Web Driver is commonly used to assess the quality of web applications. However, the resulting test scripts may require long execution times, due to their interaction with the browser GUI and backend services. To avoid repeated and costly setup of the Web application state, testers tend to build test suites whose test scripts depend on each other (i.e., one test case sets up the application state expected by another test case). In this paper we present Stile, a tool for the parallel execution of Web test scripts that ensures the compliance of all execution schedules with the dependencies among the involved test scripts, while at the same time minimizing the execution time and the computation time required for such parallel execution. Experimental results show that execution times can be approximately halved thanks to Stile. Dario Olianas, Maurizio Leotta, Filippo Ricca, Matteo Biagiola, Paolo Tonella |
ICST | 4 |
| 2020 | Dependency-Aware Web Test GenerationabstractWeb crawlers can perform long running in-depth explorations of a web application, achieving high coverage of the navigational structure. However, a crawling trace cannot be easily turned into a minimal test suite that achieves the same coverage. In fact, when the crawling trace is segmented into test cases, two problems arise: (1) test cases are dependent on each other, therefore they may raise errors when executed in isolation, and (2) test cases are redundant, since the same targets are covered multiple times by different test cases. In this paper, we propose DANTE, a novel web test generator that computes the test dependencies associated with the test cases obtained from a crawling session, and uses them to eliminate redundant tests and produce executable test schedules. DANTE can effectively turn a web crawler into a test case generator that produces minimal test suites, composed only of feasible tests that contribute to achieve the final coverage. Experimental results show that DANTE, on average, (1) reduces the error rate of the test cases obtained by crawling traces from 85% to zero, (2) produces minimized test suites that are 84% smaller than the initial ones, and (3) outperforms two competing crawling-based and model-based techniques in terms of coverage and breakage rate. Matteo Biagiola, Andrea Stocco 0001, Filippo Ricca, Paolo Tonella |
ICST | 1 |
| 2020 | A Family of Experiments to Assess the Impact of Page Object Pattern in Web Test Suite DevelopmentabstractAutomated web testing is an appealing option, especially when continuous testing practices are adopted. However, web test cases are known to be fragile and to break easily when a web application evolves. The Page Object (PO) design pattern addresses such problem by providing a layer of indirection that decouples test cases from the internals of the web page, where web page elements are located and triggered by the web tests. However, PO development could potentially introduce an additional burden to the already strictly constrained testing activities. This paper reports an empirical investigation of costs and benefits due to the introduction of the PO pattern in web test suite development. In particular, we conducted a family of controlled experiments in which test cases were developed with and without the PO pattern. While the benefits of POs did not compensate for the extra development effort they require in the limited experimental setting of our study, results indicate that when the test suite to be developed is at least 10× larger, test development becomes more efficient with than without POs. Maurizio Leotta, Matteo Biagiola, Filippo Ricca, Mariano Ceccato, Paolo Tonella |
ICST | 2 |
| 2019 | Web test dependency detectionabstractE2E web test suites are prone to test dependencies due to the heterogeneous multi-tiered nature of modern web apps, which makes it difficult for developers to create isolated program states for each test case. In this paper, we present the first approach for detecting and validating test dependencies present in E2E web test suites. Our approach employs string analysis to extract an approximated set of dependencies from the test code. It then filters potential false dependencies through natural language processing of test names. Finally, it validates all dependencies, and uses a novel recovery algorithm to ensure no true dependencies are missed in the final test dependency graph. Our approach is implemented in a tool called TEDD and evaluated on the test suites of six open-source web apps. Our results show that TEDD can correctly detect and validate test dependencies up to 72% faster than the baseline with the original test ordering in which the graph contains all possible dependencies. The test dependency graphs produced by TEDD enable test execution parallelization, with a speed-up factor of up to 7×. Matteo Biagiola, Andrea Stocco 0001, Ali Mesbah 0001, Filippo Ricca, Paolo Tonella |
ESEC/SIGSOFT FSE | 1 |
| 2019 | Diversity-based web test generationabstractExisting web test generators derive test paths from a navigational model of the web application, completed with either manually or randomly generated input values. However, manual test data selection is costly, while random generation often results in infeasible input sequences, which are rejected by the application under test. Random and search-based generation can achieve the desired level of model coverage only after a large number of test execution at- tempts, each slowed down by the need to interact with the browser during test execution. In this work, we present a novel web test generation algorithm that pre-selects the most promising candidate test cases based on their diversity from previously generated tests. As such, only the test cases that explore diverse behaviours of the application are considered for in-browser execution. We have implemented our approach in a tool called DIG. Our empirical evaluation on six real-world web applications shows that DIG achieves higher coverage and fault detection rates significantly earlier than crawling-based and search-based web test generators. Matteo Biagiola, Andrea Stocco 0001, Filippo Ricca, Paolo Tonella |
ESEC/SIGSOFT FSE | 1 |
| 2017 | Search Based Path and Input Data Generation for Web Application Testing
Matteo Biagiola, Filippo Ricca, Paolo Tonella |
SSBSE | 1 |