VLDB 2026 Research / reviewers in the wild / expert
Dario Olianas
dblp:227/6965
· DBLP profile ↗
12ranked-venue papers
7as first author
9since 2021 · last 2026
0000-0002-6618-4186ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 11 · 7 first-author · 8 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BEWT: Extended Benchmarking for End-to-End Web TestingabstractWeb applications are essential in modern society and require thorough testing to ensure reliability and dependability. End-to-End (E2E) testing is key to evaluating both front-end and back-end components of complex web applications. Recent research has focused on improving E2E test suites by reducing maintenance costs, mitigating flakiness, and enhancing test resilience. Yet, a major challenge remains the lack of a publicly available benchmark for comparing techniques. Developing test suites for research is complex and prone to bias, highlighting the need for a shared benchmark to facilitate comparisons and experimental validation. We address this gap by providing a set of test suites designed as a benchmark for E2E testing studies. They come with scripts and automated installers for seamless deployment of the application under test in Docker containers, enhancing their usability. Our benchmark consists of 36 Java Selenium WebDriver-based E2E test suites for 8 different web applications. Each application includes multiple test suites with varying characteristics, such as the use of the Page Object pattern and advanced waiting mechanisms. Additionally, for 4 out of the 8 web applications, we provide test suites for two different versions, enabling studies on test suite evolution. Our benchmark includes 1166 test scripts, featuring 492 Page Objects and 8284 locators. To manage asynchronous behavior and reduce flakiness, the test suites incorporate 718 thread sleeps and 815 explicit waits. With a total of 51838 lines of code (LOC) this dataset can potentially be used as a resource, for example, to evaluate test automation strategies, study test suite evolution, or explore flakiness mitigation techniques. Dario Olianas, Maurizio Leotta, Filippo Ricca |
J. Syst. Softw. | 1 |
| 2025 | BEWT: A Benchmark for End-to-End Web Testing
Dario Olianas, Maurizio Leotta, Filippo Ricca |
SEAA (3) | 1 |
| 2025 | Leveraging Large Language Models for Explicit Wait Management in End-to-End Web TestingabstractEnd-to-end (E2E) testing is an approach in which an application is automatically tested through scripts that simulate the actions a user would perform. Properly managing asynchronous interactions is crucial in this approach to avoid test failures and flakiness. In the Selenium WebDriver framework, this is typically addressed by using thread sleeps (which pause the test for a fixed time) or explicit waits (function calls that pause the test execution until a specified condition is met). Explicit waits require the selection of both a condition to wait for (e.g., element visibility, element clickability) and an element on which that condition applies. Since thread sleeps are unreliable and replacing them with appropriate explicit waits is a time consuming task, in this work, we leverage a Large Language Model (LLM) to assist testers in selecting the most appropriate explicit waits. We defined a structured procedure (a series of prompts) for engaging with the LLM and validated this approach empirically on three test suites affected by asynchronous waiting issues, as well as on 12 synthetic examples. Additionally, we compared our approach with SleepReplacer, the current state-of-the-art tool for replacing thread sleeps with explicit waits in E2E web test suites. The results show that the LLM-based approach can automatically replace the majority of thread sleeps in a test suite on the first attempt, outperforming SleepReplacer. Dario Olianas, Maurizio Leotta, Filippo Ricca |
ICST | 1 |
| 2025 | STILE: A tool for optimizing E2E web test scripts parallelizationabstractWeb applications quality is commonly assessed by executing End-to-End (E2E) test scripts interacting with those systems as a human tester would. To avoid setting up the web application state for each test script, testers usually create test scripts that may depend on others previously executed. However, the presence of dependencies prevents parallelization, a fundamental technique for speedup the execution of large test suites. In this paper, we present Stile , a tool for parallelizing the execution of E2E web test scripts that generates and executes a set of test schedules satisfying two important constraints: (1) every schedule respects existing test dependencies, and (2) all test scripts in the test suite are executed at least once. Moreover, Stile optimizes the execution by running only once the test scripts that are shared among the schedules. We empirically evaluated Stile on eight E2E test suites by comparing the execution time of Stile both with the sequential execution and with the parallel execution based on Selenium Grid. Our results show that Stile can reduce the execution time up to 80% w.r.t. the sequential execution and up to 50% w.r.t. Grid. Moreover, Stile provides a reduction in the CPUs usage (i.e., overall CPU-time) up to 75%. Dario Olianas, Maurizio Leotta, Filippo Ricca, Matteo Biagiola, Paolo Tonella |
J. Syst. Softw. | 1 |
| 2024 | An empirical study to compare three web test automation approaches: NLP-based, programmable, and capture&replayabstractAbstract A new advancement in test automation is the use of natural language processing (NLP) to generate test cases (or test scripts) from natural language text. NLP is innovative in this context and promises of reducing test cases creation time and simplifying understanding for “non‐developer” software testers as well. Recently, many vendors have launched on the market many proposals of NLP‐based tools and testing frameworks but their superiority has never been empirically validated. This paper investigates the adoption of NLP‐based test automation in the web context with a series of case studies conducted to compare the costs of the NLP testing approach—measured in terms of test cases development and test cases evolution—with respect to more consolidated approaches, that is, programmable (or script‐based) testing and capture&replay testing. The results of our study show that NLP‐based test automation appears to be competitive for small‐ to medium‐sized test suites such as those considered in our empirical study. It minimizes the total cumulative cost (development and evolution) and does not require software testers with programming skills. Maurizio Leotta, Filippo Ricca, Alessandro Marchetto 0001, Dario Olianas |
J. Softw. Evol. Process. | 4 |
| 2022 | A large experimentation to analyze the effects of implementation bugs in machine learning algorithms
Maurizio Leotta, Dario Olianas, Filippo Ricca |
Future Gener. Comput. Syst. | 2 |
| 2022 | MATTER: A tool for generating end-to-end IoT test scriptsabstractAbstract In the last few years, Internet of Things (IoT) systems have drastically increased their relevance in many fundamental sectors. For this reason, assuring their quality is of paramount importance, especially in safety-critical contexts. Unfortunately, few quality assurance proposals for assuring the quality of these complex systems are present in the literature. In this paper, we extended and improved our previous approach for semi-automated model-based generation of executable test scripts. Our proposal is oriented to system-level acceptance testing of IoT systems. We have implemented a prototype tool taking in input a UML model of the system under test and some additional artefacts, and producing in output a test suite that checks if the system’s behaviour is compliant with such a model. We empirically evaluated our tool employing two IoT systems: a mobile health IoT system for diabetic patients and a smart park management system part of a smart city project. Both systems involve sensors or actuators, smartphones, and a remote cloud server. Results show that the test suites generated with our tool have been able to kill 91% of the overall 260 generated mutants (i.e. artificial bugged versions of the two considered systems). Moreover, the optimisation introduced in this novel version of our prototype, based on a minimisation post-processing step, allowed to reduce the time required for executing the entire test suites (about -20/25%) with no adverse effect on the bug-detection capability. Dario Olianas, Maurizio Leotta, Filippo Ricca |
Softw. Qual. J. | 1 |
| 2022 | SleepReplacer: a novel tool-based approach for replacing thread sleeps in selenium WebDriver test codeabstractAbstract Assuring quality of web applications is fundamental, given their relevance in the today’s world. A possible way to reach this goal is through end-to-end (E2E) testing, an approach in which a web application is automatically tested by performing the actions that a user would do. With modern web applications (for example, single-page applications), it is of great importance to properly handle asynchronous calls in the test suite. In E2E Selenium WebDriver test suites, asynchronous calls are usually managed in two ways: using thread sleeps or explicit waits. The first is easier to use, but is inefficient and can lead to instability (also called flakiness, a problem often present in test suites that makes us lose confidence in the testing phase), while the second is usually more efficient but harder to use because, if the correct kind of wait is not carefully selected, it can introduce flakiness too. To help Testers, who often opt for the first strategy, we present in this work a tool-based approach to automatically replace thread sleeps with explicit waits in an E2E Selenium WebDriver test suite without introducing new flakiness. We empirically validated our tool named SleepReplacer on four different test suites, and we found that it can correctly replace in an automatic way from 81 to 100% of thread sleeps, leading to a significant reduction of the total execution time of the test suite (i.e., from 13 to 71%). Dario Olianas, Maurizio Leotta, Filippo Ricca |
Softw. Qual. J. | 1 |
| 2021 | STILE: a Tool for Parallel Execution of E2E Web Test ScriptsabstractAutomated end-to-end (E2E) Web testing relying on frameworks such as Selenium Web Driver is commonly used to assess the quality of web applications. However, the resulting test scripts may require long execution times, due to their interaction with the browser GUI and backend services. To avoid repeated and costly setup of the Web application state, testers tend to build test suites whose test scripts depend on each other (i.e., one test case sets up the application state expected by another test case). In this paper we present Stile, a tool for the parallel execution of Web test scripts that ensures the compliance of all execution schedules with the dependencies among the involved test scripts, while at the same time minimizing the execution time and the computation time required for such parallel execution. Experimental results show that execution times can be approximately halved thanks to Stile. Dario Olianas, Maurizio Leotta, Filippo Ricca, Matteo Biagiola, Paolo Tonella |
ICST | 1 |
| 2020 | Two experiments for evaluating the impact of Hamcrest and AssertJ on assertion development
Maurizio Leotta, Maura Cerioli, Dario Olianas, Filippo Ricca |
Softw. Qual. J. | 3 |
| 2019 | Comparing Testing and Runtime Verification of IoT Systems: A Preliminary Evaluation based on a Case StudyabstractAssuring the quality of Internet of Things (IoT) systems is of paramount importance, and guaranteeing their reliability and compliance with the requirements is mandatory, but few attempts have been made so far. In previous works, we proposed two approaches for acceptance testing and runtime verification of IoT systems. Both works rely on a UML state machine to specify the system expected behaviour. In the acceptance testing approach, the interesting paths to exercise are identified and translated into executable test scripts. In the runtime verification approach, the relevant events during the system execution are monitored and compared against a formal specification derived from the UML state machine. In this paper, we compare the effectiveness of our two approaches, by applying them to a mobile health IoT system for the management of diabetic patients, employing over 100 mutated versions of the original system and analysing more than 1000 different executions. Results show that both approaches are effective in different ways in detecting bugs. While the acceptance testing approach is more effective to detect the bugs affecting the user interface, the runtime verification approach tracks better the subtle deviations from the system expected behaviour, in particular those concerning network issues. Maurizio Leotta, Diego Clerissi, Luca Franceschini, Dario Olianas, Davide Ancona, Filippo Ricca, Marina Ribaudo |
ENASE | 4 |
| 2018 | An acceptance testing approach for Internet of Things systemsabstractInternet of things (IoT) systems are becoming ubiquitous and assuring their quality is fundamental. Unfortunately, a few proposals for testing these complex, and often safety‐critical, systems are present in the literature. The authors propose an approach for acceptance testing of IoT systems adopting graphical user interfaces as a principal way of interaction. Acceptance testing is a type of black box testing based on test scenarios, i.e. sequences of steps/actions performed by the user or the system. In their approach, test scenarios are derived from a state machine that expresses the behaviour of the system under test, and test cases are derived from them by specifying the actual data and assertions and made executable by implementing the corresponding test scripts. As a case study, they selected a mobile health IoT system for diabetes management composed of local sensors/actuators, smartphones, and a remote cloud‐based system. The effectiveness of the approach has been evaluated by measuring the capability of two test suites implemented using different localisation strategies (visual and structure‐based) in detecting mutants of the original m‐health system. Results show the effectiveness of the test suites implemented by following the proposed approach since 93% of the generated mutants have been detected. Maurizio Leotta, Diego Clerissi, Dario Olianas, Filippo Ricca, Davide Ancona, Giorgio Delzanno, Luca Franceschini, Marina Ribaudo |
IET Softw. | 3 |