EDBT 2026 Demo / reviewers in the wild / expert
Vincenzo Riccio
dblp:177/1943
· DBLP profile ↗
21ranked-venue papers
5as first author
16since 2021 · last 2027
0000-0002-6229-8231ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 21 · 5 first-author · 16 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Does road diversity really matter in testing automated driving systems?abstractAbstract Context The use of automated driving systems (ADSs) in the real world requires rigorous testing to ensure safety. To increase trust, ADSs should be tested on a large set of diverse road scenarios. Literature suggests that if a vehicle is driven along a set of geometrically diverse roads—measured using various diversity measures (DMs)—it will react in a wide range of behaviours, thereby increasing the chances of observing failures, or strengthening the confidence in its safety, if no failures are observed. However, this assumption has never been tested before, nor have road DMs been assessed for their properties. Objective Our goal was to perform an exploratory study on 53 currently used and new, potentially promising road DMs. Specifically, our research questions looked into the road DMs themselves, to analyse their properties (e.g. monotonicity , computation efficiency ), and to test correlation between DMs. Furthermore, we investigated the use of road DMs to determine whether the assumption that diverse test suites of roads expose diverse driving behaviour holds. Method Our empirical analysis relies on a state-of-the-art, open-source ADS testing infrastructure and uses a data set containing over 97,000 individual road geometries and matching simulation data that were collected using two driving agents. By considering test suites of various sizes and measuring their roads’ geometric diversity, we studied road DM properties, the correlation between road DMs, and the correlation between road DMs and the observed behaviour. Results Our findings reveal a strong correlation between road diversity and behavioural diversity, confirming that geometrically diverse test suites systematically exercise diverse driving behaviours. We identified and aggregations as most effective, with achieving the strongest correlation of 0.95 while requiring minimal computation time. The analysed measures maintain robust correlation with behavioural diversity across test suites containing roads of varying lengths, eliminating the need for length normalisation. Conclusions These results empirically validate the fundamental assumption underlying diversity-driven ADS testing: road geometry diversity serves as a reliable proxy for behavioural diversity. For practitioners, we recommend or as optimal choices, whilst -based measures should be avoided entirely. The near-identical correlation patterns observed across architecturally different driving agents indicate that our findings generalise beyond specific ADS implementations, providing a solid foundation for diversity-driven test generation and selection. Stefan Klikovits, Vincenzo Riccio, Ezequiel Castellano, Ahmet Cetinkaya, Alessio Gambi, Paolo Arcaini |
Empir. Softw. Eng. | 2 |
| 2026 | DeepNaqqal: Human-Aligned Automated Validation of Test Inputs for Deep LearningabstractTest input generators (TIGs) are widely used to assess the robustness of Deep Learning (DL) image classifiers, yet they often produce invalid inputs that fall outside the semantic domain of the task, misleading quality assessment. While several automated validators have been proposed, there is a critical mismatch between automated and human validation criteria and, thus, automated validators are merely a proxy of domain validity, as perceived by human testers. We introduce DeepNaqqal, a supervised test input validator that learns validity directly from human-annotated labels using transfer learning on deep vision models. Our empirical study on automated validation of misclassification-inducing inputs compares DeepNaqqal against six state-of-the-art validators across three image classification tasks and multiple TIG families, using independent human assessment as ground truth. Our results show that DeepNaqqal consistently achieves the highest agreement with human judgments, while generalizing to unseen TIGs and remaining effective with substantially reduced labeled data. Maryam, Matteo Biagiola, Paolo Tonella, Vincenzo Riccio |
ICST | 4 |
| 2026 | XMutant: XAI-based fuzzing for deep learning systemsabstractSemantic-based test generators are widely used to produce failure-inducing inputs for Deep Learning (DL) systems. They typically generate challenging test inputs by applying random perturbations to input semantic concepts until a failure is found or a timeout is reached. However, such randomness may hinder them from efficiently achieving their goal. This paper proposes XMutant, a technique that leverages explainable artificial intelligence (XAI) techniques to generate challenging test inputs. XMutant uses the local explanation of the input to inform the fuzz testing process and effectively guide it toward failures of the DL system under test. We evaluated different configurations of XMutant in triggering failures for different DL systems both for model-level (sentiment analysis, digit recognition) and system-level testing (advanced driving assistance). Our studies showed that XMutant enables more effective and efficient test generation by focusing on the most impactful parts of the input. XMutant generates up to $$125\%$$ more failure-inducing inputs compared to an existing baseline, up to 7 $$\times$$ faster. We also assessed the validity of these inputs, maintaining a validation rate above $$89\%$$ , according to automated and human validators. Xingcheng Chen, Matteo Biagiola, Vincenzo Riccio, Marcelo d'Amorim, Andrea Stocco 0001 |
Empir. Softw. Eng. | 3 |
| 2026 | GIFTbench: Generative image fuzz testing benchmarkabstractGIFTbench is a modular framework for testing Deep Learning image classifiers that combines Generative AI with genetic algorithms. Its architecture integrates pretrained generative models with a user-friendly Gradio interface, enabling automated, reproducible, and interpretable robustness testing. Supporting VAE, GAN, and Diffusion models, GIFTbench generates test inputs by perturbing latent representations to expose misbehaviors of the classifier under test. By automating test input generation and reducing the need for manual coding, GIFTbench accelerates experimentation and facilitates comparative evaluation of both classifiers and generative models. Designed for researchers and practitioners, it enables reproducible assessment of image classifiers, while supporting studies on classifier vulnerabilities, mutation strategies, and the role of generative models in robustness testing. Maryam, Matteo Biagiola, Andrea Stocco 0001, Vincenzo Riccio |
Sci. Comput. Program. | 4 |
| 2025 | Benchmarking Generative AI Models for Deep Learning Test Input GenerationabstractTest Input Generators (TIGs) are crucial to assess the ability of Deep Learning (DL) image classifiers to provide correct predictions for inputs beyond their training and test sets. Recent advancements in Generative AI(GenAI) models have made them a powerful tool for creating and manipulating synthetic images, although these advancements also imply increased complexity and resource demands for training. In this work, we benchmark and combine different GenAI models with TIGs, assessing their effectiveness, efficiency, and quality of the generated test images, in terms of domain validity and label preservation. We conduct an empirical study involving three different GenAI architectures (VAEs, GANs, Diffusion Models), five classification tasks of increasing complexity, and 364 human evaluations. Our results show that simpler architectures, such as VAEs, are sufficient for less complex datasets like MNIST. However, when dealing with feature-rich datasets, such as ImageNet, more sophisticated architectures like Diffusion Models achieve superior performance by generating a higher number of valid, misclassification-inducing inputs. Maryam, Matteo Biagiola, Andrea Stocco 0001, Vincenzo Riccio |
ICST | 4 |
| 2025 | An industrial experience report on applying search-based boundary input generation to cyber-physical systems
Vincenzo Riccio, Aitor Arrieta, Paolo Tonella, Maite Arratibel |
Empir. Softw. Eng. | 2 |
| 2024 | Two is better than one: digital siblings to improve autonomous driving testingabstractAbstract Simulation-based testing represents an important step to ensure the reliability of autonomous driving software. In practice, when companies rely on third-party general-purpose simulators, either for in-house or outsourced testing, the generalizability of testing results to real autonomous vehicles is at stake. In this paper, we enhance simulation-based testing by introducing the notion ofdigital siblings—a multi-simulator approach that tests a given autonomous vehicle on multiple general-purpose simulators built with different technologies, that operate collectively as an ensemble in the testing process. We exemplify our approach on a case study focused on testing the lane-keeping component of an autonomous vehicle. We use two open-source simulators as digital siblings, and we empirically compare such a multi-simulator approach against a digital twin of a physical scaled autonomous vehicle on a large set of test cases. Our approach requires generating and running test cases for each individual simulator, in the form of sequences of road points. Then, test cases are migrated between simulators, using feature maps to characterize the exercised driving conditions. Finally, the joint predicted failure probability is computed, and a failure is reported only in cases of agreement among the siblings. Our empirical evaluation shows that the ensemble failure predictor by the digital siblings is superior to each individual simulator at predicting the failures of the digital twin. We discuss the findings of our case study and detail how our approach can help researchers interested in automated testing of autonomous driving software. Matteo Biagiola, Andrea Stocco 0001, Vincenzo Riccio, Paolo Tonella |
Empir. Softw. Eng. | 3 |
| 2024 | Focused Test Generation for Autonomous Driving SystemsabstractTesting Autonomous Driving Systems (ADSs) is crucial to ensure their reliability when navigating complex environments. ADSs may exhibit unexpected behaviours when presented, during operation, with driving scenarios containing features inadequately represented in the training dataset. To address this shift from development to operation, developers must acquire new data with the newly observed features. This data can be then utilised to fine tune the ADS, so as to reach the desired level of reliability in performing driving tasks. However, the resource-intensive nature of testing ADSs requires efficient methodologies for generating targeted and diverse tests. In this work, we introduce a novel approach, DeepAtash-LR , that incorporates a surrogate model into the focused test generation process. This integration significantly improves focused testing effectiveness and applicability in resource-intensive scenarios. Experimental results show that the integration of the surrogate model is fundamental to the success of DeepAtash-LR . Our approach was able to generate an average of up to 60× more targeted, failure-inducing inputs compared to the baseline approach. Moreover, the inputs generated by DeepAtash-LR were useful to significantly improve the quality of the original ADS through fine tuning. Tahereh Zohdinasab, Vincenzo Riccio, Paolo Tonella |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2023 | An Empirical Study on Low- and High-Level Explanations of Deep Learning MisbehavioursabstractBackground: Most quality assessment approaches for Deep Learning (DL) focus on finding misbehaviour-inducing inputs. However, it is difficult to clearly understand the causes of misbehaviours, due to the DL software opaqueness. Recent research proposed different techniques to explain DL misbehaviours, producing input explanations either at a “low level” (raw input elements) or at a “high level” (input features). Aims: We aim to compare the similarity between different explanations and assess to what extent they are understandable. Method: We have conducted an empirical study involving 3 state-of-the-art techniques for DL explanation in 13 configurations, applied to 2 different DL tasks. We have also collected answers from 48 questionnaires submitted to SE experts. Results: Low- and high-level techniques provide dissimilar explanations for the same inputs. However, experts deemed none of the explanations as useful in 28% of the cases. Conclusion: Despite the complementarity of existing explanations, further research is needed to produce better explanations. Tahereh Zohdinasab, Vincenzo Riccio, Paolo Tonella |
ESEM | 2 |
| 2023 | When and Why Test Generators for Deep Learning Produce Invalid Inputs: an Empirical StudyabstractTesting Deep Learning (DL) based systems inherently requires large and representative test sets to evaluate whether DL systems generalise beyond their training datasets. Diverse Test Input Generators (TIGs) have been proposed to produce artificial inputs that expose issues of the DL systems by triggering misbehaviours. Unfortunately, such generated inputs may be invalid, i.e., not recognisable as part of the input domain, thus providing an unreliable quality assessment. Automated validators can ease the burden of manually checking the validity of inputs for human testers, although input validity is a concept difficult to formalise and, thus, automate. In this paper, we investigate to what extent TIGs can generate valid inputs, according to both automated and human validators. We conduct a large empirical study, involving 2 different automated validators, 220 human assessors, 5 different TIGs and 3 classification tasks. Our results show that 84% artificially generated inputs are valid, according to automated validators, but their expected label is not always preserved. Automated validators reach a good consensus with humans (78% accuracy), but still have limitations when dealing with feature-rich datasets. Vincenzo Riccio, Paolo Tonella |
ICSE | 1 |
| 2023 | Engineering Self-adaptive Microservice Applications: An Experience Report
Vincenzo Riccio, Giancarlo Sorrentino, Matteo Camilli, Raffaela Mirandola, Patrizia Scandurra |
ICSOC (1) | 1 |
| 2023 | DeepAtash: Focused Test Generation for Deep Learning SystemsabstractWhen deployed in the operation environment, Deep Learning (DL) systems often experience the so-called development to operation (dev2op) data shift, which causes a lower prediction accuracy on field data as compared to the one measured on the test set during development. To address the dev2op shift, developers must obtain new data with the newly observed features, as these are under-represented in the train/test set, and must use them to fine tune the DL model, so as to reach the desired accuracy level. In this paper, we address the issue of acquiring new data with the specific features observed in operation, which caused a dev2op shift, by proposing DeepAtash, a novel search-based focused testing approach for DL systems. DeepAtash targets a cell in the feature space, defined as a combination of feature ranges, to generate misbehaviour-inducing inputs with predefined features. Experimental results show that DeepAtash was able to generate up to 29X more targeted, failure-inducing inputs than the baseline approach. The inputs generated by DeepAtash were useful to significantly improve the quality of the original DL systems through fine tuning not only on data with the targeted features, but quite surprisingly also on inputs drawn from the original distribution. Tahereh Zohdinasab, Vincenzo Riccio, Paolo Tonella |
ISSTA | 2 |
| 2023 | Software testing in the machine learning era
Andrea Stocco 0001, Onn Shehory, Gunel Jahangirova, Vincenzo Riccio, Guy Barash, Eitan Farchi, Diptikalyan Saha |
Empir. Softw. Eng. | 4 |
| 2023 | Efficient and Effective Feature Space Exploration for Testing Deep Learning SystemsabstractAssessing the quality of Deep Learning (DL) systems is crucial, as they are increasingly adopted in safety-critical domains. Researchers have proposed several input generation techniques for DL systems. While such techniques can expose failures, they do not explain which features of the test inputs influenced the system’s (mis-) behaviour. DeepHyperion was the first test generator to overcome this limitation by exploring the DL systems’ feature space at large. In this article, we propose DeepHyperion-CS , a test generator for DL systems that enhances DeepHyperion by promoting the inputs that contributed more to feature space exploration during the previous search iterations. We performed an empirical study involving two different test subjects (i.e., a digit classifier and a lane-keeping system for self-driving cars). Our results proved that the contribution-based guidance implemented within DeepHyperion-CS outperforms state-of-the-art tools and significantly improves the efficiency and the effectiveness of DeepHyperion . DeepHyperion-CS exposed significantly more misbehaviours for five out of six feature combinations and was up to 65% more efficient than DeepHyperion in finding misbehaviour-inducing inputs and exploring the feature space. DeepHyperion-CS was useful for expanding the datasets used to train the DL systems, populating up to 200% more feature map cells than the original training set. Tahereh Zohdinasab, Vincenzo Riccio, Alessio Gambi, Paolo Tonella |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2021 | DeepHyperion: exploring the feature space of deep learning-based systems through illumination searchabstractDeep Learning (DL) has been successfully applied to a wide range of application domains, including safety-critical ones. Several DL testing approaches have been recently proposed in the literature but none of them aims to assess how different interpretable features of the generated inputs affect the system's behaviour. Tahereh Zohdinasab, Vincenzo Riccio, Alessio Gambi, Paolo Tonella |
ISSTA | 2 |
| 2021 | DeepMetis: Augmenting a Deep Learning Test Set to Increase its Mutation ScoreabstractDeep Learning (DL) components are routinely integrated into software systems that need to perform complex tasks such as image or natural language processing. The adequacy of the test data used to test such systems can be assessed by their ability to expose artificially injected faults (mutations) that simulate real DL faults.In this paper, we describe an approach to automatically generate new test inputs that can be used to augment the existing test set so that its capability to detect DL mutations increases. Our tool DeepMetis implements a search based input generation strategy. To account for the non-determinism of the training and the mutation processes, our fitness function involves multiple instances of the DL model under test. Experimental results show that DeepMetis is effective at augmenting the given test set, increasing its capability to detect mutants by 63% on average. A leave-one-out experiment shows that the augmented test set is capable of exposing unseen mutants, which simulate the occurrence of yet undetected faults. Vincenzo Riccio, Nargiz Humbatova, Gunel Jahangirova, Paolo Tonella |
ASE | 1 |
| 2020 | Taxonomy of real faults in deep learning systemsabstractThe growing application of deep neural networks in safety-critical domains makes the analysis of faults that occur in such systems of enormous importance. In this paper we introduce a large taxonomy of faults in deep learning (DL) systems. We have manually analysed 1059 artefacts gathered from GitHub commits and issues of projects that use the most popular DL frameworks (TensorFlow, Keras and PyTorch) and from related Stack Overflow posts. Structured interviews with 20 researchers and practitioners describing the problems they have encountered in their experience have enriched our taxonomy with a variety of additional faults that did not emerge from the other two sources. Our final taxonomy was validated with a survey involving an additional set of 21 developers, confirming that almost all fault categories (13/15) were experienced by at least 50% of the survey participants. Nargiz Humbatova, Gunel Jahangirova, Gabriele Bavota, Vincenzo Riccio, Andrea Stocco 0001, Paolo Tonella |
ICSE | 4 |
| 2020 | Model-based exploration of the frontier of behaviours for deep learning system testingabstractWith the increasing adoption of Deep Learning (DL) for critical tasks, such as autonomous driving, the evaluation of the quality of systems that rely on DL has become crucial. Once trained, DL systems produce an output for any arbitrary numeric vector provided as input, regardless of whether it is within or outside the validity domain of the system under test. Hence, the quality of such systems is determined by the intersection between their validity domain and the regions where their outputs exhibit a misbehaviour. Vincenzo Riccio, Paolo Tonella |
ESEC/SIGSOFT FSE | 1 |
| 2020 | Testing machine learning based systems: a systematic mappingabstractAbstract Context: A Machine Learning based System (MLS) is a software system including one or more components that learn how to perform a task from a given data set. The increasing adoption of MLSs in safety critical domains such as autonomous driving, healthcare, and finance has fostered much attention towards the quality assurance of such systems. Despite the advances in software testing, MLSs bring novel and unprecedented challenges, since their behaviour is defined jointly by the code that implements them and the data used for training them. Objective: To identify the existing solutions for functional testing of MLSs, and classify them from three different perspectives: (1) the context of the problem they address, (2) their features, and (3) their empirical evaluation. To report demographic information about the ongoing research. To identify open challenges for future research. Method: We conducted a systematic mapping study about testing techniques for MLSs driven by 33 research questions. We followed existing guidelines when defining our research protocol so as to increase the repeatability and reliability of our results. Results: We identified 70 relevant primary studies, mostly published in the last years. We identified 11 problems addressed in the literature. We investigated multiple aspects of the testing approaches, such as the used/proposed adequacy criteria, the algorithms for test input generation, and the test oracles. Conclusions: The most active research areas in MLS testing address automated scenario/input generation and test oracle creation. MLS testing is a rapidly growing and developing research area, with many open challenges, such as the generation of realistic inputs and the definition of reliable evaluation metrics and benchmarks. Vincenzo Riccio, Gunel Jahangirova, Andrea Stocco 0001, Nargiz Humbatova, Michael Weiss 0004, Paolo Tonella |
Empir. Softw. Eng. | 1 |
| 2019 | Combining Automated GUI Exploration of Android apps with Capture and Replay through Machine Learning
Domenico Amalfitano, Vincenzo Riccio, Nicola Amatucci, Vincenzo De Simone, Anna Rita Fasolino |
Inf. Softw. Technol. | 2 |
| 2018 | Why does the orientation change mess up my Android application? From GUI failures to code faultsabstractSummary This paper investigates the failures exposed in mobile apps by the mobile‐specific event of changing the screen orientation. We focus on GUI failures resulting in unexpected GUI states that should be avoided to improve the apps quality and to ensure better user experience. We propose a classification framework that distinguishes 3 main classes of GUI failures due to orientation changes and exploit it in 2 studies that investigate the impact of such failures in Android apps. The studies involved both open‐source and apps from Google Play that were specifically tested exposing them to orientation change events. The results showed that more than 88% of these apps were affected by GUI failures, some classes of GUI failures were more common than others, and some GUI objects were more frequently involved. The app source code analysis allowed us to identify 6 classes of common faults causing specific GUI failures. Domenico Amalfitano, Vincenzo Riccio, Ana C. R. Paiva, Anna Rita Fasolino |
Softw. Test. Verification Reliab. | 2 |