VLDB 2026 Research / reviewers in the wild / expert
Valerio Terragni
dblp:167/0208
· DBLP profile ↗
38ranked-venue papers
12as first author
27since 2021 · last 2025
0000-0001-5885-9297ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 34 · 12 first-author · 25 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Metamorphic Testing of Large Language Models for Natural Language ProcessingabstractUsing Large Language Models (LLMs) to perform Natural Language Processing (NLP) tasks has been becoming increasingly pervasive in recent times. The versatile nature of LLMs makes them applicable to a wide range of such tasks. While the performance of recent LLMs is generally outstanding, several studies have shown that LLMs can often produce incorrect results. Automatically identifying these faulty behaviors is extremely useful for improving the effectiveness of LLMs. One obstacle to this is the limited availability of labeled datasets, necessitating an oracle to determine the correctness of LLM behaviors. Metamorphic Testing (MT) is a popular testing approach that alleviates this oracle problem. At the core of MT are Metamorphic Relations (MRs), defining the relationship between the outputs of related inputs. MT can expose faulty behaviors without the need for explicit oracles (e.g., labeled datasets). This paper presents the most comprehensive study of MT for LLMs to date. We conducted a literature review and collected 191 MRs for NLP tasks. We implemented a representative subset (36 MRs) to conduct a series of experiments with three popular LLMs, running$\sim 560 ~\mathrm{K}$metamorphic tests. The results shed light on the capabilities and opportunities of MT for LLMs, as well as its limitations. Steven Cho, Stefano Ruberto, Valerio Terragni |
ICSME | 3 |
| 2025 | Adoption and Evolution of Code Style and Best Programming Practices in Open-Source ProjectsabstractFollowing code style conventions in software projects is essential for maintaining overall code quality. Adhering to these conventions improves maintainability, understandability, and extensibility. Additionally, following best practices during software development enhances performance and reduces the likelihood of errors. This paper analyzes 1, 036 popular open-source JAVA projects on GitHub to study how code style and programming practices are adopted and evolve over time, examining their prevalence and the most common violations. Additionally, we study a subset of active repositories on a monthly basis to track changes in adherence to coding standards over time. We found widespread violations across repositories, with Javadoc and Naming violations being the most common. We also found a significant number of violations of the Google Java Style Guide in categories often missed by modern static analysis tools. Furthermore, repositories claiming to follow code-style practices exhibited slightly higher overall adherence to code-style and best-practices. The results provide valuable insights into the adoption of code style and programming practices, highlighting key areas for improvement in the open-source development community. Furthermore, the paper identifies important lessons learned and suggests future directions for improving code quality in JAVA projects. Alvari Kupari, Nasser Giacaman, Valerio Terragni |
ICSME | 3 |
| 2025 | LLMLOOP: Improving LLM-Generated Code and Tests Through Automated Iterative Feedback LoopsabstractLarge Language Models (LLMs) are showing remarkable performance in generating source code, yet the generated code often has issues like compilation errors or incorrect code. Researchers and developers often face wasted effort in implementing checks and refining LLM-generated code, frequently duplicating their efforts. This paper presents LLMLOOP, a framework that automates the refinement of both source code and test cases produced by LLMs. LLMLOOP employs five iterative loops: resolving compilation errors, addressing static analysis issues, fixing test case failures, and improving test quality through mutation analysis. These loops ensure the generation of high-quality test cases that serve as both a validation mechanism and a regression test suite for the generated code. We evaluated llmloop on HumanEval-X, a recent benchmark of programming tasks. Results demonstrate the tool effectiveness in refining LLM-generated outputs. A demonstration video of the tool is available at https://youtu.be/2CLG9x1fsNI. Ravin Ravi, Dylan Bradshaw, Stefano Ruberto, Gunel Jahangirova, Valerio Terragni |
ICSME | 5 |
| 2025 | Towards Cross-Build Differential TestingabstractRecent concerns about software supply chain security have led to the emergence of different binaries built from the same source code. This will sometimes result in binaries that are not identical and therefore have different cryptographic hashes. The question arises whether those binaries are still equivalent, i.e., whether they have the same behaviour. We explore whether differential testing can be used to provide evidence for non-equivalence. We study this for 3,541 pairs of binaries built for the same Maven artifact version, distributed on Maven Central, Google Assured Open Source Software and/or Oracle Build-From-Source. We use EVOSUITE to generate tests for the baseline binary from Maven Central, run these tests against this baseline binary and any available alternately built binaries, and compare the results for consistency. We argue that any differences may indicate variations in program behaviour and could, therefore, be used to detect compromised binaries or failures at runtime. Although our preliminary experiments did not reveal any compromised builds, our approach successfully identified three build configuration errors that caused changes in runtime behaviour. These findings underscore the potential of our method to uncover subtle build differences, highlighting opportunities for improvement. Jens Dietrich 0001, Tim White, Valerio Terragni, Behnaz Hassanshahi |
ICST | 3 |
| 2025 | Differential Testing of Concurrent ClassesabstractConcurrent programs are pervasive, yet difficult to write. The inherent complexity of thread synchronization makes the evolution of concurrent programs prone to concurrency faults. Previous work on regression testing concurrent programs focused on reducing the cost of re-run the existing tests. However, existing tests may not be able to expose the regression faults in the modified program. In this paper, we present Condiff a differential testing technique that generates concurrent tests and oracles to expose behavioral differences between two versions of a given concurrent class. Since concurrent programs are non-deterministic, this involves exploring all possible non-deterministic thread interleavings of each generated test on both versions. However, we can afford to analyze only a few concurrent tests due to the high cost of exhaustive interleaving exploration. To address the challenge, Condiff leverages the information of code changes and trace analysis to analyze only those concurrent tests that are likely to expose behavioral differences (if they exist). We evaluated Condiff on a set of Java classes. Our results show that Condiff can effectively generate concurrent tests that expose behavioral differences. Valerio Terragni, Shing-Chi Cheung |
ICST | 1 |
| 2025 | A System-Level Testing Framework for Automated Assessment of Programming Assignments Allowing Students Object-Oriented Design FreedomabstractAutomated assessment of programming assignments is essential in software engineering education, especially for large classes where manual grading is impractical. While static analysis can evaluate code style and syntax correctness, it cannot assess the functional correctness of students' implementations. Dynamic analysis through software testing can verify program behavior and provide automated feedback to students. However, traditional unit and integration tests often restrict students' design freedom by requiring predefined interfaces and method declarations. In this paper, we present SYSCLI, a novel testing framework for system-level testing of JAVA-based command-Line interface applications. SYSCLI enables test suites that evaluate the functional correctness of students' implementations without limiting their design choices. We also share our experience using SYSCLI in a second-year programming course at the University of Auckland, which focuses on object-oriented programming and design patterns and enrolls over 300 students each offering. Analysis of student assignments from 2023 and 2024 shows that SYSCLI is effective in automating grading, allows software design flexibility, and provides actionable feedback to students. Our experience report offers valuable insights into assessing students' implementation of object-oriented concepts and design patterns. Valerio Terragni, Nasser Giacaman |
ICST | 1 |
| 2025 | MDPMorph: An MDP-Based Metamorphic Testing Framework for Deep Reinforcement Learning AgentsabstractDeep Reinforcement Learning (DRL) systems are widely used across various domains. However, testing these systems presents significant challenges. The DRL agent, which serves as the core decision-maker, generates continuous value estimates rather than discrete labels and operates under nonstationary policies within complex and stochastic environments. Consequently, there is no definitive “correct answer” for each state-action pair, complicating automated test generation due to the known oracle problem. To address this challenge, we propose a Metamorphic Testing (MT) framework (MDPMORPH) specifically designed for validating DRL agents. Our framework is based on Markov Decision Processes (MDP) and focuses on the core reasoning properties of agents to automatically uncover potential faults. To support MDPMORPH, we introduce a Metamorphic Relation (MR) design methodology tailored for DRL agents, based on the temporal characteristics of MDP. Using this method, we define nine generic MRs that encapsulate common and expected properties of an agent’s reasoning process. Furthermore, based on established assumptions and definitions within the context of MDPs, we theoretically demonstrate the soundness of these MRs. Finally, we specialize these generic MRs into environment-specific MRs by determining appropriate thresholds through training on three classic DRL environments. Our experimental results demonstrate that MDPMORPH and the proposed MRs are highly effective in automatically detecting mutants within these studied environments, with a 0.84 average mutation detection rate. Yuning Xing, Daixu Ren, Steven Cho, Valerio Terragni |
ISSRE | 6 |
| 2025 | LLMorph: Automated Metamorphic Testing of Large Language ModelsabstractAutomated testing is essential for evaluating and improving the reliability of Large Language Models (LLMs), yet the lack of automated oracles for verifying output correctness remains a key challenge. We present LLMorph, an automated testing tool specifically designed for LLMs performing NLP tasks, which leverages Metamorphic Testing (MT) to uncover faulty behaviors without relying on human-labeled data. MT uses Metamorphic Relations (MRs) to generate follow-up inputs from source test input, enabling detection of inconsistencies in model outputs without the need of expensive labelled data. LLMorph is aimed at researchers and developers who want to evaluate the robustness of LLM-based NLP systems. In this paper, we detail the design, implementation, and practical usage of LLMorph, demonstrating how it can be easily extended to any LLM, NLP task, and set of MRs. In our evaluation, we applied 36 MRs across four NLP benchmarks, testing three state-of-the-art LLMs: GPT-4, LLAMA3, and HERMES 2. This produced over 561,000 test executions. The results demonstrate LLMorph’s effectiveness in automatically exposing incorrect model behaviors at scale.The tool source code is available at https://github.com/steven-b-cho/llmorph. A screencast demo is available at https://youtu.be/sHmqdieCfw4. Steven Cho, Stefano Ruberto, Valerio Terragni |
ASE | 3 |
| 2025 | Metamorphic Testing of Deep Reinforcement Learning Agents with MDPMorphabstractWe present MDPMorph, a tool for metamorphic testing of Deep Reinforcement Learning (DRL) agents. MDPMorph is based on the Markov Decision Process (MDP) and targets the core reasoning properties of DRL agents to automatically uncover potential faults. It can generate metamorphic test suites and corresponding mutants directly from the DRL system under test. MDPMorph uses a subset of the metamorphic test suite and models to train the thresholds of the nine proposed Metamorphic Relations (MRs) using stochastic gradient descent. These MRs are based on the temporal characteristics of the MDP, and the training aims to determine the optimal threshold for each MR. After obtaining the optimal threshold, MDPMorph leverages the MRs to compare the execution results of different metamorphic test suites on the model under test and reports whether each test passes or fails. Finally, by collecting the execution results, MDPMorph calculates the mutant detection rate of MR to validate its effectiveness. Experimental results show that MDPMorph and the proposed MRs are highly effective in automatically detecting seeded faults (mutants). Yuning Xing, Daixu Ren, Steven Cho, Valerio Terragni |
ASE | 6 |
| 2025 | LspFuzz: Hunting Bugs in Language ServersabstractThe Language Server Protocol (LSP) has revolutionized the integration of code intelligence in modern software development. There are approximately 300 LSP server implementations for various languages and 50 editors offering LSP integration. However, the reliability of LSP servers is a growing concern, as crashes can disable all code intelligence features and significantly impact productivity, while vulnerabilities can put developers at risk even when editing untrusted source code. Despite the widespread adoption of LSP, no existing techniques specifically target LSP server testing. To bridge this gap, we present LspFuzz, a grey-box hybrid fuzzer for systematic LSP server testing. Our key insight is that effective LSP server testing requires holistic mutation of source code and editor operations, as bugs often manifest from their combinations. To satisfy the sophisticated constraints of LSP and effectively explore the input space, we employ a two-stage mutation pipeline: syntax-aware mutations to source code, followed by context-aware dispatching of editor operations. We evaluated LspFuzz on four widely used LSP servers. LspFuzz demonstrated superior performance compared to baseline fuzzers, and uncovered previously unknown bugs in real-world LSP servers. Of the 51 bugs we reported, 42 have been confirmed, 26 have been fixed by developers, and two have been assigned CVE numbers. Our work advances the quality assurance of LSP servers, providing both a practical tool and foundational insights for future research in this domain. Hengcheng Zhu 0001, Songqiang Chen, Valerio Terragni, Lili Wei 0001, Yepang Liu 0001, Shing-Chi Cheung |
ASE | 3 |
| 2025 | An extended study of syntactic breaking changes in the wildabstractAbstract Libraries assist in accelerating the development of software applications by providing reusable functionalities. Libraries and applications that declare these libraries as dependencies become their clients. However, as libraries evolve, maintaining the dependencies in client projects can be challenging if the new version contains breaking changes. Yet, limited research focuses on analyzing the impact of breaking changes on client projects when updating dependencies in the wild. Hence, we conduct an empirical analysis using Java projects built using Maven to investigate the impact of breaking changes introduced between two library versions. Our dataset included 18,415 Maven artifacts, declaring 142,355 direct dependencies, out of which 71.60% were not up-to-date. We automatically updated these dependencies and discovered that 11.58% of the dependency updates resulted in breaking changes that affected the client, and almost half of them were introduced during a non-major update. We analyzed the changes in the libraries that contributed towards these breaking changes, and our results indicate that changes in transitive dependencies were a significant factor in introducing breaking changes. We further investigated if it was common for clients to use functionalities of transitive dependencies directly without declaring them. This showed that over half of the clients use transitive functionality. Therefore, we analyzed actions suggested to resolve these breaking changes introduced by transitive dependencies under the discussions on open-source platforms, and the frequently suggested action was to exclude the transitive dependency from the project configuration. Dhanushka Jayasuriya, Samuel Ou, Saakshi Hegde, Valerio Terragni, Jens Dietrich 0001, Kelly Blincoe |
Empir. Softw. Eng. | 4 |
| 2025 | The Future of AI-Driven Software EngineeringabstractA paradigm shift is underway in Software Engineering, with AI systems such as LLMs playing an increasingly important role in boosting software development productivity. This trend is anticipated to persist. In the next years, we expect a growing symbiotic partnership between human software developers and AI. The Software Engineering research community cannot afford to overlook this trend; we must address the key research challenges posed by the integration of AI into the software development process. In this article, we present our vision of the future of software development in an AI-driven world and explore the key challenges that our research community should address to realize this vision. Valerio Terragni, Annie Vella, Partha S. Roop, Kelly Blincoe |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2024 | MR-Adopt: Automatic Deduction of Input Transformation Function for Metamorphic TestingabstractWhile a recent study reveals that many developer-written test cases can encode a reusable Metamorphic Relation (MR), over 70% of them directly hard-code the source input and follow-up input in the encoded relation. Such encoded MRs, which do not contain an explicit input transformation to transform the source inputs to corresponding follow-up inputs, cannot be reused with new source inputs to enhance test adequacy. Congying Xu, Songqiang Chen, Shing-Chi Cheung, Valerio Terragni, Hengcheng Zhu 0001, Jialun Cao |
ASE | 5 |
| 2024 | Semantic matching in GUI test reuseabstractReusing test cases across apps that share similar functionalities reduces both the effort required to produce useful test cases and the time to offer reliable apps to the market. The main approaches to reuse test cases across apps combine different semantic matching and test generation algorithms to migrate test cases across Android apps. In this paper we define a general framework to evaluate the impact and effectiveness of different choices of semantic matching with Test Reuse approaches on migrating test cases across Android apps. We offer a thorough comparative evaluation of the many possible choices for the components of test migration processes. We propose an approach that combines the most effective choices for each component of the test migration process to obtain an effective approach. We report the results of an experimental evaluation on 8,099 GUI events from 337 test configurations. The results attest the prominent impact of semantic matching on test reuse. They indicate that sentence level perform better than word level embedding techniques. They surprisingly suggest a negligible impact of the corpus of documents used for building the word embedding model for the Semantic Matching Algorithm. They provide evidence that semantic matching of events of selected types perform better than semantic matching of events of all types. They show that the effectiveness of overall Test Reuse approach depends on the characteristics of the test suites and apps. The replication package that we make publicly available online (https://star.inf.usi.ch/#/software-data/11) allows researchers and practitioners to refine the results with additional experiments and evaluate other choices for test reuse components. Farideh Khalili, Leonardo Mariani, Ali Mohebbi 0003, Mauro Pezzè, Valerio Terragni |
Empir. Softw. Eng. | 5 |
| 2024 | MR-Scout: Automated Synthesis of Metamorphic Relations from Existing Test CasesabstractMetamorphic Testing (MT) alleviates the oracle problem by defining oracles based on metamorphic relations (MRs) that govern multiple related inputs and their outputs. However, designing MRs is challenging, as it requires domain-specific knowledge. This hinders the widespread adoption of MT. We observe that developer-written test cases can embed domain knowledge that encodes MRs. Such encoded MRs could be synthesized for testing not only their original programs but also other programs that share similar functionalities. In this article, we propose MR-Scout to automatically synthesize MRs from test cases in open-source software (OSS) projects. MR-Scout first discovers MR-encoded test cases (MTCs), and then synthesizes the encoded MRs into parameterized methods (called codified MRs ), and filters out MRs that demonstrate poor quality for new test case generation. MR-Scout discovered over 11,000 MTCs from 701 OSS projects. Experimental results show that over 97% of codified MRs are of high quality for automated test case generation, demonstrating the practical applicability of MR-Scout . Furthermore, codified-MRs-based tests effectively enhance the test adequacy of programs with developer-written tests, leading to 13.52% and 9.42% increases in line coverage and mutation score, respectively. Our qualitative study shows that 55.76% to 76.92% of codified MRs are easily comprehensible for developers. Congying Xu, Valerio Terragni, Hengcheng Zhu 0001, Shing-Chi Cheung |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2024 | StubCoder: Automated Generation and Repair of Stub Code for Mock ObjectsabstractMocking is an essential unit testing technique for isolating the class under test from its dependencies. Developers often leverage mocking frameworks to develop stub code that specifies the behaviors of mock objects. However, developing and maintaining stub code is labor-intensive and error-prone. In this article, we present StubCoder to automatically generate and repair stub code for regression testing. StubCoder implements a novel evolutionary algorithm that synthesizes test-passing stub code guided by the runtime behavior of test cases. We evaluated our proposed approach on 59 test cases from 13 open source projects. Our evaluation results show that StubCoder can effectively generate stub code for incomplete test cases without stub code and repair obsolete test cases with broken stub code. Hengcheng Zhu 0001, Lili Wei 0001, Valerio Terragni, Yepang Liu 0001, Shing-Chi Cheung, Qin Sheng, Lihong Song |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2024 | GenMorph: Automatically Generating Metamorphic Relations via Genetic ProgrammingabstractMetamorphic testing is a popular approach that aims to alleviate the oracle problem in software testing. At the core of this approach are Metamorphic Relations (MRs), specifying properties that hold among multiple test inputs and corresponding outputs. Deriving MRs is mostly a manual activity, since their automated generation is a challenging and largely unexplored problem. This paper presentsGenMorph, a technique to automatically generate MRs for Java methods that involve inputs and outputs that are boolean, numerical, or ordered sequences.GenMorphuses an evolutionary algorithm to search foreffectivetest oracles, i.e., oracles that trigger no false alarms and expose software faults in the method under test. The proposed search algorithm is guided by two fitness functions that measure the number of false alarms and the number of missed faults for the generated MRs. Our results show thatGenMorphgenerates effective MRs for 18 out of 23 methods (mutation score >20%). Furthermore, it can increaseRandoop’s fault detection capability in 7 out of 23 methods, andEvosuite’s in 14 out of 23 methods. When compared with AUTOMR, a state-of-the-art MR generator,GenMorphalso outperformed its fault detection capability in 9 out of 10 methods. Jon Ayerdi, Valerio Terragni, Gunel Jahangirova, Aitor Arrieta, Paolo Tonella |
IEEE Trans. Software Eng. | 2 |
| 2023 | Understanding Breaking Changes in the WildabstractModern software applications rely heavily on the usage of libraries, which provide reusable functionality, to accelerate the development process. As libraries evolve and release new versions, the software systems that depend on those libraries (the clients) should update their dependencies to use these new versions as the new release could, for example, include critical fixes for security vulnerabilities. However, updating is not always a smooth process, as it can result in software failures in the clients if the new version includes breaking changes. Yet, there is little research on how these breaking changes impact the client projects in the wild. To identify if changes between two library versions cause breaking changes at the client end, we perform an empirical study on Java projects built using Maven. For the analysis, we used 18,415 Maven artifacts, which declared 142,355 direct dependencies, of which 71.60% were not up-to-date. We updated these dependencies and found that 11.58% of the dependency updates contain breaking changes that impact the client. We further analyzed these changes in the library which impact the client projects and examine if libraries have adhered to the semantic versioning scheme when introducing breaking changes in their releases. Our results show that changes in transitive dependencies were a major factor in introducing breaking changes during dependency updates and almost half of the detected client impacting breaking changes violate the semantic versioning scheme by introducing breaking changes in non-Major updates. Dhanushka Jayasuriya, Valerio Terragni, Jens Dietrich 0001, Samuel Ou, Kelly Blincoe |
ISSTA | 2 |
| 2023 | Evolving a Programming CS2 Course: A Decade-Long Experience ReportabstractDespite instructors' best efforts in designing and delivering any given course, changes are likely required from time to time. This experience report presents the changes made in a second-year programming course for non-computing engineering majors over a decade's worth of effort, and the reasons behind those changes. The changes were often reactive--in response to student feedback. However, many other changes were inspired by the desire to trial new interventions in the hope of strengthening the students' positive experience. In addition to personnel and course content changes, the gradual evolvement included how labs, assignments, and activities were structured and executed. Teaching delivery evolved, along with a number of small-scale interventions that eventually became integral elements of the course. When COVID-19 demanded a sudden shift to online learning, the course was prepared to adapt quickly and successfully. The contributions here come in the form of lessons learned over the past decade: what worked, and what did not. We present the large range of changes---and their rationales--that are particularly relevant and applicable to programming courses targeting engineering students where the luxury of pedagogically-friendlier programming languages is not possible. Nasser Giacaman, Partha S. Roop, Valerio Terragni |
SIGCSE (1) | 3 |
| 2022 | The ineffectiveness of domain-specific word embedding models for GUI test reuseabstractReusing test cases across similar applications can significantly reduce testing effort. Some recent test reuse approaches successfully exploit word embedding models to semantically match GUI events across Android apps. It is a common understanding that word embedding models trained on domain-specific corpora perform better on specialized tasks. Our recent study confirms this understanding in the context of Android test reuse. It shows that word embedding models trained with a corpus of the English descriptions of apps in the Google Play Store lead to a better semantic matching of Android GUI events. Motivated by this result, we hypothesize that we can further increase the effectiveness of semantic matching by partitioning the corpus of app descriptions into domain-specific corpora. Our experiments do not confirm our hypothesis. This paper sheds light on this unexpected negative result that contradicts the common understanding. Farideh Khalili, Ali Mohebbi 0003, Valerio Terragni, Mauro Pezzè, Leonardo Mariani, Abbas Heydarnoori |
ICPC | 3 |
| 2022 | Detect, Fix, and Verify TensorFlow API MisusesabstractThe growing application of DL makes detecting and fixing defective DL programs of paramount importance. Recent studies on DL defects report that TensorFlow API misuses represent a common class of DL defects. However to effectively detect, fix, and verify them remains an understudied problem. This paper presents the TensorFlow API misuses Detector And Fixer (TADAF) technique, which relies on 11 common API misuses patterns and corresponding fixes that we extracted from StackOverftow. TADAF statically analyses a TensorFlow program for identifying matches of any of the 11 patterns. If it finds a match, it automatically generates a fixed version of the program. To verify that the misuse brings a tangible negative effect, TADAF reports functional, accuracy, or efficiency differences when training and testing (with the same data) the original and fixed versions of the program. Our preliminary evaluation on five GitHub projects shows that TADAF detected and fixed all the API misuses. Wilson Baker, Michael O'Connor, Seyed Reza Shahamiri, Valerio Terragni |
SANER | 4 |
| 2021 | An Evolutionary Approach to Adapt Tests Across Mobile AppsabstractAutomatic generators of GUI tests often fail to generate semantically relevant test cases, and thus miss important test scenarios. To address this issue, test adaptation techniques can be used to automatically generate semantically meaningful GUI tests from test cases of applications with similar functionalities.In this paper, we present ADAPTDROID, a technique that approaches the test adaptation problem as a search-problem, and uses evolutionary testing to adapt GUI tests (including oracles) across similar Android apps. In our evaluation with 32 popular Android apps, ADAPTDROID successfully adapted semantically relevant test cases in 11 out of 20 cross-app adaptation scenarios. Leonardo Mariani, Mauro Pezzè, Valerio Terragni, Daniele Zuddas |
AST | 3 |
| 2021 | Towards effective GP multi-class classification based on dynamic targetsabstractIn the multi-class classification problem GP plays an important role when combined with other non-GP classifiers. However, when GP performs the actual classification (without relying on other classifiers) its classification accuracy is low. This is especially true when the number of classes is high. In this paper, we present DTC, a GP classifier that leverages the effectiveness of the dynamic target approach to evolve a set of discriminant functions (one for each class). Notably, DTC is the first GP classifier that defines the fitness of individuals by using the synergistic combination of linear scaling and the hinge-loss function (commonly used by SVM). Differently, most previous GP classifiers use the number of correct classifications to drive the evolution. We compare DTC with eight state-of-art multi-class classification techniques (e.g., RF, RS, MLP, and SVM) on eight popular datasets. The results show that DTC achieves competitive classification accuracy even with 15 classes, without relying on other classifiers. Stefano Ruberto, Valerio Terragni, Jason H. Moore |
GECCO | 2 |
| 2021 | Semantic matching of GUI events for test reuse: are we there yet?abstractGUI testing is an important but expensive activity. Recently, research on test reuse approaches for Android applications produced interesting results. Test reuse approaches automatically migrate human-designed GUI tests from a source app to a target app that shares similar functionalities. They achieve this by exploiting semantic similarity among textual information of GUI widgets. Semantic matching of GUI events plays a crucial role in these approaches. In this paper, we present the first empirical study on semantic matching of GUI events. Our study involves 253 configurations of the semantic matching, 337 unique queries, and 8,099 distinct GUI events. We report several key findings that indicate how to improve semantic matching of test reuse approaches, propose SemFinder a novel semantic matching algorithm that outperforms existing solutions, and identify several interesting research directions. Leonardo Mariani, Ali Mohebbi 0003, Mauro Pezzè, Valerio Terragni |
ISSTA | 4 |
| 2021 | APIzation: Generating Reusable APIs from StackOverflow Code SnippetsabstractDeveloper forums like StackOverflow have become essential resources to modern software development practices. However, many code snippets lack a well-defined method declaration, and thus they are often incomplete for immediate reuse. Developers must adapt the retrieved code snippets by parameterizing the variables involved and identifying the return value. This activity, which we call APIzation of a code snippet, can be tedious and time-consuming. In this paper, we present APIZATOR to perform APIzations of JAVA code snippets automatically. APIZATOR is grounded by four common patterns that we extracted by studying real APIzations in GitHub. APIZATOR presents a static analysis algorithm that automatically extracts the method parameters and return statements. We evaluated APIZATOR with a ground-truth of 200 APIzations collected from 20 developers. For 113 (56.50 %) and 115 (57.50 %) APIzations, APIZATOR and the developers extracted identical parameters and return statements, respectively. For 163 (81.50 %) APIzations, either the parameters or the return statements were identical. Valerio Terragni, Pasquale Salza |
ASE | 1 |
| 2021 | Generating metamorphic relations for cyber-physical systems with genetic programming: an industrial case studyabstractOne of the major challenges in the verification of complex industrial Cyber-Physical Systems is the difficulty of determining whether a particular system output or behaviour is correct or not, the so-called test oracle problem. Metamorphic testing alleviates the oracle problem by reasoning on the relations that are expected to hold among multiple executions of the system under test, which are known as Metamorphic Relations (MRs). However, the development of effective MRs is often challenging and requires the involvement of domain experts. In this paper, we present a case study aiming at automating this process. To this end, we implemented GAssertMRs, a tool to automatically generate MRs with genetic programming. We assess the cost-effectiveness of this tool in the context of an industrial case study from the elevation domain. Our experimental results show that in most cases GAssertMRs outperforms the other baselines, including manually generated MRs developed with the help of domain experts. We then describe the lessons learned from our experiments and we outline the future work for the adoption of this technique by industrial practitioners. Jon Ayerdi, Valerio Terragni, Aitor Arrieta, Paolo Tonella, Goiuria Sagardui Mendieta, Maite Arratibel |
ESEC/SIGSOFT FSE | 2 |
| 2021 | Statically driven generation of concurrent tests for thread-safe classesabstractSummary Concurrency testing is an important activity to expose concurrency faults in thread‐safe classes. A concurrent test for a thread‐safe class is a set of method call sequences that exercise the public interface of the class from multiple threads. Automatically generating fault‐revealing concurrent tests within an affordable time budget is difficult due to the huge search space of possible concurrent tests. In this paper, we present DepCon+, a novel approach that reduces the search space of concurrent tests by leveraging statically computed dependencies among public methods. DepCon+ exploits the intuition that concurrent tests can expose thread‐safety violations that manifest exceptions or deadlocks, only if they exercise some specific method dependencies. DepCon+ provides an efficient way to identify such dependencies by statically analysing the code and relies on the computed dependencies to steer the test generation towards those concurrent tests that exhibit the computed dependencies. We developed a prototype DepCon+ implementation for Java and evaluated the approach on 19 known concurrency faults of thread‐safe classes that lead to thread‐safety violations of either exception or deadlock type. The results presented in this paper show that DepCon+ is more effective than state‐of‐the‐art approaches in exposing the concurrency faults. The search space pruning of DepCon+ dramatically reduces the search space of possible concurrent tests, without missing any thread‐safety violations. Valerio Terragni, Mauro Pezzè |
Softw. Test. Verification Reliab. | 1 |
| 2020 | SGP-DT: Semantic Genetic Programming Based on Dynamic Targets
Stefano Ruberto, Valerio Terragni, Jason H. Moore |
EuroGP | 2 |
| 2020 | Measuring Software Testability Modulo Test QualityabstractComprehending the degree to which software components support testing is important to accurately schedule testing activities, train developers, and plan effective refactoring actions. Software testability estimates such property by relating code characteristics to the test effort. The main studies of testability reported in the literature investigate the relation between class metrics and test effort in terms of the size and complexity of the associated test suites. They report a moderate correlation of some class metrics to test-effort metrics, but suffer from two main limitations: (i) the results hardly generalize due to the small empirical evidence (datasets with no more than eight software projects); and (ii) mostly ignore the quality of the tests. However, considering the quality of the tests is important. Indeed, a class may have a low test effort because the associated tests are of poor quality, and not because the class is easier to test. In this paper, we propose an approach to measure testability that normalizes the test effort with respect to the test quality, which we quantify in terms of code coverage and mutation score. We present the results of a set of experiments on a dataset of 9,861 Java classes, belonging to 1,186 open source projects, with around 1.5 million of lines of code overall. The results confirm that normalizing the test effort with respect to the test quality largely improves the correlation between class metrics and the test effort. Better correlations result in better prediction power and thus better prediction of the test effort. Valerio Terragni, Pasquale Salza, Mauro Pezzè |
ICPC | 1 |
| 2020 | Image Feature Learning with Genetic Programming
Stefano Ruberto, Valerio Terragni, Jason H. Moore |
PPSN (2) | 2 |
| 2020 | Evolutionary improvement of assertion oraclesabstractAssertion oracles are executable boolean expressions placed inside the program that should pass (return true) for all correct executions and fail (return false) for all incorrect executions. Because designing perfect assertion oracles is difficult, assertions often fail to distinguish between correct and incorrect executions. In other words, they are prone to false positives and false negatives. In this paper, we propose GAssert (Genetic ASSERTion improvement), the first technique to automatically improve assertion oracles. Given an assertion oracle and evidence of false positives and false negatives, GAssert implements a novel co-evolutionary algorithm that explores the space of possible assertions to identify one with fewer false positives and false negatives. Our empirical evaluation on 34 Java methods from 7 different Java code bases shows that GAssert effectively improves assertion oracles. GAssert outperforms two baselines (random and invariant-based oracle improvement), and is comparable with and in some cases even outperformed human-improved assertions. Valerio Terragni, Gunel Jahangirova, Paolo Tonella, Mauro Pezzè |
ESEC/SIGSOFT FSE | 1 |
| 2019 | Coverage-Driven Test Generation for Thread-Safe Classes via Parallel and Conflict DependenciesabstractThread-safe classes are common in concurrent object-oriented programs. Testing such classes is important to ensure the reliability of the concurrent programs that rely on them. Recently, researchers have proposed the automated generation of concurrent (multi-threaded) tests to expose concurrency faults in thread-safe classes (thread-safety violations). However, generating fault-revealing concurrent tests within an affordable time-budget is difficult due to the huge search space of possible concurrent tests. In this paper, we present DepCon, an approach to effectively reduce the search space of concurrent tests by means of both parallel and conflict dependency analyses. DepCon is based on the intuition that only methods that can both interleave (parallel dependent) and access the same shared memory locations (conflict dependent) can lead to thread-safety violations when concurrently executed. DepCon implements an efficient static analysis to compute the parallel and conflict dependencies among the methods of a class and uses the computed dependencies to steer the generation of tests towards concurrent tests that exhibit the computed dependencies. We evaluated DepCon by experimenting with a prototype implementation for Java programs on a set of thread-safe classes with known concurrency faults. The experimental results show that DepCon is more effective in exposing concurrency faults than state-of-the-art techniques. Valerio Terragni, Mauro Pezzè, Francesco A. Bianchi |
ICST | 1 |
| 2018 | Effectiveness and challenges in generating concurrent tests for thread-safe classesabstractDeveloping correct and efficient concurrent programs is difficult and error-prone, due to the complexity of thread synchronization. Often, developers alleviate such problem by relying on thread-safe classes, which encapsulate most synchronization-related challenges. Thus, testing such classes is crucial to ensure the reliability of the concurrency aspects of programs. Some recent techniques and corresponding tools tackle the problem of testing thread-safe classes by automatically generating concurrent tests. In this paper, we present a comprehensive study of the state-of-the-art techniques and an independent empirical evaluation of the publicly available tools. We conducted the study by executing all tools on the JaConTeBe benchmark that contains 47 well-documented concurrency faults. Our results show that 8 out of 47 faults (17%) were detected by at least one tool. By studying the issues of the tools and the generated tests, we derive insights to guide future research on improving the effectiveness of automated concurrent test generation. Valerio Terragni, Mauro Pezzè |
ASE | 1 |
| 2017 | Reproducing concurrency failures from crash stacksabstractReproducing field failures is the first essential step for understanding, localizing and removing faults. Reproducing concurrency field failures is hard due to the need of synthesizing a test code jointly with a thread interleaving that induce the failure in the presence of limited information from the field. Current techniques for reproducing concurrency failures focus on identifying failure-inducing interleavings, leaving largely open the problem of synthesizing the test code that manifests such interleavings. In this paper, we present ConCrash, a technique to automatically generate test codes that reproduce concurrency failures that violate thread-safety from crash stacks, which commonly summarize the conditions of field failures. ConCrash efficiently explores the huge space of possible test codes to identify a failure-inducing one by using a suitable set of search pruning strategies. Combined with existing techniques for exploring interleavings, ConCrash automatically reproduces a given concurrency failure that violates the thread-safety of a class by identifying both a failure-inducing test code and corresponding interleaving. In the paper, we define the ConCrash approach, present a prototype implementation of ConCrash, and discuss the experimental results that we obtained on a known set of ten field failures that witness the effectiveness of the approach. Francesco A. Bianchi, Mauro Pezzè, Valerio Terragni |
ESEC/SIGSOFT FSE | 3 |
| 2016 | Coverage-driven test code generation for concurrent classesabstractPrevious techniques on concurrency testing have mainly focused on exploring the interleaving space of manually written test code to expose faulty interleavings of shared memory accesses. These techniques assume the availability of failure-inducing tests. In this paper, we present AutoConTest, a coverage-driven approach to generate effective concurrent test code that achieve high interleaving coverage. AutoConTest consists of three components. First, it computes the coverage requirements dynamically and iteratively during sequential test code generation, using a coverage metric that captures the execution context of shared memory accesses. Second, it smartly selects these sequential codes based on the computed result and assembles them for concurrent tests, achieving increased context-sensitive interleaving coverage. Third, it explores the newly covered interleavings. We have implemented AutoConTest as an automated tool and evaluated it using 6 real-world concurrent Java subjects. The results show that AutoConTest is able to generate effective concurrent tests that achieve high interleaving coverage and expose concurrency faults quickly. AutoConTest took less than 65 seconds (including program analysis, test generation and execution) to expose the faults in the program subjects. Valerio Terragni, Shing-Chi Cheung |
ICSE | 1 |
| 2016 | CSNIPPEX: automated synthesis of compilable code snippets from Q&A sitesabstractPopular Q&A sites like StackOverflow have collected numerous code snippets. However, many of them do not have complete type information, making them uncompilable and inapplicable to various software engineering tasks. This paper analyzes this problem, and proposes a technique CSNIPPEX to automatically convert code snippets into compilable Java source code files by resolving external dependencies, generating import declarations, and fixing syntactic errors. We implemented CSNIPPEX as a plug-in for Eclipse and evaluated it with 242,175 StackOverflow posts that contain code snippets. CSNIPPEX successfully synthesized compilable Java files for 40,410 of them. It was also able to effectively recover import declarations for each post with a precision of 91.04% in a couple of seconds. Valerio Terragni, Yepang Liu 0001, Shing-Chi Cheung |
ISSTA | 1 |
| 2016 | Understanding and detecting wake lock misuses for Android applicationsabstractWake locks are widely used in Android apps to protect critical computations from being disrupted by device sleeping. Inappropriate use of wake locks often seriously impacts user experience. However, little is known on how wake locks are used in real-world Android apps and the impact of their misuses. To bridge the gap, we conducted a large-scale empirical study on 44,736 commercial and 31 open-source Android apps. By automated program analysis and manual investigation, we observed (1) common program points where wake locks are acquired and released, (2) 13 types of critical computational tasks that are often protected by wake locks, and (3) eight patterns of wake lock misuses that commonly cause functional and non-functional issues, only three of which had been studied by existing work. Based on our findings, we designed a static analysis technique, Elite, to detect two most common patterns of wake lock misuses. Our experiments on real-world subjects showed that Elite is effective and can outperform two state-of-the-art techniques. Yepang Liu 0001, Chang Xu 0001, Shing-Chi Cheung, Valerio Terragni |
SIGSOFT FSE | 4 |
| 2015 | RECONTEST: Effective Regression Testing of Concurrent ProgramsabstractConcurrent programs proliferate as multi-core technologies advance. The regression testing of concurrent programs often requires running a failing test for weeks before catching a faulty interleaving, due to the myriad of possible interleavings of memory accesses arising from concurrent program executions. As a result, the conventional approach that selects a sub-set of test cases for regression testing without considering interleavings is insufficient. In this paper we present RECONTEST to address the problem by selecting the new interleavings that arise due to code changes. These interleavings must be explored in order to uncover regression bugs. RECONTEST efficiently selects new interleavings by first identifying shared memory accesses that are affected by the changes, and then exploring only those problematic interleavings that contain at least one of these accesses. We have implemented RECONTEST as an automated tool and evaluated it using 13 real-world concurrent program subjects. Our results show that RECONTEST can significantly reduce the regression testing cost without missing any faulty interleavings induced by code changes. Valerio Terragni, Shing-Chi Cheung, Charles Zhang 0001 |
ICSE (1) | 1 |