VLDB 2026 Research / reviewers in the wild / expert
Pietro Liguori
dblp:230/2388
· DBLP profile ↗
21ranked-venue papers
4as first author
18since 2021 · last 2026
0000-0001-5579-1696ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 15 · 3 first-author · 13 since 2021Security and privacy · 4 · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Will It Break in Production? Metric-Driven Prediction of Residual Defects in Python Systems
Giuseppe De Rosa, Pietro Liguori |
DSN | 2 |
| 2026 | cAPTure dataset: How fast can you detect APT threats?abstractHigh-quality datasets are essential for machine learning-based intrusion detection systems, which are considered a promising defense for cyber-physical systems against Advanced Persistent Threats (APTs). However, existing datasets often are not built to capture the long, multi-stage nature of real APT campaigns, and they are labeled as general cyber-attacks rather than explicitly as APTs. To address this gap, we propose a methodology for creating a semi-synthetic, labeled dataset that reflects the complex attack paths typical of APTs targeting cyber-physical environments. Our approach integrates realistic network traffic gathered from a real testbed with multi-step APT attack scenarios modeled on the well-established MITRE ATT&CK framework and CVE exploits repository. The cAPTure dataset provides a rich basis for evaluating intrusion detection systems, enabling an evaluation methodology that relates false positive rate and time-to-detection, two metrics that are crucial for practical, real-world NIDS deployment. Tommaso Puccetti, Simona De Vivo, Davide Zhang, Pietro Liguori, Roberto Natella, Andrea Ceccarelli |
Comput. Networks | 4 |
| 2026 | Editorial of the special issue in the journal of systems and software on reliable and secure large language models for software engineering
Antonio Mastropaolo, Pietro Liguori, Gabriele Bavota |
J. Syst. Softw. | 2 |
| 2025 | Human-Written vs. AI-Generated Code: A Large-Scale Study of Defects, Vulnerabilities, and ComplexityabstractAs AI code assistants become increasingly integrated into software development workflows, understanding how their code compares to human-written programs is critical for ensuring reliability, maintainability, and security. In this paper, we present a large-scale comparison of code authored by human developers and three state-of-the-art LLMs, i.e., ChatGPT, DeepSeek-Coder, and Qwen-Coder, on multiple dimensions of software quality: code defects, security vulnerabilities, and structural complexity. Our evaluation spans over 500k code samples in two widely used languages, Python and Java, classifying defects via Orthogonal Defect Classification and security vulnerabilities using the Common Weakness Enumeration. We find that AI-generated code is generally simpler and more repetitive, yet more prone to unused constructs and hardcoded debugging, while humanwritten code exhibits greater structural complexity and a higher concentration of maintainability issues. Notably, AI-generated code also contains more high-risk security vulnerabilities. These findings highlight the distinct defect profiles of AI-and humanauthored code and underscore the need for specialized quality assurance practices in AI-assisted programming. Domenico Cotroneo, Cristina Improta, Pietro Liguori |
ISSRE | 3 |
| 2025 | Quality In, Quality Out: Investigating Training Data's Role in AI Code GenerationabstractDeep Learning (DL)-based code generators have seen significant advancements in recent years. Tools such as GitHub Copilot are used by thousands of developers with the main promise of a boost in productivity. However, researchers have recently questioned their impact on code quality showing, for example, that code generated by DL-based tools may be affected by security vulnerabilities. Since DL models are trained on large code corpora, one may conjecture that low-quality code they output is the result of low-quality code they have seen during training. However, there is very little empirical evidence documenting this phenomenon. Indeed, most of previous work look at the frequency with which commercial code generators (e.g., Copilot, ChatGPT) recommend low-quality code without the possibility of relating this to their (publicly unavailable) training set. In this paper, we investigate the extent to which low-quality code instances seen during training affect the quality of the code generated at inference time. We start by fine-tuning a pre-trained DL model on a large-scale dataset ($>4.4 \mathrm{M}$functions) being representative of those usually adopted in the training of code generators. We show that 4.98 % of functions in this dataset exhibit one or more quality issues related to security, maintainability, coding practices, etc. We use the fine-tuned model to generate 551 k Python functions, showing that 5.85 % of them are affected by at least one quality issue. We then remove from the training set the low-quality functions, and use the cleaned dataset to fine-tune a second model which has been used to generate the same 551k Python functions. We show that the model trained on the cleaned dataset exhibits similar performance in terms of functional correctness as compared to the original model (i.e., the one trained on the whole dataset) while, however, generating a statistically significant lower number of low-quality functions (2.16 %). Our study empirically documents the importance of high-quality training data for code generators. Cristina Improta, Rosalia Tufano, Pietro Liguori, Domenico Cotroneo, Gabriele Bavota |
ICPC | 3 |
| 2025 | Creation and Use of a Representative Dataset for Advanced Persistent Threats Detection
Tommaso Puccetti, Simona De Vivo, Davide Zhang, Pietro Liguori, Roberto Natella, Andrea Ceccarelli |
SAFECOMP | 4 |
| 2025 | Enhancing robustness of AI offensive code generators via data augmentation
Cristina Improta, Pietro Liguori, Roberto Natella, Bojan Cukic, Domenico Cotroneo |
Empir. Softw. Eng. | 2 |
| 2025 | DeVAIC: A tool for security assessment of AI-generated code
Domenico Cotroneo, Roberta De Luca, Pietro Liguori |
Inf. Softw. Technol. | 3 |
| 2025 | CGP-Tuning: Structure-Aware Soft Prompt Tuning for Code Vulnerability DetectionabstractLarge language models (LLMs) have been proposed as powerful tools for detecting software vulnerabilities, where task-specific fine-tuning is typically employed to provide vulnerability-specific knowledge to the LLMs. However, existing fine-tuning techniques often treat source code as plain text, losing the graph-based structural information inherent in code.Graph-enhanced soft prompt tuning addresses this by translating the structural information into contextual cues that the LLM can understand. However, current methods are primarily designed for general graph-related tasks and focus more on adjacency information, they fall short in preserving the rich semantic information (e.g., control/data flow) within code graphs. They also fail to ensure computational efficiency while capturing graph-text interactions in their cross-modal alignment module.This paper presents CGP-Tuning, a new code graph-enhanced, structure-aware soft prompt tuning method for vulnerability detection. CGP-Tuning introduces type-aware embeddings to capture the rich semantic information within code graphs, along with an efficient cross-modal alignment module that achieves linear computational costs while incorporating graph-text interactions. It is evaluated on the latestDiverseVuldataset and three advanced open-source code LLMs, CodeLlama, CodeGemma, and Qwen2.5-Coder. Experimental results show that CGP-Tuning delivers model-agnostic improvements and maintains practical inference speed, surpassing the best graph-enhanced soft prompt tuning baseline by an average of four percentage points and outperforming non-tuned zero-shot prompting by 15 percentage points. Ruijun Feng, Hammond A. Pearce, Pietro Liguori, Yulei Sui |
IEEE Trans. Software Eng. | 3 |
| 2024 | Enhancing AI-based Generation of Software Exploits with Contextual InformationabstractThis practical experience report explores Neural Machine Translation (NMT) models’ capability to generate offensive security code from natural language (NL) descriptions, highlighting the significance of contextual understanding and its impact on model performance. Our study employs a dataset comprising real shellcodes to evaluate the models across various scenarios, including missing information, necessary context, and unnecessary context. The experiments are designed to assess the models’ resilience against incomplete descriptions, their proficiency in leveraging context for enhanced accuracy, and their ability to discern irrelevant information. The findings reveal that the introduction of contextual data significantly improves performance. However, the benefits of additional context diminish beyond a certain point, indicating an optimal level of contextual information for model training. Moreover, the models demonstrate an ability to filter out unnecessary context, maintaining high levels of accuracy in the generation of offensive security code. This study paves the way for future research on optimizing context use in AI-driven code generation, particularly for applications requiring a high degree of technical precision such as the generation of offensive code. Pietro Liguori, Cristina Improta, Roberto Natella, Bojan Cukic, Domenico Cotroneo |
ISSRE | 1 |
| 2024 | Vulnerabilities in AI Code Generators: Exploring Targeted Data Poisoning AttacksabstractAI-based code generators have become pivotal in assisting developers in writing software starting from natural language (NL). However, they are trained on large amounts of data, often collected from unsanitized online sources (e.g., GitHub, HuggingFace). As a consequence, AI models become an easy target for data poisoning, i.e., an attack that injects malicious samples into the training data to generate vulnerable code. Domenico Cotroneo, Cristina Improta, Pietro Liguori, Roberto Natella |
ICPC | 3 |
| 2024 | Automating the correctness assessment of AI-generated code for security contextsabstractEvaluating the correctness of code generated by AI is a challenging open problem. In this paper, we propose a fully automated method, named ACCA, to evaluate the correctness of AI-generated code for security purposes. The method uses symbolic execution to assess whether the AI-generated code behaves as a reference implementation. We use ACCA to assess four state-of-the-art models trained to generate security-oriented assembly code and compare the results of the evaluation with different baseline solutions, including output similarity metrics, widely used in the field, and the well-known ChatGPT, the AI-powered language model developed by OpenAI. Our experiments show that our method outperforms the baseline solutions and assesses the correctness of the AI-generated code similar to the human-based evaluation, which is considered the ground truth for the assessment in the field. Moreover, ACCA has a very strong correlation with the human evaluation (Pearson’s correlation coefficient r=0.84 on average). Finally, since it is a full y automated solution that does not require any human intervention, the proposed method performs the assessment of every code snippet in ∼0.17 s on average, which is definitely lower than the average time required by human analysts to manually inspect the code, based on our experience. Domenico Cotroneo, Alessio Foggia, Cristina Improta, Pietro Liguori, Roberto Natella |
J. Syst. Softw. | 4 |
| 2023 | Who evaluates the evaluators? On automatic metrics for assessing AI-based offensive code generatorsabstractAI-based code generators are an emerging solution for automatically writing programs starting from descriptions in natural language, by using deep neural networks (Neural Machine Translation, NMT). In particular, code generators have been used for ethical hacking and offensive security testing by generating proof-of-concept attacks. Unfortunately, the evaluation of code generators still faces several issues. The current practice uses output similarity metrics, i.e., automatic metrics that compute the textual similarity of generated code with ground-truth references. However, it is not clear what metric to use, and which metric is most suitable for specific contexts. This work analyzes a large set of output similarity metrics on offensive code generators. We apply the metrics on two state-of-the-art NMT models using two datasets containing offensive assembly and Python code with their descriptions in the English language. We compare the estimates from the automatic metrics with human evaluation and provide practical insights into their strengths and limitations. Pietro Liguori, Cristina Improta, Roberto Natella, Bojan Cukic, Domenico Cotroneo |
Expert Syst. Appl. | 1 |
| 2023 | Run-time failure detection via non-intrusive event analysis in a large-scale cloud computing platformabstractCloud computing systems fail in complex and unforeseen ways due to unexpected combinations of events and interactions among hardware and software components. These failures are especially problematic when they are silent, i.e., not accompanied by any explicit failure notification, hindering the timely detection and recovery. In this work, we propose an approach to run-time failure detection tailored for monitoring multi-tenant and concurrent cloud computing systems. The approach uses a non-intrusive form of event tracing, without manual changes to the system’s internals to propagate session identifiers (IDs), and builds a set of lightweight monitoring rules from fault-free executions. We evaluated the effectiveness of the approach in detecting failures in the context of the OpenStack cloud computing platform, a complex and “off-the-shelf” distributed system, by executing a campaign of fault injection experiments in a multi-tenant scenario. Our experiments show that the approach detects the failure with an F1 score (0.85) and accuracy (0.77) higher than the ones provided by the OpenStack failure logging mechanisms (0.53 and 0.50) and two non-session-aware run-time verification approaches (both lower than 0.15). Moreover, the approach significantly decreases the average time to detect failures at run-time (∼114 seconds) compared to the OpenStack logging mechanisms. Domenico Cotroneo, Luigi De Simone, Pietro Liguori, Roberto Natella |
J. Syst. Softw. | 3 |
| 2022 | Can we generate shellcodes via natural language? An empirical studyabstractAbstract Writing software exploits is an important practice for offensive security analysts to investigate and prevent attacks. In particular, shellcodes are especially time-consuming and a technical challenge, as they are written in assembly language. In this work, we address the task of automatically generating shellcodes, starting purely from descriptions in natural language, by proposing an approach based on Neural Machine Translation (NMT). We then present an empirical study using a novel dataset ( Shellcode_IA32 ), which consists of 3200 assembly code snippets of real Linux/x86 shellcodes from public databases, annotated using natural language. Moreover, we propose novel metrics to evaluate the accuracy of NMT at generating shellcodes. The empirical analysis shows that NMT can generate assembly code snippets from the natural language with high accuracy and that in many cases can generate entire shellcodes with no errors. Pietro Liguori, Erfan Al-Hossami, Domenico Cotroneo, Roberto Natella, Bojan Cukic, Samira Shaikh |
Autom. Softw. Eng. | 1 |
| 2022 | Fault Injection Analytics: A Novel Approach to Discover Failure Modes in Cloud-Computing SystemsabstractCloud computing systems fail in complex and unexpected ways due to unexpected combinations of events and interactions between hardware and software components. Fault injection is an effective means to bring out these failures in a controlled environment. However, fault injection experiments produce massive amounts of data, and manually analyzing these data is inefficient and error-prone, as the analyst can miss severe failure modes that are yet unknown. This article introduces a new paradigm (fault injection analytics) that applies unsupervised machine learning on execution traces of the injected system, to ease the discovery and interpretation of failure modes. We evaluated the proposed approach in the context of fault injection experiments on the OpenStack cloud computing platform, where we show that the approach can accurately identify failure modes with a low computational cost. Domenico Cotroneo, Luigi De Simone, Pietro Liguori, Roberto Natella |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2021 | EVIL: Exploiting Software via Natural LanguageabstractWriting exploits for security assessment is a challenging task. The writer needs to master programming and obfuscation techniques to develop a successful exploit. To make the task easier, we propose an approach (EVIL) to automatically generate exploits in assembly/Python language from descriptions in natural language. The approach leverages Neural Machine Translation (NMT) techniques and a dataset that we developed for this work. We present an extensive experimental study to evaluate the feasibility of EVIL, using both automatic and manual analysis, and both at generating individual statements and entire exploits. The generated code achieved high accuracy in terms of syntactic and semantic correctness. Pietro Liguori, Erfan Al-Hossami, Vittorio Orbinato, Roberto Natella, Samira Shaikh, Domenico Cotroneo, Bojan Cukic |
ISSRE | 1 |
| 2021 | Enhancing the analysis of software failures in cloud computing systems with deep learning
Domenico Cotroneo, Luigi De Simone, Pietro Liguori, Roberto Natella |
J. Syst. Softw. | 3 |
| 2020 | ProFIPy: Programmable Software Fault Injection as-a-ServiceabstractIn this paper, we present a new fault injection tool (ProFIPy) for Python software. The tool is designed to be programmable, in order to enable users to specify their software fault model, using a domain-specific language (DSL) for fault injection. Moreover, to achieve better usability, ProFIPy is provided as software-as-a-service and supports the user through the configuration of the faultload and workload, failure data analysis, and full automation of the experiments using container- based virtualization and parallelization. Domenico Cotroneo, Luigi De Simone, Pietro Liguori, Roberto Natella |
DSN | 3 |
| 2019 | Enhancing Failure Propagation Analysis in Cloud Computing SystemsabstractIn order to plan for failure recovery, the designers of cloud systems need to understand how their system can potentially fail. Unfortunately, analyzing the failure behavior of such systems can be very difficult and time-consuming, due to the large volume of events, non-determinism, and reuse of third-party components. To address these issues, we propose a novel approach that joins fault injection with anomaly detection to identify the symptoms of failures. We evaluated the proposed approach in the context of the OpenStack cloud computing platform. We show that our model can significantly improve the accuracy of failure analysis in terms of false positives and negatives, with a low computational cost. Domenico Cotroneo, Luigi De Simone, Pietro Liguori, Roberto Natella, Nematollah Bidokhti |
ISSRE | 3 |
| 2019 | How bad can a bug get? an empirical analysis of software failures in the OpenStack cloud computing platformabstractCloud management systems provide abstractions and APIs for programmatically configuring cloud infrastructures. Unfortunately, residual software bugs in these systems can potentially lead to high-severity failures, such as prolonged outages and data losses. In this paper, we investigate the impact of failures in the context widespread OpenStack cloud management system, by performing fault injection and by analyzing the impact of the resulting failures in terms of fail-stop behavior, failure detection through logging, and failure propagation across components. The analysis points out that most of the failures are not timely detected and notified; moreover, many of these failures can silently propagate over time and through components of the cloud management system, which call for more thorough run-time checks and fault containment. Domenico Cotroneo, Luigi De Simone, Pietro Liguori, Roberto Natella, Nematollah Bidokhti |
ESEC/SIGSOFT FSE | 3 |