VLDB 2026 Research / reviewers in the wild / expert
Aurora Papotti
dblp:329/4857
· DBLP profile ↗
5ranked-venue papers
4as first author
5since 2021 · last 2026
0000-0003-3207-7662ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 5 · 4 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Is common sense all you need? Using expert defined rules to identify vulnerability patches instead of machine learningabstractGoal: Machine learning (ML) has been proposed to identify security fixing commits with mixed success. We evaluate a different alternative in which expert-defined common sense rules with power-law weights are used to identify security fixing commit. Experiments: We first evaluated how those rules perform against ML models trained on the same features on which rules are based on the ProjectKB real world dataset of Java security commits. We then ran a think aloud protocol with seven senior analysts to analyze whether the selected rules are consistent with usage. Lastly, we ran the experiment with Master students ([Formula: see text]) to check whether the last ranking or the actual mention of which rules were applicable is beneficial to users with less experience. We found that the common-sense rules defined by security experts have a similar performance to classical ML approaches. Findings: Using Shap values to explain earned features, we find that ML models have the same chance corrected accuracy of expert defined rules, and they learned essentially the same top features identified by the experts and only differs on minor features (either among themselves and between the expert features). We observed that the juniors performance are comparable to the experts' one (CI [0.62, 0.76]). When asked to junior analysts to identify the fixing commits (in the top 10 selected by the tool) showing them which features was responsible for the selection do not seem to help (when compared with a control group with no support). A conclusion of our study is that one might just use ML to first analyze the data and then distill common sense rules that are still effective to apply. More experiments are needed to make features useful for the final decision. Aurora Papotti, Serena Elisa Ponta, Antonino Sabetta, Fabio Massacci |
Empir. Softw. Eng. | 1 |
| 2025 | On the effects of program slicing for vulnerability detection during code inspectionabstractAbstract Slicing is a fault localization technique that has been proposed to support debugging and program comprehension. Yet, its empirical effectiveness during code inspection by humans has received limited attention. The goal of our study is two-fold. First, we aim to define what it means for a code reviewer to identify the vulnerable lines correctly. Second, we investigate whether reducing the number of to-be-inspected lines by method-level slicing supports code reviewers in detecting security vulnerabilities. We propose a novel approach based on the notion of a $$\delta $$ δ -neighborhood (intuitively based on the idea of the context size of the command ) to define correctly identified lines. Then, we conducted a multi-year controlled experiment (2017-2023) in which MSc students attending security courses ( $$n=236$$ n = 236 ) were tasked with identifying vulnerable lines in original or sliced Java files from Apache Tomcat. We provide perfect seed lines for a slicing algorithm to control for confounding factors. Each treatment differs in the pair (Vulnerability, Original/Sliced) with a balanced design with vulnerabilities from the OWASP Top 10 2017: A1 (Injection), A5 (Broken Access Control), A6 (Security Misconfiguration), and A7 (Cross-Site Scripting). To generate smaller slices for human consumption, we used a variant of intra-procedural thin slicing. We report the results for $$\delta = 0$$ δ = 0 which corresponds to exactly matching the vulnerable ground truth lines, and $$\delta = 3$$ δ = 3 which represents the scenario of identifying the vulnerable area. For both cases, we found that slicing helps in ‘finding something’ (the participant has found at least some vulnerable lines) as opposed to ‘finding nothing’. For the case of $$\delta = 0$$ δ = 0 analyzing a slice and analyzing the original file are statistically equivalent from the perspective of lines found by those who found something. With $$\delta = 3$$ δ = 3 slicing helps to find more vulnerabilities compared to analyzing an original file, as we would normally expect. Given the type of population, additional experiments are necessary to be generalized to experienced developers. Aurora Papotti, Katja Tuma, Fabio Massacci |
Empir. Softw. Eng. | 1 |
| 2024 | On the acceptance by code reviewers of candidate security patches suggested by Automated Program Repair toolsabstractAbstract Objective We investigated whether (possibly wrong) security patches suggested by Automated Program Repairs (APR) for real world projects are recognized by human reviewers. We also investigated whether knowing that a patch was produced by an allegedly specialized tool does change the decision of human reviewers. Method We perform an experiment with $$n= 72$$ n = 72 Master students in Computer Science. In the first phase, using a balanced design, we propose to human reviewers a combination of patches proposed by APR tools for different vulnerabilities and ask reviewers to adopt or reject the proposed patches. In the second phase, we tell participants that some of the proposed patches were generated by security-specialized tools (even if the tool was actually a ‘normal’ APR tool) and measure whether the human reviewers would change their decision to adopt or reject a patch. Results It is easier to identify wrong patches than correct patches, and correct patches are not confused with partially correct patches. Also patches from APR Security tools are adopted more often than patches suggested by generic APR tools but there is not enough evidence to verify if ‘bogus’ security claims are distinguishable from ‘true security’ claims. Finally, the number of switches to the patches suggested by security tool is significantly higher after the security information is revealed irrespective of correctness. Limitations The experiment was conducted in an academic setting, and focused on a limited sample of popular APR tools and popular vulnerability types. Aurora Papotti, Ranindya Paramitha, Fabio Massacci |
Empir. Softw. Eng. | 1 |
| 2024 | Addressing combinatorial experiments and scarcity of subjects by provably orthogonal and crossover experimental designsabstractContext: Experimentation in Software and Security Engineering is a common research practice, in particular with human subjects.Problem: The combinatorial nature of software configurations and the difficulty of recruiting experienced subjects or running complex and expensive experiments make the use of full factorial experiments unfeasible to obtain statistically significant results.Contribution: Provide comprehensive alternative Designs of Experiments (DoE) based on orthogonal designs or crossover designs that provably meet desired requirements such as balanced pair-wise configurations or balanced ordering of scenarios to mitigate bias or learning effects.We also discuss and formalize the statistical implications of these design choices, in particular for crossover designs.Artifact: We made available the algorithmic construction of the design for 𝓁 = 2, 3, 4, 5 levels for arbitrary 𝐾 factors and illustrated their use with examples from security and software engineering research. Fabio Massacci, Aurora Papotti, Ranindya Paramitha |
J. Syst. Softw. | 2 |
| 2022 | Assessment of Automated (Intelligent) Toolchainsabstract[Background:] Automated Intelligent Toolchains, which are a composition of different tools that use AI or static analysis, are widely used in software engineering to deploy automated program repair techniques, or in software security to identify vulnerabilities. [Overall Research Problem:] Most studies with automated intelligent toolchains report uncertainty and evaluations only of the individual components of the chain. How do we calculate the uncertainty and error propagation on the overall automated toolchain? [Approach:] I plan to replicate research case studies to collect data and design a methodology to reconstruct the overall correctness metrics of the toolchains, or identifying missing variables. Further confirmatory experiments with humans will be performed. Finally, I will implement an artifact to automate the overall assessment of automated toolchains. [Current Status:] A preliminary validation of published studies showed promising results. Aurora Papotti |
ASE | 1 |