Fahad Al Debeyan

dblp:332/6815 · DBLP profile ↗
← Back
2ranked-venue papers
2as first author
2since 2021 · last 2026
0000-0002-3981-2722ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 How vulnerability explanations help software practitioners confirm and fix code vulnerabilities
abstract
Context: Most current code vulnerability detection tools provide only a binary classification (vulnerable/non-vulnerable) with little to no additional context. This paper explores the impact of providing explanations for vulnerabilities alongside code labelled as vulnerable. Objective: We investigate the influence of explanations on the ability of software practitioners to confirm such labelled code as actually vulnerable (i.e., a true positive vulnerability) and to fix such vulnerable code correctly. Method: We surveyed 99 software practitioners to establish their use of code-vulnerability detection tools and to evaluate the impact of explanations on their behaviour towards code labelled as vulnerable in a series of coding exercises. Participants were presented with four forms of explanation: vulnerable lines , vulnerability type , short-form text , and long-form text . Results: Software practitioners performed better at confirming and fixing code vulnerabilities when presented with any of the four forms of explanation. Although practitioners stated a preference for long-form text explanations, they achieved the highest confirmation and fixing performance with short-form text explanations. Practitioners also indicated willingness to accept modest drops in detection precision and recall if richer explanations were provided, and their preferences for explanation types and performance trade-offs varied according to where a detection tool is used in the software-development pipeline. Conclusions: Vulnerability-detection and prediction tools should provide explanatory output and allow different explanation types tailored to their deployment stage in the development workflow. Few current tools provide any explanations, and none identified in this study provide text-based explanations.
Fahad Al Debeyan, Tracy Hall, Lech Madeyski, Emily Winter 0001
Inf. Softw. Technol.1
2024 The impact of hard and easy negative training data on vulnerability prediction performance
abstract
Vulnerability prediction models have been shown to perform poorly in the real world. We examine how the composition of negative training data influences vulnerability prediction model performance. Inspired by other disciplines (e.g. image processing), we focus on whether distinguishing between negative training data that is ‘easy’ to recognise from positive data (very different from positive data) and negative training data that is ‘hard’ to recognise from positive data (very similar to positive data) impacts on vulnerability prediction performance. We use a range of popular machine learning algorithms, including deep learning, to build models based on vulnerability patch data curated by Reis and Abreu, as well as the MSR dataset. Our results suggest that models trained on higher ratios of easy negatives perform better, plateauing at 15 easy negatives per positive instance. We also report that different ML algorithms work better based on the negative sample used. Overall, we found that the negative sampling approach used significantly impacts model performance, potentially leading to overly optimistic results. The ratio of ‘easy’ versus ‘hard’ negative training data should be explicitly considered when building vulnerability prediction models for the real world.
Fahad Al Debeyan, Lech Madeyski, Tracy Hall, David Bowes
J. Syst. Softw.1