VLDB 2026 Research / reviewers in the wild / expert
Krishna Kanth Arumugam
dblp:382/4759
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2024
0009-0000-2309-0203ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Network and information security
1 paper |
Blockchain and cryptocurrency security · 50% Systems and software security · 50% | |
| Software engineering, system software, and programming languages
1 paper |
Empirical software engineering · 100% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Systems and software security › vulnerability discovery › machine-learning-based vulnerability detection
deep learning-based vulnerability detection |
0.8 | 1 | 2024 | Revisiting the Performance of Deep Learning-Based Vulnerability Detection on Realistic Datasets · IEEE Trans. Software Eng. 2024 |
Blockchain and cryptocurrency security › smart contract security
vulnerability detection |
0.8 | 1 | 2024 | Revisiting the Performance of Deep Learning-Based Vulnerability Detection on Realistic Datasets · IEEE Trans. Software Eng. 2024 |
Empirical software engineering
mining software repositories |
0.8 | 1 | 2024 | Revisiting the Performance of Deep Learning-Based Vulnerability Detection on Realistic Datasets · IEEE Trans. Software Eng. 2024 |
Methods — techniques the papers use, named apart from their topics
deep learning · 1.5data augmentation · 1.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Revisiting the Performance of Deep Learning-Based Vulnerability Detection on Realistic DatasetsabstractThe impact of software vulnerabilities on everyday software systems is concerning. Although deep learning-based models have been proposed for vulnerability detection, their reliability remains a significant concern. While prior evaluation of such models reports impressive recall/F1 scores of up to 99%, we find that these models underperform in practical scenarios, particularly when evaluated on the entire codebases rather than only the fixing commit. In this paper, we introduce a comprehensive dataset (Real-Vul) designed to accurately represent real-world scenarios for evaluating vulnerability detection models. We evaluate DeepWukong, LineVul, ReVeal, and IVDetect vulnerability detection approaches and observe a surprisingly significant drop in performance, with precision declining by up to 95 percentage points and F1 scores dropping by up to 91 percentage points. A closer inspection reveals a substantial overlap in the embeddings generated by the models for vulnerable and uncertain samples (non-vulnerable or vulnerability not reported yet), which likely explains why we observe such a large increase in the quantity and rate of false positives. Additionally, we observe fluctuations in model performance based on vulnerability characteristics (e.g., vulnerability types and severity). For example, the studied models achieve 26 percentage points better F1 scores when vulnerabilities are related to information leaks or code injection rather than when vulnerabilities are related to path resolution or predictable return values. Our results highlight the substantial performance gap that still needs to be bridged before deep learning-based vulnerability detection is ready for deployment in practical settings. We dive deeper into why models underperform in realistic settings and our investigation revealed overfitting as a key issue. We address this by introducing an augmentation technique, potentially improving performance by up to 30%. We contribute (a) an approach to creating a dataset that future research can use to improve the practicality of model evaluation; (b)Real-Vul– a comprehensive dataset that adheres to this approach; and (c) empirical evidence that the deep learning-based models struggle to perform in a real-world setting. Partha Chakraborty, Krishna Kanth Arumugam, Mahmoud Alfadel, Meiyappan Nagappan, Shane McIntosh |
IEEE Trans. Software Eng. | 2 |