Jake Norton

dblp:429/6360 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
1since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Network and information security
1 paper
Blockchain and cryptocurrency security · 100%
Software engineering, system software, and programming languages
1 paper
Empirical software engineering · 100%
Artificial intelligence
1 paper
Deep learning architectures and training · 100%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Blockchain and cryptocurrency security › smart contract security
vulnerability detection
1.012026
VulnBench: A Comprehensive Benchmark for Transformer-Based Vulnerability Detection · AAAI 2026
Empirical software engineering
benchmarking
1.012026
VulnBench: A Comprehensive Benchmark for Transformer-Based Vulnerability Detection · AAAI 2026
Machine learning › Deep learning architectures and training
transformer
0.312026
VulnBench: A Comprehensive Benchmark for Transformer-Based Vulnerability Detection · AAAI 2026

Methods — techniques the papers use, named apart from their topics

threshold optimization · 3.0natgen · 3.0codet5 · 3.0GraphCodeBERT · 3.0CodeBERT · 3.0
YearPublicationVenuePosition
2026 VulnBench: A Comprehensive Benchmark for Transformer-Based Vulnerability Detection
abstract
Reproducible benchmarking of tools that automatically detect vulnerabilities in source code remains challenging due to inconsistent implementations, varying data preprocessing, and methodological flaws that compromise fair model comparison. In a recent study, 9 in 10 vulnerability detection studies were found to use inappropriate evaluation approaches, with models achieving high scores through spurious correlations rather than actual vulnerability detection. We present VulnBench, an extensible, open-source benchmarking tool that enables fair comparison across models and datasets. Our systematic evaluation of CodeBERT, GraphCodeBERT, CodeT5 (encoder-only and full), and NatGen across eight mostly C/C++ source code datasets reveals that proper threshold optimization can improve F1-scores by up to 54%, as well as wide variation in F1-scores showing the large gap in the difficulty of the vulnerability dataset field. By standardising evaluation protocols, VulnBench enables researchers to distinguish between genuine model improvements and methodological artifacts as well as reducing wasteful duplication of effort spent on reproducing results.
Jake Norton, David M. Eyers, Veronica Liesaputra
AAAI1