EDBT 2026 Demo / reviewers in the wild / expert
Yu Luo 0020
dblp:45/6469-20
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2026
0009-0008-3317-4935ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 4 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LLM-Based Data Generation and Augmentation for Rust Vulnerability Detection
Irfan Ali Khan, Yu Luo 0020, Dianxiang Xu |
COMPSAC | 2 |
| 2025 | C2RustTV: An LLM-based Framework for C to Rust Translation and ValidationabstractTransitioning legacy C codebases to Rust has emerged as a promising approach to addressing memory safety issues inherent in C. However, existing methods often produce non-idiomatic or semantically inconsistent Rust code. This paper introduces C2RustTV, a framework that leverages large language models (LLMs) to translate C programs to idiomatic Rust code while assuring translation quality through automated conformance testing. C2RustTV integrates test case generation, automated translation of both production and test code, and conformance validation via test execution. Experiments across multiple datasets demonstrate the effectiveness of C2RustTV, achieving higher rates of compilation success and functional conformance compared to state-of-the-art techniques. Yu Luo 0020, Mengtao Zhang, Dianxiang Xu |
COMPSAC | 2 |
| 2025 | From Text to STIX: Reducing Hallucinations through Fine-Tuned LLMsabstractThe Structured Threat Information Expression (STIX) standard has become essential for representing and sharing cyber threat intelligence (CTI). However, automatically generating valid and semantically consistent STIX objects remains challenging. Large language models (LLMs) often succeed at extracting low-level indicators but struggle with higher-level abstractions, leading to missing fields, invalid properties, and structurally inconsistent outputs. To address these issues, we propose a methodology that integrates instruction fine-tuning with a novel evaluation framework. Specifically, we fine-tune LLMs using curated text–STIX pairs and domain-specific instructions to internalize schema alignment and reduce hallucinations. We further introduce the Object and Properties Similarity Evaluation (OPSE), which measures both structural validity and semantic fidelity of generated STIX objects against expert ground truth. Experimental results show that our fine-tuned model substantially improves object coverage, property completeness, and structural accuracy compared with baseline prompting and existing STIX generation tools. Ruoyao Xiao, Yu Luo 0020, Dianxiang Xu |
TrustCom | 2 |
| 2024 | Predicting Code Vulnerability Types via Heterogeneous GNN Learning
Yu Luo 0020, Dianxiang Xu |
ESORICS (3) | 1 |
| 2024 | GNN-Based Transfer Learning and Tuning for Detecting Code Vulnerabilities across Different Programming LanguagesabstractMachine learning has emerged as a promising method for detecting code vulnerabilities. In existing research, models trained on available samples are applied to predict potential vulnerabilities in new code within the same language. This approach may not be directly applicable when dealing with languages that have limited vulnerable samples available. This paper leverages transfer learning and k-shot tuning to address this issue, aiming to classify various Common Weakness Enumeration (CWE) categories in code written in a target language by utilizing a model trained on data from a different base language. We build upon a foundational model that enhances a state-of-the-art Graph Neural Network (GNN) framework designed for code vulnerability detection. Our experiments using extensive datasets of 42 common CWEs in C, Java, and C# have yielded valuable findings. Specifically, when models trained in one language are transferred to another, there is a noticeable improvement in detection performance. The magnitude of performance enhancements varies across different CWEs. k-shot tuning further refines the transferred models. However, the extent of performance gains gradually diminishes as the value of k increases. Once k exceeds a certain threshold, the transferred knowledge may be overwritten. Furthermore, our research unveiled an asymmetry in the two transfer learning directions within each language pair, even when the languages share the same programming paradigm. Transferred models in both directions often exhibit differing performance levels across the majority of CWEs. Irfan Ali Khan, Yu Luo 0020, Dianxiang Xu |
QRS | 2 |
| 2024 | Analyzing Relationship Consistency in Digital Forensic Knowledge Graphs with Graph Learning
Ruoyao Xiao, Yu Luo 0020, Harshmeet Lamba, Dianxiang Xu |
TrustCom | 2 |
| 2022 | Compact Abstract Graphs for Detecting Code Vulnerability with GNN ModelsabstractSource code representation is critical to the machine-learning-based approach to detecting code vulnerability. This paper proposes Compact Abstract Graphs (CAGs) of source code in different programming languages for predicting a broad range of code vulnerabilities with Graph Neural Network (GNN) models. CAGs make the source code representation aligned with the task of vulnerability classification and reduce the graph size to accelerate model training with minimum impact on the prediction performance. We have applied CAGs to six GNN models and large Java/C datasets with 114 vulnerability types in Java programs and 106 vulnerability types in C programs. The experiment results show that the GNN models have performed well, with accuracy ranging from 94.7% to 96.3% on the Java dataset and from 91.6% to 93.2% on the C dataset. The resultant GNN models have achieved promising performance when applied to more than 2,500 vulnerabilities collected from real-world software projects. The results also show that using CAGs for GNN models is significantly better than ASTs, CFGs (Control Flow Graphs), and PDGs (Program Dependence Graphs). A comparative study has demonstrated that the CAG-based GNN models can outperform the existing methods for machine learning-based vulnerability detection. Yu Luo 0020, Dianxiang Xu |
ACSAC | 1 |
| 2021 | Detecting Integer Overflow Errors in Java Source Code via Machine LearningabstractInteger overflow is a common cause of software failure and security vulnerability. Existing approaches to detecting integer overflow errors rely on traditional static code analysis and dynamic testing. This paper presents a novel machine learning-based approach that predicts integer overflow errors by treating source code as text. It exploits text classifiers to determine whether each method in a given Java program contains an integer overflow error. As the training data is essential, we have constructed a comprehensive dataset to accounts for (a) integer overflow errors of all integer types and operations in Java (i.e., positive samples); (b) various programming techniques for preventing integer overflow errors (i.e., negative samples); and (c) malicious scenarios that may mislead text classifiers (i.e., adversarial samples). We have trained three classifiers, BERT, fastText, and NBSVM, that represent different text embedding techniques. BERT, as a representative deep-learning transformer, has achieved the highest performance scores and remained robust even when tested with the adversarial samples. Yu Luo 0020, Dianxiang Xu |
ICTAI | 1 |