VLDB 2026 Research / reviewers in the wild / expert
Khang Mai
dblp:386/4634
· DBLP profile ↗
4ranked-venue papers
3as first author
4since 2021 · last 2026
0009-0000-2488-6043ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 3 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CodeEnhancer: LLM-generated Python code enhancement through SAST integration and fine-tuningabstractDespite the rapid adoption of Large Language Models (LLMs) for automatic code generation, their output often exhibits syntax errors, security vulnerabilities, and functional inconsistencies. To address these issues, we present CodeEnhancer, a two-stage framework that tightly integrates LLMs with static application security testing (SAST) tools and targeted fine-tuning. The goal is to produce more secure and functionally correct Python code. In the first stage, our iterative validation pipeline couples LLM-generated code with tools such as Pylint and Bandit. These tools automatically identify and remediate issues through structured feedback loops. When applied to the GPT-4o model, this process eliminated 82.8% of the initial vulnerabilities and resolved all the detected functional correctness issues when tested on the LLMSecEval dataset. In the second stage, we fine-tune the LLMs using two types of secure code examples: expert-written samples and code refined by our framework. Comparative experiments demonstrate that the framework-tuned model outperforms the baseline and expert-tuned models. The framework-tuned model generates only 18.4% vulnerable code snippets on the LLMSecEval dataset, whereas the baseline and expert-tuned models produce 43.6% and 54.7% vulnerable code snippets, respectively. The framework-tuned model reduces final vulnerability rates to 6.7% on LLMSecEval and 3.5% on the SecurityEval dataset. Our results highlight the synergistic effect of integrating static analysis with feedback-informed fine-tuning. They also reveal limitations in current evaluation metrics and dataset representativeness. These findings suggest a scalable, robust approach to achieving more secure, trustworthy, and practical AI-assisted code generation. • Combines language models with SAST Tools to enhance Syntax, security and functional correctness Python code. • First approach to address syntax, security, and functional correctness in LLM-generated code. • Automated feedback and learning process helps LLMs generate more secure, correct code. • Fine-tuning on framework-refined code leads to better security than training on expert-written code. • Scalable approach enables robust and trustworthy AI-assisted code generation and refinement with minimal manual effort. Khang Mai, Nakul Ghate, Tomohiko Yagyu, Razvan Beuran, Yasuo Tan |
Knowl. Based Syst. | 2 |
| 2025 | CyLLM-DAP: Cybersecurity Domain-Adaptive Pre-Training Framework of Large Language Models
Khang Mai, Razvan Beuran, Naoya Inoue |
ICISSP (2) | 1 |
| 2025 | LLM-Based Fine-Grained ABAC Policy Generation
Khang Mai, Nakul Ghate, Razvan Beuran |
ICISSP (2) | 1 |
| 2025 | RAF-AG: Report analysis framework for attack path generationabstractInformation sharing is a key practice in cybersecurity for coping with the ever-changing cyberattacks that are targeting computer systems. Thus, when cyber incidents happen, cyber threat intelligence (CTI) reports are prepared and shared among cybersecurity practitioners to help them get up-to-date information about those incidents. However, reading and analyzing the report text to comprehend the included information is a cumbersome process. Although techniques based on deep learning were proposed to speed up report analysis in order to obtain the enclosed essential information, such as attack path, training data insufficiency makes these methods inefficient in practical circumstances. This paper presents RAF-AG, a report analysis framework for attack path generation. To analyze CTI reports, RAF-AG utilizes the sentence dependency tree for entity and relation extraction, and a weak supervision approach for entity labeling. This is followed by graph building and graph alignment for generating the attack paths. Our approach resolves the data insufficiency problem in the cybersecurity domain by lowering the need for expert involvement. We evaluated RAF-AG by comparing the generated attack paths with those produced by AttacKG, a state-of-the-art automatic report analysis framework. RAF-AG was able to identify cyberattack steps by matching their appearance order inside the report, and link them with techniques from the MITRE ATT&CK knowledge base with an improved F1 score compared to AttacKG (0.708 versus 0.393). Khang Mai, Razvan Beuran, Ryosuke Hotchi, Ooi Sian En, Takayuki Kuroda, Yasuo Tan |
Comput. Secur. | 1 |