VLDB 2026 Research / reviewers in the wild / expert
M. Ali Babar
dblp:358/5031
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 5 · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DDPT: Diffusion-Driven Prompt Tuning for Large Language Model Code GenerationabstractLarge Language Models (LLMs) have demonstrated remarkable capabilities in code generation. However, the quality of the generated code is heavily dependent on the structure and composition of the prompts used. Crafting high-quality prompts is a challenging task that requires significant knowledge and skills of prompt engineering. To advance the automation support for the prompt engineering for LLM-based code generation, we propose a novel solution Diffusion-Driven Prompt Tuning (DDPT) that learns how to generate optimal prompt embedding from Gaussian Noise to automate the prompt engineering for code generation. We evaluate the feasibility of diffusion-based optimization and abstract the optimal prompt embedding as a directional vector toward the optimal embedding. We use the code generation loss given by the LLMs to help the diffusion model capture the distribution of optimal prompt embedding during training. The trained diffusion model can build a path from the noise distribution to the optimal distribution at the sampling phrase, the evaluation result demonstrates that DDPT helps improve the prompt optimization for code generation. Sangwon Hyun, M. Ali Babar |
CAIN | 3 |
| 2025 | Toward Realistic Evaluations of Just-In-Time Vulnerability PredictionabstractModern software systems are increasingly complex, presenting significant challenges in quality assurance. Just-intime vulnerability prediction (JIT-VP) is a proactive approach to identifying vulnerable commits and providing early warnings about potential security risks. However, we observe that current JIT-VP evaluations rely on an idealized setting, where the evaluation datasets are artificially balanced, consisting exclusively of vulnerability-introducing and vulnerability-fixing commits. To address this limitation, this study assesses the effectiveness of JIT-VP techniques under a more realistic setting that includes both vulnerability-related and vulnerability-neutral commits. To enable a reliable evaluation, we introduce a large-scale public dataset comprising over one million commits from FFmpeg and the Linux kernel. Our empirical analysis of eight state-of-theart JIT-VP techniques reveals a significant decline in predictive performance when applied to real-world conditions; for example, the average PR-AUC on Linux drops 98 % from 0.805 to 0.016. This discrepancy is mainly attributed to the severe class imbalance in real-world datasets, where vulnerability-introducing commits constitute only a small fraction of all commits. To mitigate this issue, we explore the effectiveness of widely adopted techniques for handling dataset imbalance, including customized loss functions, oversampling, and undersampling. Surprisingly, our experimental results indicate that these techniques are ineffective in addressing the imbalance problem in JIT-VP. These findings underscore the importance of realistic evaluations of JIT-VP and the need for domain-specific techniques to address data imbalance in such scenarios. Thanh Le-Cong, Triet Huynh Minh Le, M. Ali Babar, Huynh Quyet Thang |
ICSME | 4 |
| 2025 | VulGuard: An Unified Tool for Evaluating Just-In-Time Vulnerability Prediction ModelsabstractWe present VulGuard, an automated tool designed to streamline the extraction, processing, and analysis of commits from GitHub repositories for Just-In-Time vulnerability prediction (JIT-VP) research. VulGuard automatically mines commit histories, extracts fine-grained code changes, commit messages, and software engineering metrics, and formats them for downstream analysis. In addition, it integrates several state-of-the-art vulnerability prediction models, allowing researchers to train, evaluate, and compare models with minimal setup. By supporting both repository-scale mining and model-level experimentation within a unified framework, VulGuard addresses key challenges in reproducibility and scalability in software security research. VulGuard can also be easily integrated into the CI/CD pipeline. We demonstrate the effectiveness of the tool in two influential open-source projects, FFmpeg and the Linux kernel, highlighting its potential to accelerate real-world JIT-VP research and promote standardized benchmarking. A demo video is available at: https://youtu.be/j96096-pxbs. Manh Tran-Duc, Thanh Le-Cong, Triet Huynh Minh Le, M. Ali Babar, Huynh Quyet Thang |
ICSME | 5 |
| 2025 | SAGELY - Context-Aware Holistic Service Policy Enforcement Across Swarm-Edge ContinuumabstractSwarm-edge-cloud service-based applications (SESA) utilize UAV swarm nodes and edge-cloud computing infrastructures to perform complex tasks. These applications leverage distributed services across UAVs and edge resources, transitioning to cloud resources as needed, to support continuum computing, adapting to dynamic workloads, network conditions, and mission requirements. However, existing solutions lack robust mechanisms for dynamic policy enforcement in such environments. This paper introduces SAGELY (SwArm-edGE PoLicY), a framework for secure, efficient, and adaptive policy enforcement in dynamic SESA environments. SAGELY incorporates: (1) context-aware policy adaptation to adjust enforcement dynamically, (2) flexible policy enforcement across centralized and decentralized models, and (3) a pluggable service architecture for continuum context management. We present a testbed to enforce continuum policies and conduct experiments. Tri Nguyen 0001, Anh-Dung Nguyen, M. Ali Babar, Hong Linh Truong 0001 |
ICWS | 4 |
| 2025 | LLMSecConfig: An LLM-Based Approach for Fixing Software Container MisconfigurationsabstractSecurity misconfigurations in Container Orchestrators (COs) can pose serious threats to software systems. While Static Analysis Tools (SATs) can effectively detect these security vulnerabilities, the industry currently lacks automated solutions capable of fixing these misconfigurations. The emergence of Large Language Models (LLMs), with their proven capabilities in code understanding and generation, presents an opportunity to address this limitation. This study introduces LLMSecConfig, an innovative framework that bridges this gap by combining SATs with LLMs. Our approach leverages advanced prompting techniques and Retrieval-Augmented Generation (RAG) to automatically repair security misconfigurations while preserving operational functionality. Evaluation of 1,000 real-world Kubernetes configurations achieved a 94% success rate while maintaining a low rate of introducing new misconfigurations.Our work makes a promising step towards automated container security management, reducing the manual effort required for configuration maintenance. Triet Huynh Minh Le, M. Ali Babar |
MSR | 3 |