VLDB 2026 Research / reviewers in the wild / expert
Anh M. T. Bui
dblp:348/7948
· DBLP profile ↗
6ranked-venue papers
0as first author
6since 2021 · last 2026
0000-0001-7877-9438ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 6 · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Larger is Not Always Better: Leveraging Structured Code Diffs for Comment Inconsistency Detection
Hoang Vinh-Phong Nguyen, Anh M. T. Bui, Phuong T. Nguyen |
SANER | 2 |
| 2025 | Bake Two Cakes with One Oven: RL for Defusing Popularity Bias and Cold-start in Third-Party Library RecommendationsabstractThird-party libraries (TPLs) are an integral part of modern software development, enhancing developer productivity and accelerating time-to-market. However, identifying suitable candidates from a rapidly growing and continuously evolving collection of TPLs remains a challenging task. TPL recommender systems have been developed to address this issue. They typically rely on collaborative filtering (CF) which exploits a two-dimensional project-library matrix (user-item in general context of recommendation) when making recommendations. In fact, CF-based approaches often encounter two challenges: (i) a tendency to recommend popular items more frequently, making them even more dominant, a phenomenon known as popularity bias, and (ii) difficulty in generating recommendations for new users or items due to limited user-item interactions, commonly referred to as the cold-start problem. In this paper, we propose a reinforcement learning (RL)-based approach to address popularity bias and the cold-start problem in TPL recommendation. We conducted experiments on benchmark datasets for TPL recommendation, demonstrating that our proposed approach outperforms state-of-the-art models in cold-start scenarios while effectively mitigating the impact of popularity bias. Vuong Hoang Minh, Anh M. T. Bui, Phuong T. Nguyen 0001, Davide Di Ruscio |
EASE | 2 |
| 2025 | When Retriever Meets Generator: A Joint Model for Code Comment GenerationabstractBackground. Automatically generating concise, informative comments for source code can lighten documentation effort and accelerate program comprehension. Retrievalaugmented approaches first fetch code snippets with existing comments and then synthesize a new comment, yet retrieval and generation are typically optimized in isolation, allowing irrelevant neighbors to propagate noise downstream. Aims. To tackle the issue, we propose a novel approach named RAGSum with the aim of both effectiveness and efficiency in recommendations. Method. RAGSum is built on top of fuse retrieval and generation using a single CodeT5 backbone. Results. We report preliminary results on a unified retrievalgeneration framework built on CodeT5. A contrastive pretraining phase shapes code embeddings for nearest-neighbor search; these weights then seed end-to-end training with a composite loss that (i) rewards accurate top-k retrieval; and (ii) minimizes comment-generation error. More importantly, a lightweight self-refinement loop is deployed to polish the final output. We evaluated the framework on three cross-language benchmarks (Java, Python, C), and compared it with three well-established baselines. The results show that our approach substantially outperforms the baselines with respect to BLEU, METEOR, and ROUTE-L. Conclusions. These findings indicate that tightly coupling retrieval and generation can raise the ceiling for comment automation and motivate forthcoming replications and qualitative developer studies. Tien P. T. Le, Anh M. T. Bui, Huy N. D. Pham, Alessio Bucaioni, Phuong T. Nguyen 0001 |
ESEM | 2 |
| 2025 | EnseSmells : Deep ensemble and programming language models for automated code smells detectionabstractA smell in software source code denotes an indication of suboptimal design and implementation decisions, potentially hindering the code understanding and, in turn, raising the likelihood of being prone to changes and faults. Identifying these code issues at an early stage in the software development process can mitigate these problems and enhance the overall quality of the software. Current research primarily focuses on the utilization of deep learning-based models to investigate the contextual information concealed within source code instructions to detect code smells, with limited attention given to the importance of structural and design-related features. This paper proposes a novel approach to code smell detection, constructing a deep learning architecture that places importance on the fusion of structural features and statistical semantics derived from pre-trained models for programming languages. We further provide a thorough analysis of how different source code embedding models affect the detection performance with respect to different code smell types. Using four widely-used code smells from well-designed datasets, our empirical study shows that incorporating design-related features significantly improves detection accuracy, outperforming state-of-the-art methods on the MLCQ dataset with improvements ranging from 5.98% to 28.26%, depending on the type of code smell. Anh Ho, Anh M. T. Bui, Phuong T. Nguyen 0001, Amleto Di Salle, Bach Le 0001 |
J. Syst. Softw. | 2 |
| 2024 | LEGION: Harnessing Pre-trained Language Models for GitHub Topic Recommendations with Distribution-Balance LossabstractOpen-source development has revolutionized the software industry by promoting collaboration, transparency, and community-driven innovation. Today, a vast amount of various kinds of open-source software, which form networks of repositories, is often hosted on GitHub – a popular software development platform. To enhance the discoverability of the repository networks, i.e., groups of similar repositories, GitHub introduced repository topics in 2017 that enable users to more easily explore relevant projects by type, technology, and more. It is thus crucial to accurately assign topics for each GitHub repository. Current methods for automatic topic recommendation rely heavily on TF-IDF for encoding textual data, presenting challenges in understanding semantic nuances. Yen-Trang Dang, Thanh Le-Cong, Phuc-Thanh Nguyen, Anh M. T. Bui, Phuong T. Nguyen 0001, Bach Le 0001, Huynh Quyet Thang |
EASE | 4 |
| 2023 | Fusion of deep convolutional and LSTM recurrent neural networks for automated detection of code smellsabstractCode smells is the term used to signal certain patterns or structures in software code that may contain a potential design or architecture problem, leading to maintainability or other software quality issues. Detecting code smells early in the software development process helps prevent these problems and improve the overall software quality. Existing research concentrates on the process of collecting and handling dataset, then exploring the potential of utilizing deep learning models to detect smells, while ignoring extensive feature engineering. Though these approaches obtained promising results, the following issues need to be tackled: (i) extracting both structural and semantic features from the software units; (ii) mitigating the effects of imbalanced data distribution on the performance.In this paper, we propose DeepSmells as a novel approach to code smells detection. To learn the complex hierarchical representations of the code fragment, we apply a deep convolutional neural network (CNN). Then, in order to improve the quality of the context encoding and preserve semantic information, long short-term memory networks (LSTM) is placed immediately after the CNN. The final classification is conducted by deep neural networks with weighted loss function to reduce the impact of skewed data distribution. We performed an empirical study using the existing code smell benchmark datasets to assess the performance of our proposed approach, and compare it with state-of-the-art baselines. The results demonstrate the effectiveness of our proposed method for all kinds of code smells with outperformed evaluation metrics in terms of F1 score and MCC. Anh Ho, Anh M. T. Bui, Phuong T. Nguyen 0001, Amleto Di Salle |
EASE | 2 |