VLDB 2026 Research / reviewers in the wild / expert
Weihuan Min
dblp:362/8701
· DBLP profile ↗
6ranked-venue papers
0as first author
6since 2021 · last 2025
0009-0004-4668-0192ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 5 · 5 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | EvoAPR: Enhancing Large Language Models for Automatic Program Repair with Genetic Algorithm and Dynamic LoRAabstractAutomated Program Repair (APR) aims to automate the patch generation for buggy code and is vital in software devel-opment and maintenance. While large language models (LLMs) excel in various tasks, our empirical study shows they still face challenges in APR. LLMs take infilling templates with different qualities as input, and low-quality templates may misguide LLMs in constantly generating incorrect patches. Additionally, LLMs lack project-specific knowledge and struggle to leverage bug context fully. Therefore, we propose EvoAPR, which integrates the genetic algorithm and dynamic LoRA technique to enhance LLMs for better APR. First, to generate high-quality infilling templates, buggy codes are encoded to genetic representations, and genetic operators and multi-angle evaluation are designed to produce better templates. Then, to effectively utilize bug context, a bug-context-aware LoRA fine-tuning method is proposed to fuse various bug contexts by dynamically activating certain blocks of the LoRA module according to the bug index. Finally, these high-quality infilling templates are fed the fine-tuned LLMs for patch generation. Experimental results demonstrate that EvoAPR significantly improves LLMs' performance, surpassing state-of-the-art methods. Qingyang Yan, Weihuan Min, Li Kuang, Yingjie Xia |
ICWS | 3 |
| 2025 | DLCoG: A Novel Framework for Dual-Level Code Comment Generation Based on Semantic Segmentation and In-Context LearningabstractIn large software projects with collaborative development, comprehensive code comments are crucial for code readability and maintainability. Code comments mainly include method comments and inline comments, where the former describes the functionality globally, and the latter describes the implementation details locally. Existing methods typically generate these two kinds of comments with specific locations independently, which results in weak correlations between comments and code context, as well as high model inference costs due to long token inputs. To address these issues, we define the combination of inline comments and method comments as Dual-Level Code Comments. We formulate the novel task of automatically generate dual-level code comments based on given code and propose an approach named DLCoG (DualLevel Code Comment Generation) to automate this task. First, a Semantic Segmentation and Identification multi-task model based on CodeBERT, termed Se2Iden (Semantic Segmentation and Identification model), is proposed to identify code segments requiring inline comments. Next, we retrieve similar samples to adopting the in-context learning paradigm, which can enhance the generation quality of large language models (LLMs) in specific domains. Finally, the LLM is guided to generate duallevel code comments using Chain-of-Thought (CoT) prompts that first produce inline comments, followed by method comments. We manually constructed a high-quality clean Java dataset consisting of*> based on open-source Java projects by (i) determining comments type and (ii) manually associating inline comments with their corresponding code. Then, we trained a multi-task learning model based on CodeBERT to automatically take the two steps needed, termed ICSA (Inline Comment Classification and Scope Association), thus to expand to a dataset containing 80k dual-level code comments. Experimental results on clean and extended datasets show that DLCoG outperforms all baselines by substantial margins. The contextual information provided by DLCoG can effectively improve the inline comments generated by LLM. Coordinated generation of dual-level comment also brings effective improvements to method comments, which is particularly significant when there are few contextual examples. Our work fills the long-standing gap in the dual-level code comment generation field, and can provide insights for future research in this direction. We provide open-source datasets and source code for future research. Haiyang Yang, Qingyang Yan, Weihuan Min, Zhao Wei, Li Kuang, Yingjie Xia |
ICPC | 5 |
| 2024 | RegGPT: A Tool for Cross-Domain Service Regulation Language ConversionabstractDigital services have become essential to the modern service industry, offering great convenience to consumers. Because of the lack of regulation, it has posed an unprecedented challenge to the regulation of digital services. With the development of Regulatory Technology, many regulatory platforms have emerged. However, most of these platforms focus on a single service domain and are difficult to migrate to other domains to meet cross-domain regulatory demands. By extracting common elements from multi-domain regulatory rules, we propose a regulatory language called Cross-Domain Service Regulation Language (CDSRL), which aims to improve the comprehension of rules for machines, so the automated regulation can be achieved. Meanwhile, we construct the fine-tuned datasets for the regulatory domain and train the RegGPT based on the large language model, which can identify and classify natural language rules and automatically convert them into CDSRL. Experiments show that the language is highly comprehensible, scalable, and suitable for expressing rules in different domains and categories. The RegGPT shows a strong ability in the conversion process and improves regulatory efficiency. It provides a new scheme for automatically converting regulatory rule into rule language. Qi Xie 0010, Weihuan Min, Li Kuang |
ICWS | 4 |
| 2024 | Improving AST-Level Code Completion with Graph Retrieval and Multi-Field AttentionabstractCode completion, which provides code suggestions by generating code snippets or structures, has become an essential feature of integrated development environments (IDEs). Recently, some studies have begun to use graph neural networks to complete AST-level code, and shown that it is promising to introduce GNNs into ASTlevel completion. However, these methods do not fully exploit the potential of reference codes with similar structures nor solve out-of-vocabulary (OOV). We propose Retrieval-Assisted Graph Code Completion (ReGCC) to enhance AST-level code completion further. ReGCC integrates a retrieval model that searches for similar code graphs to generate graph nodes and a completion model that leverages information from multiple domains. The key component of both the retrieval and completion models is the Multi-field Graph Attention Block, which consists of three layers of stacked attention: (1) Neighborhood Attention: preserves the heterogeneity and local dependency of the graph, enabling nodes to exchange information within their neighborhood. (2) Global & Memory Attention: addresses the long-distance dependency problem by providing nodes with a global view and the ability to extract information from the memory domain. (3) Reference Attention: lets nodes obtain valuable information from structurally similar reference code graphs. Furthermore, we tackle the OOV issue by employing feature matching and copying values from existing nodes. Specifically, we predict edges between nodes beyond the vocabulary, enabling effective information transfer. Experimental results demonstrate the superiority of our approach over state-of-the-art AST-level completion methods and generative language models. Yu Xia 0010, Weihuan Min, Li Kuang |
ICPC | 3 |
| 2024 | A Just-in-time Software Defect Localization Method based on Code Graph RepresentationabstractTraditional software defect localization aims to locate defective files, methods, or code lines based on symptoms such as defect reports. In comparison, Just-In-Time (JIT) software defect localization focuses on identifying defective code lines when a defective code change is initially submitted. It can identify issues at the code line level before the defect becomes apparent, preventing it from adversely affecting the software. Although researchers have proposed various methods for JIT defect localization, existing methods still have the following shortcomings: (1) Most methods rely heavily on tokens from single code lines to calculate naturalness for defect localization, which makes it challenging to effectively distinguish between code lines that have the same content but different labels (defective code lines or non-defective code lines) - termed Duplicate Lines with Different Labels (DLDL). (2) Existing methods represent code in the form of sequences, neglecting the structural information of the code. Therefore, we propose a JIT defect localization method based on code graph representation. First, we construct code linelevel code graphs for code changes to distinguish DLDL explicitly. Next, to extract sequential and structural information from the code, we propose a code graph representation model with contrastive learning to generate graph feature vectors and node scores with rich semantics. Finally, we calculate the naturalness of code lines based on the graph feature vectors and node scores. Using this naturalness, we identify defective code lines. Experimental results show that our JIT defect localization method outperforms the state-of-the-art methods. Huan Zhang 0017, Weihuan Min, Zhao Wei, Li Kuang, Honghao Gao, Huaikou Miao |
ICPC | 2 |
| 2023 | Enhancing intelligent IoT services development by integrated multi-token code completion
Yu Xia 0010, Weihuan Min, Li Kuang, Honghao Gao |
Comput. Commun. | 3 |