VLDB 2026 Research / reviewers in the wild / expert
Xiaowei Zhang 0018
dblp:93/4664-18
· DBLP profile ↗
9ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0003-1481-5158ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 9 · 4 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AdaptGen: A Problem-Adaptive Solution Template Generation Technique for Online Programming PlatformsabstractOnline programming platforms that offer various programming tasks play a crucial role in helping programmers enhance their coding skills. Programming tasks posted for different application scenarios may follow the same or similar programming patterns in their solutions. This leads programmers to repeatedly write not just the core code of problem-solving, but also the same basic framework and some peripheral code (such as variable declarations and input/output handling) that are needed to make their solutions executable. Repeatedly writing boilerplate or well-mastered algorithmic frameworks wastes time and adds little value for programmers focused on skill-specific practice.Toward this, we propose to develop AdaptGen, a problemadaptive code template generation method for online programming platforms. AdaptGen analyzes and extracts solution patterns from various programming problem solutions (i.e., code accepted by online programming platforms) and generates templates tailored to each problem. More specifically, AdaptGen is built on genetic programming and uses a linear hashing sequence encoding strategy to represent solutions. It incorporates selection, crossover, and de-duplication operators to maintain diversity in the evolution process, and a fitness function tailored to generate solution templates. These templates are then abstracted and structured with a flexible core-code-hiding mechanism, enabling programmers of different experience levels to practice efficiently.We evaluated AdaptGen using two datasets from LeetCode and NowCoder, containing a total of 997 tasks and over 3,200 solution categories. Results show that AdaptGen successfully generates usable templates for 77%-84% of solution categories, with 80% of the templates performing well in manual evaluations. It also outperforms seven advanced representative large language models (LLMs), achieving the best overall performance in template quality, consistency, and generation efficiency. To validate Adapt- Gen’s effectiveness in real-world programming environments, we further conduct a user study involving live coding practice by programmers in online programming platforms, which effectively demonstrated its utility in practical application scenarios. Weiqin Zou, Xiaowei Zhang 0018, Jifeng Xuan |
IEEE Trans. Software Eng. | 3 |
| 2025 | CodeQG: Automated Multiple Question Generation for Source Code ComprehensionabstractDuring software maintenance and evolution, developers spend more than half of their time on code comprehension activities. In order to understand an unfamiliar code base, they would naturally ask different types of questions related to code snippets and try to find the answers. In this paper, we conduct an initial work to explore the possibility of automatic question generation for program comprehension. We construct a large-scale data set containing pairs of source code and questions that are automatically transformed from inline comments based on dependency analysis and semantic role labeling. We also build a comprehensive taxonomy of question types so as to generate questions concerning different aspects of code snippets, such as purpose, implementation details and so on. Then, we propose a deep learning-based prototype CodeQG to automatically generates multiple types of questions for code snippets. We evaluate CodeQG by using both typical performance metrics and manual evaluation. The results show that (1) we can achieve a value of 42.02 on BLEU4 and 60.81 on ROUGE-L for the generated questions; (2) overall, the questions are very correct in grammatical, semantic and format; (3) the questions are related to the corresponding code snippet and are helpful for developers in source code comprehension activities. Our work gives insights into automatically generating multiple types of questions for code comprehension. We expect this exploration will improve the applicability and generality of machine code comprehension. Xiaowei Zhang 0018, Lin Chen 0015, Kaiyuan Qi, Weiqin Zou, Liye Pang, Lianfa Zhang, Peng Zhang 0083, Guanqun Xu |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2024 | Efficient Construction of Practical Python Call Graphs with Entity Knowledge BaseabstractCall graphs facilitate various tasks in software engineering. However, for the dynamic language Python, the complex language features and external library dependencies pose enormous challenges for building the call graphs of real projects. Some program analysis techniques used for call graph construction in other languages are impractical for Python. In this paper, we present STAR, a practical technique for the construction of Python static call graphs. We reformulate call graph construction as an entity identification task. STAR leverages inter-module summary and cross-project dependencies to construct a fine-grained entity knowledge base to identify the possible nodes and edges of the call graph in the code, and then construct the call graph. Our evaluation of three benchmarks shows that (1) STAR improves recall in three benchmarks compared to three baseline tools. Especially, STAR improves the recall of reachable nodes and reachable edges compared with the state-of-the-art tool by 11.3% and 9.8%, respectively; (2) STAR achieves comparable performance as three baseline tools in execution time and memory usage and is more efficient in large projects; (3) STAR can be effectively used for the task of detecting vulnerability propagation with real-world cases. We expect our results will attract more exploration of practical methods and improve the application of Python call graphs. Yulu Cao, Lin Chen 0015, Zhifei Chen, Jiacheng Zhong, Xiaowei Zhang 0018, Linzhang Wang |
Int. J. Softw. Eng. Knowl. Eng. | 5 |
| 2024 | Multi-Intent Inline Code Comment Generation via Large Language ModelabstractCode comment generation typically refers to the process of generating concise natural language descriptions for a piece of code, which facilitates program comprehension activities. Inline code comments, as a part of code comments, are also crucial for program comprehension. Recently, the emergence of large language models (LLMs) has significantly boosted the performance of natural language processing tasks. This naturally inspires us to explore the performance of the LLMs in the task of inline code comment generation. To this end, we evaluate open-source LLMs on a large-scale dataset and compare the results with the current state-of-the-art methods. Specifically, we explore the model performance in the following scenarios based on the widely used evaluation metrics (i.e. BLEU, Meteor, and ROUGE-L): (1) generation with simple instruction; (2) few-shot-guided generation with random examples selected from the database; (3) few-shot-guided generation with similar examples selected from the database; and (4) adopt the re-ranking strategy for the output of LLMs. Our findings reveal that: (1) under the simple instruction scenario, LLMs could not fully show the potential in the task of inline comment generation compared to the state-of-the-art models; (2) random few-shot leads to a slight improvement; (3) similar few-shot and re-ranking strategy could significantly enhance the performance of LLMs; and (4) for inline comment and code snippet pairs with different intents, why category achieves the best performance and what category achieves relatively poorer performance. That remains consistent across all four scenarios. Our findings shed light on future research directions for using LLMs in inline comment generation tasks. Xiaowei Zhang 0018, Zhifei Chen, Yulu Cao, Lin Chen 0015, Yuming Zhou |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2024 | ICG: A Machine Learning Benchmark Dataset and Baselines for Inline Code Comments Generation TaskabstractAs a fundamental component of software documentation, code comments could help developers comprehend and maintain programs. Several datasets of method header comments have been proposed in previous studies for machine learning-based code comment generation. As part of code comments, inline code comments are also crucial for code understanding activities. However, unlike method header comments written in a standard format and describing the whole method code, inline comments are often written in arbitrary formats by developers due to timelines pressures and different aspects of code snippets in the method are described. Currently, there is no large-scale dataset used for inline comments generation considering these. Hence, this naturally inspires us to explore whether we can construct a dataset to foster machine learning research that not only performs fine-grained noise-cleaning but conducts a taxonomy of inline comments. To this end, we first collect inline comments and code snippets from 8000 Java projects on GitHub. Then, we conduct a manual review to obtain heuristic rules, which could be used to clean the data noise in a fine-grained manner. As a result, we construct a large-scale benchmark dataset named ICG with 5,740,770 pairs of inline comments and code snippets. We then build a comprehensive taxonomy and conduct a statistical and manual analysis to explore the performances of different categories of inline comments, such as helpfulness in code understanding. After that, we provide and compare several baseline models to automatically generate inline comments, such as CodeBERT, to enhance the usability of the benchmark for researchers. The availability of our benchmark and baselines can help develop and validate new inline comment generation methods, which would also further facilitate code understanding activities. Xiaowei Zhang 0018, Lin Chen 0015, Weiqin Zou, Yulu Cao, Hao Ren 0011, Yanhui Li 0001, Yuming Zhou |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2024 | Diagnosis of package installation incompatibility via knowledge base
Yulu Cao, Zhifei Chen, Xiaowei Zhang 0018, Yanhui Li 0001, Lin Chen 0015, Linzhang Wang |
Sci. Comput. Program. | 3 |
| 2024 | Just-in-time identification for cross-project correlated issuesabstractAbstract Issue tracking systems are now prevalent in software development, which would help developers submit and discuss issues to solve development problems on software projects. Most previous studies have been conducted to analyze issue relations within projects, such as recommending similar or duplicate bug issues. However, along with the popularization of co‐developing through multiple projects, many issues are cross‐project correlated (CPC), that is, one issue is associated with another issue in a different project. When developers meet with CPC issues, it may primarily increase the difficulties of solving them because they need information from not only their projects but also other related projects that developers are not familiar with. Identifying a CPC issue as early as possible is a fundamental challenge for both managers and developers to allocate the resources for software maintenance and estimate the effort to solve it. This paper proposes 11 issue metrics of two groups to describe textual summary and reporters' activity, which can be extracted just after the issue was reported. We employ these 11 issue metrics to construct just‐in‐time (JIT) prediction models to identify CPC issues. To evaluate the effect of CPC issue prediction models, we conduct experiments on 16 open‐source data science and deep learning projects and compare our prediction model with two baseline models based on textual features (i.e., Term Frequency‐Inverse Document Frequency [TF‐IDF] and Word Embedding), which are commonly adopted by previous studies on issue prediction. The results show that the JIT prediction model based on issue metrics has significantly improved the performance of CPC issue prediction under two evaluation indicators, Matthew's correlation coefficient (MCC) and F1. In addition, we find that the prediction model is more suitable for large‐scale complex core projects in the open‐source ecosystem. Hao Ren 0011, Yanhui Li 0001, Lin Chen 0015, Yulu Cao, Xiaowei Zhang 0018, Changhai Nie |
J. Softw. Evol. Process. | 5 |
| 2023 | Effective Recommendation of Cross-Project Correlated Issues based on Issue MetricsabstractThe calling relationship between projects becomes complicated as the number of open-source projects increases. Different issues across projects can also be related, referred to as cross-project correlated issues (CPCIs), and bring new challenges for developers to fix these issues. When solving these CPCIs, developers have to accurately locate the source code that causes it in the current project and also needs to know the related issues in other projects. However, few studies have proposed specific methods to help developers effectively address these CPCIs, i.e., find related issues for CPCIs. Hao Ren 0011, Mingliang Ma, Xiaowei Zhang 0018, Yulu Cao, Changhai Nie |
Internetware | 3 |
| 2023 | Towards the Analysis and Completion of Syntactic Structure Ellipsis for Inline CommentsabstractThe ellipsis of the syntactic structure is a common phenomenon in ordinary textual documents. Existing studies have found that despite syntactic ellipsis could help avoid repetition of normative documents, it could also, for example, lead to ambiguity and hamper the understandability of document contents. As a fundamental component of software, code comments are generally written by developers in a non-structured way just like normative documents. This naturally inspires us to explore whether syntactic ellipsis is also a common phenomenon in code comments and what potential negative effects would such ellipsis have on software tasks such as code/comments comprehension activities. Such explorations, in our opinion, are expected to facilitate the research on code comments and comments-related software tasks. To this end, we conduct the first large-scale study to explore the syntactic structure ellipsis problem of code comments, with a focus on Java inline comments. Specifically, we construct a data set of 1,000 Java projects with 1,307,457 inline comments and associated codes. Based on this data set, we first study the prevalence of syntactic structure ellipsis in inline comments. We find that syntactic structure ellipsis is quite common in inline comments where 83.6% comments have structure ellipsis (such as subject/predicate omissions). Then, we investigate the effects of syntactic structure ellipsis on code/comment understanding activities. As a result, we find that there indeed exists a negative relationship between them, with a medium effect size. Based on these findings, we further propose neural network based approaches to complete the ellipsis parts for the inline comments. With our approach, we could achieve: 1) a medium improvement in assisting code/comment understanding activities, and 2) a substantial improvement of 11.3% in comment-assisted code abbreviation extension task. Xiaowei Zhang 0018, Weiqin Zou, Lin Chen 0015, Yanhui Li 0001, Yuming Zhou |
IEEE Trans. Software Eng. | 1 |