Zhang Zhang 0005

dblp:94/2468-5 · DBLP profile ↗
← Back
16ranked-venue papers
7as first author
12since 2021 · last 2025
0000-0002-9914-026XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 14 · 6 first-author · 10 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Instruct or Interact? Exploring and Eliciting LLMs' Capability in Code Snippet Adaptation Through Prompt Engineering
abstract
Code snippet adaptation is a fundamental activity in the software development process. Unlike code generation, code snippet adaptation is not a “free creation”, which requires developers to tailor a given code snippet in order to fit specific requirements and the code context. Recently, large language models (LLMs) have confirmed their effectiveness in the code generation task with promising results. However, their performance on code snippet adaptation, a reuse-oriented and context-dependent code change prediction task, is still unclear. To bridge this gap, we conduct an empirical study to investigate the performance and issues of LLMs on the adaptation task. We first evaluate the adaptation performances of three popular LLMs and compare them to the code generation task. Our result indicates that their adaptation ability is weaker than generation, with a nearly 15% decrease on pass@1 and more context-related errors. By manually inspecting 200 cases, we further investigate the causes of LLMs' sub-optimal performance, which can be classified into three categories, i.e., Unclear Requirement, Requirement Misalignment and Context Misapplication. Based on the above empirical research, we propose an interactive prompting approach to eliciting LLMs' ability on the adaptation task. Specifically, we enhance the prompt by enriching the context and decomposing the task, which alleviates context misapplication and improves requirement understanding. Besides, we enable LLMs' reflection by requiring them to interact with a human or a LLM counselor, compensating for unclear requirement. Our experimental result reveals that our approach greatly improve LLMs' adaptation performance. The best-performing Human-LLM interaction successfully solves 159 out of the 202 identified defects and improves the pass@1 and pass@5 by over 40% compared to the initial instruction-based prompt. Considering human efforts, we suggest multi-agent interaction as a trade-off, which can achieve comparable performance with excellent generalization ability. We deem that our approach could provide methodological assistance for autonomous code snippet reuse and adaptation with LLMs.
Tanghaoran Zhang, Yue Yu 0001, Xinjun Mao, Shangwen Wang, Kang Yang 0001, Yao Lu 0003, Zhang Zhang 0005
ICSE7
2025 Large Language Models Are Qualified Benchmark Builders: Rebuilding Pre-Training Datasets for Advancing Code Intelligence Tasks
abstract
Pre-trained code models are essential for various code intelligence tasks. Yet, their effectiveness is heavily influenced by the quality of the pre-training dataset, particularly human-written reference comments, which usually serve as a bridge between the programming language and natural language. One significant challenge is that such comments could become inconsistent with the corresponding code as the software evolves, leading to suboptimal model performance. Large language models (LLMs) have demonstrated superior capabilities in generating high-quality code comments. This work investigates whether substituting original human-written comments with LLM-generated ones can improve pre-training datasets for more effective pretrained code models. As existing reference-based metrics cannot evaluate the quality of human-written reference comments themselves, to enable direct comparison between LLM-generated and human reference comments, we introduce two auxiliary tasks as novel reference-free metrics, including code-comment inconsistency detection and semantic code search. Experimental results show that LLM-generated comments exhibit superior semantic consistency with the code compared to human-written reference comments. Our manual evaluation also corroborates this conclusion, which indicates the potential of utilizing LLMs to enhance the quality of the pre-training dataset. Based on this finding, we rebuilt the CodeSearchNet dataset with LLM-generated comments and re-pre-trained the CodeT5 model. Evaluations on multiple code intelligence tasks demonstrate that models pretrained by LLM-enhanced data outperform their counterparts (pre-trained by original human reference comments data) on code summarization, code generation, and code translation tasks. This research validates the feasibility of rebuilding the pre-training dataset by LLMs to advance code intelligence tasks. It advocates rethinking the reliance on human reference comments for coderelated tasks.
Kang Yang 0001, Xinjun Mao, Shangwen Wang, Yanlin Wang 0001, Tanghaoran Zhang, Bo Lin 0011, Yihao Qin, Zhang Zhang 0005, Yao Lu 0003, Kamal Al-Sabahi
ICPC8
2025 AdaptEval: A Benchmark for Evaluating Large Language Models on Code Snippet Adaptation
abstract
Recent advancements in large language models (LLMs) have automated various software engineering tasks, with benchmarks emerging to evaluate their capabilities. However, for adaptation, a critical activity during code reuse, there is no benchmark to assess LLMs’ performance, leaving their practical utility in this area unclear. To fill this gap, we propose AdaptEval, a benchmark designed to evaluate LLMs on code snippet adaptation. Unlike existing benchmarks, AdaptEval incorporates the following three distinctive features: First, practical context. Tasks in AdaptEval are derived from developers’ practices, preserving rich contextual information from Stack Overflow and GitHub communities. Second, multi-granularity annotation. Each task is annotated with requirements at both task and adaptation levels, supporting the evaluation of LLMs across diverse adaptation scenarios. Third, fine-grained evaluation. AdaptEval includes a two-tier testing framework combining adaptation-level and function-level tests, which enables evaluating LLMs’ performance across various individual adaptations. Based on AdaptEval, we conduct the first empirical study to evaluate six instruction-tuned LLMs and especially three reasoning LLMs on code snippet adaptation. Experimental results demonstrate that AdaptEval enables the assessment of LLMs’ adaptation capabilities from various perspectives. It also provides critical insights into their current limitations, particularly their struggle to follow explicit instructions. We hope AdaptEval can facilitate further investigation and enhancement of LLMs’ capabilities in code snippet adaptation, supporting their real-world applications.
Tanghaoran Zhang, Xinjun Mao, Shangwen Wang, Yao Lu 0003, Zhang Zhang 0005, Kang Yang 0001, Yue Yu 0001
ASE7
2025 Improving API Knowledge Comprehensibility: A Context-Dependent Entity Detection and Context Completion Approach Using LLM
abstract
Extracting API knowledge from Stack Overflow has become a crucial way to assist developers in using APIs. Existing research has primarily focused on extracting relevant API-related knowledge at the sentence level to enhance API documentation. However, this level of extraction can lead to a loss of crucial context, especially when sentences contain context-dependent entities (i.e., whose understanding requires reference to the surrounding context) that may hinder developers' understanding. To investigate this issue, we conducted an empirical study of 384 Stack Overflow posts and found that (1) approximately one-third of API functionality sentences contain context-dependent entities, and (2) these entities fall into two categories: Referential ContextDependent Entities and Local Variable Context-Dependent Entities. In response, we developed a novel method, CEDCC, which combines an entity filtering strategy informed by insights from our empirical study, with a large language model (LLM) to construct coreference chains for detecting context-dependent entities. Additionally, it employs a step-by-step approach with the LLM to complete the necessary context for understanding these entities. To evaluate CEDCC, we constructed a dataset of 1,023 API knowledge sentences, including 567 context-dependent entities and their required contexts. The results demonstrate the effectiveness of CEDCC in accurately detecting contextdependent entities and completing context tasks, achieving an F1score of 0.865 and a BERTScore of 0.373, significantly surpassing the baseline methods. Human evaluations further confirmed that CEDCC effectively improves the comprehensibility of API knowledge sentences.
Zhang Zhang 0005, Xinjun Mao, Shangwen Wang, Kang Yang 0001, Tanghaoran Zhang, Xunhui Zhang
SANER1
2025 ConflictLens: an LLM-Based Method for Detecting Semantic Merge Conflicts
abstract
Semantic conflicts in branch merging occur when merged code violates specifications from one or both branches.These conflicts are often subtle and can lead to serious runtime errors such as crashes or data corruption.Existing detection methods fail to achieve both high precision and recall: Static analysis-based methods ensure high recall but lack precision, whereas dynamic execution-based methods provide better precision but struggle with recall due to limited test coverage.To better understand such conflicts, we first conduct an empirical study on a real-world merge dataset and identify four common conflict patterns.These patterns reveal key characteristics of semantic conflicts and serve as guidance for automated detection.Based on these insights, we propose ConflictLens, a two-stage LLM-based method that combines static analysis and dynamic execution to balance precision and recall.First, LLMs are guided by few-shot and chain-of-thought prompting using the patterns to localize conflicts statically.Then, conflicts are dynamically verified with LLM-generated targeted tests, refined through execution feedback.Evaluated on 85 real-world merge scenarios, ConflictLens achieves 0.91 precision and 0.76 recall, outperforming static and dynamic baselines.Ablation studies demonstrate the contribution and synergy of each component.Cross-LLM evaluations confirm robustness, with DeepSeek-R1 performing best and cost-efficient models like GPT-4o-Mini still competitive.
Longfei Sun, Yao Lu 0003, Xinjun Mao, Tanghaoran Zhang, Zhang Zhang 0005
SEKE5
2025 ROS package search for robot software development: a knowledge graph-based approach
Xinjun Mao, Shuo Yang 0005, Menghan Wu 0001, Zhang Zhang 0005
Frontiers Comput. Sci.5
2025 CARLDA: An Approach for Stack Overflow API Mention Recognition Driven by Context and LLM-Based Data Augmentation
abstract
ABSTRACT The recognition of Application Programming Interface (API) mentions in software‐related texts is vital for extracting API‐related knowledge, providing deep insights into API usage and enhancing productivity efficiency. Previous research identifies two primary technical challenges in this task: (1) differentiating APIs from common words and (2) identifying morphological variants of standard APIs. While deep learning‐based methods have demonstrated advancements in addressing these challenges, they rely heavily on high‐quality labeled data, leading to another significant data‐related challenge: (3) the lack of such high‐quality data due to the substantial effort required for labeling. To overcome these challenges, this paper proposes a context‐aware API recognition method named CARLDA. This approach utilizes two key components, namely, Bidirectional Encoder Representations from Transformers (BERT) and Bidirectional Long Short‐Term Memory (BiLSTM), to extract context at both the word and sequence levels, capturing syntactic and semantic information to address the first challenge. For the second challenge, it incorporates a character‐level BiLSTM with an attention mechanism to grasp global character‐level context, enhancing the recognition of morphological features of APIs. To address the third challenge, we developed specialized data augmentation techniques using large language models (LLMs) to tackle both in‐library and cross‐library data shortages. These techniques generate a variety of labeled samples through targeted transformations (e.g., replacing tokens and restructuring sentences) and hybrid augmentation strategies (e.g., combining real‐world and generated data while applying style rules to replicate authentic programming contexts). Given the uncertainty about the quality of LLM‐generated samples, we also developed sample selection algorithms to filter out low‐quality samples (i.e., incomplete or incorrectly labeled samples). Moreover, specific datasets have been constructed to evaluate CARLDA's ability to address the aforementioned challenges. Experimental results demonstrate that (1) CARLDA significantly enhances F1 by 11.0% and the Matthews correlation coefficient (MCC) by 10.0% compared to state‐of‐the‐art methods, showing superior overall performance and effectively tackling the first two challenges, and (2) LLM‐based data augmentation techniques successfully yield high‐quality labeled data and effectively alleviate the third challenge.
Zhang Zhang 0005, Xinjun Mao, Shangwen Wang, Kang Yang 0001, Tanghaoran Zhang, Yao Lu 0003
J. Softw. Evol. Process.1
2024 CAREER: Context-Aware API Recognition with Data Augmentation for API Knowledge Extraction
abstract
The recognition of Application Programming Interface (API) mentions in the software-related texts is a prerequisite task for extracting API-related knowledge. Previous studies have demonstrated the superiority of deep learning-based methods in accomplishing this task. However, such techniques still meet their bottlenecks due to their inability to effectively handle the following three challenges: (1) differentiating APIs from common words; (2) identifying APIs in morphological variants of the standard APIs; and (3) the lack of high-quality labeled data for training. To overcome these challenges, this paper proposes a context-aware API recognition method named CAREER. This approach utilizes two key components, namely Bidirectional Encoder Representations from Transformers (BERT) and Bi-directional Long Short-Term Memory (BiLSTM), to extract context information at both the word-level and sequence-level. This strategic combination empowers the method to dynamically capture both syntactic and semantic information, effectively addressing the first challenge. To tackle the second challenge, CAREER introduces a character-level BiLSTM component, enriched with an attention mechanism. This enables the model to grasp character-level global context information, thereby enhancing the recognition of morphological attributes within API mentions. Furthermore, to address the third challenge, the paper introduces three data augmentation techniques aimed at generating new data samples. Accompanying these techniques is a novel sample selection algorithm designed to screen out high-quality instances. This dual-pronged approach effectively mitigates the requirement for data labeling. Experiments demonstrate that CAREER significantly improves F1-score by 11.0% compared with state-of-the-art methods. We also construct specific datasets to assess CAREER's capacity to tackle the aforementioned challenges. Results confirm that (1) CAREER significantly outperforms baseline methods in addressing the first and second challenges, and (2) with the aid of data augmentation techniques and sample selection algorithms, high-quality samples can be generated to improve the performance, and alleviate the third challenge.
Zhang Zhang 0005, Xinjun Mao, Shangwen Wang, Kang Yang 0001, Yao Lu 0003
ICPC1
2023 MUSE: A Multi-Feature Semantic Fusion Method for ROS Node Search Based on Knowledge Graph
abstract
Reusing ROS components, specifically ROS Nodes, is crucial for improving the efficiency and quality of robotic software development. However, developers face challenges in finding the desired ROS Nodes for reuse due to scattered organization of ROS Nodes and the ambiguity in their ROS Node name. To address these challenges, this paper proposes a MUlti-feature SEmantic fusion method (MUSE) that leverages a domain-specific ROS knowledge graph for searching ROS Nodes. Firstly, a large dataset is constructed, comprising code files and textual descriptions related to ROS Nodes obtained from GitHub and ROS Wiki. Secondly, an in-depth analysis of user queries regarding the reuse of ROS Nodes is conducted, leading to the selection of multiple features that provide a comprehensive representation of ROS Node semantics, including Function, Hardware, Input, and Output. Subsequently, a knowledge graph of ROS Nodes is developed based on the dataset, incorporating the selected features. This knowledge graph effectively organizes scattered knowledge and resolves the issue of diverse mentions through entity disambiguation and resolution. To eliminate the semantic gap between the descriptions of features mentioned in user queries and the entities in the knowledge graph, a pretrained transformer-based model was used to measure the multi-feature semantic similarity between user queries and ROS Nodes knowledge. Finally, we employ a linear regression model to integrate the multi-feature knowledge between user queries and ROS Nodes knowledge. The proposed method has shown a 20% improvement in performance on NDCG@1 compared to other ROS Node search methods. Further evaluations highlight the effectiveness of each feature incorporated in the knowledge graph, as well as the significance of each parameter within the regression model. These findings underscore the robustness of this research in optimizing the reuse of ROS Nodes and facilitating the development of robotics software.
Xinjun Mao, Tanghaoran Zhang, Zhang Zhang 0005
APSEC4
2022 The Maintenance of Top Frameworks and Libraries Hosted on GitHub: An Empirical Study
abstract
The number of repositories on GitHub is huge and growing rapidly.However, most repositories are inactive, while active maintenance is essential in choosing a project.In this paper, we study the maintenance of top (i.e., most starred) frameworks and libraries hosted on GitHub, for they can be widely reused and in critical positions in the dependency network, so their maintenance status is significant.Furthermore, their maintenance practices may inspire other projects to thrive on the collaborative development platform.By investigating their adoption of recommended Open Source Software (OSS) maintenance practices and recent maintenance activities, and the association between maintenance status and usage, we find that: (1) more than 20% of the top frameworks and libraries have no commit for more than one year; (2) Some maintenance practices (e.g., codes of conduct) have relatively low adoption rates, while continuous integration has a high adoption rate of around 80%; (3) the maintenance status may have an effect on the usage frequency.
Xinjun Mao, Zhang Zhang 0005
SEKE3
2022 On the Way to Microservices: Exploring Problems and Solutions from Online Q&A Community
abstract
Microservice architecture is a dominant architectural style in SaaS industry, which helps to develop a single application as a collection of independent, well-defined, and inter-communicating services. The number of microservice-related questions in Q&Awebsites, such as Stack Overflow, has expanded substantially over the last years. Due to its increasing popularity, it is essential to understand the existing problems that microservice developers face in practices as well as the potential solutions to these problems. Such an investigation of problems and solutions is vital for long-term, impactful, and qualified research and practices in microservice community. Unfortunately, we currently know relatively little about such knowledge. To fill this gap, we conduct a large-scale in-depth empirical study on 17,522 Stack Overflow microservice-related posts. Our analysis leads to the first taxonomy of microservice-related topics based on the software development process. By analyzing the characteristics of the accepted answers, we find that there are fewer experts in the microservice than other domains, and such a phenomenon is most significant with respect to the microservice design phase. Furthermore, we perform manual analysis on 6,013 answers accepted by developers and distill 47 general solution strategies for different microservice-related problems, 22 of which are proposed for the first time. For instance, several problems inherent in the delivery phase can be lessened by referring to external sources like GitHub code examples. Our findings can therefore facilitate research and development on emerging microservice systems.
Menghan Wu 0001, Yang Zhang 0026, Shangwen Wang, Zhang Zhang 0005, Xin Xia 0001, Xinjun Mao
SANER5
2022 ForkXplorer: an approach of fork summary generation
Zhang Zhang 0005, Xinjun Mao, Yao Lu 0003
Frontiers Comput. Sci.1
2020 Understanding the Non-Repairability Factors of Automated Program Repair Techniques
abstract
Automated Program Repair (APR) is becoming a hot topic in Software Engineering community with many approaches being proposed and experiments being performed over the years. The results obtained from different experiments can be used as practical guidance to advance APR techniques. However, researchers have generally ignored the biases with respect to the unexpected results generated by various APR techniques, in which case the repair process cannot be finished normally and is terminated with unexpected exceptions (referred to as the non-repairability factors). In this paper, we aim to thoroughly understand the reasons for such non-repairability factors of various APR techniques, thus to provide practical insights for diverse stakeholders to establish an unbiased evaluation of APR techniques. To achieve so, we performed a systematic study on the existing execution logs that are ended with unexpected exceptions collected from different APR studies. Specifically, we investigated different types of exceptions with their frequencies of occurrence, the behind reasons of such occurrences, as well as the impact of such exceptions on the repairability of APR techniques. Our experimental results reveal that: 1) non-repairability factors happen in 25.7% of our studied logs and are widespread among diverse combinations of APR tools with FL strategies; 2) Inherent defect of APR tools is the most common reason for the occurrence of the non-repairability factors; 3) the impact of the non-repairability factors on the performance of APR tools can be rather significant. Our empirical study indicates that it is of great importance to eliminate the biases from the non-repairability factors. We also highlight several implications for actions that we can take to eliminate such biases.
Bo Lin 0011, Shangwen Wang, Ming Wen 0001, Zhang Zhang 0005, Yihao Qin, Xiaoguang Mao
APSEC4
2020 An Empirical Study on the Influence of Social Interactions for the Acceptance of Answers in Stack Overflow
abstract
In knowledge-sharing communities like Stack Overflow (SO), users can post questions, give answers and choose one answer as an accepted answer. The accepted answers will be important references for users when they encounter similar questions. Essentially, posting questions and giving answers is an interactive process occurring among community users, and choosing accepted answers is actually a decision-making process involving multiple factors. Previous works examined the impact on this decision process from the user, question and answer viewpoints. Social interactions between the questioners and answerers, although being popular according to our pre-analysis, have never been considered as a factor that can influence the decisions. To fill this gap, this paper first proposes a comprehensive answer acceptance model that integrates the answer features established by social interactions as well as information of users, questions and answers. We then divide social interactions into two stages and propose a method to calculate the relationship between the questioner and the answerer by analyzing these social interactions. Finally, we investigate the influence of social interactions for the acceptance of answers by performing logistic regression analysis. The results reveal several findings: (1) social-based features explain 16.6 % of the variance explained together, indicating that social interactions have significant and important effects on the acceptance of answers; (2) social interactions that occur after the answer is posted are more influential than these occur before the answer is posted. Based on the findings, we further conduct an online study of 132 SO users, and the respondents report that social interactions have a greater impact on the acceptance of answers than other judgments of answers such as upvotes, downvotes and not accepting answers.
Zhang Zhang 0005, Xinjun Mao, Yao Lu 0003, Shangwen Wang, Jinyu Lu
APSEC1
2020 Who Should Close the Questions: Recommending Voters for Closing Questions Based on Tags
Zhang Zhang 0005, Xinjun Mao, Yao Lu 0003, Jinyu Lu
SEKE1
2020 Automatic Voter Recommendation Method for Closing Questions in Stack Overflow
abstract
Stack Overflow is the most popular programming question and answer community that continuously receives a large number of questions every day. To ensure the quality of questions, the community grants privileges for the moderators and a group of experienced users to review the quality of questions and close the low-quality ones (e.g. duplicate or irrelevant questions). The review process is a typical crowdsourcing job that relies on users’ volunteer participation, and the current practices of closing questions in Stack Overflow face two aspects of challenges: (1) an obvious increase in both the absolute number and the percentage of “closed” questions; (2) a considerable decrease in participation willingness of experienced users to close questions. In order to solve the problem, we present a novel model of user willingness for reviewing and voting questions by incorporating four types of user activity history, including questions, answers, comments and votes of closing questions. Then we propose an automatic recommendation method based on the model to assign experienced users proper questions, to utilize the forces of them to close questions. The evaluation shows that the successful recommendation probability in the top 5, top 10, top 20, top 30, top 40, top 50 users are 48.23%, 58.93%, 68.83%, 74.27%, 78.13% and 81%, respectively.
Zhang Zhang 0005, Xinjun Mao, Yao Lu 0003, Jinyu Lu, Yue Yu 0001
Int. J. Softw. Eng. Knowl. Eng.1