EDBT 2026 Demo / reviewers in the wild / expert
Tanghaoran Zhang
dblp:329/3958
· DBLP profile ↗
15ranked-venue papers
4as first author
14since 2021 · last 2027
0000-0001-7241-9730ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 14 · 4 first-author · 14 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Evaluating LLMs in ROS robotic software code generation
Xinjun Mao, Tanghaoran Zhang, Zhiqun Xiao |
Empir. Softw. Eng. | 3 |
| 2025 | Instruct or Interact? Exploring and Eliciting LLMs' Capability in Code Snippet Adaptation Through Prompt EngineeringabstractCode snippet adaptation is a fundamental activity in the software development process. Unlike code generation, code snippet adaptation is not a “free creation”, which requires developers to tailor a given code snippet in order to fit specific requirements and the code context. Recently, large language models (LLMs) have confirmed their effectiveness in the code generation task with promising results. However, their performance on code snippet adaptation, a reuse-oriented and context-dependent code change prediction task, is still unclear. To bridge this gap, we conduct an empirical study to investigate the performance and issues of LLMs on the adaptation task. We first evaluate the adaptation performances of three popular LLMs and compare them to the code generation task. Our result indicates that their adaptation ability is weaker than generation, with a nearly 15% decrease on pass@1 and more context-related errors. By manually inspecting 200 cases, we further investigate the causes of LLMs' sub-optimal performance, which can be classified into three categories, i.e., Unclear Requirement, Requirement Misalignment and Context Misapplication. Based on the above empirical research, we propose an interactive prompting approach to eliciting LLMs' ability on the adaptation task. Specifically, we enhance the prompt by enriching the context and decomposing the task, which alleviates context misapplication and improves requirement understanding. Besides, we enable LLMs' reflection by requiring them to interact with a human or a LLM counselor, compensating for unclear requirement. Our experimental result reveals that our approach greatly improve LLMs' adaptation performance. The best-performing Human-LLM interaction successfully solves 159 out of the 202 identified defects and improves the pass@1 and pass@5 by over 40% compared to the initial instruction-based prompt. Considering human efforts, we suggest multi-agent interaction as a trade-off, which can achieve comparable performance with excellent generalization ability. We deem that our approach could provide methodological assistance for autonomous code snippet reuse and adaptation with LLMs. Tanghaoran Zhang, Yue Yu 0001, Xinjun Mao, Shangwen Wang, Kang Yang 0001, Yao Lu 0003, Zhang Zhang 0005 |
ICSE | 1 |
| 2025 | Understanding the Faults in Serverless Computing Based Applications: An Empirical StudyabstractServerless computing is a novel cloud computing paradigm that enables developers to develop, deploy, and run applications in the cloud without complex and error-prone cloud resource management. However, its characteristics also introduce new types of faults (e.g., faults due to insufficient computing resource allocation) and challenges to serverless computing-based applications (abbreviated as serverless applications). While prior studies have highlighted that serverless developers encounter various challenges, no attempts have been made to understand the faults in serverless applications. These faults may cause catastrophic consequences such as application crash, thereby hindering the further spread of serverless computing. We aim in this paper to understand the symptoms, root causes, and fix patterns of faults in serverless applications. To this end, we conduct an empirical study investigating developers' issues on GitHub and posts on Stack Overflow (SO). We first identify 546 real-world serverless-related faults from GitHub and SO. Then, we manually analyze and construct taxonomies of the symptoms, root causes, and fix patterns for these faults, respectively. Our study leads to the first taxonomy for symptoms of serverlessrelated faults, covering 5 categories and 21 subcategories. The findings of our study inform that the Permission Denied error is the most common type ($\mathbf{1 0. 8 1 \%}$) of faults. Furthermore, the Incorrect Code Logic is the main cause ($\mathbf{1 7. 9 5 \%}$) behind the faults. Furthermore, we summarize 15 fix patterns that can resolve$\mathbf{7 3. 6 3 \%}$of faults in this study. Based on the results, we provide actionable implications that can potentially facilitate research and assist developers in improving the development of serverless applications. Finally, we implement a knowledge-based Q&A tool named SafHelper to help developers understand and fix faults. Changrong Xie, Yang Zhang 0026, Xinjun Mao, Kang Yang 0001, Tanghaoran Zhang |
ICSME | 5 |
| 2025 | Large Language Models Are Qualified Benchmark Builders: Rebuilding Pre-Training Datasets for Advancing Code Intelligence TasksabstractPre-trained code models are essential for various code intelligence tasks. Yet, their effectiveness is heavily influenced by the quality of the pre-training dataset, particularly human-written reference comments, which usually serve as a bridge between the programming language and natural language. One significant challenge is that such comments could become inconsistent with the corresponding code as the software evolves, leading to suboptimal model performance. Large language models (LLMs) have demonstrated superior capabilities in generating high-quality code comments. This work investigates whether substituting original human-written comments with LLM-generated ones can improve pre-training datasets for more effective pretrained code models. As existing reference-based metrics cannot evaluate the quality of human-written reference comments themselves, to enable direct comparison between LLM-generated and human reference comments, we introduce two auxiliary tasks as novel reference-free metrics, including code-comment inconsistency detection and semantic code search. Experimental results show that LLM-generated comments exhibit superior semantic consistency with the code compared to human-written reference comments. Our manual evaluation also corroborates this conclusion, which indicates the potential of utilizing LLMs to enhance the quality of the pre-training dataset. Based on this finding, we rebuilt the CodeSearchNet dataset with LLM-generated comments and re-pre-trained the CodeT5 model. Evaluations on multiple code intelligence tasks demonstrate that models pretrained by LLM-enhanced data outperform their counterparts (pre-trained by original human reference comments data) on code summarization, code generation, and code translation tasks. This research validates the feasibility of rebuilding the pre-training dataset by LLMs to advance code intelligence tasks. It advocates rethinking the reliance on human reference comments for coderelated tasks. Kang Yang 0001, Xinjun Mao, Shangwen Wang, Yanlin Wang 0001, Tanghaoran Zhang, Bo Lin 0011, Yihao Qin, Zhang Zhang 0005, Yao Lu 0003, Kamal Al-Sabahi |
ICPC | 5 |
| 2025 | AdaptEval: A Benchmark for Evaluating Large Language Models on Code Snippet AdaptationabstractRecent advancements in large language models (LLMs) have automated various software engineering tasks, with benchmarks emerging to evaluate their capabilities. However, for adaptation, a critical activity during code reuse, there is no benchmark to assess LLMs’ performance, leaving their practical utility in this area unclear. To fill this gap, we propose AdaptEval, a benchmark designed to evaluate LLMs on code snippet adaptation. Unlike existing benchmarks, AdaptEval incorporates the following three distinctive features: First, practical context. Tasks in AdaptEval are derived from developers’ practices, preserving rich contextual information from Stack Overflow and GitHub communities. Second, multi-granularity annotation. Each task is annotated with requirements at both task and adaptation levels, supporting the evaluation of LLMs across diverse adaptation scenarios. Third, fine-grained evaluation. AdaptEval includes a two-tier testing framework combining adaptation-level and function-level tests, which enables evaluating LLMs’ performance across various individual adaptations. Based on AdaptEval, we conduct the first empirical study to evaluate six instruction-tuned LLMs and especially three reasoning LLMs on code snippet adaptation. Experimental results demonstrate that AdaptEval enables the assessment of LLMs’ adaptation capabilities from various perspectives. It also provides critical insights into their current limitations, particularly their struggle to follow explicit instructions. We hope AdaptEval can facilitate further investigation and enhancement of LLMs’ capabilities in code snippet adaptation, supporting their real-world applications. Tanghaoran Zhang, Xinjun Mao, Shangwen Wang, Yao Lu 0003, Zhang Zhang 0005, Kang Yang 0001, Yue Yu 0001 |
ASE | 1 |
| 2025 | Improving API Knowledge Comprehensibility: A Context-Dependent Entity Detection and Context Completion Approach Using LLMabstractExtracting API knowledge from Stack Overflow has become a crucial way to assist developers in using APIs. Existing research has primarily focused on extracting relevant API-related knowledge at the sentence level to enhance API documentation. However, this level of extraction can lead to a loss of crucial context, especially when sentences contain context-dependent entities (i.e., whose understanding requires reference to the surrounding context) that may hinder developers' understanding. To investigate this issue, we conducted an empirical study of 384 Stack Overflow posts and found that (1) approximately one-third of API functionality sentences contain context-dependent entities, and (2) these entities fall into two categories: Referential ContextDependent Entities and Local Variable Context-Dependent Entities. In response, we developed a novel method, CEDCC, which combines an entity filtering strategy informed by insights from our empirical study, with a large language model (LLM) to construct coreference chains for detecting context-dependent entities. Additionally, it employs a step-by-step approach with the LLM to complete the necessary context for understanding these entities. To evaluate CEDCC, we constructed a dataset of 1,023 API knowledge sentences, including 567 context-dependent entities and their required contexts. The results demonstrate the effectiveness of CEDCC in accurately detecting contextdependent entities and completing context tasks, achieving an F1score of 0.865 and a BERTScore of 0.373, significantly surpassing the baseline methods. Human evaluations further confirmed that CEDCC effectively improves the comprehensibility of API knowledge sentences. Zhang Zhang 0005, Xinjun Mao, Shangwen Wang, Kang Yang 0001, Tanghaoran Zhang, Xunhui Zhang |
SANER | 5 |
| 2025 | ConflictLens: an LLM-Based Method for Detecting Semantic Merge ConflictsabstractSemantic conflicts in branch merging occur when merged code violates specifications from one or both branches.These conflicts are often subtle and can lead to serious runtime errors such as crashes or data corruption.Existing detection methods fail to achieve both high precision and recall: Static analysis-based methods ensure high recall but lack precision, whereas dynamic execution-based methods provide better precision but struggle with recall due to limited test coverage.To better understand such conflicts, we first conduct an empirical study on a real-world merge dataset and identify four common conflict patterns.These patterns reveal key characteristics of semantic conflicts and serve as guidance for automated detection.Based on these insights, we propose ConflictLens, a two-stage LLM-based method that combines static analysis and dynamic execution to balance precision and recall.First, LLMs are guided by few-shot and chain-of-thought prompting using the patterns to localize conflicts statically.Then, conflicts are dynamically verified with LLM-generated targeted tests, refined through execution feedback.Evaluated on 85 real-world merge scenarios, ConflictLens achieves 0.91 precision and 0.76 recall, outperforming static and dynamic baselines.Ablation studies demonstrate the contribution and synergy of each component.Cross-LLM evaluations confirm robustness, with DeepSeek-R1 performing best and cost-efficient models like GPT-4o-Mini still competitive. Longfei Sun, Yao Lu 0003, Xinjun Mao, Tanghaoran Zhang, Zhang Zhang 0005 |
SEKE | 4 |
| 2025 | CARLDA: An Approach for Stack Overflow API Mention Recognition Driven by Context and LLM-Based Data AugmentationabstractABSTRACT The recognition of Application Programming Interface (API) mentions in software‐related texts is vital for extracting API‐related knowledge, providing deep insights into API usage and enhancing productivity efficiency. Previous research identifies two primary technical challenges in this task: (1) differentiating APIs from common words and (2) identifying morphological variants of standard APIs. While deep learning‐based methods have demonstrated advancements in addressing these challenges, they rely heavily on high‐quality labeled data, leading to another significant data‐related challenge: (3) the lack of such high‐quality data due to the substantial effort required for labeling. To overcome these challenges, this paper proposes a context‐aware API recognition method named CARLDA. This approach utilizes two key components, namely, Bidirectional Encoder Representations from Transformers (BERT) and Bidirectional Long Short‐Term Memory (BiLSTM), to extract context at both the word and sequence levels, capturing syntactic and semantic information to address the first challenge. For the second challenge, it incorporates a character‐level BiLSTM with an attention mechanism to grasp global character‐level context, enhancing the recognition of morphological features of APIs. To address the third challenge, we developed specialized data augmentation techniques using large language models (LLMs) to tackle both in‐library and cross‐library data shortages. These techniques generate a variety of labeled samples through targeted transformations (e.g., replacing tokens and restructuring sentences) and hybrid augmentation strategies (e.g., combining real‐world and generated data while applying style rules to replicate authentic programming contexts). Given the uncertainty about the quality of LLM‐generated samples, we also developed sample selection algorithms to filter out low‐quality samples (i.e., incomplete or incorrectly labeled samples). Moreover, specific datasets have been constructed to evaluate CARLDA's ability to address the aforementioned challenges. Experimental results demonstrate that (1) CARLDA significantly enhances F1 by 11.0% and the Matthews correlation coefficient (MCC) by 10.0% compared to state‐of‐the‐art methods, showing superior overall performance and effectively tackling the first two challenges, and (2) LLM‐based data augmentation techniques successfully yield high‐quality labeled data and effectively alleviate the third challenge. Zhang Zhang 0005, Xinjun Mao, Shangwen Wang, Kang Yang 0001, Tanghaoran Zhang, Yao Lu 0003 |
J. Softw. Evol. Process. | 5 |
| 2024 | An Empirical Study of Cross-Project Pull Request Recommendation in GitHubabstractAs a core contribution merge mechanism in distributed collaborative development, pull requests contain valuable knowledge of code evolution and issue resolution. With the co-evolution of multiple projects in a software ecosystem, relevant and similar issues can arise across different projects. Leveraging existing solutions in pull requests (PRs) through cross-project pull request recommendation (CPR) can enrich context knowledge and improve the efficiency of issue resolution. However, the characteristics of CPR and its effectiveness in the process of issue resolution still remain unclear. To bridge this gap, we conduct an empirical study of the CPR on GitHub. We first extract 4,445 CPRs from 2,500 open source projects and quantitatively analyze the characteristics of CPR. Then we conduct a qualitative analysis of sampled CPR cases to understand the influence of CPR. We also use a regression model to explore the impact of CPRs on issue resolution. Our main findings are as follows: (1) Experienced contributors in target projects make most of the CPRs and their CPRs are more timely than inexperienced contributors; (2) In CPR dataset, bugs constitute the largest proportion of target issue types, followed by enhancements, features and questions; (3) Nearly half of the CPRs are accepted by issue participants; (4) A greater number of the CPRs contribute indirectly to solving the target issue by offering solutions and contextual information, rather than providing appropriate code that can be directly applied to the issue; (5) Most of CPR-related factors have a significant impact on issue resolution delay. Among these, recommendation latency has the most significant impact, followed by the type of recommender. Our work has important insights into CPR and offers important guidance for developers on recommending cross-project PRs to resolve the mushrooming issues. Wenyu Xu, Yao Lu 0003, Xunhui Zhang, Tanghaoran Zhang, Bo Lin 0011, Xinjun Mao |
APSEC | 4 |
| 2024 | How Do Developers Adapt Code Snippets to Their Contexts? An Empirical Study of Context-Based Code Snippet AdaptationsabstractReusing code snippets from online programming Q&A communities has become a common development practice, in which developers often need to adapt code snippets to their code contexts to satisfy their own programming needs. However, how developers make these code adaptations based on contexts is still unclear. To bridge this gap, we first conduct a semi-structured interview of 21 developers to investigate their adaptation practices and perceived challenges during this process. The result suggests that code snippet adaptation is a challenging and exhausting task for developers, as they should tailor the snippets to guarantee their correctness and quality with laborious work. We also note that developers all resort to their intra-file context to complete adaptations, which motivates us to further study how developers performed context-based adaptations (CAs) in real scenarios. To this end, we conduct a quantitative study on an adaptation dataset comprising 300 code snippet reuse cases with 1,384 adaptations from Stack Overflow to GitHub. For each adaptation, we manually annotate its intention and relationship with the context. Based on our annotated data, we employ frequent itemset mining to obtain four CA patterns from our dataset, includingFortification,Code Wiring,Attribute-izationandParameterization. Our main findings reveal that: (1) more than half of the code snippet reuse cases include CAs and 23.3% of the adaptations are CAs; (2) more than half of the CAs are corrective adaptations and variable is the primary adapted language construct; (3) attribute is the most frequently utilized context and 88% of the local contexts are within the nearest 10 LOCs; and (4) CAs towards different intentions are repetitive, which are useful for automatic adaptation. Overall, our study provides valuable insights into code snippet adaptation and has important implications for research, practice, and tool design. Tanghaoran Zhang, Yao Lu 0003, Yue Yu 0001, Xinjun Mao, Yang Zhang 0026 |
IEEE Trans. Software Eng. | 1 |
| 2023 | MUSE: A Multi-Feature Semantic Fusion Method for ROS Node Search Based on Knowledge GraphabstractReusing ROS components, specifically ROS Nodes, is crucial for improving the efficiency and quality of robotic software development. However, developers face challenges in finding the desired ROS Nodes for reuse due to scattered organization of ROS Nodes and the ambiguity in their ROS Node name. To address these challenges, this paper proposes a MUlti-feature SEmantic fusion method (MUSE) that leverages a domain-specific ROS knowledge graph for searching ROS Nodes. Firstly, a large dataset is constructed, comprising code files and textual descriptions related to ROS Nodes obtained from GitHub and ROS Wiki. Secondly, an in-depth analysis of user queries regarding the reuse of ROS Nodes is conducted, leading to the selection of multiple features that provide a comprehensive representation of ROS Node semantics, including Function, Hardware, Input, and Output. Subsequently, a knowledge graph of ROS Nodes is developed based on the dataset, incorporating the selected features. This knowledge graph effectively organizes scattered knowledge and resolves the issue of diverse mentions through entity disambiguation and resolution. To eliminate the semantic gap between the descriptions of features mentioned in user queries and the entities in the knowledge graph, a pretrained transformer-based model was used to measure the multi-feature semantic similarity between user queries and ROS Nodes knowledge. Finally, we employ a linear regression model to integrate the multi-feature knowledge between user queries and ROS Nodes knowledge. The proposed method has shown a 20% improvement in performance on NDCG@1 compared to other ROS Node search methods. Further evaluations highlight the effectiveness of each feature incorporated in the knowledge graph, as well as the significance of each parameter within the regression model. These findings underscore the robustness of this research in optimizing the reuse of ROS Nodes and facilitating the development of robotics software. Xinjun Mao, Tanghaoran Zhang, Zhang Zhang 0005 |
APSEC | 3 |
| 2023 | An Extensive Study of the Structure Features in Transformer-based Code Semantic SummarizationabstractTransformers are now widely utilized in code intelligence tasks. To better fit highly structured source code, various structure information is passed into Transformer, such as positional encoding and abstract syntax tree (AST) based structures. However, it is still not clear how these structural features affect code intelligence tasks, such as code summarization. Addressing this problem is of vital importance for designing Transformer-based code models. Existing works are keen to introduce various structural information into Transformers while lacking persuasive analysis to reveal their contributions and interaction effects. In this paper, we conduct an empirical study of frequently-used code structure features for code representation, including two types of position encoding features and AST-based structure features. We propose a couple of probing tasks to detect how these structure features perform in Transformer and conduct comprehensive ablation studies to investigate how these structural features affect code semantic summarization tasks. To further validate the effectiveness of code structure features in code summarization tasks, we assess Transformer models equipped with these code structure features on a structural dependent summarization dataset. Our experimental results reveal several findings that may inspire future study: (1) there is a conflict between the influence of the absolute positional embeddings and relative positional embeddings in Transformer; (2) AST-based code structure features and relative position encoding features show a strong correlation and much contribution overlap for code semantic summarization tasks indeed exists between them; (3) Transformer models still have space for further improvement in explicitly understanding code structure information. Kang Yang 0001, Xinjun Mao, Shangwen Wang, Yihao Qin, Tanghaoran Zhang, Yao Lu 0003, Kamal Al-Sabahi |
ICPC | 5 |
| 2023 | An Effective Method for Constructing Knowledge Graph to Search Reusable ROS Nodes (S)abstractDeveloping robot software is difficult for most software engineers as it requires multi-discipline knowledge such as robotics, AI, and software engineering.Robot Operating Systems (ROS) provides a software development framework and lots of reusable ROS Nodes that encapsulate various robotics functions, which can simplify robot software development in terms of software reuse.However, searching and reusing required ROS Nodes from thousands of ROS Nodes is still challenging due to the scattered distribution of ROS Node information and the need for adequate search methods.In this paper, we present an effective method to construct a ROS Node knowledge graph in support of searching and reusing ROS Nodes.Our method uses multiple data sources, including open-source ROS software in Github and ROS wiki community.We extract two-tuple functional information and task-related noun phrases from the ROS Node description and ROS communication interactions from the ROS Node source code.The constructed ROS Node knowledge graph (RNKG) contains 14,065 entities and 15,767 relations.It provides rich semantic information to comprehensively and precisely describe ROS Nodes, their services, and related interaction topics and messages. Xinjun Mao, Sun Bo, Tanghaoran Zhang, Shuo Yang 0005 |
SEKE | 4 |
| 2022 | FENSE: A feature-based ensemble modeling approach to cross-project just-in-time defect prediction
Tanghaoran Zhang, Yue Yu 0001, Xinjun Mao, Yao Lu 0003, Huaimin Wang 0001 |
Empir. Softw. Eng. | 1 |
| 2020 | Verifying ReLU Neural Networks from a Model Checking Perspective
Wanwei Liu, Fu Song, Tanghaoran Zhang, Ji Wang 0001 |
J. Comput. Sci. Technol. | 3 |