Kang Yang 0001

dblp:86/8501-1 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
11since 2021 · last 2025
0000-0003-2313-7141ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 9 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Instruct or Interact? Exploring and Eliciting LLMs' Capability in Code Snippet Adaptation Through Prompt Engineering
abstract
Code snippet adaptation is a fundamental activity in the software development process. Unlike code generation, code snippet adaptation is not a “free creation”, which requires developers to tailor a given code snippet in order to fit specific requirements and the code context. Recently, large language models (LLMs) have confirmed their effectiveness in the code generation task with promising results. However, their performance on code snippet adaptation, a reuse-oriented and context-dependent code change prediction task, is still unclear. To bridge this gap, we conduct an empirical study to investigate the performance and issues of LLMs on the adaptation task. We first evaluate the adaptation performances of three popular LLMs and compare them to the code generation task. Our result indicates that their adaptation ability is weaker than generation, with a nearly 15% decrease on pass@1 and more context-related errors. By manually inspecting 200 cases, we further investigate the causes of LLMs' sub-optimal performance, which can be classified into three categories, i.e., Unclear Requirement, Requirement Misalignment and Context Misapplication. Based on the above empirical research, we propose an interactive prompting approach to eliciting LLMs' ability on the adaptation task. Specifically, we enhance the prompt by enriching the context and decomposing the task, which alleviates context misapplication and improves requirement understanding. Besides, we enable LLMs' reflection by requiring them to interact with a human or a LLM counselor, compensating for unclear requirement. Our experimental result reveals that our approach greatly improve LLMs' adaptation performance. The best-performing Human-LLM interaction successfully solves 159 out of the 202 identified defects and improves the pass@1 and pass@5 by over 40% compared to the initial instruction-based prompt. Considering human efforts, we suggest multi-agent interaction as a trade-off, which can achieve comparable performance with excellent generalization ability. We deem that our approach could provide methodological assistance for autonomous code snippet reuse and adaptation with LLMs.
Tanghaoran Zhang, Yue Yu 0001, Xinjun Mao, Shangwen Wang, Kang Yang 0001, Yao Lu 0003, Zhang Zhang 0005
ICSE5
2025 Understanding the Faults in Serverless Computing Based Applications: An Empirical Study
abstract
Serverless computing is a novel cloud computing paradigm that enables developers to develop, deploy, and run applications in the cloud without complex and error-prone cloud resource management. However, its characteristics also introduce new types of faults (e.g., faults due to insufficient computing resource allocation) and challenges to serverless computing-based applications (abbreviated as serverless applications). While prior studies have highlighted that serverless developers encounter various challenges, no attempts have been made to understand the faults in serverless applications. These faults may cause catastrophic consequences such as application crash, thereby hindering the further spread of serverless computing. We aim in this paper to understand the symptoms, root causes, and fix patterns of faults in serverless applications. To this end, we conduct an empirical study investigating developers' issues on GitHub and posts on Stack Overflow (SO). We first identify 546 real-world serverless-related faults from GitHub and SO. Then, we manually analyze and construct taxonomies of the symptoms, root causes, and fix patterns for these faults, respectively. Our study leads to the first taxonomy for symptoms of serverlessrelated faults, covering 5 categories and 21 subcategories. The findings of our study inform that the Permission Denied error is the most common type ($\mathbf{1 0. 8 1 \%}$) of faults. Furthermore, the Incorrect Code Logic is the main cause ($\mathbf{1 7. 9 5 \%}$) behind the faults. Furthermore, we summarize 15 fix patterns that can resolve$\mathbf{7 3. 6 3 \%}$of faults in this study. Based on the results, we provide actionable implications that can potentially facilitate research and assist developers in improving the development of serverless applications. Finally, we implement a knowledge-based Q&A tool named SafHelper to help developers understand and fix faults.
Changrong Xie, Yang Zhang 0026, Xinjun Mao, Kang Yang 0001, Tanghaoran Zhang
ICSME4
2025 Large Language Models Are Qualified Benchmark Builders: Rebuilding Pre-Training Datasets for Advancing Code Intelligence Tasks
abstract
Pre-trained code models are essential for various code intelligence tasks. Yet, their effectiveness is heavily influenced by the quality of the pre-training dataset, particularly human-written reference comments, which usually serve as a bridge between the programming language and natural language. One significant challenge is that such comments could become inconsistent with the corresponding code as the software evolves, leading to suboptimal model performance. Large language models (LLMs) have demonstrated superior capabilities in generating high-quality code comments. This work investigates whether substituting original human-written comments with LLM-generated ones can improve pre-training datasets for more effective pretrained code models. As existing reference-based metrics cannot evaluate the quality of human-written reference comments themselves, to enable direct comparison between LLM-generated and human reference comments, we introduce two auxiliary tasks as novel reference-free metrics, including code-comment inconsistency detection and semantic code search. Experimental results show that LLM-generated comments exhibit superior semantic consistency with the code compared to human-written reference comments. Our manual evaluation also corroborates this conclusion, which indicates the potential of utilizing LLMs to enhance the quality of the pre-training dataset. Based on this finding, we rebuilt the CodeSearchNet dataset with LLM-generated comments and re-pre-trained the CodeT5 model. Evaluations on multiple code intelligence tasks demonstrate that models pretrained by LLM-enhanced data outperform their counterparts (pre-trained by original human reference comments data) on code summarization, code generation, and code translation tasks. This research validates the feasibility of rebuilding the pre-training dataset by LLMs to advance code intelligence tasks. It advocates rethinking the reliance on human reference comments for coderelated tasks.
Kang Yang 0001, Xinjun Mao, Shangwen Wang, Yanlin Wang 0001, Tanghaoran Zhang, Bo Lin 0011, Yihao Qin, Zhang Zhang 0005, Yao Lu 0003, Kamal Al-Sabahi
ICPC1
2025 AdaptEval: A Benchmark for Evaluating Large Language Models on Code Snippet Adaptation
abstract
Recent advancements in large language models (LLMs) have automated various software engineering tasks, with benchmarks emerging to evaluate their capabilities. However, for adaptation, a critical activity during code reuse, there is no benchmark to assess LLMs’ performance, leaving their practical utility in this area unclear. To fill this gap, we propose AdaptEval, a benchmark designed to evaluate LLMs on code snippet adaptation. Unlike existing benchmarks, AdaptEval incorporates the following three distinctive features: First, practical context. Tasks in AdaptEval are derived from developers’ practices, preserving rich contextual information from Stack Overflow and GitHub communities. Second, multi-granularity annotation. Each task is annotated with requirements at both task and adaptation levels, supporting the evaluation of LLMs across diverse adaptation scenarios. Third, fine-grained evaluation. AdaptEval includes a two-tier testing framework combining adaptation-level and function-level tests, which enables evaluating LLMs’ performance across various individual adaptations. Based on AdaptEval, we conduct the first empirical study to evaluate six instruction-tuned LLMs and especially three reasoning LLMs on code snippet adaptation. Experimental results demonstrate that AdaptEval enables the assessment of LLMs’ adaptation capabilities from various perspectives. It also provides critical insights into their current limitations, particularly their struggle to follow explicit instructions. We hope AdaptEval can facilitate further investigation and enhancement of LLMs’ capabilities in code snippet adaptation, supporting their real-world applications.
Tanghaoran Zhang, Xinjun Mao, Shangwen Wang, Yao Lu 0003, Zhang Zhang 0005, Kang Yang 0001, Yue Yu 0001
ASE8
2025 Improving API Knowledge Comprehensibility: A Context-Dependent Entity Detection and Context Completion Approach Using LLM
abstract
Extracting API knowledge from Stack Overflow has become a crucial way to assist developers in using APIs. Existing research has primarily focused on extracting relevant API-related knowledge at the sentence level to enhance API documentation. However, this level of extraction can lead to a loss of crucial context, especially when sentences contain context-dependent entities (i.e., whose understanding requires reference to the surrounding context) that may hinder developers' understanding. To investigate this issue, we conducted an empirical study of 384 Stack Overflow posts and found that (1) approximately one-third of API functionality sentences contain context-dependent entities, and (2) these entities fall into two categories: Referential ContextDependent Entities and Local Variable Context-Dependent Entities. In response, we developed a novel method, CEDCC, which combines an entity filtering strategy informed by insights from our empirical study, with a large language model (LLM) to construct coreference chains for detecting context-dependent entities. Additionally, it employs a step-by-step approach with the LLM to complete the necessary context for understanding these entities. To evaluate CEDCC, we constructed a dataset of 1,023 API knowledge sentences, including 567 context-dependent entities and their required contexts. The results demonstrate the effectiveness of CEDCC in accurately detecting contextdependent entities and completing context tasks, achieving an F1score of 0.865 and a BERTScore of 0.373, significantly surpassing the baseline methods. Human evaluations further confirmed that CEDCC effectively improves the comprehensibility of API knowledge sentences.
Zhang Zhang 0005, Xinjun Mao, Shangwen Wang, Kang Yang 0001, Tanghaoran Zhang, Xunhui Zhang
SANER4
2025 GTE: learning code AST representation efficiently and effectively
Yihao Qin, Shangwen Wang, Bo Lin 0011, Kang Yang 0001, Xiaoguang Mao
Sci. China Inf. Sci.4
2025 CARLDA: An Approach for Stack Overflow API Mention Recognition Driven by Context and LLM-Based Data Augmentation
abstract
ABSTRACT The recognition of Application Programming Interface (API) mentions in software‐related texts is vital for extracting API‐related knowledge, providing deep insights into API usage and enhancing productivity efficiency. Previous research identifies two primary technical challenges in this task: (1) differentiating APIs from common words and (2) identifying morphological variants of standard APIs. While deep learning‐based methods have demonstrated advancements in addressing these challenges, they rely heavily on high‐quality labeled data, leading to another significant data‐related challenge: (3) the lack of such high‐quality data due to the substantial effort required for labeling. To overcome these challenges, this paper proposes a context‐aware API recognition method named CARLDA. This approach utilizes two key components, namely, Bidirectional Encoder Representations from Transformers (BERT) and Bidirectional Long Short‐Term Memory (BiLSTM), to extract context at both the word and sequence levels, capturing syntactic and semantic information to address the first challenge. For the second challenge, it incorporates a character‐level BiLSTM with an attention mechanism to grasp global character‐level context, enhancing the recognition of morphological features of APIs. To address the third challenge, we developed specialized data augmentation techniques using large language models (LLMs) to tackle both in‐library and cross‐library data shortages. These techniques generate a variety of labeled samples through targeted transformations (e.g., replacing tokens and restructuring sentences) and hybrid augmentation strategies (e.g., combining real‐world and generated data while applying style rules to replicate authentic programming contexts). Given the uncertainty about the quality of LLM‐generated samples, we also developed sample selection algorithms to filter out low‐quality samples (i.e., incomplete or incorrectly labeled samples). Moreover, specific datasets have been constructed to evaluate CARLDA's ability to address the aforementioned challenges. Experimental results demonstrate that (1) CARLDA significantly enhances F1 by 11.0% and the Matthews correlation coefficient (MCC) by 10.0% compared to state‐of‐the‐art methods, showing superior overall performance and effectively tackling the first two challenges, and (2) LLM‐based data augmentation techniques successfully yield high‐quality labeled data and effectively alleviate the third challenge.
Zhang Zhang 0005, Xinjun Mao, Shangwen Wang, Kang Yang 0001, Tanghaoran Zhang, Yao Lu 0003
J. Softw. Evol. Process.4
2025 A Little Help Goes a Long Way: Tutoring LLMs in Solving Competitive Programming Through Hints
abstract
Code generation has advanced with large language models (LLMs), but LLMs still struggle with complex tasks, especially in competitive programming. These tasks require understanding complex problems, generating correct code that passes numerous test cases, and meeting tight time and memory limits. We observed that there are some critical hints provided by competition platforms, which often point to the most critical information needed to solve the problem, thus guiding participants to accurate solutions. Inspired by these observations, we propose TEACH1, an approach that tutors LLMs in solving competitive programming by combining critical hints with a structured Chain of Thought (CoT). The key insight of TEACH is to employ a domain-specialized hint generator that is fine-tuned on curated data from competitive programming platforms, enabling it to produce concise and targeted algorithmic hints. By integrating these hints into the reasoning process of LLMs, TEACH helps LLMs bridge the gap between complex tasks and solutions. Furthermore, TEACH simulates human problem-solving through a structured CoT that covers problem understanding, analysis, algorithm selection, and coding. We extensively evaluate TEACH on both proprietary (GPT-3.5, GPT-4o, Claude-3.5-Sonnet, Gemini-2.5-Flash) and open-source (DeepSeek-V3) LLMs. TEACH achieves up to 6.56 absolute (17.4% relative) gain in pass@1 on LeetCode, and demonstrates strong generalization to APPS and ASAC, with maximum pass@1 relative improvements of 17.6% and 26.9%, respectively. Furthermore, existing CoT methods with the hints generated from TEACH yield additional gains, demonstrating its compatibility and extensibility across models and prompting strategies.
Wei Dong 0006, Shangwen Wang, Deze Wang, Tiecheng Ma, Yiwei Li 0006, Kang Yang 0001
IEEE Trans. Software Eng.8
2024 CAREER: Context-Aware API Recognition with Data Augmentation for API Knowledge Extraction
abstract
The recognition of Application Programming Interface (API) mentions in the software-related texts is a prerequisite task for extracting API-related knowledge. Previous studies have demonstrated the superiority of deep learning-based methods in accomplishing this task. However, such techniques still meet their bottlenecks due to their inability to effectively handle the following three challenges: (1) differentiating APIs from common words; (2) identifying APIs in morphological variants of the standard APIs; and (3) the lack of high-quality labeled data for training. To overcome these challenges, this paper proposes a context-aware API recognition method named CAREER. This approach utilizes two key components, namely Bidirectional Encoder Representations from Transformers (BERT) and Bi-directional Long Short-Term Memory (BiLSTM), to extract context information at both the word-level and sequence-level. This strategic combination empowers the method to dynamically capture both syntactic and semantic information, effectively addressing the first challenge. To tackle the second challenge, CAREER introduces a character-level BiLSTM component, enriched with an attention mechanism. This enables the model to grasp character-level global context information, thereby enhancing the recognition of morphological attributes within API mentions. Furthermore, to address the third challenge, the paper introduces three data augmentation techniques aimed at generating new data samples. Accompanying these techniques is a novel sample selection algorithm designed to screen out high-quality instances. This dual-pronged approach effectively mitigates the requirement for data labeling. Experiments demonstrate that CAREER significantly improves F1-score by 11.0% compared with state-of-the-art methods. We also construct specific datasets to assess CAREER's capacity to tackle the aforementioned challenges. Results confirm that (1) CAREER significantly outperforms baseline methods in addressing the first and second challenges, and (2) with the aid of data augmentation techniques and sample selection algorithms, high-quality samples can be generated to improve the performance, and alleviate the third challenge.
Zhang Zhang 0005, Xinjun Mao, Shangwen Wang, Kang Yang 0001, Yao Lu 0003
ICPC4
2024 Multi-head sequence tagging model for Grammatical Error Correction
Kamal Al-Sabahi, Kang Yang 0001, Wangwang Liu, Guanyu Jiang
Eng. Appl. Artif. Intell.2
2023 An Extensive Study of the Structure Features in Transformer-based Code Semantic Summarization
abstract
Transformers are now widely utilized in code intelligence tasks. To better fit highly structured source code, various structure information is passed into Transformer, such as positional encoding and abstract syntax tree (AST) based structures. However, it is still not clear how these structural features affect code intelligence tasks, such as code summarization. Addressing this problem is of vital importance for designing Transformer-based code models. Existing works are keen to introduce various structural information into Transformers while lacking persuasive analysis to reveal their contributions and interaction effects. In this paper, we conduct an empirical study of frequently-used code structure features for code representation, including two types of position encoding features and AST-based structure features. We propose a couple of probing tasks to detect how these structure features perform in Transformer and conduct comprehensive ablation studies to investigate how these structural features affect code semantic summarization tasks. To further validate the effectiveness of code structure features in code summarization tasks, we assess Transformer models equipped with these code structure features on a structural dependent summarization dataset. Our experimental results reveal several findings that may inspire future study: (1) there is a conflict between the influence of the absolute positional embeddings and relative positional embeddings in Transformer; (2) AST-based code structure features and relative position encoding features show a strong correlation and much contribution overlap for code semantic summarization tasks indeed exists between them; (3) Transformer models still have space for further improvement in explicitly understanding code structure information.
Kang Yang 0001, Xinjun Mao, Shangwen Wang, Yihao Qin, Tanghaoran Zhang, Yao Lu 0003, Kamal Al-Sabahi
ICPC1
2019 EcForest: Extractive document summarization through enhanced sentence embedding and cascade forest
abstract
Summary We present EcForest, an extractive summarization model through Enhanced Sentence Embedding and Cascade Forest. Sentence representation is of great significance for many summarization methods. Bag‐of‐words mostly fails to grasp the semantics, and typical embedding models cannot capture more complex semantic features, such as polysemy and the meaning of a phrase, which is usually ignored by simply averaging the word embeddings included in a sentence. To this end, we propose Enhanced Sentence Embedding (ESE) model to solve such drawbacks via mapping several valid features to dense vectors. Essentially, the enhanced sentence embedding is a novel model for improving the distributed representation of sentence. Our sentence embedding model is universally applicable and it can be adapted to other NLP tasks. Moreover, deep forest is used as a sentence extraction algorithm for its robustness to the hyper‐parameters and its efficient training algorithm compared to deep neural network. The evaluation of variant models proposed in this work proves the validation of the enhanced sentence embedding. The comparison results between EcForest and several baselines on two different datasets demonstrate that the proposed summarization model performs better than or with high competitiveness to the state‐of‐the‐art.
Kang Yang 0001, Hongye He, Kamal Al-Sabahi, Zuping Zhang 0001
Concurr. Comput. Pract. Exp.1