EDBT 2026 Demo / reviewers in the wild / expert
Yao Lu 0003
dblp:26/5662-3
· DBLP profile ↗
24ranked-venue papers
5as first author
14since 2021 · last 2025
0000-0002-3520-5829ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 21 · 4 first-author · 12 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Instruct or Interact? Exploring and Eliciting LLMs' Capability in Code Snippet Adaptation Through Prompt EngineeringabstractCode snippet adaptation is a fundamental activity in the software development process. Unlike code generation, code snippet adaptation is not a “free creation”, which requires developers to tailor a given code snippet in order to fit specific requirements and the code context. Recently, large language models (LLMs) have confirmed their effectiveness in the code generation task with promising results. However, their performance on code snippet adaptation, a reuse-oriented and context-dependent code change prediction task, is still unclear. To bridge this gap, we conduct an empirical study to investigate the performance and issues of LLMs on the adaptation task. We first evaluate the adaptation performances of three popular LLMs and compare them to the code generation task. Our result indicates that their adaptation ability is weaker than generation, with a nearly 15% decrease on pass@1 and more context-related errors. By manually inspecting 200 cases, we further investigate the causes of LLMs' sub-optimal performance, which can be classified into three categories, i.e., Unclear Requirement, Requirement Misalignment and Context Misapplication. Based on the above empirical research, we propose an interactive prompting approach to eliciting LLMs' ability on the adaptation task. Specifically, we enhance the prompt by enriching the context and decomposing the task, which alleviates context misapplication and improves requirement understanding. Besides, we enable LLMs' reflection by requiring them to interact with a human or a LLM counselor, compensating for unclear requirement. Our experimental result reveals that our approach greatly improve LLMs' adaptation performance. The best-performing Human-LLM interaction successfully solves 159 out of the 202 identified defects and improves the pass@1 and pass@5 by over 40% compared to the initial instruction-based prompt. Considering human efforts, we suggest multi-agent interaction as a trade-off, which can achieve comparable performance with excellent generalization ability. We deem that our approach could provide methodological assistance for autonomous code snippet reuse and adaptation with LLMs. Tanghaoran Zhang, Yue Yu 0001, Xinjun Mao, Shangwen Wang, Kang Yang 0001, Yao Lu 0003, Zhang Zhang 0005 |
ICSE | 6 |
| 2025 | Large Language Models Are Qualified Benchmark Builders: Rebuilding Pre-Training Datasets for Advancing Code Intelligence TasksabstractPre-trained code models are essential for various code intelligence tasks. Yet, their effectiveness is heavily influenced by the quality of the pre-training dataset, particularly human-written reference comments, which usually serve as a bridge between the programming language and natural language. One significant challenge is that such comments could become inconsistent with the corresponding code as the software evolves, leading to suboptimal model performance. Large language models (LLMs) have demonstrated superior capabilities in generating high-quality code comments. This work investigates whether substituting original human-written comments with LLM-generated ones can improve pre-training datasets for more effective pretrained code models. As existing reference-based metrics cannot evaluate the quality of human-written reference comments themselves, to enable direct comparison between LLM-generated and human reference comments, we introduce two auxiliary tasks as novel reference-free metrics, including code-comment inconsistency detection and semantic code search. Experimental results show that LLM-generated comments exhibit superior semantic consistency with the code compared to human-written reference comments. Our manual evaluation also corroborates this conclusion, which indicates the potential of utilizing LLMs to enhance the quality of the pre-training dataset. Based on this finding, we rebuilt the CodeSearchNet dataset with LLM-generated comments and re-pre-trained the CodeT5 model. Evaluations on multiple code intelligence tasks demonstrate that models pretrained by LLM-enhanced data outperform their counterparts (pre-trained by original human reference comments data) on code summarization, code generation, and code translation tasks. This research validates the feasibility of rebuilding the pre-training dataset by LLMs to advance code intelligence tasks. It advocates rethinking the reliance on human reference comments for coderelated tasks. Kang Yang 0001, Xinjun Mao, Shangwen Wang, Yanlin Wang 0001, Tanghaoran Zhang, Bo Lin 0011, Yihao Qin, Zhang Zhang 0005, Yao Lu 0003, Kamal Al-Sabahi |
ICPC | 9 |
| 2025 | AdaptEval: A Benchmark for Evaluating Large Language Models on Code Snippet AdaptationabstractRecent advancements in large language models (LLMs) have automated various software engineering tasks, with benchmarks emerging to evaluate their capabilities. However, for adaptation, a critical activity during code reuse, there is no benchmark to assess LLMs’ performance, leaving their practical utility in this area unclear. To fill this gap, we propose AdaptEval, a benchmark designed to evaluate LLMs on code snippet adaptation. Unlike existing benchmarks, AdaptEval incorporates the following three distinctive features: First, practical context. Tasks in AdaptEval are derived from developers’ practices, preserving rich contextual information from Stack Overflow and GitHub communities. Second, multi-granularity annotation. Each task is annotated with requirements at both task and adaptation levels, supporting the evaluation of LLMs across diverse adaptation scenarios. Third, fine-grained evaluation. AdaptEval includes a two-tier testing framework combining adaptation-level and function-level tests, which enables evaluating LLMs’ performance across various individual adaptations. Based on AdaptEval, we conduct the first empirical study to evaluate six instruction-tuned LLMs and especially three reasoning LLMs on code snippet adaptation. Experimental results demonstrate that AdaptEval enables the assessment of LLMs’ adaptation capabilities from various perspectives. It also provides critical insights into their current limitations, particularly their struggle to follow explicit instructions. We hope AdaptEval can facilitate further investigation and enhancement of LLMs’ capabilities in code snippet adaptation, supporting their real-world applications. Tanghaoran Zhang, Xinjun Mao, Shangwen Wang, Yao Lu 0003, Zhang Zhang 0005, Kang Yang 0001, Yue Yu 0001 |
ASE | 5 |
| 2025 | ConflictLens: an LLM-Based Method for Detecting Semantic Merge ConflictsabstractSemantic conflicts in branch merging occur when merged code violates specifications from one or both branches.These conflicts are often subtle and can lead to serious runtime errors such as crashes or data corruption.Existing detection methods fail to achieve both high precision and recall: Static analysis-based methods ensure high recall but lack precision, whereas dynamic execution-based methods provide better precision but struggle with recall due to limited test coverage.To better understand such conflicts, we first conduct an empirical study on a real-world merge dataset and identify four common conflict patterns.These patterns reveal key characteristics of semantic conflicts and serve as guidance for automated detection.Based on these insights, we propose ConflictLens, a two-stage LLM-based method that combines static analysis and dynamic execution to balance precision and recall.First, LLMs are guided by few-shot and chain-of-thought prompting using the patterns to localize conflicts statically.Then, conflicts are dynamically verified with LLM-generated targeted tests, refined through execution feedback.Evaluated on 85 real-world merge scenarios, ConflictLens achieves 0.91 precision and 0.76 recall, outperforming static and dynamic baselines.Ablation studies demonstrate the contribution and synergy of each component.Cross-LLM evaluations confirm robustness, with DeepSeek-R1 performing best and cost-efficient models like GPT-4o-Mini still competitive. Longfei Sun, Yao Lu 0003, Xinjun Mao, Tanghaoran Zhang, Zhang Zhang 0005 |
SEKE | 2 |
| 2025 | CARLDA: An Approach for Stack Overflow API Mention Recognition Driven by Context and LLM-Based Data AugmentationabstractABSTRACT The recognition of Application Programming Interface (API) mentions in software‐related texts is vital for extracting API‐related knowledge, providing deep insights into API usage and enhancing productivity efficiency. Previous research identifies two primary technical challenges in this task: (1) differentiating APIs from common words and (2) identifying morphological variants of standard APIs. While deep learning‐based methods have demonstrated advancements in addressing these challenges, they rely heavily on high‐quality labeled data, leading to another significant data‐related challenge: (3) the lack of such high‐quality data due to the substantial effort required for labeling. To overcome these challenges, this paper proposes a context‐aware API recognition method named CARLDA. This approach utilizes two key components, namely, Bidirectional Encoder Representations from Transformers (BERT) and Bidirectional Long Short‐Term Memory (BiLSTM), to extract context at both the word and sequence levels, capturing syntactic and semantic information to address the first challenge. For the second challenge, it incorporates a character‐level BiLSTM with an attention mechanism to grasp global character‐level context, enhancing the recognition of morphological features of APIs. To address the third challenge, we developed specialized data augmentation techniques using large language models (LLMs) to tackle both in‐library and cross‐library data shortages. These techniques generate a variety of labeled samples through targeted transformations (e.g., replacing tokens and restructuring sentences) and hybrid augmentation strategies (e.g., combining real‐world and generated data while applying style rules to replicate authentic programming contexts). Given the uncertainty about the quality of LLM‐generated samples, we also developed sample selection algorithms to filter out low‐quality samples (i.e., incomplete or incorrectly labeled samples). Moreover, specific datasets have been constructed to evaluate CARLDA's ability to address the aforementioned challenges. Experimental results demonstrate that (1) CARLDA significantly enhances F1 by 11.0% and the Matthews correlation coefficient (MCC) by 10.0% compared to state‐of‐the‐art methods, showing superior overall performance and effectively tackling the first two challenges, and (2) LLM‐based data augmentation techniques successfully yield high‐quality labeled data and effectively alleviate the third challenge. Zhang Zhang 0005, Xinjun Mao, Shangwen Wang, Kang Yang 0001, Tanghaoran Zhang, Yao Lu 0003 |
J. Softw. Evol. Process. | 6 |
| 2024 | An Empirical Study of Cross-Project Pull Request Recommendation in GitHubabstractAs a core contribution merge mechanism in distributed collaborative development, pull requests contain valuable knowledge of code evolution and issue resolution. With the co-evolution of multiple projects in a software ecosystem, relevant and similar issues can arise across different projects. Leveraging existing solutions in pull requests (PRs) through cross-project pull request recommendation (CPR) can enrich context knowledge and improve the efficiency of issue resolution. However, the characteristics of CPR and its effectiveness in the process of issue resolution still remain unclear. To bridge this gap, we conduct an empirical study of the CPR on GitHub. We first extract 4,445 CPRs from 2,500 open source projects and quantitatively analyze the characteristics of CPR. Then we conduct a qualitative analysis of sampled CPR cases to understand the influence of CPR. We also use a regression model to explore the impact of CPRs on issue resolution. Our main findings are as follows: (1) Experienced contributors in target projects make most of the CPRs and their CPRs are more timely than inexperienced contributors; (2) In CPR dataset, bugs constitute the largest proportion of target issue types, followed by enhancements, features and questions; (3) Nearly half of the CPRs are accepted by issue participants; (4) A greater number of the CPRs contribute indirectly to solving the target issue by offering solutions and contextual information, rather than providing appropriate code that can be directly applied to the issue; (5) Most of CPR-related factors have a significant impact on issue resolution delay. Among these, recommendation latency has the most significant impact, followed by the type of recommender. Our work has important insights into CPR and offers important guidance for developers on recommending cross-project PRs to resolve the mushrooming issues. Wenyu Xu, Yao Lu 0003, Xunhui Zhang, Tanghaoran Zhang, Bo Lin 0011, Xinjun Mao |
APSEC | 2 |
| 2024 | CAREER: Context-Aware API Recognition with Data Augmentation for API Knowledge ExtractionabstractThe recognition of Application Programming Interface (API) mentions in the software-related texts is a prerequisite task for extracting API-related knowledge. Previous studies have demonstrated the superiority of deep learning-based methods in accomplishing this task. However, such techniques still meet their bottlenecks due to their inability to effectively handle the following three challenges: (1) differentiating APIs from common words; (2) identifying APIs in morphological variants of the standard APIs; and (3) the lack of high-quality labeled data for training. To overcome these challenges, this paper proposes a context-aware API recognition method named CAREER. This approach utilizes two key components, namely Bidirectional Encoder Representations from Transformers (BERT) and Bi-directional Long Short-Term Memory (BiLSTM), to extract context information at both the word-level and sequence-level. This strategic combination empowers the method to dynamically capture both syntactic and semantic information, effectively addressing the first challenge. To tackle the second challenge, CAREER introduces a character-level BiLSTM component, enriched with an attention mechanism. This enables the model to grasp character-level global context information, thereby enhancing the recognition of morphological attributes within API mentions. Furthermore, to address the third challenge, the paper introduces three data augmentation techniques aimed at generating new data samples. Accompanying these techniques is a novel sample selection algorithm designed to screen out high-quality instances. This dual-pronged approach effectively mitigates the requirement for data labeling. Experiments demonstrate that CAREER significantly improves F1-score by 11.0% compared with state-of-the-art methods. We also construct specific datasets to assess CAREER's capacity to tackle the aforementioned challenges. Results confirm that (1) CAREER significantly outperforms baseline methods in addressing the first and second challenges, and (2) with the aid of data augmentation techniques and sample selection algorithms, high-quality samples can be generated to improve the performance, and alleviate the third challenge. Zhang Zhang 0005, Xinjun Mao, Shangwen Wang, Kang Yang 0001, Yao Lu 0003 |
ICPC | 5 |
| 2024 | How Do Developers Adapt Code Snippets to Their Contexts? An Empirical Study of Context-Based Code Snippet AdaptationsabstractReusing code snippets from online programming Q&A communities has become a common development practice, in which developers often need to adapt code snippets to their code contexts to satisfy their own programming needs. However, how developers make these code adaptations based on contexts is still unclear. To bridge this gap, we first conduct a semi-structured interview of 21 developers to investigate their adaptation practices and perceived challenges during this process. The result suggests that code snippet adaptation is a challenging and exhausting task for developers, as they should tailor the snippets to guarantee their correctness and quality with laborious work. We also note that developers all resort to their intra-file context to complete adaptations, which motivates us to further study how developers performed context-based adaptations (CAs) in real scenarios. To this end, we conduct a quantitative study on an adaptation dataset comprising 300 code snippet reuse cases with 1,384 adaptations from Stack Overflow to GitHub. For each adaptation, we manually annotate its intention and relationship with the context. Based on our annotated data, we employ frequent itemset mining to obtain four CA patterns from our dataset, includingFortification,Code Wiring,Attribute-izationandParameterization. Our main findings reveal that: (1) more than half of the code snippet reuse cases include CAs and 23.3% of the adaptations are CAs; (2) more than half of the CAs are corrective adaptations and variable is the primary adapted language construct; (3) attribute is the most frequently utilized context and 88% of the local contexts are within the nearest 10 LOCs; and (4) CAs towards different intentions are repetitive, which are useful for automatic adaptation. Overall, our study provides valuable insights into code snippet adaptation and has important implications for research, practice, and tool design. Tanghaoran Zhang, Yao Lu 0003, Yue Yu 0001, Xinjun Mao, Yang Zhang 0026 |
IEEE Trans. Software Eng. | 2 |
| 2023 | An Extensive Study of the Structure Features in Transformer-based Code Semantic SummarizationabstractTransformers are now widely utilized in code intelligence tasks. To better fit highly structured source code, various structure information is passed into Transformer, such as positional encoding and abstract syntax tree (AST) based structures. However, it is still not clear how these structural features affect code intelligence tasks, such as code summarization. Addressing this problem is of vital importance for designing Transformer-based code models. Existing works are keen to introduce various structural information into Transformers while lacking persuasive analysis to reveal their contributions and interaction effects. In this paper, we conduct an empirical study of frequently-used code structure features for code representation, including two types of position encoding features and AST-based structure features. We propose a couple of probing tasks to detect how these structure features perform in Transformer and conduct comprehensive ablation studies to investigate how these structural features affect code semantic summarization tasks. To further validate the effectiveness of code structure features in code summarization tasks, we assess Transformer models equipped with these code structure features on a structural dependent summarization dataset. Our experimental results reveal several findings that may inspire future study: (1) there is a conflict between the influence of the absolute positional embeddings and relative positional embeddings in Transformer; (2) AST-based code structure features and relative position encoding features show a strong correlation and much contribution overlap for code semantic summarization tasks indeed exists between them; (3) Transformer models still have space for further improvement in explicitly understanding code structure information. Kang Yang 0001, Xinjun Mao, Shangwen Wang, Yihao Qin, Tanghaoran Zhang, Yao Lu 0003, Kamal Al-Sabahi |
ICPC | 6 |
| 2022 | Towards a behavior tree-based robotic software architecture with adjoint observation schemes for robotic software development
Shuo Yang 0005, Xinjun Mao, Yao Lu 0003 |
Autom. Softw. Eng. | 3 |
| 2022 | FENSE: A feature-based ensemble modeling approach to cross-project just-in-time defect prediction
Tanghaoran Zhang, Yue Yu 0001, Xinjun Mao, Yao Lu 0003, Huaimin Wang 0001 |
Empir. Softw. Eng. | 4 |
| 2022 | ForkXplorer: an approach of fork summary generation
Zhang Zhang 0005, Xinjun Mao, Yao Lu 0003 |
Frontiers Comput. Sci. | 4 |
| 2022 | Motivation Under Gamification: An Empirical Study of Developers' Motivations and Contributions in Stack OverflowabstractTo encourage developers' volunteer contributions, modern programming question and answer (Q&A) sites like Stack Overflow (SO) employ gamified incentive mechanisms such as reputation and badges. Understanding developers' motivations in the presence of gamification and the relationship between their motivations and behavioral outcomes is crucial for community building and designing good incentive mechanisms. Grounded on self-determination theory, we conducted a survey with 938 developers who participate in SO to understand their participation motivations and incentive perceptions. By connecting the survey responses with the SO data, we quantitatively analyzed how the developers' motivations and satisfaction of needs relate to their effort and contribution quality. Our main findings are as follows: (1) despite the presence of gamified incentive mechanisms, developers are mainly motivated by intrinsic motivation to participate in SO; (2) developers who have strong motivations to gain gamification rewards are associated with higher intrinsic and integrated motivations, while developers with more development experiences are less motivated by the gamified incentives; (3) both extrinsic motivations (in terms of career prospects) and intrinsic motivations (regarding self-improvement and helping others) can motivate developers to make high-quantity and high-quality contributions; and (4) high-level satisfaction of needs for competency and autonomy has a positive effect on developers making high-quantity and high-quality contributions and addressing difficult problems. Based on these findings, we discuss implications for developer motivation and gamification in the crowdsourcing context and for the mechanism design of gamified crowdsourced platforms. Yao Lu 0003, Xinjun Mao, Minghui Zhou 0001, Yang Zhang 0026, Zude Li, Tao Wang 0006, Gang Yin, Huaimin Wang 0001 |
IEEE Trans. Software Eng. | 1 |
| 2021 | TagNN: A Code Tag Generation Technology for Resource Retrieval from Open-Source Big DataabstractWith the vigorous development of open‐source software, a huge number of open‐source projects and open‐source codes have been accumulated in open‐source big data, which contains a wealth of code resources. However, effectively and efficiently retrieving the relevant code snippets in such a large amount of open‐source big data is an extremely difficult problem. There are usually large gaps between the user’s natural language description and the open‐source code snippets. In this paper, we propose a novel code tag generation and code retrieval approach named TagNN, which combines software engineering empirical knowledge and a deep learning algorithm. The experimental results show that our method has good effects on code tag generation and code snippet retrieval. Lingbin Zeng, Cheng Yang 0004, Yao Lu 0003, Xiao Li 0039 |
Wirel. Commun. Mob. Comput. | 4 |
| 2020 | An Empirical Study on the Influence of Social Interactions for the Acceptance of Answers in Stack OverflowabstractIn knowledge-sharing communities like Stack Overflow (SO), users can post questions, give answers and choose one answer as an accepted answer. The accepted answers will be important references for users when they encounter similar questions. Essentially, posting questions and giving answers is an interactive process occurring among community users, and choosing accepted answers is actually a decision-making process involving multiple factors. Previous works examined the impact on this decision process from the user, question and answer viewpoints. Social interactions between the questioners and answerers, although being popular according to our pre-analysis, have never been considered as a factor that can influence the decisions. To fill this gap, this paper first proposes a comprehensive answer acceptance model that integrates the answer features established by social interactions as well as information of users, questions and answers. We then divide social interactions into two stages and propose a method to calculate the relationship between the questioner and the answerer by analyzing these social interactions. Finally, we investigate the influence of social interactions for the acceptance of answers by performing logistic regression analysis. The results reveal several findings: (1) social-based features explain 16.6 % of the variance explained together, indicating that social interactions have significant and important effects on the acceptance of answers; (2) social interactions that occur after the answer is posted are more influential than these occur before the answer is posted. Based on the findings, we further conduct an online study of 132 SO users, and the respondents report that social interactions have a greater impact on the acceptance of answers than other judgments of answers such as upvotes, downvotes and not accepting answers. Zhang Zhang 0005, Xinjun Mao, Yao Lu 0003, Shangwen Wang, Jinyu Lu |
APSEC | 3 |
| 2020 | Gathering GitHub OSS Requirements from Q&A Community: an Empirical StudyabstractCross-community collaboration can exploit the expertise and knowledges of crowds in different communities. Recently increasing users in open source software (OSS) community like GitHub attempt to gather software requirements from question and answer (Q&A) communities such as Stack Overflow (SO). In order to investigate this emerging cross-community collaboration phenomenon, the paper presents an exploratory study on cross-community requirements gathering of OSS projects in GitHub. We manually sample 3266 practice cases and quantitatively analyze the popularity of the phenomenon, the characteristics of the gathered requirements, and cross-community collaboration behaviors of users. Some important findings are obtained: more than half of the requirements gathered from SO are enhancements and the majority of the gathered requirements are non-functional requirements. In addition, OSS developers can directly obtain related solutions and contributions of the gathered requirements from SO in the gathering process. Yao Lu 0003, Xinjun Mao |
ICECCS | 2 |
| 2020 | Haste Makes Waste: An Empirical Study of Fast Answers in Stack OverflowabstractModern programming question & answer (Q&A) sites such as Stack Overflow (SO) employ gamified mechanisms to stimulate volunteers' contributions. To maximize the chances of winning gamification rewards such as reputation and badges, a portion of users race to post answers as quickly as possible (i.e., fast answers or FAs), which makes SO the fastest Q&A site; however, this behavior may affect the contribution quality as well. In this paper, we report on a large-scale, mixed-methods empirical study of the gamification-influenced FA phenomenon in SO. We first quantitatively investigate the popularity of the phenomenon and user behaviors regarding FAs. Then, we study the quality of FAs by using regression modeling and qualitatively analyzing 300 instances of FAs. Our main findings reveal that more than 70% and 90% of FAs are not edited by the answerers and other users, respectively, and that later incoming answers have lower chances of being voted on and accepted. Notably, we find that the answer length, code snippets length, and readability of FAs are significantly lower than those of non-fast answers. Although FAs have higher crowd assessment scores, they have no relationship with acceptance from the perspective of asker assessment, and a considerable portion of FAs solve the problem by interacting with the asker in the comments. These results help us better understand the effects of reward-based gamification on crowdsourced software engineering communitites and provide implications for designers of gamified systems. Yao Lu 0003, Xinjun Mao, Minghui Zhou 0001, Yang Zhang 0026, Tao Wang 0006, Zude Li |
ICSME | 1 |
| 2020 | Exploring the Dependency Network of Docker Containers: Structure, Diversity, and RelationshipabstractContainer technologies are being widely used in large scale production cloud environments, of which Docker has become the de-facto industry standard. As a key step, containers need to define their dependent base image, which makes complex dependencies exist in a large number of containers. Prior studies have shown that references between software packages could form technical dependencies, thus forming a dependency network. However, little is known about the details of docker container dependency networks. In this paper, we perform an empirical study on the dependency network of docker containers from more than 120,000 dockerfiles. We construct the container dependency network and analyze its network structure. Further, we focus on the Top-100 dominant containers and investigate their subnetworks, including diversity and relationships. Our findings help to characterize and understand the container dependencies in the docker community and motivate the need for developing container dependency management tools. Yinyuan Zhang, Yang Zhang 0026, Yiwen Wu 0001, Yao Lu 0003, Tao Wang 0006, Xinjun Mao |
Internetware | 4 |
| 2020 | Who Should Close the Questions: Recommending Voters for Closing Questions Based on Tags
Zhang Zhang 0005, Xinjun Mao, Yao Lu 0003, Jinyu Lu |
SEKE | 3 |
| 2020 | Exploring CQA User Contributions and Their Influence on Answer Distribution
Yi Yang 0004, Xinjun Mao, Zixi Xu, Yao Lu 0003 |
SEKE | 4 |
| 2020 | Improving students' programming quality with the continuous inspection process: a social coding perspective
Yao Lu 0003, Xinjun Mao, Tao Wang 0006, Gang Yin, Zude Li |
Frontiers Comput. Sci. | 1 |
| 2020 | Automatic Voter Recommendation Method for Closing Questions in Stack OverflowabstractStack Overflow is the most popular programming question and answer community that continuously receives a large number of questions every day. To ensure the quality of questions, the community grants privileges for the moderators and a group of experienced users to review the quality of questions and close the low-quality ones (e.g. duplicate or irrelevant questions). The review process is a typical crowdsourcing job that relies on users’ volunteer participation, and the current practices of closing questions in Stack Overflow face two aspects of challenges: (1) an obvious increase in both the absolute number and the percentage of “closed” questions; (2) a considerable decrease in participation willingness of experienced users to close questions. In order to solve the problem, we present a novel model of user willingness for reviewing and voting questions by incorporating four types of user activity history, including questions, answers, comments and votes of closing questions. Then we propose an automatic recommendation method based on the model to assign experienced users proper questions, to utilize the forces of them to close questions. The evaluation shows that the successful recommendation probability in the top 5, top 10, top 20, top 30, top 40, top 50 users are 48.23%, 58.93%, 68.83%, 74.27%, 78.13% and 81%, respectively. Zhang Zhang 0005, Xinjun Mao, Yao Lu 0003, Jinyu Lu, Yue Yu 0001 |
Int. J. Softw. Eng. Knowl. Eng. | 3 |
| 2018 | Internal quality assurance for external contributions in GitHub: An empirical investigationabstractAbstract For popular open‐source software projects, there are always a large number of worldwide developers who have been glued to making code contributions, while most of these developers play the role of casual contributors because of their very limited code commits. The frequent turnover of such a group of developers and the wide variations in their coding experiences challenge the project management on code and quality. This paper aims to investigate the status quo of internal quality assurance for external contributions in social coding sites. We first conducted a case study of 21 popular GitHub projects to estimate the code quality of the casual contributors. The quantitative results show that the casual contributors introduced greater quantity and severity of code quality issues than the main contributors; the developers who contribute to different projects as main and casual contributors did not perform significantly differently in terms of their code quality. On the basis of these findings, we further conducted a survey of 81 developers on GitHub to understand their practices on internal quality assurance. The qualitative results expose some limitations of present internal quality control for external contributions in GitHub. Finally, we discuss an alternative quality management paradigm: Continuous Inspection for industrial practices. Yao Lu 0003, Xinjun Mao, Zude Li, Yang Zhang 0026, Tao Wang 0006, Gang Yin |
J. Softw. Evol. Process. | 1 |
| 2016 | Does the Role Matter? An Investigation of the Code Quality of Casual Contributors in GitHubabstractFor popular Open Source Software (OSS) projects there are always a large number of worldwide developers who have been glued to making code contributions, while most of these developers play the role of casual contributors due to their very limited code commits (for fixing defects and enhancing features, casually). The frequent turnover of such group of casual developers and the wide variations among their coding experiences challenge the project management on code and quality.This paper describes a case study which aims to estimate the quality of code made by casual contributors in 21 popular GitHub projects. The results of this case study show that: (1) casual contributors introduced greater quantity and severity of Code Quality Issues (CQIs) than main contributors; (2) developers who contribute in different projects as main and casual contributors didn't perform statistically differently in terms of code quality; (3) casual contributors who have few project stars introduced more CQIs than those who have many. Furthermore, the paper lists the CQI categories which are most frequently introduced by casual contributors in the investigated projects. These findings provide valuable insights into code quality in the OSS context, and can guide OSS developers in improving the quality of the code contributions. Yao Lu 0003, Xinjun Mao, Zude Li, Yang Zhang 0026, Tao Wang 0006, Gang Yin |
APSEC | 1 |