Xunhui Zhang

dblp:183/9165 · DBLP profile ↗
← Back
23ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0002-4027-9443ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 13 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-authorHuman-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Security and privacy · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 Divergence or Convergence? A Deep Insight into the Crowd Collaboration and its Productivity in Open Source Software based on Entropy
abstract
The Fork and Pull-Request model is widely used in collaborative development of open source software (OSS), fostering innovation through independent repository copies, but it can also lead to inefficiencies and fragmentation. A key underexplored aspect is the integration effectiveness—the degree to which distributed original commits across forks are effectively integrated back into the main repository. It plays a critical role in OSS project productivity but remains poorly understood. In response, we introduce convergence entropy, a novel metric that quantifies the integration effectiveness by measuring the similarity between distributions of original and merged commits across forks, adjusted for integration ratio. This metric highlights not only the volume of contributions but also their diversity and coordination, offering a unique lens to understand forking practices. Moreover, we explore the relationship between convergence entropy and three dimensions of OSS project productivity, showing significant correlations. We also observe that other factors can alter this dynamic.
Tao Wang 0158, Xunhui Zhang, Yang Zhang 0026, Cheng Yang 0004, Bo Ding 0001, Huaimin Wang 0001
CHI3
2025 Improving API Knowledge Comprehensibility: A Context-Dependent Entity Detection and Context Completion Approach Using LLM
abstract
Extracting API knowledge from Stack Overflow has become a crucial way to assist developers in using APIs. Existing research has primarily focused on extracting relevant API-related knowledge at the sentence level to enhance API documentation. However, this level of extraction can lead to a loss of crucial context, especially when sentences contain context-dependent entities (i.e., whose understanding requires reference to the surrounding context) that may hinder developers' understanding. To investigate this issue, we conducted an empirical study of 384 Stack Overflow posts and found that (1) approximately one-third of API functionality sentences contain context-dependent entities, and (2) these entities fall into two categories: Referential ContextDependent Entities and Local Variable Context-Dependent Entities. In response, we developed a novel method, CEDCC, which combines an entity filtering strategy informed by insights from our empirical study, with a large language model (LLM) to construct coreference chains for detecting context-dependent entities. Additionally, it employs a step-by-step approach with the LLM to complete the necessary context for understanding these entities. To evaluate CEDCC, we constructed a dataset of 1,023 API knowledge sentences, including 567 context-dependent entities and their required contexts. The results demonstrate the effectiveness of CEDCC in accurately detecting contextdependent entities and completing context tasks, achieving an F1score of 0.865 and a BERTScore of 0.373, significantly surpassing the baseline methods. Human evaluations further confirmed that CEDCC effectively improves the comprehensibility of API knowledge sentences.
Zhang Zhang 0005, Xinjun Mao, Shangwen Wang, Kang Yang 0001, Tanghaoran Zhang, Xunhui Zhang
SANER7
2025 Open source oriented cross-platform survey
Simeng Yao, Xunhui Zhang, Yang Zhang 0026, Tao Wang 0006
Inf. Softw. Technol.2
2025 Are External Contributions Important to Project Productivity in Open Source Software? A Deep Insight based on Issue Entropy
abstract
In the realm of open source software (OSS) development, the resolution of issues is not just a technical task but a pivotal activity that drives ongoing enhancement and secures a project's sustainability. Contributing to issues is the majority form for external contributors to take part in OSS projects. Although the significance of external contributors is recognized, their contributions in issue process still lacks full quantification and clarity, and the correlation between their contributions and the project productivity remains unclear. In response, we propose issue entropy, a novel metric that quantifies the complexity of event sequence in issue process applied to study the external contributions. The metric applies principles of information theory to scrutinize granular details within a project's issue-related activities, providing a unique lens to understand and assess external contributions. To explore the correlation between external contributions and project productivity, we employed issue entropy as a novel way to examine external contributions, and analyzed its correlation with new bugs, commits, and bug fix time, serving as proxies for OSS project productivity. Our findings reveal a strong positive correlation with new bugs, variable relationship with commit volume based on company sponsorship, and significant negative correlation with bug fix time. We also observe the significant interactions between external contributions and other factors, such as project age and the number of files. Moreover, issue entropy offers a new perspective on the health of the OSS project ecosystem, potentially supporting further research and practical applications.
Tao Wang 0158, Xunhui Zhang, Yang Zhang 0026, Cheng Yang 0004, Yue Yu 0001, Huaimin Wang 0001
Proc. ACM Hum. Comput. Interact.3
2024 An Empirical Study of Cross-Project Pull Request Recommendation in GitHub
abstract
As a core contribution merge mechanism in distributed collaborative development, pull requests contain valuable knowledge of code evolution and issue resolution. With the co-evolution of multiple projects in a software ecosystem, relevant and similar issues can arise across different projects. Leveraging existing solutions in pull requests (PRs) through cross-project pull request recommendation (CPR) can enrich context knowledge and improve the efficiency of issue resolution. However, the characteristics of CPR and its effectiveness in the process of issue resolution still remain unclear. To bridge this gap, we conduct an empirical study of the CPR on GitHub. We first extract 4,445 CPRs from 2,500 open source projects and quantitatively analyze the characteristics of CPR. Then we conduct a qualitative analysis of sampled CPR cases to understand the influence of CPR. We also use a regression model to explore the impact of CPRs on issue resolution. Our main findings are as follows: (1) Experienced contributors in target projects make most of the CPRs and their CPRs are more timely than inexperienced contributors; (2) In CPR dataset, bugs constitute the largest proportion of target issue types, followed by enhancements, features and questions; (3) Nearly half of the CPRs are accepted by issue participants; (4) A greater number of the CPRs contribute indirectly to solving the target issue by offering solutions and contextual information, rather than providing appropriate code that can be directly applied to the issue; (5) Most of CPR-related factors have a significant impact on issue resolution delay. Among these, recommendation latency has the most significant impact, followed by the type of recommender. Our work has important insights into CPR and offers important guidance for developers on recommending cross-project PRs to resolve the mushrooming issues.
Wenyu Xu, Yao Lu 0003, Xunhui Zhang, Tanghaoran Zhang, Bo Lin 0011, Xinjun Mao
APSEC3
2024 Think Before Acting: The Necessity of Endowing Robot Terminals With the Ability to Fine-Tune Reinforcement Learning Policies
abstract
Goal-Conditioned Reinforcement Learning (GCRL) has gained widespread application in robotics. A typical application of GCRL is to pre-train policies in a development environment and then deploy them to robots. In this approach of application, we find that the policies trained by GCRL exhibit discontinuity in goal space, indicating that we cannot effectively estimate the policy’s performance in the actual production environment based on its evaluation results in the development environment. To ensure that the robot terminals can complete tasks effectively, we propose the Think Before Acting (TBA) framework, which evaluates and fine-tunes policies on the robot terminal. Within the TBA framework, for a goal to be executed, the policy’s performance is first evaluated. If the performance does not meet the requirements, the policy is fine-tuned based on this goal. We conduct experiments on velocity vector control of fixed-wing UAVs to validate the effectiveness of TBA. The results show that for a well-pre-trained policy, less than 105samples and less than 2 minutes of fine-tuning time are required to achieve satisfactory performance on the target goal.
Xudong Gong, Xunhui Zhang
ISPA3
2023 Pull Request Decisions Explained: An Empirical Overview
abstract
Context: The pull-based development model is widely used in open source projects, leading to the emergence of trends in distributed software development. One aspect that has garnered significant attention concerning pull request decisions is the identification of explanatory factors.Objective: This study builds on a decade of research on pull request decisions and provides further insights. We empirically investigate how factors influence pull request decisions and the scenarios that change the influence of such factors.Method: We identify factors influencing pull request decisions on GitHub through a systematic literature review and infer them by mining archival data. We collect a total of 3,347,937 pull requests with 95 features from 11,230 diverse projects on GitHub. Using these data, we explore the relations among the factors and build mixed effects logistic regression models to empirically explain pull request decisions.Results: Our study shows that a small number of factors explain pull request decisions, with that concerning whether the integrator is the same as or different from the submitter being the most important factor. We also note that the influence of factors on pull request decisions change with a change in context; e.g., the area hotness of pull request is important only in the early stage of project development, however it becomes unimportant for pull request decisions as projects become mature.
Xunhui Zhang, Yue Yu 0001, Georgios Gousios, Ayushi Rastogi
IEEE Trans. Software Eng.1
2022 Who, What, Why and How? Towards the Monetary Incentive in Crowd Collaboration: A Case Study of Github's Sponsor Mechanism
abstract
While many forms of financial support are currently available, there are still many complaints about inadequate financing from software maintainers. In May 2019, GitHub, the world’s most active social coding platform, launched the Sponsor mechanism as a step toward more deeply integrating open source development and financial support. This paper collects data on 8,028 maintainers, 13,555 sponsors, and 22,515 sponsorships and conducts a comprehensive analysis. We explore the relationship between the Sponsor mechanism and developers along four dimensions using a combination of qualitative and quantitative analysis, examining why developers participate, how the mechanism affects developer activity, who obtains more sponsorships, and what mechanism flaws developers have encountered in the process of using it. We find a long-tail effect in the act of sponsorship, with most maintainers’ expectations remaining unmet, and sponsorship has only a short-term, slightly positive impact on development activity but is not sustainable. While sponsors participate in this mechanism mainly as a means of thanking the developers of OSS that they use, in practice, the social status of developers is the primary influence on the number of sponsorships. We find that both the Sponsor mechanism and open source donations have certain shortcomings and need further improvements to attract more participants.
Xunhui Zhang, Tao Wang 0006, Yue Yu 0001, Qiubing Zeng, Huaimin Wang 0001
CHI1
2022 Teegraph: trusted execution environment and directed acyclic graph-based consensus algorithm for IoT blockchains
Xiang Fu 0002, Huaimin Wang 0001, Peichang Shi, Xingkong Ma, Xunhui Zhang
Sci. China Inf. Sci.5
2022 Pull request latency explained: an empirical overview
Xunhui Zhang, Yue Yu 0001, Tao Wang 0006, Ayushi Rastogi, Huaimin Wang 0001
Empir. Softw. Eng.1
2022 Teegraph: A Blockchain consensus algorithm based on TEE and DAG for data sharing in IoT
Xiang Fu 0002, Huaimin Wang 0001, Peichang Shi, Xunhui Zhang
J. Syst. Archit.4
2021 Ladder: A Blockchain Model of Low-Overhead Storage
Xunhui Zhang, Liangliang Xiang, Peichang Shi
BlockSys2
2021 Jointgraph: A DAG-based efficient consensus algorithm for consortium blockchains
abstract
Summary The blockchain is a distributed ledger that records all transactions and operations in a shared manner. Public blockchains such as Bitcoin realize decentralization at the cost of mining overhead, which is not suitable for real‐life scenarios requiring high throughput. Techniques such as the consortium blockchain improve efficiency through partial decentralization. However, the consensus algorithms used in the existing state‐of‐the‐art consortium blockchains face many challenges when dealing with commercial applications. For example, the high communication overhead hinders the scalability of PBFT‐based consensus algorithms even though they are efficient at small scale. Hashgraph, one of the most popular Directed Acyclic Graph‐based (DAG‐based) consensus algorithms, achieves good performance in scalability; however, it does not allow users' dynamic participation. To deal with these challenges, we propose Jointgraph, a Byzantine fault‐tolerance consensus algorithm for consortium blockchains based on DAG. In Jointgraph, transactions are packed into events and validated by no less than 2/3 of all members. A supervisor is introduced in our design, who monitors member behaviors and improves consensus efficiency. Simulation results demonstrate that Jointgraph outperforms Hashgraph in both throughput and latency.
Xiang Fu 0002, Huaimin Wang 0001, Peichang Shi, Xue Ouyang 0003, Xunhui Zhang
Softw. Pract. Exp.5
2020 On the Shoulders of Giants: A New Dataset for Pull-based Development Research
abstract
Pull-based development is a widely adopted paradigm for collaboration in distributed software development, attracting eyeballs from both academic and industry. To better study pull-based development model, this paper presents a new dataset containing 96 features collected from 11,230 projects and 3,347,937 pull requests. We describe the creation process and explain the features in details. To the best of our knowledge, our dataset is the most comprehensive and largest one toward a complete picture for pull-based development research.
Xunhui Zhang, Ayushi Rastogi, Yue Yu 0001
MSR1
2019 BBCPS: A Blockchain Based Open Source Contribution Protection System
Qiubing Zeng, Xunhui Zhang, Tao Wang 0006, Peichang Shi, Xiang Fu 0002, Chenhui Feng
BlockSys2
2019 A Neural-Network based Code Summarization Approach by Using Source Code and its Call Dependencies
abstract
Code summarization aims at generating natural language abstraction for source code, and it can be of great help for program comprehension and software maintenance. The current code summarization approaches have made progress with neural-network. However, most of these methods focus on learning the semantic and syntax of source code snippets, ignoring the dependency of codes. In this paper, we propose a novel method based on neural-network model using the knowledge of the call dependency between source code and its related codes. We extract call dependencies from the source code, transform it as a token sequence of method names, and leverage the Seq2Seq model for code summarization using the combination of source code and call dependency information. About 100,000 code data is collected from 1,000 open source Java proejects on github for experiment. The large-scale code experiment shows that by considering not only the code itself but also the codes it called, the code summarization model can be improved with the BLEU score to 33.08.
Bohong Liu, Tao Wang 0006, Xunhui Zhang, Gang Yin, Jinsheng Deng
Internetware3
2019 RepoLike: amulti-feature-based personalized recommendation approach for open-source repositories
abstract
With the deep integration of software collaborative development and social networking, social coding represents a new style of software production and creation paradigm. Because of their good flexibility and openness, a large number of external contributors have been attracted to the open-source communities. They are playing a significant role in open-source development. However, the open-source development online is a globalized and distributed cooperative work. If left unsupervised, the contribution process may result in inefficiency. It takes contributors a lot of time to find suitable projects or tasks from thousands of open-source projects in the communities to work on. In this paper, we propose a new approach called “RepoLike,” to recommend repositories for developers based on linear combination and learning to rank. It uses the project popularity, technical dependencies among projects, and social connections among developers to measure the correlations between a developer and the given projects. Experimental results show that our approach can achieve over 25% of hit ratio when recommending 20 candidates, meaning that it can recommend closely correlated repositories to social developers.
Cheng Yang 0004, Tao Wang 0006, Gang Yin, Xunhui Zhang, Yue Yu 0001, Huaimin Wang 0001
Frontiers Inf. Technol. Electron. Eng.5
2018 Cross-Project Issue Classification Based on Ensemble Modeling in a Social Coding World
Yarong Zeng, Yue Yu 0001, Xunhui Zhang, Tao Wang 0006, Gang Yin, Huaimin Wang 0001
ICONIP (4)4
2018 Improving code summarization by combining deep learning and empirical knowledge (S)
abstract
Code summaries are human-readable text that describes the functionality of code blocks.Software developers use code summaries to understand the specification of API while code retrieve system relies on code summaries for effective code search.However, code summaries are often written by software developers.Writing good code summaries usually requires great effort.It could be helpful if developers use automatic code summarization system to generate code summaries.Recently, some works have applied deep learning methods to generate code summaries for code snippets.However, those deep learning methods treat code snippets as streams of text tokens while ignoring the inherent code structure information.In this paper, we propose a novel code summarization method named the CDE-Model (Code summarization by Deep learning and Empirical knowledge) that combines inherent code structure information with deep learning models.The CDE-Model proposes several empirical strategies to transform code snippets to refined code representation and feeds them into an encoder-decoder neural network for text generation.We conduct large-scale experiments on 1500 popular Java projects on GitHub 1 with 396,184 pairs of code snippets and summaries.Experimental results show that the quality of code summaries generated by our CDE-Model is better than other two methods.To the best of our knowledge, this paper is the first to combine code structure information with deep learning.
Lingbin Zeng, Xunhui Zhang, Tao Wang 0006, Xiao Li 0039, Huaimin Wang 0001
SEKE2
2017 DevRec: A Developer Recommendation System for Open Source Repositories
Xunhui Zhang, Tao Wang 0006, Gang Yin, Cheng Yang 0004, Yue Yu 0001, Huaimin Wang 0001
ICSR1
2017 An Empirical Study of Reviewer Recommendation in Pull-based Development Model
abstract
Code review is an important process to reduce code defects and improve software quality. However, in social coding communities using the pull-based model, everyone can submit code changes, which increases the required code review efforts. Therefore, there is a great need of knowing the process of code review and analyzing the pre-existing reviewer recommendation algorithms. In this paper, we do an empirical study about the PRs and their reviewers in Rails project. Moreover, we reproduce a popular and effective IR-based code reviewer recommendation algorithm and validate it on our dataset which contains 16,049 PRs. We find that the inactive reviewers are very important to code reviewing process, however, the pre-existing method's recommendation result strongly depends on the activeness of reviewers.
Cheng Yang 0004, Xunhui Zhang, Lingbin Zeng, Gang Yin, Huaimin Wang 0001
Internetware2
2017 Who Will be Interested in? A Contributor Recommendation Approach for Open Source Projects
abstract
The crowds' continuous participation and contribution are the key factors for the success of open source projects.However, among the massive competitors, it is difficult for a project to attract enough contributors by just passively waiting for enthusiasts to join in.Instead, it should actively seek gifted developers.Most of the current studies mainly focus on recommending experts inside a repository for some specific development tasks.In this paper, we propose a novel approach ConRec to recommend potential contributors across the entire open source community for given projects.It leverages the developers' historical activities in projects to analyze their technical interests and technical connections with others.Thereafter, it combines collaborative filtering algorithm with text matching algorithm to recommend proper developers.We conducted extensive experiments on 5,995 open source projects and 2,938,620 developers in GitHub.The results show that the proposed algorithm can recommend contributors to open source projects with the best performance of 63% in accuracy, and solve the cold start problem as well.
Xunhui Zhang, Tao Wang 0006, Gang Yin, Cheng Yang 0004, Huaimin Wang 0001
SEKE1
2016 A Novel Open Source Software Ecosystem: From a Graphic Point of View and Its Application
abstract
With the rapid development of open source software, various elements such as OSS, developers, users and online posts, across different communities and their interactions constitute a novel software ecosystem.Most of the current researches about software ecosystems care the connections between software, and few of them consider the relationship across communities from an overall perspective, and fail to cover the users and their activities which should be an indispensable part of the OSS ecosystem.This paper model the OSS ecosystem as a graph, which combines different types of OSS communities as a whole.Based on this graph model, we analyze the characteristics of ecosystem, like the evolution, competition and symbiosis.In addition, we build a recommendation system as well, and the experiment results suggest the validation of our approach.
Chenxi Song, Tao Wang 0006, Gang Yin, Xunhui Zhang, Cheng Yang 0004
SEKE4