Baihui Sang

dblp:320/7975 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Mining and Assessing Issue Resolution Processes in Open Source Software Repositories
abstract
Resolving issues is a key activity in open-source software (OSS) development, but the ad hoc nature of issue handling on platforms like GitHub makes it challenging to understand the underlying processes of issue resolution, how well such processes are organized, and how to effectively improve issue resolution efficiency. In this work, we quantitatively examine the relationship between issues' organization levels, modeled as process uncertainties, and resolution efficiency measured by issue lifetime and event transition time. We propose an information-theoretical approach to assess the uncertainty lies in issue processing by calculating the entropy of Direct Follow Graphs (DFGs) that represent issue processes. Instead of performing analysis on a project's all issues as a whole, which may lead to over complex models, we find concise yet representative DFG-based models for a project's issues with an entropy-guided KMeans++ clustering algorithm. Findings from real-world OSS projects' issues suggest that higher process uncertainty of issues is associated with longer issue lifetime and extended event transition time. The result also indicates that enhancing issue process organization can potentially improve issue resolution efficiency.
Baihui Sang, Liang Wang 0006, XianPing Tao
CSCWD1
2025 Attributed Multiplex Learning for Analogical Third-Party Library Recommendation and Retrieval
abstract
Third-party libraries (TPLs) play a critical role in modern software development by providing reusable code that accelerates project development. However, the vast number of TPLs available makes selecting the appropriate library for a given task or finding replacements for deprecated libraries a challenging task. Existing methods are limited, relying only on miningbased approaches or feature-based solutions. In this study, we propose an innovative attributed multiplex learning approach that combines both textual and relational data across multiple layers to perform effective analogical library recommendation and retrieval. By representing libraries as nodes with attributes and modeling cross-library relationships as graph edges, our method constructs an attributed multiplex network for TPL representation embeddings. Our approach uses a unified, concise model to include different aspects of information. The proposed inductive model can also address cold-start issues. Moreover, our model is scalable and can adapt to a large number of libraries. To validate our approach, we conduct experiments including an ablation study within the NPM ecosystem. By using a ground-truth data set of 8,308 libraries, the results demonstrate a recommendation precision of 89.8 % at Hit@10. Additionally, we contribute a new data set extracted from deprecation messages containing 4,070 migration rules, enriching the relatively small existing data sets in the NPM ecosystem. In summary, our approach is efficient and promising for supporting real-world, large-scale TPL recommendation and retrieval.
Baihui Sang, Liang Wang 0006, Jierui Zhang, XianPing Tao
ICPC1
2025 An entropy-based measure of fork diversity and its correlations with open source software projects' received contributions
Xiangchen Wu, Liang Wang 0006, Baihui Sang, Jierui Zhang, XianPing Tao
Empir. Softw. Eng.4
2023 Fork Entropy: Assessing the Diversity of Open Source Software Projects' Forks
abstract
On open source software (OSS) platforms such as GitHub, forking and accepting pull-requests is an important approach for OSS projects to receive contributions, especially from external contributors who cannot directly commit into the source repositories. Having a large number of forks is often considered as an indicator of a project being popular. While extensive studies have been conducted to understand the reasons of forking, communications between forks, features and impacts of forks, there are few quantitative measures that can provide a simple yet informative way to gain insights about an OSS project's forks besides their count. Inspired by studies on biodiversity and OSS team diversity, in this paper, we propose an approach to measure the diversity of an OSS project's forks (i.e., its fork population). We devise a novel fork entropy metric based on Rao's quadratic entropy to measure such diversity according to the forks' modifications to project files. With properties including symmetry, continuity, and monotonicity, the proposed fork entropy metric is effective in quantifying the diversity of a project's fork population. To further examine the usefulness of the proposed metric, we conduct empirical studies with data retrieved from fifty projects on GitHub. We observe significant correlations between a project's fork entropy and different outcome variables including the project's external productivity measured by the number of external contributors' commits, acceptance rate of external contributors' pull-requests, and the number of reported bugs. We also observe significant interactions between fork entropy and other factors such as the number of forks. The results suggest that fork entropy effectively enriches our understanding of OSS projects' forks beyond the simple number of forks, and can potentially support further research and applications.
Liang Wang 0006, Xiangchen Wu, Baihui Sang, Jierui Zhang, XianPing Tao
ASE4