Minghui Zhou 0001

dblp:07/4732-1 · also Ming-Hui Zhou 0001 · DBLP profile ↗
← Back
67ranked-venue papers
9as first author
32since 2021 · last 2026
0000-0001-6324-3964ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 61 · 9 first-author · 31 since 2021Applied, interdisciplinary, general and emerging computing · 6Databases, data management, data science and information retrieval · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Correction to: A comprehensive analysis of challenges and strategies for software release notes on github
Hao He 0012, Kai Gao 0008, Wenxin Xiao, Jingyue Li, Minghui Zhou 0001
Empir. Softw. Eng.6
2026 Facilitating Wise Decision-Making for Bounty Backers in Open Source Software Communities
abstract
Bounty programs have become a pivotal incentive mechanism in open-source software (OSS) communities, attracting contributors by offering monetary rewards for task completion. Despite their long-standing implementation, the optimal utilization of this mechanism from the perspective of backers (individuals or entities funding bounties) remains insufficiently understood, hindering its refinement and broader adoption. To bridge this gap, we conduct a mixed-methods study analyzing 10,561 bounty issues fromGitcoin, their linkedGitHubdevelopment data, and surveys from 46 bounty backers. We investigate three core decision-making dimensions: (1) why backers use bounties and the actual outcomes, (2) what issues backers prioritize, and (3) how bounty amounts are set. Our findings reveal that backers primarily seek to enhance developer engagement, project visibility, and task efficiency. However, the actual outcomes often diverge from expectations: although bounty issues have a higher resolution rate (+12%) than non-bounty issues, they also introduce systemic challenges, such as delayed resolutions (+33 days) and difficulties in engaging new developers. Notably, backers tend to prioritize feature-related, intermediate-complexity tasks with short completion timelines, while showing relatively less interest in overly simplistic or highly specialized work. Reward allocation follows a nuanced approach: lower bounties target beginner-friendly tasks, while higher rewards are reserved for advanced skills or multi-week commitments. However, backers often lack systematic methods to calibrate rewards, leading to frequent bounty adjustments. To enable data-driven decision-making, we propose a bounty recommendation predictor that uses empirical factors to predict appropriate bounty amount. By synthesizing these insights, our study offers OSS communities actionable strategies to refine bounty programs, balancing short-term productivity with long-term ecosystem sustainability.
Xin Tan 0003, Xianjun Ni, Yuxia Zhang, Jing Jiang 0005, Minghui Zhou 0001, Li Zhang 0029
IEEE Trans. Software Eng.6
2025 Licoeval: Evaluating LLMs on License Compliance in Code Generation
abstract
Recent advances in Large Language Models (LLMs) have revolutionized code generation, leading to widespread adoption of AI coding tools by developers. However, LLMs can generate license-protected code without providing the necessary license information, leading to potential intellectual property violations during software production. This paper addresses the critical, yet underexplored, issue of license compliance in LLM-generated code by establishing a benchmark to evaluate the ability of LLMs to provide accurate license information for their generated code. To establish this benchmark, we conduct an empirical study to identify a reasonable standard for “striking similarity” that excludes the possibility of independent creation, indicating a copy relationship between the LLM output and certain opensource code. Based on this standard, we propose LiCoEval, to evaluate the license compliance capabilities of LLMs, i.e., the ability to provide accurate license or copyright information when they generate code with striking similarity to already existing copyrighted code. Using LiCoEval, we evaluate 14 popular LLMs, finding that even top-performing LLMs produce a non-negligible proportion (0.88 % to 2.01 %) of code strikingly similar to existing open-source implementations. Notably, most LLMs fail to provide accurate license information, particularly for code under copyleft licenses. These findings underscore the urgent need to enhance LLM compliance capabilities in code generation tasks. Our study provides a foundation for future research and development to improve license compliance in AIassisted software development, contributing to both the protection of open-source software copyrights and the mitigation of legal risks for LLM users.
Weiwei Xu 0001, Kai Gao 0008, Hao He 0012, Minghui Zhou 0001
ICSE4
2025 A First Look at Package-to-Group Mechanism: An Empirical Study of the Linux Distributions
abstract
Reusing third-party software packages is a common practice in software development. As the scale and complexity of open-source software (OSS) projects continue to grow (e.g., Linux distributions), the number of reused third-party packages has significantly increased. Therefore, effective package management is essential for the development and evolution of the OSS project. To achieve this, a package-to-group mechanism (P2G) is used to enable the unified installation, uninstallation, and updates of multiple packages at once. To better understand the mechanism, this paper takes Linux distributions as a case study and presents an empirical study focusing on its application trends, evolution patterns, group quality, and group tendency. By analyzing 11,746 groups and 193,548 packages from 89 versions of 5 popular Linux distributions and conducting questionnaire surveys with Linux practitioners and researchers, we derive several key insights. Our findings show that P2G is increasingly being adopted, particularly in popular Linux distributions. P2G follows six evolutionary patterns (e.g., splitting and merging groups). Interestingly, packages no longer managed through P2G are more likely to remain in Linux distributions rather than being directly removed. In addition, we propose a metric called GValue to evaluate the quality of groups and identify issues such as inadequate group descriptions and insufficient group sizes. We also summarize five types of packages that tend to adopt P2G, including graphical desktops and networks. To our knowledge, this is the first study to focus on P2G mechanisms. We hope that our study can assist in the efficient management of packages and reduce the burden on practitioners in the rapidly growing Linux distributions and other open-source software projects.
Dongming Jin, Nianyu Li, Kai Yang 0053, Minghui Zhou 0001, Zhi Jin 0001
SANER4
2025 Characterising Open Source Co-opetition in Company-hosted Open Source Software Projects: The Cases of PyTorch, TensorFlow, and Transformers
abstract
Companies, including market rivals, have long collaborated on open source software (OSS) development, resulting in a tangle of co-operation and competition known as "open source co-opetition". While prior work investigates open source co-opetition in OSS projects that are hosted by vendor-neutral foundations, we have a limited understanding thereof in OSS projects that are hosted and governed by one company. Given their prevalence, it is timely to investigate open source co-opetition in such contexts. Towards this end, we conduct a mixed-methods analysis of three company-hosted OSS projects in the artificial intelligence (AI) industry: Meta's PyTorch prior to its donation to the Linux Foundation, Google's TensorFlow, and Hugging Face's Transformers. We contribute three key findings. First, while the projects exhibit similar code authorship patterns between host and external companies (~80%/20% of commits), collaborations are structured differently (e.g. decentralised vs. hub-and-spoke networks). Second, host and external companies engage in strategic, non-strategic, and contractual collaborations, with varying incentives and collaboration practices. Some of the observed collaborations are specific to the AI industry (e.g. AI model integrations), while others are typical of the broader software industry (e.g. bug fixing or task outsourcing). Third, single-vendor governance creates a power imbalance that influences open source co-opetition practices and possibilities, from the host company's singular decision-making power (e.g. the risk of license changes) to their community involvement strategy (e.g. from over-control to over-delegation). We conclude with recommendations for future research.
Cailean Osborne, Farbod Daneshyan, Runzhi He, Hengzhi Ye, Yuxia Zhang, Minghui Zhou 0001
Proc. ACM Hum. Comput. Interact.6
2025 Systematic Literature Review of Commercial Participation in Open Source Software
abstract
Open source software (OSS) has been playing a fundamental role in not only information technology but also our social lives. Attracted by various advantages of OSS, increasing commercial companies are participating extensively in open source development, and this has had a broad impact. Enormous research efforts have been devoted to understanding this phenomenon and trying to pursue a win-win result. To characterize the current research achievement and identify challenges, this article provides a comprehensive systematic literature review (SLR) of existing research on company participation in OSS. We collected 105 papers and organized them based on their research topics, which cover three main directions, i.e., participation motivation, contribution model, and impact on OSS development. We found that companies have diverse motivations from economic, technological, and social aspects, and no one study covered all the motivation categories. Existing studies categorize five main companies’ contribution models in OSS projects through their objectives and how they shape OSS communities. Researchers also explored how commercial participation affects OSS development, including companies, developers, and OSS projects. This study contributes to a comprehensive understanding of commercial participation in OSS development. Based on our findings, we present a set of research challenges and promising directions for companies’ better participation in OSS.
Xuetao Li, Yuxia Zhang, Cailean Osborne, Minghui Zhou 0001, Zhi Jin 0001, Hui Liu 0003
ACM Trans. Softw. Eng. Methodol.4
2025 Developers' Views on Commercial Involvement in OSS: A Survey From Three Projects
abstract
Given the well-established merits of open source software (OSS), many profit-oriented companies actively participate in OSS communities, making significant contributions. Existing studies have predominantly focused on several advanced and specific questions regarding this phenomenon, mainly from the companies’ perspective, such as companies’ domination and withdrawal. A more basic and comprehensive understanding is missing, i.e., how OSS developers perceive such corporate engagement. Individual developers, including both volunteers and developers assigned by companies, are directly impacted by and have personal experiences with the consequences of commercial participation in OSS projects. This paper aims to bridge this gap by amplifying the voices of individual developers and providing valuable insights that have the potential to enhance companies’ participation in OSS projects. We conducted a survey involving developers from three OSS projects, i.e., Rust, OpenStack, and the Linux kernel, focusing on their attitudes and expectations regarding corporate involvement. We received 84 meaningful responses and analyzed their open-ended responses through thematic analysis. The results suggest that regardless of whether developers were paid or voluntary contributors, a prevailing attitude emerged – 67.9% of developers expressed a positive view of companies’ participation in OSS. The key idea behind their positive attitudes is perceiving commercial participation as a win-win for both the OSS community and companies. The Rust community remains more neutral when compared with the other two communities. We also surveyed and analyzed developers’ expectations of companies’ better participation, which can shed light on how OSS ecosystems can sustainably evolve with the companies involved.
Yuxia Zhang, Minghui Zhou 0001, Haoyang Li 0008, Hui Liu 0003
IEEE Trans. Software Eng.3
2024 How Are Paid and Volunteer Open Source Developers Different? A Study of the Rust Project
abstract
It is now commonplace for organizations to pay developers to work on specific open source software (OSS) projects to pursue their business goals. Such paid developers work alongside voluntary contributors, but given the different motivations of these two groups of developers, conflict may arise, which may pose a threat to a project's sustainability. This paper presents an empirical study of paid developers and volunteers in Rust, a popular open source programming language project. Rust is a particularly interesting case given considerable concerns about corporate participation. We compare volunteers and paid developers through contribution characteristics and long-term participation, and solicit volunteers' perceptions on paid developers. We find that core paid developers tend to contribute more frequently; commits contributed by onetime paid developers have bigger sizes; peripheral paid developers implement more features; and being paid plays a positive role in becoming a long-term contributor. We also find that volunteers do have some prejudices against paid developers. This study suggests that the dichotomous view of paid vs. volunteer developers is too simplistic and that further subgroups can be identified. Companies should become more sensitive to how they engage with OSS communities, in certain ways as suggested by this study.
Yuxia Zhang, Klaas-Jan Stol, Minghui Zhou 0001, Hui Liu 0003
ICSE4
2024 A comprehensive analysis of challenges and strategies for software release notes on GitHub
Hao He 0012, Kai Gao 0008, Wenxin Xiao, Jingyue Li, Minghui Zhou 0001
Empir. Softw. Eng.6
2024 Characterizing Deep Learning Package Supply Chains in PyPI: Domains, Clusters, and Disengagement
abstract
Deep learning (DL) frameworks have become the cornerstone of the rapidly developing DL field. Through installation dependencies specified in the distribution metadata, numerous packages directly or transitively depend on DL frameworks, layer after layer, forming DL package supply chains (SCs), which are critical for DL frameworks to remain competitive. However, vital knowledge on how to nurture and sustain DL package SCs is still lacking. Achieving this knowledge may help DL frameworks formulate effective measures to strengthen their SCs to remain competitive and shed light on dependency issues and practices in the DL SC for researchers and practitioners. In this paper, we explore the domains, clusters, and disengagement of packages in two representative PyPI DL package SCs to bridge this knowledge gap. We analyze the metadata of nearly six million PyPI package distributions and construct version-sensitive SCs for two popular DL frameworks: TensorFlow and PyTorch. We find that popular packages (measured by the number of monthly downloads) in the two SCs cover 34 domains belonging to eight categories. Applications , Infrastructure , and Sciences categories account for over 85% of popular packages in either SC and TensorFlow and PyTorch SC have developed specializations on Infrastructure and Applications packages, respectively. We employ the Leiden community detection algorithm and detect 131 and 100 clusters in the two SCs. The clusters mainly exhibit four shapes: Arrow, Star, Tree, and Forest with increasing dependency complexity. Most clusters are Arrow or Star, while Tree and Forest clusters account for most packages (Tensorflow SC: 70.7%, PyTorch SC: 92.9%). We identify three groups of reasons why packages disengage from the SC (i.e., remove the DL framework and its dependents from their installation dependencies): dependency issues, functional improvements, and ease of installation. The most common reason in TensorFlow SC is dependency incompatibility and in PyTorch SC is to simplify functionalities and reduce installation size. Our study provides rich implications for DL framework vendors, researchers, and practitioners on the maintenance and dependency management practices of PyPI DL SCs.
Kai Gao 0008, Runzhi He, Minghui Zhou 0001
ACM Trans. Softw. Eng. Methodol.4
2023 Is It Enough to Recommend Tasks to Newcomers? Understanding Mentoring on Good First Issues
abstract
Newcomers are critical for the success and continuity of open source software (OSS) projects. To attract newcomers and facilitate their onboarding, many OSS projects recommend tasks for newcomers, such as good first issues (GFIs). Previous studies have preliminarily investigated the effects of GFIs and techniques to identify suitable GFIs. However, it is still unclear whether just recommending tasks is enough and how significant mentoring is for newcomers. To better understand mentoring in OSS communities, we analyze the resolution process of 48,402 GFIs from 964 repositories through a mix-method approach. We investigate the extent, the mentorship structures, the discussed topics, and the relevance of expert involvement. We find that ~70% of GFIs have expert participation, with each GFI usually having one expert who makes two comments. Half of GFIs will receive their first expert comment within 8.5 hours after a newcomer comment. Through analysis of the collaboration networks of newcomers and experts, we observe that community mentorship presents four types of structure: centralized mentoring, decentralized mentoring, collaborative mentoring, and distributed mentoring. As for discussed topics, we identify 14 newcomer challenges and 18 expert mentoring content. By fitting the generalized linear models, we find that expert involvement positively correlates with newcomers' successful contributions but negatively correlates with newcomers' retention. Our study manifests the status and significance of mentoring in the OSS projects, which provides rich practical implications for optimizing the mentoring process and helping newcomers contribute smoothly and suecessfully,
Xin Tan 0003, Yiran Chen 0005, Haohua Wu, Minghui Zhou 0001, Li Zhang 0029
ICSE4
2023 Personalized First Issue Recommender for Newcomers in Open Source Projects
abstract
Many open source projects provide good first issues (GFIs) to attract and retain newcomers. Although several automated GFI recommenders have been proposed, existing recommenders are limited to recommending generic GFIs without considering differences between individual newcomers. However, we observe mismatches between generic GFIs and the diverse background of newcomers, resulting in failed attempts, discouraged onboarding, and delayed issue resolution. To address this problem, we assume that personalized first issues (PFIs) for newcomers could help reduce the mismatches. To justify the assumption, we empirically analyze 37 newcomers and their first issues resolved across multiple projects. We find that the first issues resolved by the same newcomer share similarities in task type, programming language, and project domain. These findings underscore the need for a PFI recommender to improve over state-of-the-art approaches. For that purpose, we identify features that influence newcomers' personalized selection of first issues by analyzing the relationship between possible features of the newcomers and the characteristics of the newcomers' chosen first issues. We find that the expertise preference, OSS experience, activeness, and sentiment of newcomers drive their personalized choice of the first issues. Based on these findings, we propose a Personalized First Issue Recommender (PFIRec), which employs LamdaMART to rank candidate issues for a given newcomer by leveraging the identified influential features. We evaluate PFIRec using a dataset of 68,858 issues from 100 GitHub projects. The evaluation results show that PFIRec outperforms existing first issue recommenders, potentially doubling the probability that the top recommended issue is suitable for a specific newcomer and reducing one-third of a newcomer's unsuccessful attempts to identify suitable first issues, in the median. We provide a replication package at https://zenodo.org/record/7915841.
Wenxin Xiao, Jingyue Li, Hao He 0012, Ruiqiao Qiu, Minghui Zhou 0001
ASE5
2023 Understanding and Remediating Open-Source License Incompatibilities in the PyPI Ecosystem
abstract
The reuse and distribution of open-source software must be in compliance with its accompanying open-source license. In modern packaging ecosystems, maintaining such compliance is challenging because a package may have a complex multi-layered dependency graph with many packages, any of which may have an incompatible license. Although prior research finds that license incompatibilities are prevalent, empirical evidence is still scarce in some modern packaging ecosystems (e.g., PyPI). It also remains unclear how developers remediate the license incompatibilities in the dependency graphs of their packages (including direct and transitive dependencies), let alone any automated approaches. To bridge this gap, we conduct a large-scale empirical study of license incompatibilities and their remediation practices in the PyPI ecosystem. We find that 7.27% of the PyPI package releases have license incompatibilities and 61.3 % of them are caused by transitive dependencies, causing challenges in their remediation; for remediation, developers can apply one of the five strategies: migration, removal, pinning versions, changing their own licenses, and negotiation. Inspired by our findings, we propose Silence, an SMT-solver-based approach to recommend license incompatibility remediations with minimal costs in package dependency graph. Our evaluation shows that the remediations proposed by Silencecan match 19 historical real-world cases (except for migrations not covered by an existing knowledge base) and have been accepted by five popular PyPI packages whose developers were previously unaware of their license incompatibilities.
Weiwei Xu 0001, Hao He 0012, Kai Gao 0008, Minghui Zhou 0001
ASE4
2023 AutoML from Software Engineering Perspective: Landscapes and Challenges
abstract
Machine learning (ML) has been widely adopted in modern software, but the manual configuration of ML (e.g., hyper-parameter configuration) poses a significant challenge to software developers. Therefore, automated ML (AutoML), which seeks the optimal configuration of ML automatically, has received increasing attention from the software engineering community. However, to date, there is no comprehensive understanding of how AutoML is used by developers and what challenges developers encounter in using AutoML for software development. To fill this knowledge gap, we conduct the first study on understanding the use and challenges of AutoML from software developers’ perspective. We collect and analyze 1,554 AutoML downstream repositories, 769 AutoML-related Stack Overflow questions, and 1,437 relevant GitHub issues. The results suggest the increasing popularity of AutoML in a wide range of topics, but also the lack of relevant expertise. We manually identify specific challenges faced by developers for AutoML-enabled software. Based on the results, we derive a series of implications for AutoML framework selection, framework development, and research.
Zhenpeng Chen 0001, Minghui Zhou 0001
MSR3
2023 How Early Participation Determines Long-Term Sustained Activity in GitHub Projects?
abstract
Although the open source model bears many advantages in software development, open source projects are always hard to sustain. Previous research on open source sustainability mainly focuses on projects that have already reached a certain level of maturity (e.g., with communities, releases, and downstream projects). However, limited attention is paid to the development of (sustainable) open source projects in their infancy, and we believe an understanding of early sustainability determinants is crucial for project initiators, incubators, newcomers, and users.
Wenxin Xiao, Hao He 0012, Weiwei Xu 0001, Yuxia Zhang, Minghui Zhou 0001
ESEC/SIGSOFT FSE5
2023 Self-Admitted Library Migrations in Java, JavaScript, and Python Packaging Ecosystems: A Comparative Study
abstract
Reusing open-source software libraries has become the norm in modern software development, but libraries can fail due to various reasons, e.g., security vulnerabilities, lacking features, and end of maintenance. In some cases, developers need to replace a library with another competent library with similar functionalities, i.e., library migration. Previous studies have leveraged library migrations as a unique lens of observation to reveal insights into library selection and dependency management in general. However, they are heavily biased toward Java while the generalizability of their findings remains unknown.In this paper, we present a comparative study on self-admitted library migrations (SALMs) from three packaging ecosystems: Java/Maven, JavaScript/npm, and Python/PyPI. For this study, we design a set of semi-automatic methods that accurately locate SALMs, their domains, and their rationales from git repositories. We reveal that SALMs are prevalent and highly unidirectional in all three ecosystems, and the underlying rationales can be well covered by a previous theoretical framework. Also, SALMs in these ecosystems present domain similarity (testing frameworks, web frameworks, HTTP clients, and serialization). However, we observe differences in the longitudinal trends, the distributions of rationales, the ecosystem-specific domains, and the levels of unidirectionality, all of which indicate that Python/PyPI sees increasingly intense competition between libraries and deserves more research on library recommendation and migration.
Haiqiao Gu, Hao He 0012, Minghui Zhou 0001
SANER3
2023 Characterize Software Release Notes of GitHub Projects: Structure, Writing Style, and Content
abstract
Release notes (RNs) summarize important changes between two successive software releases to facilitate software upgrades and serve as means of communication between software and its users. However, existing research has shown that many users cannot extract the information they want from RNs effectively and efficiently due to poor structure and insufficient content. Many efforts have been devoted to categorizing documented information in RNs, however, how exactly RNs are organized, in what way RNs are written, and what is written in RNs with respect to project domains and release types remain under investigation. To bridge this knowledge gap, we manually analyzed 612 RNs from 233 top popular GitHub projects to characterize their Structure, Writing Style, and Content. We find 64.54% of RNs organize changes into hierarchical structures following three strategies, i.e., by Change Type, Affected Module, or Change Priority. And 11.60% of RNs adopt multiple strategies to present the changes. Among the three strategies to organize changes, by Change Type is mostly adopted. RNs of major releases and System Software are more likely to organize changes by Affected Module. We also find three types of Writing Styles: Expository, Descriptive, and Persuasive with increasing explanation information and manual effort, taking 30.07%, 34.80%, and 35.13% of RNs, respectively. 83.10% of RNs in System Software projects adopt the Descriptive and Persuasive writing style. For Content, we find System Software and Libraries & Frameworks projects are more likely to record Breaking Changes, while Software Tools projects emphasize Enhancements. Besides, Fixed Bugs and Security Changes are less common in RNs of major releases. Our findings not only serve practitioners with a roadmap for customizing high-quality RNs but also shed light on future research on automating RN.
Weiwei Xu 0001, Kai Gao 0008, Jingyue Li, Minghui Zhou 0001
SANER5
2023 Suboptimal Comments in Java Projects: From Independent Comment Changes to Commenting Practices
abstract
High-quality source code comments are valuable for software development and maintenance, however, code often contains low-quality comments or lacks them altogether. We name such source code comments as suboptimal comments. Such suboptimal comments create challenges in code comprehension and maintenance. Despite substantial research on low-quality source code comments, empirical knowledge about commenting practices that produce suboptimal comments and reasons that lead to suboptimal comments are lacking. We help bridge this knowledge gap by investigating (1) independent comment changes (ICCs)—comment changes committed independently of code changes—which likely address suboptimal comments, (2) commenting guidelines, and (3) comment-checking tools and comment-generating tools, which are often employed to help commenting practice—especially to prevent suboptimal comments. We collect 24M+ comment changes from 4,392 open-source GitHub Java repositories and find that ICCs widely exist. TheICC ratio—proportion of ICCs among all comment changes—is ~15.5%, with 98.7% of the repositories having ICC. Our thematic analysis of 3,533 randomly sampled ICCs provides a three-dimensional taxonomy forwhatis changed (four comment categories and 13 subcategories),howit changed (six commenting activity categories), andwhat factorsare associated with the change (three factors). We investigate 600 repositories to understand the prevalence, content, impact, and violations of commenting guidelines. We find that only 15.5% of the 600 sampled repositories have any commenting guidelines. We provide the first taxonomy for elements in commenting guidelines: where and what to comment are particularly important. The repositories without such guidelines have a statistically significantly higher ICC ratio, indicating the negative impact of the lack of commenting guidelines. However, commenting guidelines are not strictly followed: 85.5% of checked repositories have violations. We also systematically study how developers use two kinds of tools, comment-checking tools and comment-generating tools, in the 4,392 repositories. We find that the use ofJavadoctool is negatively correlated with the ICC ratio, while the use ofCheckstylehas no statistically significant correlation; the use of comment-generating tools leads to a higher ICC ratio. To conclude, we reveal issues and challenges in current commenting practice, which help understand how suboptimal comments are introduced. We propose potential research directions on comment location prediction, comment generation, and comment quality assessment; suggest how developers can formulate commenting guidelines and enforce rules with tools; and recommend how to enhance current comment-checking and comment-generating tools.
Hao He 0012, Uma Pal, Darko Marinov, Minghui Zhou 0001
ACM Trans. Softw. Eng. Methodol.5
2023 On the Variability of Software Engineering Needs for Deep Learning: Stages, Trends, and Application Types
abstract
The wide use of Deep learning (DL) has not been followed by the corresponding advances in software engineering (SE) for DL. Research shows that developers writing DL software have specific development stages (i.e., SE4DL stages) and face new DL-specific problems. Despite substantial research, it is unclear how DL developers’ SE needs for DL vary over stages, application types, or if they change over time. To help focus research and development efforts on DL-development challenges, we analyze 92,830 Stack Overflow (SO) questions and 227,756 READMEs of public repositories related to DL. Latent Dirichlet Allocation (LDA) reveals 27 topics for the SO questions where 19 (70.4%) topics mainly relate to a single SE4DL stage, and eight topics span multiple stages. Most questions concernData PreparationandModel Setupstages. The relative rates of questions for 11 topics have increased, for eight topics decreased over time. Questions for the former 11 topics had a lower percentage of accepting an answer than the remaining questions. LDA on README files reveals 16 distinct application types for the 227k repositories. We apply the LDA model fitted on READMEs to the 92,830 SO questions and find that 27% of the questions are related to the 16 DL application types. The most asked question topic varies across application types, with half primarily relating to the second and third stages. Specifically, developers ask the most questions about topics primarily relating toData Preparation(2nd) stage for four mature application types such as${{\sf Image\ Segmentation}}$, and topics primarily relating toModel Setup(3rd) stage for four application types concerning emerging methods such as${{\sf Transfer\ Learning}}$. Based on our findings, we distill several actionable insights for SE4DL research, practice, and education, such as better support for using trained models, application-type specific tools, and teaching materials.
Kai Gao 0008, Audris Mockus, Minghui Zhou 0001
IEEE Trans. Software Eng.4
2023 Automating Dependency Updates in Practice: An Exploratory Study on GitHub Dependabot
abstract
Dependency management bots automatically open pull requests to update software dependencies on behalf of developers. Early research shows that developers are suspicious of updates performed by dependency management bots and feel tired of overwhelming notifications from these bots. Despite this, dependency management bots are becoming increasingly popular. Such contrast motivates us to investigate Dependabot, currently the most visible bot on GitHub, to reveal the effectiveness and limitations of state-of-art dependency management bots. We use exploratory data analysis and a developer survey to evaluate the effectiveness of Dependabot in keeping dependencies up-to-date, interacting with developers, reducing update suspicion, and reducing notification fatigue. We obtain mixed findings. On the positive side, projects do reduce technical lag after Dependabot adoption and developers are highly receptive to its pull requests. On the negative side, its compatibility scores are too scarce to be effective in reducing update suspicion; developers tend to configure Dependabot toward reducing the number of notifications; and 11.3% of projects have deprecated Dependabot in favor of other alternatives. The survey confirms our findings and provides insights into the key missing features of Dependabot. Based on our findings, we derive and summarize the key characteristics of an ideal dependency management bot which can be grouped into four dimensions: configurability, autonomy, transparency, and self-adaptability.
Runzhi He, Hao He 0012, Yuxia Zhang, Minghui Zhou 0001
IEEE Trans. Software Eng.4
2023 Understanding Mentors' Engagement in OSS Communities via Google Summer of Code
abstract
A constant influx of newcomers is essential for the sustainability and success of open source software (OSS) projects. However, successful onboarding is always challenging because newcomers face various initial contributing barriers. To support newcomer onboarding, OSS communities widely adopt the mentoring approach. Despite its significance, previous mentoring studies tend to focus on the newcomer's perspective, leaving the mentor's perspective relatively under-studied. To better support mentoring, we study the popular Google Summer of Code (GSoC). It is a well-established global program that offers stipends and mentors to students aiming to bring more student developers into OSS development. We combine online data analysis, an email survey, and semi-structured interviews with the GSoC mentors to understand their motivations, challenges, strategies, and gains. We propose a taxonomy of GSoC mentors’ engagement with four themes, ten categories, 34 sub-categories, and 118 codes, as well as the mentors’ attitudes toward the codes. In particular, we find that mentors participating in GSoC are primarily intrinsically motivated, and some new motivators emerge adapting to the contemporary challenges, e.g., sustainability and advertisement of projects. Forty-one challenges and 52 strategies associated with the program timeline are identified, most of which are first time revealed. Although almost all the challenges are agreed upon by specific mentors, some mentors believe that several challenges are reasonable and even have a positive effect. For example, the cognitive differences between mentors and mentees can stimulate new perspectives. Most of the mentors agreed that they had adopted these strategies during the mentoring process, but a few strategies recommended by the GSoC administration were not agreed upon. Self-satisfaction, different skills, and peer recognition are the main gains of mentors to participate in GSoC. Eventually, we discuss practical implications for mentors, students, OSS communities, GSoC programs, and researchers.
Xin Tan 0003, Minghui Zhou 0001, Li Zhang 0029
IEEE Trans. Software Eng.2
2022 An Exploratory Study of Deep learning Supply Chain
abstract
Deep learning becomes the driving force behind many contemporary technologies and has been successfully applied in many fields. Through software dependencies, a multi-layer supply chain (SC) with a deep learning framework as the core and substantial down-stream projects as the periphery has gradually formed and is constantly developing. However, basic knowledge about the structure and characteristics of the SC is lacking, which hinders effective support for its sustainable development. Previous studies on software SC usually focus on the packages in different registries without paying attention to the SCs derived from a single project. We present an empirical study on two deep learning SCs: TensorFlow and PyTorch SCs. By constructing and analyzing their SCs, we aim to understand their structure, application domains, and evolutionary factors. We find that both SCs exhibit a short and sparse hierarchy structure. Overall, the relative growth of new projects increases month by month. Projects have a tendency to attract downstream projects shortly after the release of their packages, later the growth becomes faster and tends to stabilize. We propose three criteria to identify vulnerabilities and identify 51 types of packages and 26 types of projects involved in the two SCs. A comparison reveals their similarities and differences, e.g., TensorFlow SC provides a wealth of packages in experiment result analysis, while PyTorch SC contains more specific framework packages. By fitting the GAM model, we find that the number of dependent packages is significantly negatively associated with the number of downstream projects, but the relationship with the number of authors is nonlinear. Our findings can help further open the "black box" of deep learning SCs and provide insights for their healthy and sustainable development.
Xin Tan 0003, Kai Gao 0008, Minghui Zhou 0001, Li Zhang 0029
ICSE3
2022 Recommending Good First Issues in GitHub OSS Projects
abstract
Attracting and retaining newcomers is vital for the sustainability of an open-source software project. However, it is difficult for newcomers to locate suitable development tasks, while existing "Good First Issues" (GFI) in GitHub are often insufficient and inappropriate. In this paper, we propose RecGFI, an effective practical approach for the recommendation of good first issues to newcomers, which can be used to relieve maintainers' burden and help newcomers onboard. RecGFI models an issue with features from multiple dimensions (content, background, and dynamics) and uses an XGBoost classifier to generate its probability of being a GFI. To evaluate RecGFI, we collect 53,510 resolved issues among 100 GitHub projects and carefully restore their historical states to build ground truth datasets. Our evaluation shows that RecGFI can achieve up to 0.853 AUC in the ground truth dataset and outperforms alternative models. Our interpretable analysis of the trained model further reveals interesting observations about GFI characteristics. Finally, we report latest issues (without GFI-signaling labels but recommended as GFI by our approach) to project maintainers among which 16 are confirmed as real GFIs and five have been resolved by a newcomer.
Wenxin Xiao, Hao He 0012, Weiwei Xu 0001, Xin Tan 0003, Jinhao Dong, Minghui Zhou 0001
ICSE6
2022 Demystifying software release note issues on GitHub
abstract
Release notes (RNs) summarize main changes between two consecutive software versions and serve as a central source of information when users upgrade software. While producing high quality RNs can be hard and poses a variety of challenges to developers, a comprehensive empirical understanding of these challenges is still lacking. In this paper, we bridge this knowledge gap by manually analyzing 1,731 latest GitHub issues to build a comprehensive taxonomy of RN issues with four dimensions: Content, Presentation, Accessibility, and Production. Among these issues, nearly half (48.47%) of them focus on Production; Content, Accessibility, and Presentation take 25.61%, 17.65%, and 8.27%, respectively. We find that: 1) RN producers are more likely to miss information than to include incorrect information, especially for breaking changes; 2) improper layout may bury important information and confuse users; 3) many users find RNs inaccessible due to link deterioration, lack of notification, and obfuscate RN locations; 4) automating and regulating RN production remains challenging despite the great needs of RN producers. Our taxonomy not only pictures a roadmap to improve RN production in practice but also reveals interesting future research directions for automating RN production.
Hao He 0012, Wenxin Xiao, Kai Gao 0008, Minghui Zhou 0001
ICPC5
2022 GFI-bot: automated good first issue recommendation on GitHub
abstract
To facilitate newcomer onboarding, GitHub recommends the use of "good first issue" (GFI) labels to signal issues suitable for newcomers to resolve. However, previous research shows that manually labeled GFIs are scarce and inappropriate, showing a need for automated recommendations. In this paper, we present GFI-Bot (accessible at https://gfibot.io), a proof-of-concept machine learning powered bot for automated GFI recommendation in practice. Project maintainers can configure GFI-Bot to discover and label possible GFIs so that newcomers can easily locate issues for making their first contributions. GFI-Bot also provides a high-quality, up-to-date dataset for advancing GFI recommendation research.
Hao He 0012, Haonan Su, Wenxin Xiao, Runzhi He, Minghui Zhou 0001
ESEC/SIGSOFT FSE5
2022 Corporate dominance in open source ecosystems: a case study of OpenStack
abstract
Corporate participation plays an increasing role in Open Source Software (OSS) development. Unlike volunteers in OSS projects, companies are driven by business objectives. To pursue corporate interests, companies may try to dominate the development direction of OSS projects. One company's domination in OSS may 'crowd out' other contributors, changing the nature of the project, and jeopardizing the sustainability of the OSS ecosystem. Prior studies of corporate involvement in OSS have primarily focused on predominately positive aspects such as business strategies, contribution models, and collaboration patterns. However, there is a scarcity of research on the potential drawbacks of corporate engagement. In this paper, we investigate corporate dominance in OSS ecosystems. We draw on the field of Economics and quantify company domination using a dominance measure; we investigate the prevalence, patterns, and impact of domination in the evolution of the OpenStack ecosystem. We find evidence of company domination in over 73% of the repositories in OpenStack, and approximately 25% of companies dominate one or more repositories per version. We identify five patterns of corporate dominance: Early incubation, Full-time hosting, Growing domination, Occasional domination, and Last remaining. We find that domination has a significantly negative relationship with the survival probability of OSS projects. This study provides insights for building sustainable relationships between companies and the OSS ecosystems in which they seek to get involved.
Yuxia Zhang, Klaas-Jan Stol, Hui Liu 0003, Minghui Zhou 0001
ESEC/SIGSOFT FSE4
2022 Turnover of Companies in OpenStack: Prevalence and Rationale
abstract
To achieve commercial goals, companies have made substantial contributions to large open-source software (OSS) ecosystems such as OpenStack and have become the main contributors. However, they often withdraw their employees for a variety of reasons, which may affect the sustainability of OSS projects. While the turnover of individual contributors has been extensively investigated, there is a lack of knowledge about the nature of companies’ withdrawal. To this end, we conduct a mixed-methods empirical study on OpenStack to reveal how common company withdrawals were, to what degree withdrawn companies made contributions, and what the rationale behind withdrawals was. By analyzing the commit data of 18 versions of OpenStack, we find that the number of companies that have left is increasing and even surpasses the number of companies that have joined in later versions. Approximately 12% of the companies in each version have exited by the next version. Compared to the sustaining companies that joined in the same version, the withdrawn companies tend to have a weaker contribution intensity but contribute to a similar scope of repositories in OpenStack. Through conducting a developer survey, we find four aspects of reasons for companies’ withdrawal from OpenStack: company, community, developer, and project. The most common reasons lie in the company aspect, i.e., the company either achieved its goals or failed to do so. By fitting the survival analysis model, we find that commercial goals are associated with the probability of the company’s withdrawal, and that a company’s contribution intensity and scale are positively correlated with its retention. Maintaining good retention is important but challenging for OSS ecosystems, and our results may shed light on potential approaches to improve company retention and reduce the negative impact of company withdrawal.
Yuxia Zhang, Hui Liu 0003, Xin Tan 0003, Minghui Zhou 0001, Zhi Jin 0001
ACM Trans. Softw. Eng. Methodol.4
2022 Redundancy, Context, and Preference: An Empirical Study of Duplicate Pull Requests in OSS Projects
abstract
OSS projects are being developed by globally distributed contributors, who often collaborate through the pull-based model today. While this model lowers the barrier to entry for OSS developers by synthesizing, automating and optimizing the contribution process, coordination among an increasing number of contributors remains as a challenge due to the asynchronous and self-organized nature of distributed development. In particular, duplicate contributions, where multiple different contributors unintentionally submit duplicate pull requests to achieve the same goal, are an elusive problem that may waste effort in automated testing, code review and software maintenance. While the issue of duplicate pull requests has been highlighted, to what extent duplicate pull requests affect the development in OSS communities has not been well investigated. In this paper, we conduct a mixed-approach study to bridge this gap. Based on a comprehensive dataset constructed from 26 popular GitHub projects, we obtain the following findings: (a) Duplicate pull requests result in redundant human and computing resources, exerting a significant impact on the contribution and evaluation process. (b) Contributors’ inappropriate working patterns and the drawbacks of their collaborating environment might result in duplicate pull requests. (c) Compared to non-duplicate pull requests, duplicate pull requests have significantly different features, e.g., being submitted by inexperienced contributors, being fixing bugs, touching cold files, and solving tracked issues. (d) Integrators choosing between duplicate pull requests prefer to accept those with early submission time, accurate and high-quality implementation, broad coverage, test code, high maturity, deep discussion, and active response. Finally, actionable suggestions and implications are proposed for OSS practitioners.
Yue Yu 0001, Minghui Zhou 0001, Tao Wang 0006, Gang Yin, Long Lan, Huaimin Wang 0001
IEEE Trans. Software Eng.3
2022 Motivation Under Gamification: An Empirical Study of Developers' Motivations and Contributions in Stack Overflow
abstract
To encourage developers' volunteer contributions, modern programming question and answer (Q&A) sites like Stack Overflow (SO) employ gamified incentive mechanisms such as reputation and badges. Understanding developers' motivations in the presence of gamification and the relationship between their motivations and behavioral outcomes is crucial for community building and designing good incentive mechanisms. Grounded on self-determination theory, we conducted a survey with 938 developers who participate in SO to understand their participation motivations and incentive perceptions. By connecting the survey responses with the SO data, we quantitatively analyzed how the developers' motivations and satisfaction of needs relate to their effort and contribution quality. Our main findings are as follows: (1) despite the presence of gamified incentive mechanisms, developers are mainly motivated by intrinsic motivation to participate in SO; (2) developers who have strong motivations to gain gamification rewards are associated with higher intrinsic and integrated motivations, while developers with more development experiences are less motivated by the gamified incentives; (3) both extrinsic motivations (in terms of career prospects) and intrinsic motivations (regarding self-improvement and helping others) can motivate developers to make high-quantity and high-quality contributions; and (4) high-level satisfaction of needs for competency and autonomy has a positive effect on developers making high-quantity and high-quality contributions and addressing difficult problems. Based on these findings, we discuss implications for developer motivation and gamification in the crowdsourcing context and for the mechanism design of gamified crowdsourced platforms.
Yao Lu 0003, Xinjun Mao, Minghui Zhou 0001, Yang Zhang 0026, Zude Li, Tao Wang 0006, Gang Yin, Huaimin Wang 0001
IEEE Trans. Software Eng.3
2021 A large-scale empirical study on Java library migrations: prevalence, trends, and rationales
abstract
With the rise of open-source software and package hosting platforms, reusing 3rd-party libraries has become a common practice. Due to various failures during software evolution, a project may remove a used library and replace it with another library, which we call library migration. Despite substantial research on dependency management, the understanding of how and why library migrations occur is still lacking. Achieving this understanding may help practitioners optimize their library selection criteria, develop automated approaches to monitor dependencies, and provide migration suggestions for their libraries or software projects. In this paper, through a fine-grained commit-level analysis of 19,652 Java GitHub projects, we extract the largest migration dataset to-date (1,194 migration rules, 3,163 migration commits). We show that 8,065 (41.04%) projects having at least one library removal, 1,564 (7.96%, lower-bound) to 5,004 (25.46%, upper-bound) projects have at least one migration, and a median project with migrations has 2 to 4 migrations in total. We discover that library migrations are dominated by several domains (logging, JSON, testing and web service) presenting a long tail distribution. Also, migrations are highly unidirectional in that libraries are either mostly abandoned or mostly chosen in our project corpus. A thematic analysis on related commit messages, issues, and pull requests identifies 14 frequently mentioned migration reasons (e.g., lack of maintenance, usability, integration, etc), 7 of which are not discussed in previous work. Our findings can be operationalized into actionable insights for package hosting platforms, project maintainers, and library developers. We provide a replication package at https://doi.org/10.5281/zenodo.4816752.
Hao He 0012, Runzhi He, Haiqiao Gu, Minghui Zhou 0001
ESEC/SIGSOFT FSE4
2021 A Multi-Metric Ranking Approach for Library Migration Recommendations
abstract
The wide adoption of third-party libraries in software projects is beneficial but also risky. An already-adopted third-party library may be abandoned by its maintainers, may have license incompatibilities, or may no longer align with current project requirements. Under such circumstances, developers need to migrate the library to another library with similar functionalities, but the migration decisions are often opinion-based and sub-optimal with limited information at hand. Therefore, several filtering-based approaches have been proposed to mine library migrations from existing software data to leverage "the wisdom of crowd," but they suffer from either low precision or low recall with different thresholds, which limits their usefulness in supporting migration decisions. In this paper, we present a novel approach that utilizes multiple metrics to rank and therefore recommend library migrations. Given a library to migrate, our approach first generates candidate target libraries from a large corpus of software repositories, and then ranks them by combining the following four metrics to capture different dimensions of evidence from development histories: Rule Support, Message Support, Distance Support, and API Support. We evaluate the performance of our approach with 773 migration rules (190 source libraries) that we borrow from previous work and recover from 21,358 Java GitHub projects. The experiments show that our metrics are effective to help identify real migration targets, and our approach significantly outperforms existing works, with MRR of 0.8566, top-1 precision of 0.7947, top-10 NDCG of 0.7702, and top-20 recall of 0.8939. To demonstrate the generality of our approach, we manually verify the recommendation results of 480 popular libraries not included in prior work, and we confirm 661 new migration rules from 231 of the 480 libraries with comparable performance. The source code, data, and supplementary materials are provided at: https://github.com/hehao98/MigrationHelper.
Hao He 0012, Guangtai Liang, Minghui Zhou 0001
SANER6
2021 Companies' Participation in OSS Development-An Empirical Study of OpenStack
abstract
Commercial participation continues to grow in open source software (OSS) projects and novel arrangements appear to emerge in company-dominated projects and ecosystems. What is the nature of these novel arrangements? Does volunteers’ participation remain critical for these ecosystems? Despite extensive research on commercial participation in OSS, the exact nature and extent of company contributions to OSS development, and the impact of this engagement may have on the volunteer community have not been clarified. To bridge the gap, we perform an exploratory study of OpenStack: a large OSS ecosystem with intense commercial participation. We quantify companies’ contributions via the developers that they provide and the commits made by those developers. We find that companies made far more contributions than volunteers and the distribution of the contributions made by different companies is also highly unbalanced. We observe eight unique contribution models based on companies’ commercial objectives and characterize each model according to three dimensions: contribution intensity, extent, and focus. Companies providing full cloud solutions tend to make both intensive (more than other companies) and extensive (involving a wider variety of projects) contributions. Usage-oriented companies make extensive but less intense contributions. Companies driven by particular business needs focus their contributions on the specific projects addressing these needs. Minor contributors include community players (e.g., the Linux Foundation) and research groups. A model relating the number of volunteers to the diversity of contribution shows a strong positive association between them.
Yuxia Zhang, Minghui Zhou 0001, Audris Mockus, Zhi Jin 0001
IEEE Trans. Software Eng.2
2020 Scaling open source communities: an empirical study of the Linux kernel
abstract
Large-scale open source communities, such as the Linux kernel, have gone through decades of development, substantially growing in scale and complexity. In the traditional workflow, maintainers serve as "gatekeepers" for the subsystems that they maintain. As the number of patches and authors significantly increases, maintainers come under considerable pressure, which may hinder the operation and even the sustainability of the community. A few subsystems have begun to use new workflows to address these issues. However, it is unclear to what extent these new workflows are successful, or how to apply them. Therefore, we conduct an empirical study on the multiple-committer model (MCM) that has provoked extensive discussion in the Linux kernel community. We explore the effect of the model on the i915 subsystem with respect to four dimensions: pressure, latency, complexity, and quality assurance. We find that after this model was adopted, the burden of the i915 maintainers was significantly reduced. Also, the model scales well to allow more committers. After analyzing the online documents and interviewing the maintainers of i915, we propose that overloaded subsystems which have trustworthy candidate committers are suitable for adopting the model. We further suggest that the success of the model is closely related to a series of measures for risk mitigation---sufficient precommit testing, strict review process, and the use of tools to simplify work and reduce errors. We employ a network analysis approach to locate candidate committers for the target subsystems and validate this approach and contextual success factors through email interviews with their maintainers. To the best of our knowledge, this is the first study focusing on how to scale open source communities. We expect that our study will help the rapidly growing Linux kernel and other similar communities to adapt to changes and remain sustainable.
Xin Tan 0003, Minghui Zhou 0001, Brian Fitzgerald 0001
ICSE2
2020 How do companies collaborate in open source ecosystems?: an empirical study of OpenStack
abstract
Open Source Software (OSS) has come to play a critical role in the software industry. Some large ecosystems enjoy the participation of large numbers of companies, each of which has its own focus and goals. Indeed, companies that otherwise compete, may become collaborators within the OSS ecosystem they participate in. Prior research has largely focused on commercial involvement in OSS projects, but there is a scarcity of research focusing on company collaborations within OSS ecosystems. Some of these ecosystems have become critical building blocks for organizations worldwide; hence, a clear understanding of how companies collaborate within large ecosystems is essential. This paper presents the results of an empirical study of the OpenStack ecosystem, in which hundreds of companies collaborate on thousands of project repositories to deliver cloud distributions. Based on a detailed analysis, we identify clusters of collaborations, and identify four strategies that companies adopt to engage with the OpenStack ecosystem. We alsofind that companies may engage in intentional or passive collaborations, or may work in an isolated fashion. Further, wefi nd that a company's position in the collaboration network is positively associated with its productivity in OpenStack. Our study sheds light on how large OSS ecosystems work, and in particular on the patterns of collaboration within one such large ecosystem.
Yuxia Zhang, Minghui Zhou 0001, Klaas-Jan Stol, Zhi Jin 0001
ICSE2
2020 Haste Makes Waste: An Empirical Study of Fast Answers in Stack Overflow
abstract
Modern programming question & answer (Q&A) sites such as Stack Overflow (SO) employ gamified mechanisms to stimulate volunteers' contributions. To maximize the chances of winning gamification rewards such as reputation and badges, a portion of users race to post answers as quickly as possible (i.e., fast answers or FAs), which makes SO the fastest Q&A site; however, this behavior may affect the contribution quality as well. In this paper, we report on a large-scale, mixed-methods empirical study of the gamification-influenced FA phenomenon in SO. We first quantitatively investigate the popularity of the phenomenon and user behaviors regarding FAs. Then, we study the quality of FAs by using regression modeling and qualitatively analyzing 300 instances of FAs. Our main findings reveal that more than 70% and 90% of FAs are not edited by the answerers and other users, respectively, and that later incoming answers have lower chances of being voted on and accepted. Notably, we find that the answer length, code snippets length, and readability of FAs are significantly lower than those of non-fast answers. Although FAs have higher crowd assessment scores, they have no relationship with acceptance from the perspective of asker assessment, and a considerable portion of FAs solve the problem by interacting with the asker in the comments. These results help us better understand the effects of reward-based gamification on crowdsourced software engineering communitites and provide implications for designers of gamified systems.
Yao Lu 0003, Xinjun Mao, Minghui Zhou 0001, Yang Zhang 0026, Tao Wang 0006, Zude Li
ICSME3
2020 A first look at good first issues on GitHub
abstract
Keeping a good influx of newcomers is critical for open source software projects' survival, while newcomers face many barriers to contributing to a project for the first time. To support newcomers onboarding, GitHub encourages projects to apply labels such as good first issue (GFI) to tag issues suitable for newcomers. However, many newcomers still fail to contribute even after many attempts, which not only reduces the enthusiasm of newcomers to contribute but makes the efforts of project members in vain. To better support the onboarding of newcomers, this paper reports a preliminary study on this mechanism from its application status, effect, problems, and best practices. By analyzing 9,368 GFIs from 816 popular GitHub projects and conducting email surveys with newcomers and project members, we obtain the following results. We find that more and more projects are applying this mechanism in the past decade, especially the popular projects. Compared to common issues, GFIs usually need more days to be solved. While some newcomers really join the projects through GFIs, almost half of GFIs are not solved by newcomers. We also discover a series of problems covering mechanism (e.g., inappropriate GFIs), project (e.g., insufficient GFIs) and newcomer (e.g., uneven skills) that makes this mechanism ineffective. We discover the practices that may address the problems, including identifying GFIs that have informative description and available support, and require limited scope and skill, etc. Newcomer onboarding is an important but challenging question in open source projects and our work enables a better understanding of GFI mechanism and its problems, as well as highlights ways in improving them.
Xin Tan 0003, Minghui Zhou 0001, Zeyu Sun 0004
ESEC/SIGSOFT FSE2
2020 Pandemic programming
abstract
Abstract Context As a novel coronavirus swept the world in early 2020, thousands of software developers began working from home. Many did so on short notice, under difficult and stressful conditions. Objective This study investigates the effects of the pandemic on developers’ wellbeing and productivity. Method A questionnaire survey was created mainly from existing, validated scales and translated into 12 languages. The data was analyzed using non-parametric inferential statistics and structural equation modeling. Results The questionnaire received 2225 usable responses from 53 countries. Factor analysis supported the validity of the scales and the structural model achieved a good fit (CFI = 0.961, RMSEA = 0.051, SRMR = 0.067). Confirmatory results include: (1) the pandemic has had a negative effect on developers’ wellbeing and productivity; (2) productivity and wellbeing are closely related; (3) disaster preparedness, fear related to the pandemic and home office ergonomics all affect wellbeing or productivity. Exploratory analysis suggests that: (1) women, parents and people with disabilities may be disproportionately affected; (2) different people need different kinds of support. Conclusions To improve employee productivity, software companies should focus on maximizing employee wellbeing and improving the ergonomics of employees’ home offices. Women, parents and disabled persons may require extra support.
Paul Ralph, Sebastian Baltes, Gianisa Adisaputri, Richard Torkar, Vladimir Kovalenko, Marcos Kalinowski, Nicole Novielli, Shin Yoo, Xavier Devroey, Xin Tan 0003, Minghui Zhou 0001, Burak Turhan, Rashina Hoda, Hideaki Hata, Gregorio Robles, Amin Milani Fard, Rana Alkadhi
Empir. Softw. Eng.11
2019 How to Communicate when Submitting Patches: An Empirical Study of the Linux Kernel
abstract
Communication when submitting patches (CSP) is critical in software development, because ineffective communication wastes the time of developers and reviewers, and is even harmful to future product release and maintenance. In distributed software development, CSP is usually in the form of computer-mediated communication (CMC), which is a crucial topic concerned by the CSCW community. However, what to say and how to say in communication including CSP has been rarely studied. To bridge this knowledge gap and provide relevant guidance for developers to conduct CSP, in this study, we conducted an empirical study on the Linux kernel. We found four themes involving 17 expression elements that characterize what to express when submitting patches and three themes of contextual factors that determine how these elements are applied. Considering both expression elements and context, combined with an online survey, we obtained 17 practices for communication. Among them, four practices, such as "provide sufficient reasons" and "provide a trade-off" are the most important but difficult practices, for which we provide specific instructions. We also found that the "individual factors" plays a special role in communication, which may lead to potential problems even in accepted patches. Based on these findings, we discuss the recommendations for different practitioners, including patch submitters, reviewers, and tool designers, and the implications for open source software (OSS) communities and the CSCW researchers.
Xin Tan 0003, Minghui Zhou 0001
Proc. ACM Hum. Comput. Interact.2
2018 A neural framework for retrieval and summarization of source code
abstract
Code retrieval and summarization are two tasks often employed by software developers to reuse code that spreads over online repositories. In this paper, we present a neural framework that allows bidirectional mapping between source code and natural language to improve these two tasks. Our framework, BVAE, is designed to have two Variational AutoEncoders (VAEs) to model bimodal data: C-VAE for source code and L-VAE for natural language. Both VAEs are trained jointly to reconstruct their input as much as possible with regularization that captures the closeness between the latent variables of code and description. BVAE could learn semantic vector representations for both code and description and generate completely new descriptions for arbitrary code snippets. We design two instance models of BVAE for retrieval and summarization tasks respectively and evaluate their performance on a benchmark which involves two programming languages: C# and SQL. Experiments demonstrate BVAE’s potential on the two tasks.
Qingying Chen, Minghui Zhou 0001
ASE2
2018 A multi-level dataset of linux kernel patchwork
abstract
In many open source software projects (e.g., the Linux kernel), people contribute by sending code patches to the community. The community evaluates these contributions and decides whether to integrate the changes. To improve the efficiency of code contributions, substantial effort has been devoted to analyzing how patches are submitted and processed. Patch data are critical for this type of analysis, while retrieving and cleaning the data is a non-trivial job. To facilitate these studies, we share a multi-level dataset of a Linux kernel patchwork covering a nine-year history of patches and related discussion recorded by the Linux kernel mailing list (LKML). The data and scripts are provided at: https://zenodo.org/record/1165576
Minghui Zhou 0001
MSR2
2018 Be careful of when: an empirical study on time-related misuse of issue tracking data
abstract
Issue tracking data have been used extensively to aid in predicting or recommending software development practices. Issue attributes typically change over time, but users may use data from a separate time of data collection rather than the time of their application scenarios. We, therefore, investigate data leakage, which results from ignoring the chronological order in which the data were produced. Information leaked from the "future" makes prediction models misleadingly optimistic. We examine existing literature to confirm the existence of data leakage and reproduce three typical studies (detecting duplicate issues, localizing issues, and predicting issue-fix time) adjusted for appropriate data to quantify the impact of the data leakage. We confirm that 11 out of 58 studies have leakage problem, while 44 are suspected. We observe biased results caused by data leakage while the extent is not striking. Attributes of summary, component, and assignee have the largest impact on the results. Our findings suggest that data users are often unaware of the context of the data being produced. We recommend researchers and practitioners who attempt to utilize issue tracking data to address software development problems to have a full understanding of their application scenarios, the origin and change of the data, and the influential issue attributes to manage data leakage risks.
Feifei Tu, Qimu Zheng, Minghui Zhou 0001
ESEC/SIGSOFT FSE4
2017 Understanding the Variation of Software Development Tasks: a Qualitative Study
abstract
In order to reduce cost, get to market faster and utilize global talents, large companies often organize their software development globally (distributed over Internet), a paradigm advocated by Internetware. Considering the complexity of distributed software development, it is important to understand the variation of various tasks, so the projects could work more efficiently on plan formulation, personnel organization, or task allocation. The main goal of this paper is to understand the variations of software development tasks. We conduct an interview with 47 interviewees and a survey with 148 people from 15 projects with different size and domain.Through the analysis of interviews and surveys, we find that a software task could be characterized through three aspects: value, difficulty and centrality. Among them, task value is influenced by the role of stakeholders and project context, and is related to the task difficulty and task centrality. Task difficulty is reflected by technology, domain difference, working relationships, customer related issues, and it is also related to the developers' personalities. Task centrality can be described by customer impact, system-wide impact, team impact and future impact. We believe our results can help project managers to optimize task allocation, or adjust project plan, and thus achieve efficient development.
Xin Tan 0003, Hanmin Qin, Minghui Zhou 0001
Internetware3
2017 On the scalability of Linux kernel maintainers' work
abstract
Open source software ecosystems evolve ways to balance the workload among groups of participants ranging from core groups to peripheral groups. As ecosystems grow, it is not clear whether the mechanisms that previously made them work will continue to be relevant or whether new mechanisms will need to evolve. The impact of failure for critical ecosystems such as Linux is enormous, yet the understanding of why they function and are effective is limited. We, therefore, aim to understand how the Linux kernel sustains its growth, how to characterize the workload of maintainers, and whether or not the existing mechanisms are scalable. We quantify maintainers' work through the files that are maintained, and the change activity and the numbers of contributors in those files. We find systematic differences among modules; these differences are stable over time, which suggests that certain architectural features, commercial interests, or module-specific practices lead to distinct sustainable equilibria. We find that most of the modules have not grown appreciably over the last decade; most growth has been absorbed by a few modules. We also find that the effort per maintainer does not increase, even though the community has hypothesized that required effort might increase. However, the distribution of work among maintainers is highly unbalanced, suggesting that a few maintainers may experience increasing workload. We find that the practice of assigning multiple maintainers to a file yields only a power of 1/2 increase in productivity. We expect that our proposed framework to quantify maintainer practices will help clarify the factors that allow rapidly growing ecosystems to be sustainable.
Minghui Zhou 0001, Qingying Chen, Audris Mockus, Fengguang Wu
ESEC/SIGSOFT FSE1
2016 Multi-extract and multi-level dataset of mozilla issue tracking history
abstract
any studies analy e issue trac ing repositories to understand and support software development To facilitate the analyses we share a o illa issue trac ing dataset covering a -year history The dataset includes three e tracts and multiple levels for each e tract The three e tracts were retrieved through two channels a front-end we user interface I and a ac -end o cial data ase dump of o illa ug illa at three di erent times The variations dynamics among e tracts provide space for researchers to reproduce and validate their studies while revealing potential opportunities for studies that otherwise could not e conducted e provide di erent data levels for each e tract ranging from raw data to standardi ed data as well as to the calculated data level for targeting speci c research uestions Data retrieving and processing scripts related to each data level are o ered too y employing the multi-level structure analysts can more e ciently start an in uiry from the standardi ed level and easily trace the data chain when necessary e g to verify if a phenomenon re ected y the data is an actual event e applied this dataset to several pu lished studies and intend to e pand the multi-level and multi-e tract feature to other software engineering datasets.
Minghui Zhou 0001, Hong Mei 0001
MSR2
2016 Effectiveness of code contribution: from patch-based to pull-request-based tools
abstract
Code contributions in Free/Libre and Open Source Software projects are controlled to maintain high-quality of software. Alternatives to patch-based code contribution tools such as mailing lists and issue trackers have been developed with the pull request systems being the most visible and widely available on GitHub. Is the code contribution process more effective with pull request systems? To answer that, we quantify the effectiveness via the rates contributions are accepted and ignored, via the time until the first response and final resolution and via the numbers of contributions. To control for the latent variables, our study includes a project that migrated from an issue tracker to the GitHub pull request system and a comparison between projects using mailing lists and pull request systems. Our results show pull request systems to be associated with reduced review times and larger numbers of contributions. However, not all the comparisons indicate substantially better accept or ignore rates in pull request systems. These variations may be most simply explained by the differences in contribution practices the projects employ and may be less affected by the type of tool. Our results clarify the importance of understanding the role of tools in effective management of the broad network of potential contributors and may lead to strategies and practices making the code contribution more satisfying and efficient from both contributors' and maintainers' perspectives.
Minghui Zhou 0001, Audris Mockus
SIGSOFT FSE2
2016 Inflow and Retention in OSS Communities with Commercial Involvement: A Case Study of Three Hybrid Projects
abstract
Motivation:Open-source projects are often supported by companies, but such involvement often affects the robust contributor inflow needed to sustain the project and sometimes prompts key contributors to leave. To capture user innovation and to maintain quality of software and productivity of teams, these projects need to attract and retain contributors.Aim:We want to understand and quantify how inflow and retention are shaped by policies and actions of companies in three application server projects.Method:We identified three hybrid projects implementing the same JavaEE specification and used published literature, online materials, and interviews to quantify actions and policies companies used to get involved. We collected project repository data, analyzed affiliation history of project participants, and used generalized linear models and survival analysis to measure contributor inflow and retention.Results:We identified coherent groups of policies and actions undertaken by sponsoring companies as three models of community involvement and quantified tradeoffs between the inflow and retention each model provides. We found that full control mechanisms and high intensity of commercial involvement were associated with a decrease of external inflow and with improved retention. However, a shared control mechanism was associated with increased external inflow contemporaneously with the increase of commercial involvement.Implications:Inspired by a natural experiment, our methods enabled us to quantify aspects of the balance between community and private interests in open- source software projects and provide clear implications for the structure of future open-source communities.
Minghui Zhou 0001, Audris Mockus, Lu Zhang 0023, Hong Mei 0001
ACM Trans. Softw. Eng. Methodol.1
2015 A method to identify and correct problematic software activity data: exploiting capacity constraints and data redundancies
abstract
Mining software repositories to understand and improve software development is a common approach in research and practice. The operational data obtained from these repositories often do not faithfully represent the intended aspects of software development and, therefore, may jeopardize the conclusions derived from it. We propose an approach to identify problematic values based on the constraints of software development and to correct such values using data redundancies. We investigate the approach using issue and commit data of Mozilla project. In particular, we identified problematic data in four types of events and found the fraction of problematic values to exceed 10% and rapidly rising. We found the corrected values to be 50% closer to the most accurate estimate of task completion time. Finally, we found that the models of time until fix changed substantially when data were corrected, with the corrected data providing a 20% better fit. We discuss how the approach may be generalized to other types of operational data to increase fidelity of software measurement in practice and in research.
Qimu Zheng, Audris Mockus, Minghui Zhou 0001
ESEC/SIGSOFT FSE3
2015 Who Will Stay in the FLOSS Community? Modeling Participant's Initial Behavior
abstract
Motivation: To survive and succeed, FLOSS projects need contributors able to accomplish critical project tasks. However, such tasks require extensive project experience of long term contributors (LTCs). Aim: We measure, understand, and predict how the newcomers' involvement and environment in the issue tracking system (ITS) affect their odds of becoming an LTC. Method: ITS data of Mozilla and Gnome, literature, interviews, and online documents were used to design measures of involvement and environment. A logistic regression model was used to explain and predict contributor's odds of becoming an LTC. We also reproduced the results on new data provided by Mozilla. Results: We constructed nine measures of involvement and environment based on events recorded in an ITS. Macro-climate is the overall project environment while micro-climate is person-specific and varies among the participants. Newcomers who are able to get at least one issue reported in the first month to be fixed, doubled their odds of becoming an LTC. The macro-climate with high project popularity and the micro-climate with low attention from peers reduced the odds. The precision of LTC prediction was 38 times higher than for a random predictor. We were able to reproduce the results with new Mozilla data without losing the significance or predictive power of the previously published model. We encountered unexpected changes in some attributes and suggest ways to make analysis of ITS data more reproducible. Conclusions: The findings suggest the importance of initial behaviors and experiences of new participants and outline empirically-based approaches to help the communities with the recruitment of contributors for long-term participation and to help the participants contribute more effectively. To facilitate the reproduction of the study and of the proposed measures in other contexts, we provide the data we retrieved and the scripts we wrote at https://www.passion-lab.org/projects/developerfluency.html.
Minghui Zhou 0001, Audris Mockus
IEEE Trans. Software Eng.1
2014 Patterns of folder use and project popularity: a case study of github repositories
abstract
Context: Every software development project uses folders to organize software artifacts. Goal: We would like to understand how folders are used and what ramifications different uses may have. Method: In this paper we study the frequency of folders used by 140k Github projects and use regression analysis to model how folder use is related to project popularity, i.e., the extent of forking. Results: We find that the standard folders, such as document, testing, and examples, are not only among the most frequently used, but their presence in a project is associated with increased chances that a project's code will be forked (i.e., used by others) and an increased number of forks. Conclusions: This preliminary study of folder use suggests opportunities to quantify (and improve) file organization practices based on folder use patterns of large collections of repositories.
Minghui Zhou 0001, Audris Mockus
ESEM2
2014 Mining micro-practices from operational data
abstract
Micro-practices are actual (and usually undocumented or incorrectly documented) activity patterns used by individuals or projects to accomplish basic software development tasks, such as writing code, testing, triaging bugs, or mentoring newcomers. The operational data in software repositories presents the tantalizing possibility to discover such fine-scale behaviors and use them to understand and improve software development. We propose a large-scale evidence-based approach to accomplish this by first creating a mirror of the projects in the open source universe. The next step would involve the inductive generalization from in-depth studies of specific projects from one side and the categorization of micro-practices in the entire universe from the other side.
Minghui Zhou 0001, Audris Mockus
SIGSOFT FSE1
2013 Impact of Triage: A Study of Mozilla and Gnome
abstract
Triage is of great interest in software projects because it has the potential to reduce developer effort by involving a broader base of non-developer contributors to filter and augment reported issues. Using issue tracking data and interviews with experienced contributors we investigate ways to quantify the impact of triagers on reducing the number of issues developers need to resolve in two OSS projects: Mozilla and Gnome. We find the primary impact of triagers to involve issue filtering, filling missing information, and determining the relevant product. While triagers were good at filtering invalid issues and as accurate as developers in filling in missing issue attributes, they had more difficulty accurately pinpointing the relevant product. We expect that this work will highlight the importance of issue triage in software projects and will help design further studies on understanding and improving triage practices.
Minghui Zhou 0001, Audris Mockus
ESEM2
2013 How commercial involvement affects open source projects: three case studies on issue reporting
Minghui Zhou 0001, Dirk Riehle
Sci. China Inf. Sci.2
2012 Towards an Adaptive Service Degradation Approach for Handling Server Overload
abstract
While people get used to surfing web, managing the overload of web applications has become a critical problem for application providers. Targeting overload issue of complex, dynamic web applications, this paper presents an adaptive service degradation approach. Our approach attempts to automatically locate the bottleneck inside the applications and generate proper degradation plans in real time. This is accomplished through internally monitoring the performance and resource utilization state of the application, which is decomposed into a set of services. By dynamic controlling the bottleneck of resource utilization, the application can keep providing key services even when overload occurs, for example, by degrading only the service which consumes critical resources and has a low priority from the perspective of business logic. We implement a prototype and conduct a case study with a business application. The case study demonstrates the approach is effective when overload occurs.
Ziyou Wang, Minghui Zhou 0001, Hong Mei 0001
APSEC2
2012 Towards Online Localization and Recovery for Faulty Components in Component-Based Applications
abstract
Off-The-Shelf (COTS) software components have been extensively used by applications over the world. However, COTS components always carry issues that might and sometimes only take place at runtime, in particular, when combined with other components. This might bring serious trouble for the application's dependability. Based on the industry experiences that lots of faulty components usually occupy excessive resources, this paper proposes an approach to localize the faulty components in an application automatically through analyzing the resource usage of components. Once a faulty component interferes with the application's performance, our approach generates an anomaly report, localizes the faulty component, and enables component level recovery to remove the negative impact. We have implemented a prototype and demonstrated its effectiveness on the well-known JPetStore benchmark with performance overhead of less than 3%.
Chao You, Minghui Zhou 0001, Hongwu Lin, Zan Xiao, Hong Mei 0001
COMPSAC2
2012 What make long term contributors: Willingness and opportunity in OSS community
abstract
To survive and succeed, software projects need to attract and retain contributors. We model the individual's chances to become a valuable contributor through her capacity, willingness, and the opportunity to contribute at the time of joining. Using issue tracking data of Mozilla and Gnome, we find that the probability for a new joiner to become a Long Term Contributor (LTC) is associated with her willingness and environment. Specifically, during their first month, future LTCs tend to be more active and show more community-oriented attitude than other joiners. Joiners who start by commenting on instead of reporting an issue or ones who succeed to get at least one reported issue to be fixed, more than double their odds of becoming an LTC. The micro-climate with a productive and clustered peer group increases the odds. On the contrary, the macro-climate with high project popularity and the micro-climate with low attention from peers reduce the odds. This implies that the interaction between individual's attitude and project's climate are associated with the odds that an individual would become a valuable contributor or disengage from the project. Our findings may provide a basis for empirical approaches to design a better community architecture and to improve the experience of contributors.
Minghui Zhou 0001, Audris Mockus
ICSE1
2012 Review code evolution history in OSS universe
abstract
Software evolves all the time because of the changing requirements, in particular, in the diverse Internet environment. Evolution history recorded in software repositories, e.g., Version Control Systems, reflects people's software development practice. Exploring this history could help practitioners to reuse the best practices therefore improve productivity and software quality. Because of the difficulty of collecting and standardizing data, most existing work could only utilize small project set. In this study, we target the open source software universe to build a universal code evolution model for large-scale data. We consider code evolution from two aspects: code version changing history in a single project and code reuse history in the whole universe. In the model, files/modules are built as nodes, and relations (version change or reuse) between files/modules are built as connections. Based on the model, we design and implement a code evolution review framework, i.e., Code Evolution Reviewer (CER), which provides a series of data interfaces to review code evolution history, in particular, code version changing in single project and code reuse among projects. Further, CER could be utilized to explore best practices across large-scale project set.
Hongwu Lin, Minghui Zhou 0001, Hong Mei 0001
Internetware3
2012 Towards a degradation-based mechanism for adaptive overload control
Ziyou Wang, Minghui Zhou 0001, Hong Mei 0001
Sci. China Inf. Sci.2
2011 Does the initial environment impact the future of developers
abstract
Software developers need to develop technical and social skills to be successful in large projects. We model the relative sociality of developer as a ratio between the size of her communication network and the number of tasks she participates in. We obtain both measures from the problem tracking systems. We use her workflow peer network to represent her social learning, and the issues she has worked on to represent her technical learning. Using three open source and three traditional projects we investigate how the project environment reflected by the sociality measure at the time a developer joins, affects her future participation. We find: a) the probability that a new developer will become one of long-term and productive developers is highest when the project sociality is low; b) times of high sociality are associated with a higher intensity of new contributors joining the project; c) there are significant differences between the social learning trajectories of the developers who join in low and in high sociality environments; d) the open source and commercial projects exhibit different nature in the relationship between developer's tenure and the project's environment at the time she joins. These findings point out the importance of the initial environment in determining the future of the developers and may lead to better training and learning strategies in software organizations.
Minghui Zhou 0001, Audris Mockus
ICSE1
2010 A case study of internetware development
abstract
The open, dynamic and ever-changing characteristics of Internet have attracted much attention to carry out research on Internetware. Current researches mainly focus on the framework of the Internetware. However, there are a variety of issues facing the Internetware development today with more joint work distributed over the world, and how should we improve the efficiency of such development? In order to resolve this issue we investigate three open source projects from J2EE platform domain: JBossAS, JOnAS, and Apache Geronimo to find out that, in the sampled projects, how many people will involve the Internetware development, how they allocate the work, and how the speed to resolve the issues reported by the customer. By answering five research questions referred from the Apache study, we proposed four hypotheses: (1) Open source Interware development will have a core of developers who will create approximately 80% or more of the new functionality. The group will be no larger than 30 people; (2) In a specific server-side domain, the group who will repair defects and report issues will have the equal or even smaller number people compared to the core group; (3) Commercial support can attract more volunteers to the open source Internetware projects; (4) Open Source Internetware developments exhibit very rapid responses to customer issues.
Minghui Zhou 0001, Hong Mei 0001
Internetware2
2010 Developer fluency: achieving true mastery in software projects
abstract
Outsourcing and offshoring lead to a rapid influx of new developers in software projects. That, in turn, manifests in lower productivity and project delays. To address this common problem we study how the developers become fluent in software projects. We found that developer productivity in terms of number of tasks per month increases with project tenure and plateaus within a few months in three small and medium projects and it takes up to 12 months in a large project. When adjusted for the task difficulty, developer productivity did not plateau but continued to increase over the entire three year measurement interval. We also discovered that tasks vary according to their importance(centrality) to a project. The increase in task centrality along four dimensions: customer, system-wide, team, and future impact was approximately linear over the entire period. By studying developer fluency we contribute by determining dimensions along which developer expertise is acquired, finding ways to measure them, and quantifying the trajectories of developer learning.
Minghui Zhou 0001, Audris Mockus
SIGSOFT FSE1
2010 Self-Adaptive Resource Management for Large-Scale Shared Clusters
Yan Li 0067, Feng-Hong Chen, Minghui Zhou 0001, Wenpin Jiao, Donggang Cao, Hong Mei 0001
J. Comput. Sci. Technol.4
2009 Dual-Container: Extending the EJB2.x Container to Support EJB3.0
abstract
In order to support EJB3.0, many open-source application server vendors have implemented a new container which is independent with the existing EJB2.x one. However, the functional redundancy between both of the containers increases the maintenance cost when the software are modified because of problems or new requirements. This paper proposes a Dual-container to support EJB2.x and EJB3.0 in a unified architecture. Software engineering perspective is applied, three principles including greatest-common-divisor, lease-common-multiple, and transparent intrusion, are presented to guide the construction of the architecture. We implemented the Dual-container in an open-source J2EE application server. The evaluation results show that the Dual-container is effective in reducing the functional redundancy and increasing the maintainability. We demonstrate that in a comparative analysis.
Ziyou Wang, Minghui Zhou 0001, Donggang Cao
COMPSAC (2)2
2009 Towards a dynamic and adaptable application server
abstract
Internetware is proposed as a new software paradigm to cope with the open, dynamic and ever-changing Internet environment for applications. The incarnated characteristics of Internetware promote the operating platform to be more dynamic and adaptable. As an operating platform, PKUAS (Peking University Application Server) has been successfully applied in various fields. However, its inexplicit module boundary, insufficient lifecycle management and absent dependency management, make it hardly meet challenges. In this paper we refactor PKUAS into PKUAS II by introducing OSGi (Open Services Gateway Initiative) to achieve a better dynamic capability. First, PKUAS II adopts service component oriented model as its structure so that the boundary among modules can be explicitly explained, and the continuous lifecycle management and dynamic dependency management can get supported. PKUAS II is able to upgrade without interruption of service and extend on the fly with new services. Second, PKUAS II supports application--aware customization, which can dynamically generate a just enough application server for the application at runtime, to satisfy different applications' requirements and reduce resource costs. Last but not least, some evaluations have been done, which show that PKUAS II is more flexible and dynamic without significant performance overhead, and might support Internetware better.
Chao You, Minghui Zhou 0001, Zan Xiao, Hong Mei 0001
Internetware2
2008 Constructing Flexible Application Servers with Off-the-Shelf Middleware Services Integration Framework
Yan Li 0067, Minghui Zhou 0001, Donggang Cao, Lu Zhang 0023
ICSR2
2007 Supporting crosscutting concern modelling in software architecture design
Donggang Cao, Hong Mei 0001, Minghui Zhou 0001
Frontiers Comput. Sci. China3
2006 A Service-Oriented Trust Management Model on Application Server
abstract
In the service-oriented architecture, the components deployed on application servers are published as Web services. Though many researches focus on how to authorize at the Web service level currently, there is little work involving the authorization gap between the service and its component implementation. This paper tries to bridge the gap by proposing a service-oriented trust management model, which expands the application server's capability to deal with more complex trust relationship between service users and services, and supplies a flexible trust management mechanism to integrate authentication and authorization together. Moreover, the model provides a finer granularity access control, sustains delegation between users, and has a certain extent reasoning capability. The model has been implemented in a J2EE application server, and the experiment has demonstrated that the model has high flexibility and scalability
Minghui Zhou 0001, Hong Mei 0001
ICWS1
2005 Customizable Framework for Managing Trusted Components Deployed on Middleware
abstract
Due to the widespread trust threat under the open and dynamic Internet environment, the computer community has endeavored to engage in the studies of technologies for protecting and evaluating trustworthiness. This paper firstly defines trust from three aspects: trust relationship, trust property and trust entity, and build a uniform view for multiple trust properties. Secondly, according to abstract the common characteristics over varied trust properties, a model of trust management is elaborated, in which the trust entities, measurement model and trust policy are described. Thirdly, the trust management is implemented as a kind of public service on a J2EE-compliant middleware platform, i.e., the PKUAS.
Minghui Zhou 0001, Wenpin Jiao, Hong Mei 0001
ICECCS1