Shurui Zhou

dblp:166/5038 · DBLP profile ↗
← Back
31ranked-venue papers
4as first author
21since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 20 · 4 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 10 · 10 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Untangling the Timeline: Challenges and Opportunities in Supporting Version Control in Modern Computer-Aided Design
abstract
Version control is critical in mechanical computer-aided design (CAD) to enable traceability, manage product variation, and support collaboration. Yet, its implementation in modern CAD software as an essential information infrastructure for product development remains plagued by issues due to the complexity and interdependence of design data. This paper presents a systematic review of user-reported challenges with version control in modern CAD tools. Analyzing 170 online forum threads, we identify recurring socio-technical issues that span the management, continuity, scope, and distribution of versions. Our findings inform a broader reflection on how version control should be designed and improved for CAD and motivate opportunities for tools and mechanisms that better support articulation work, facilitate cross-boundary collaboration, and operate with infrastructural reflexivity. This study offers actionable insights for CAD software providers and highlights opportunities for researchers to rethink version control.
Yuanzhe Deng, Shutong Zhang, Kathy Cheng, Alison Olechowski, Shurui Zhou
CHI5
2026 CADModelScope: Revealing the Dependency Structure Behind Parametric Computer-Aided Design Models
abstract
Parametric computer-aided design (CAD) models are constructed by a sequence of operations, where each operation may reference geometries created by earlier operations. This network of dependencies enables efficient modelling of complex geometry but also results in fragile models, where small modifications can trigger cascading errors. These interdependencies are obscured in commercial CAD systems, leaving users to rely on trial and error when navigating, modularizing, and debugging unfamiliar and complex models. In this paper, we motivate, present, and pilot CADModelScope, a multi-level graph-based visualization of operation dependencies integrated into a commercial CAD platform. In a qualitative lab study, we observed how participants locate and interpret operations, and how CADModelScope enhances awareness of hidden interdependencies and supports more structured navigation. Our findings highlight the potential of using the network of operation dependency as an effective representation for understanding and interacting with parametric CAD models, and we discuss implications for future tool design.
Yuanzhe Deng, Zhijing Zhang, Shurui Zhou, Alison Olechowski
CHI3
2025 The Product Beyond the Model - An Empirical Study of Repositories of Open-Source ML Products
abstract
Machine learning (ML) components are increasingly incorporated into software products for end-users, but developers face challenges in transitioning from ML prototypes to products. Academics have limited access to the source of commercial ML products, hindering research progress to address these challenges. In this study, first and foremost, we contribute a dataset of 262 open-source ML products for end users (not just models), identified among more than half a million ML-related projects on GitHub. Then, we qualitatively and quantitatively analyze 30 open-source ML products to answer six broad research questions about development practices and system architecture. We find that the majority of the ML products in our sample represent more startup-style development than reported in past interview studies. We report 21 findings, including limited involvement of data scientists in many open-source ML products, unusually low modularity between ML and non-ML code, diverse architectural choices on incorporating models into products, and limited prevalence of industry best practices such as model testing, pipeline automation, and monitoring. Additionally, we discuss seven implications of this study on research, development, and education, including the need for tools to assist teams without data scientists, education opportunities, and open-source-specific research for privacy-preserving telemetry.
Nadia Nahar, Grace A. Lewis, Shurui Zhou, Christian Kästner
ICSE4
2025 MAML: Towards a Faster Web in Developing Regions
abstract
The web experience in developing regions remains subpar, primarily due to the growing complexity of modern webpages and insufficient optimization by content providers. Users in these regions typically rely on low-end devices and limited bandwidth, which results in a poor user experience as they download and parse webpages bloated with excessive third-party CSS and JavaScript (JS). To address these challenges, we introduce the Mobile Application Markup Language (MAML), a flat layout-based web specification language that reduces computational and data transmission demands, while replacing the excessive bloat from JS with a new scripting language centered on essential (and popular) web functionalities. Last but not least, MAML is backward compatible as it can be transpiled to minimal HTML/JavaScript/CSS and thus work with legacy browsers. We benchmark MAML in terms of page load times and sizes, using a translator which can automatically port any webpage to MAML. When compared to the popular Google AMP, across 100 testing webpages, MAML offers webpage speedups by tens of seconds under challenging network conditions thanks to its significant size reductions. Next, we run a competition involving 25 university students porting 50 of the above webpages to MAML using a web-based editor we developed. This experiment verifies that, with little developer effort, MAML is quite effective in maintaining the visual and functional correctness of the originating webpages.
Ayush Pandey 0002, Matteo Varvello, Syed Ishtiaque Ahmed, Shurui Zhou, Lakshminarayanan Subramanian, Yasir Zaki
WWW4
2025 Collaboration Challenges and Opportunities in Developing Scientific Open-Source Software Ecosystem: A Case Study on Astropy
abstract
Scientific open-source software (OSS) has greatly benefited research communities through its transparent and collaborative nature. Given its critical role in scientific research, ensuring the efficiency of collaboration within development teams of scientific OSS has become vital. Earlier research has identified both the challenges and opportunities associated with interdisciplinary team collaboration in developing conventional scientific software, as well as the dynamics of distributed teams in the context of OSS development. However, it remains unclear whether these challenges are still present and if their solutions can be seamlessly adapted to the context of scientific OSS and its broader ecosystem. Therefore, this study examines the challenges and opportunities for improving the collaboration efficiency in the development and maintenance of scientific OSS, focusing on interdisciplinary and multi-project collaboration within the open-source environment. We conducted a mixed-methods case study on the Astropy project, a widely-used software ecosystem in astronomy, including (1) a detailed analysis of the commit history to understand the roles and activities of each contributor; (2) an in-depth investigation of cross-referenced issues and pull requests to identify challenges and best practices for cross-project collaboration at the ecosystem level; and (3) an interview study with core contributors to complement the first two steps, examining their collaborative efforts within an interdisciplinary team and across a multi-project ecosystem. We contribute to the CSCW community by deepening the understanding of collaboration in scientific OSS ecosystems, highlighting practices and challenges both at the individual project level and across a broader ecosystem. Our findings offer insights into managing cross-project interdependencies and propose strategies to address key obstacles in scientific OSS development.
Jiayi Sun 0001, Aarya Patil, Youhai Li, Jin L. C. Guo, Shurui Zhou
Proc. ACM Hum. Comput. Interact.5
2025 It's a Complete Haystack: Understanding Dependency Management Needs in Computer-Aided Design
abstract
In today's landscape, hardware development teams face increasing demands for better quality products, greater innovation, and shorter manufacturing lead times. Despite the need for more efficient and effective processes, hardware designers continue to struggle with a lack of awareness of design changes and other collaborators' actions, a persistent issue in decades of CSCW research. One significant and unaddressed challenge is understanding and managing dependencies between 3D CAD (computer-aided design) models, especially when products can contain thousands of interconnected components. In this two-phase formative study, we explore designers' pain points of CAD dependency management through a thematic analysis of 100 online forum discussions and semi-structured interviews with 10 designers. We identify nine key challenges related to the traceability, navigation, and consistency of CAD dependencies, that harm the effective coordination of hardware development teams. To address these challenges, we propose design goals and necessary features to enhance hardware designers' awareness and management of dependencies, ultimately with the goal of improving collaborative workflows.
Kathy Cheng, Alison Olechowski, Shurui Zhou
Proc. ACM Hum. Comput. Interact.3
2025 Who is to Blame: A Comprehensive Review of Challenges and Opportunities in Designer-Developer Collaboration
abstract
Software development relies on effective collaboration between Software Development Engineers (SDEs) and User eXperience Designers (UXDs) to create software products of high quality and usability. While this collaboration issue has been explored over the past decades, anecdotal evidence continues to indicate the existence of challenges in their collaborative efforts. To understand this gap, we first conducted a systematic literature review (SLR) of 45 papers published since 2004, uncovering three key collaboration challenges and two main categories of potential best practices. We then analyzed designer and developer forums and discussions from one open-source software repository to assess how the challenges and practices manifest in the status quo. Our findings have broad applicability for collaboration in software development, extending beyond the partnership between SDEs and UXDs. The suggested best practices and interventions also act as a reference for future research, assisting in the development of dedicated collaboration tools for SDEs and UXDs.
Shutong Zhang, Jinghui Cheng 0001, Shurui Zhou
Proc. ACM Hum. Comput. Interact.4
2024 Can We Do Better with What We Have Done? Unveiling the Potential of ML Pipeline in Notebooks
abstract
Computational notebooks are widely adopted by data scientists for experimenting with machine learning (ML) models. Despite the support for exploratory programming enabled by notebooks, they fall short in the ability to manage alternatives across different stages of the ML pipeline. In this study, we conduct a qualitative analysis to examine how data scientists explore various alternatives through a series of versions of notebooks on Kaggle. The findings indicate that data scientists investigate multiple alternatives at each stage across multiple versions, yet only a limited number of combinations from different stages are explored. Next, by combining alternatives from all stages to form previously unexplored paths, we discover that certain untested combinations of alternatives can outperform the best models as identified in the original notebooks. Moreover, by substituting the hyperparameter optimization and model configuration stages with AutoML methods, we observe that only a select number of ML pipelines experience improvement via AutoML, which implies the limitation of the current AutoML techniques. In summary, our study provides insights into the systematic and effective exploration of overlooked ML pipeline configuration combinations that yield superior results. The findings shed light on future research directions such as the development of tooling support of alternative management while striking a balance between manual exploration and automated optimization.
Yuangan Zou, Xinpeng Shan, Shiqi Tan, Shurui Zhou
ICSME4
2024 "A Lot of Moving Parts": A Case Study of Open-Source Hardware Design Collaboration in the Thingiverse Community
abstract
Open-source is a decentralized and collaborative method of development that encourages open contribution from an extensive and undefined network of individuals. Although commonly associated with software development (OSS), the open-source model extends to hardware development, forming the basis of open-source hardware development (OSH). Compared to OSS, OSH is relatively nascent, lacking adequate tooling support from existing platforms and best practices for efficient collaboration. Taking a necessary step towards improving OSH collaboration, we conduct a detailed case study of DrawBot, a successful OSH project that remarkably fostered a long-term collaboration on Thingiverse - a platform not explicitly intended for complex collaborative design. Through analyzing comment threads and design changes over the course of the project, we found how collaboration occurred, the challenges faced, and how the DrawBot community managed to overcome these obstacles. Beyond offering a detailed account of collaboration practices and challenges, our work contributes best practices, design implications, and practical implications for OSH project maintainers, platform builders, and researchers, respectively. With these insights and our publicly available dataset, we aim to foster more effective and efficient collaborative design in OSH projects.
Kathy Cheng, Shurui Zhou, Alison Olechowski
Proc. ACM Hum. Comput. Interact.2
2023 A Meta-Summary of Challenges in Building Products with ML Components - Collecting Experiences from 4758+ Practitioners
abstract
Incorporating machine learning (ML) components into software products raises new software-engineering challenges and exacerbates existing ones. Many researchers have invested significant effort in understanding the challenges of industry practitioners working on building products with ML components, through interviews and surveys with practitioners. With the intention to aggregate and present their collective findings, we conduct a meta-summary study: We collect 50 relevant papers that together interacted with over 4758 practitioners using guidelines for systematic literature reviews. We then collected, grouped, and organized the over 500 mentions of challenges within those papers. We highlight the most commonly reported challenges and hope this meta-summary will be a useful resource for the research community to prioritize research and education in this field.
Nadia Nahar, Grace A. Lewis, Shurui Zhou, Christian Kästner
CAIN4
2023 Aspirations and Practice of ML Model Documentation: Moving the Needle with Nudging and Traceability
abstract
The documentation practice for machine-learned (ML) models often falls short of established practices for traditional software, which impedes model accountability and inadvertently abets inappropriate or misuse of models. Recently, model cards, a proposal for model documentation, have attracted notable attention, but their impact on the actual practice is unclear. In this work, we systematically study the model documentation in the field and investigate how to encourage more responsible and accountable documentation practice. Our analysis of publicly available model cards reveals a substantial gap between the proposal and the practice. We then design a tool named DocML aiming to (1) nudge the data scientists to comply with the model cards proposal during the model development, especially the sections related to ethics, and (2) assess and manage the documentation quality. A lab study reveals the benefit of our tool towards long-term documentation quality and accountability.
Avinash Bhat, Austin Coursey, Grace Hu, Sixian Li, Nadia Nahar, Shurui Zhou, Christian Kästner, Jin L. C. Guo
CHI6
2023 Interaction of Thoughts: Towards Mediating Task Assignment in Human-AI Cooperation with a Capability-Aware Shared Mental Model
abstract
The existing work on task assignment of human-AI cooperation did not consider the differences between individual team members regarding their capabilities, leading to sub-optimal task completion results. In this work, we propose a capability-aware shared mental model (CASMM) with the components of task grouping and negotiation, which utilize tuples to break down tasks into sets of scenarios relating to difficulties and then dynamically merge the task grouping ideas raised by human and AI through negotiation. We implement a prototype system and a 3-phase user study for the proof of concept via an image labeling task. The result shows building CASMM boosts the accuracy and time efficiency significantly through forming the task assignment close to real capabilities within few iterations. It helps users better understand the capability of AI and themselves. Our method has the potential to generalize to other scenarios such as medical diagnoses and automatic driving in facilitating better human-AI cooperation.
Ziyao He, Yunpeng Song, Shurui Zhou, Zhongmin Cai
CHI3
2023 Aligning Documentation and Q&A Forum through Constrained Decoding with Weak Supervision
abstract
Stack Overflow (SO) is a widely used question-and-answer (Q&A) forum dedicated to software development. It plays a supplementary role to official documentation (DOC for short) by offering practical examples and resolving uncertainties. However, the process of simultaneously consulting both the documentation and SO posts can be challenging and time-consuming due to their disconnected nature. In this study, we propose DOSA, a novel approach to automatically align SO and DOC, which inject domain-specific knowledge about the DOC structure into large language models (LLMs) through weak supervision and constrained decoding, thereby enhancing knowledge retrieval and streamlining task completion during the software development procedure. Our preliminary experiments find that DOSA outperforms various widely-used baselines, showing the promise of using generative retrieval models to perform low-resource software engineering tasks.
Rohith Pudari, Shiyuan Zhou, Iftekhar Ahmed 0001, Zhuyun Dai, Shurui Zhou
ICSME5
2023 User Perspectives on Branching in Computer-Aided Design
abstract
Branching is a feature of distributed version control systems that facilitates the "divide and conquer" strategy present in complex and collaborative work domains. Branching has revolutionized modern software development and has the potential to similarly transform hardware product development via CAD (computer-aided design). Yet, contrasting with its status in software, branching as a feature of commercial CAD systems is in its infancy, and little research exists to investigate its use in the digital design and development of physical products. To address this knowledge gap, in this paper, we mine and analyze 719 user-generated posts from online CAD forums to qualitatively study designers' intentions for and preliminary use of branching in CAD. Our work contributes a taxonomy of CAD branching use cases, an identification of deficiencies of existing branching capabilities in CAD, and a discussion of the untapped potential of CAD branching to support a new paradigm of collaborative mechanical design. The insights gained from this study may help CAD tool developers address design shortcomings in CAD branching tools and assist CAD practitioners by raising their awareness of CAD branching to improve design efficiency and collaborative workflows in hardware development teams.
Kathy Cheng, Phil Cuvin, Alison Olechowski, Shurui Zhou
Proc. ACM Hum. Comput. Interact.4
2023 In the Age of Collaboration, the Computer-Aided Design Ecosystem is Behind: An Interview Study of Distributed CAD Practice
abstract
Computer-aided design (CAD) has become indispensable to increasingly collaborative hardware design processes. Despite the long-standing and growing need for collaboration with CAD models and tools, anecdotal reports and ongoing researcher efforts point to a complex and unresolved set of challenges faced by designers when working with distributed CAD. We aim to close this academic-practitioner knowledge gap through the first systematic study of professional user-driven CAD collaboration challenges. In this work, we conduct semi-structured interviews with 20 CAD professionals of diverse industries, roles, and experience levels to understand their collaborative workflows with distributed CAD tools. In total, we identify 14 challenges related to collaborative design, communication, data management, and permissioning that are currently impeding effective collaboration in professional CAD teams. Our systematic classification of CAD collaboration challenges presents a guide for pressing areas of future work, highlighting important implications for CAD researchers, practitioners, and tool builders to target new advancement in CAD infrastructure, management choices, and modelling best practices. With the insights gained from this work, we hope to ultimately improve collaboration efficiency, quality, and innovation for future product design teams.
Kathy Cheng, Michal K. Davis, Xiyue Zhang 0004, Shurui Zhou, Alison Olechowski
Proc. ACM Hum. Comput. Interact.4
2023 Perceptions of open-source software developers on collaborations: An interview and survey study
abstract
Abstract With the emergence of social coding platforms, collaboration has become a key and dynamic aspect to the success of software projects. In such platforms, developers have to collaborate and deal with issues of collaboration in open‐source software development. Although collaboration is challenging, collaborative development produces better software systems than any developer could produce alone. Several approaches have investigated collaboration challenges, for instance, by proposing or evaluating models and tools to support collaborative work. Despite the undeniable importance of the existing efforts in this direction, there are few works on collaboration from perspectives of developers. In this work, we aim to investigate the perceptions of open‐source software developers on collaborations, such as motivations, techniques, and tools to support global, productive, and collaborative development. Following an ad hoc literature review, an exploratory interview study with 12 open‐source software developers fromGitHub, our novel approach for this problem also relies on an extensive survey with 121 developers to confirm or refute the interview results. We found different collaborative contributions, such as managing change requests. Besides, we observed that most collaborators prefer to collaborate with the core team instead of their peers. We also found that most collaboration happens in software development (60%) and maintenance (47%) tasks. Furthermore, despite personal preferences to work independently, developers still consider collaborating with others in specific task categories, for instance, software development. Finally, developers also expressed the importance of the social coding platforms, such asGitHub, to support maintainers, and contributors in making decisions and developing tasks of the projects. Therefore, these findings may help project leaders optimize the collaborations among developers and reduce entry barriers. Moreover, these findings may support the project collaborators in understanding the collaboration process and engaging others in the project.
Kattiana Constantino, Maurício R. de A. Souza, Shurui Zhou, Eduardo Figueiredo 0001, Christian Kästner
J. Softw. Evol. Process.3
2022 Collaboration Challenges in Building ML-Enabled Systems: Communication, Documentation, Engineering, and Process
abstract
The introduction of machine learning (ML) components in software projects has created the need for software engineers to collaborate with data scientists and other specialists. While collaboration can always be challenging, ML introduces additional challenges with its exploratory model development process, additional skills and knowledge needed, difficulties testing ML systems, need for continuous evolution and monitoring, and non-traditional quality requirements such as fairness and explainability. Through interviews with 45 practitioners from 28 organizations, we identified key collaboration challenges that teams face when building and deploying ML systems into production. We report on common collaboration points in the development of production ML systems for requirements, data, and integration, as well as corresponding team patterns and challenges. We find that most of these challenges center around communication, documentation, engineering, and process, and collect recommendations to address these challenges.
Nadia Nahar, Shurui Zhou, Grace A. Lewis, Christian Kästner
ICSE2
2022 Elevating Jupyter Notebook Maintenance Tooling by Identifying and Extracting Notebook Structures
abstract
Data analysis is an exploratory, interactive, and often collaborative process. Computational notebooks have become a popular tool to support this process, among others because of their ability to interleave code, narrative text, and results. However, notebooks in practice are often criticized as hard to maintain and being of low code quality, including problems such as unused or duplicated code and out-of-order code execution. Data scientists can benefit from better tool support when maintaining and evolving notebooks. We argue that central to such tool support is identifying the structure of notebooks. We present a lightweight and accurate approach to extract notebook structure and outline several ways such structure can be used to improve maintenance tooling for notebooks, including navigation and finding alternatives.
Christian Kästner, Shurui Zhou
ICSME3
2022 An empirical study of emoji use in software development communication
Shiyue Rong, Weisheng Wang, Umme Ayda Mannan, Eduardo Santana de Almeida, Shurui Zhou, Iftekhar Ahmed 0001
Inf. Softw. Technol.5
2021 Interactive Patch Filtering as Debugging Aid
abstract
It is widely recognized that patches generated by program repair tools have to be correct to be useful. However, it is fundamentally difficult to ensure the correctness of the patches. Many tools generate only the patches that are highly likely to be correct by taking conservative strategies which inevitably limit the recall of APR approaches. While the recall of APR can potentially be improved by relaxing the requirement on precision, more incorrect patches may also be generated. In this paper, we conjecture that reviewing incorrect patches also helps developers to understand the bug, and with proper tool support, reviewing incorrect patches would at least not reduce the repair performance. To evaluate this, we propose an interactive patch filtering approach to facilitate developers in the patch review process via effectively filtering out groups of incorrect patches. We implemented the approach as an Eclipse plugin, InPaFer, and evaluated the effectiveness and usefulness with a mixed-method evaluation. The results show that our approach improves the repair performance of developers, with 62.5% more successfully repaired bugs and 25.3% less debugging time. In particular, even if all generated patches are incorrect, the performance of developers would not be significantly reduced, and could still be improved. Our work provides a new way of thinking for the APR research.
Ruyi Ji, Jiajun Jiang, Shurui Zhou, Yiling Lou, Yingfei Xiong 0001, Gang Huang 0001
ICSME4
2021 Subtle Bugs Everywhere: Generating Documentation for Data Wrangling Code
abstract
Data scientists reportedly spend a significant amount of their time in their daily routines on data wrangling, i.e. cleaning data and extracting features. However, data wrangling code is often repetitive and error-prone to write. Moreover, it is easy to introduce subtle bugs when reusing and adopting existing code, which results in reduced model quality. To support data scientists with data wrangling, we present a technique to generate documentation for data wrangling code. We use (1) program synthesis techniques to automatically summarize data transformations and (2) test case selection techniques to purposefully select representative examples from the data based on execution information collected with tailored dynamic program analysis. We demonstrate that a JupyterLab extension with our technique can provide on-demand documentation for many cells in popular notebooks and find in a user study that users with our plugin are faster and more effective at finding realistic bugs in data wrangling code.
Chenyang Yang 0002, Shurui Zhou, Jin L. C. Guo, Christian Kästner
ASE2
2020 Understanding collaborative software development: an interview study
abstract
In globally distributed software development, many software developers have to collaborate and deal with issues of collaboration. Although collaboration is challenging, collaborative development produces better software than any developer could produce alone. Unlike previous work which focuses on the proposal and evaluation of models and tools to support collaborative work, this paper presents an interview study aiming to understand (i) the motivations, (ii) how collaboration happens, and (iii) the challenges and barriers of collaborative software development. After interviewing twelve experienced software developers from GitHub, we found different types of collaborative contributions, such as in the management of requests for changes. Our analysis also indicates that the main barriers for collaboration are related to non-technical, rather than technical issues.
Kattiana Constantino, Shurui Zhou, Maurício R. de A. Souza, Eduardo Figueiredo 0001, Christian Kästner
ICGSE2
2020 How has forking changed in the last 20 years?: a study of hard forks on GitHub
abstract
The notion of forking has changed with the rise of distributed version control systems and social coding environments, like GitHub. Traditionally forking refers to splitting off an independent development branch (which we call hard forks); research on hard forks, conducted mostly in pre-GitHub days showed that hard forks were often seen critical as they may fragment a community Today, in social coding environments, open-source developers are encouraged to fork a project in order to contribute to the community (which we call social forks), which may have also influenced perceptions and practices around hard forks. To revisit hard forks, we identify, study, and classify 15,306 hard forks on GitHub and interview 18 owners of hard forks or forked repositories. We find that, among others, hard forks often evolve out of social forks rather than being planned deliberately and that perception about hard forks have indeed changed dramatically, seeing them often as a positive noncompetitive alternative to the original project.
Shurui Zhou, Bogdan Vasilescu, Christian Kästner
ICSE1
2020 An Exploratory Study to Find Motives Behind Cross-platform Forks from Software Heritage Dataset
abstract
The fork-based development mechanism provides the flexibility and the unified processes for software teams to collaborate easily in a distributed setting without too much coordination overhead. Currently, multiple social coding platforms support fork-based development, such as GitHub, GitLab, and Bitbucket. Although these different platforms virtually share the same features, they have different emphasis. As GitHub is the most popular platform and the corresponding data is publicly available, most of the current studies are focusing on GitHub hosted projects. However, we observed anecdote evidences that people are confused about choosing among these platforms, and some projects are migrating from one platform to another, and the reasons behind these activities remain unknown. With the advances of Software Heritage Graph Dataset (SWHGD), we have the opportunity to investigate the forking activities across platforms. In this paper, we conduct an exploratory study on 10 popular open-source projects to identify cross-platform forks and investigate the motivation behind. Preliminary result shows that cross-platform forks do exist. For the 10 subject systems used in this study, we found 81,357 forks in total among which 179 forks are on GitLab. Based on our qualitative analysis, we found that most of the cross-platform forks that we identified are mirrors of the repositories on another platform, but we still find cases that were created due to preference of using certain functionalities (e.g. Continuous Integration (CI)) supported by different platforms. This study lays the foundation of future research directions, such as understanding the differences between platforms and supporting cross-platform collaboration.
Avijit Bhattacharjee, Sristy Sumana Nath, Shurui Zhou, Debasish Chakroborti, Banani Roy, Chanchal Kumar Roy, Kevin A. Schneider
MSR3
2019 How to Explain a Patch: An Empirical Study of Patch Explanations in Open Source Projects
abstract
Bugs are inevitable in software development and maintenance processes. Recently a lot of research efforts have been devoted to automatic program repair, aiming to reduce the efforts of debugging. However, since it is difficult to ensure that the generated patches meet all quality requirements such as correctness, developers still need to review the patch. In addition, current techniques produce only patches without explanation, making it difficult for the developers to understand the patch. Therefore, we believe a more desirable approach should generate not only the patch but also an explanation of the patch. To generate a patch explanation, it is important to first understand how patches were explained. In this paper, we explored how developers explain their patches by manually analyzing 300 merged bug-fixing pull requests from six projects on GitHub. Our contribution is twofold. First, we build a patch explanation model, which summarizes the elements in a patch explanation, and corresponding expressive forms. Second, we conducted a quantitative analysis to understand the distributions of elements, and the correlation between elements and their expressive forms.
Yaozong Hou, Shurui Zhou, Junjie Chen 0003, Yingfei Xiong 0001, Gang Huang 0001
ISSRE3
2019 Improving Collaboration Efficiency in Fork-Based Development
abstract
Fork-based development is a lightweight mechanism that allows developers to collaborate with or without explicit coordination. Although it is easy to use and popular, when developers each create their own fork and develop independently, their contributions are usually not easily visible to others. When the number of forks grows, it becomes very difficult to maintain an overview of what happens in individual forks, which would lead to additional problems and inefficient practices: lost contributions, redundant development, fragmented communities, and so on. Facing the problems mentioned above, we developed two complementary strategies: (1) Identifying existing best practices and suggesting evidence-based interventions for projects that are inefficient; (2) designing new interventions that could improve the awareness of a community using fork-based development, and help developers to detect redundant development to reduce unnecessary effort.
Shurui Zhou
ASE1
2019 What the fork: a study of inefficient and efficient forking practices in social coding
abstract
Forking and pull requests have been widely used in open-source communities as a uniform development and contribution mechanism, giving developers the flexibility to modify their own fork without affecting others before attempting to contribute back. However, not all projects use forks efficiently; many experience lost and duplicate contributions and fragmented communities. In this paper, we explore how open-source projects on GitHub differ with regard to forking inefficiencies. First, we observed that different communities experience these inefficiencies to widely different degrees and interviewed practitioners to understand why. Then, using multiple regression modeling, we analyzed which context factors correlate with fewer inefficiencies.We found that better modularity and centralized management are associated with more contributions and a higher fraction of accepted pull requests, suggesting specific best practices that project maintainers can adopt to reduce forking-related inefficiencies in their communities.
Shurui Zhou, Bogdan Vasilescu, Christian Kästner
ESEC/SIGSOFT FSE1
2019 Identifying Redundancies in Fork-based Development
abstract
Fork-based development is popular and easy to use, but makes it difficult to maintain an overview of the whole community when the number of forks increases. This may lead to redundant development where multiple developers are solving the same problem in parallel without being aware of each other. Redundant development wastes effort for both maintainers and developers. In this paper, we designed an approach to identify redundant code changes in forks as early as possible by extracting clues indicating similarities between code changes, and building a machine learning model to predict redundancies. We evaluated the effectiveness from both the maintainer's and the developer's perspectives. The result shows that we achieve 57-83% precision for detecting duplicate code changes from maintainer's perspective, and we could save developers' effort of 1.9-3.0 commits on average. Also, we show that our approach significantly outperforms existing state-of-art.
Luyao Ren, Shurui Zhou, Christian Kästner, Andrzej Wasowski
SANER2
2018 Adding sparkle to social coding: an empirical study of repository badges in the npm ecosystem
abstract
In fast-paced, reuse-heavy, and distributed software development, the transparency provided by social coding platforms like GitHub is essential to decision making. Developers infer the quality of projects using visible cues, known as signals, collected from personal profile and repository pages. We report on a large-scale, mixed-methods empirical study of npm packages that explores the emerging phenomenon of repository badges, with which maintainers signal underlying qualities about their projects to contributors and users. We investigate which qualities maintainers intend to signal and how well badges correlate with those qualities. After surveying developers, mining 294,941 repositories, and applying statistical modeling and time-series analyses, we find that non-trivial badges, which display the build status, test coverage, and up-to-dateness of dependencies, are mostly reliable signals, correlating with more tests, better pull requests, and fresher dependencies. Displaying such badges correlates with best practices, but the effects do not always persist.
Asher Trockman, Shurui Zhou, Christian Kästner, Bogdan Vasilescu
ICSE2
2018 Identifying features in forks
abstract
Fork-based development has been widely used both in open source communities and in industry, because it gives developers flexibility to modify their own fork without affecting others. Unfortunately, this mechanism has downsides: When the number of forks becomes large, it is difficult for developers to get or maintain an overview of activities in the forks. Current tools provide little help. We introduce Infox, an approach to automatically identify non-merged features in forks and to generate an overview of active forks in a project. The approach clusters cohesive code fragments using code and network-analysis techniques and uses information-retrieval techniques to label clusters with keywords. The clustering is effective, with 90 % accuracy on a set of known features. In addition, a human-subject evaluation shows that Infox can provide actionable insight for developers of forks.
Shurui Zhou, Stefan Stanciulescu, Olaf Leßenich, Yingfei Xiong 0001, Andrzej Wasowski, Christian Kästner
ICSE1
2013 Elastic resource management for heterogeneous applications on PaaS
abstract
Elastic resource management is one of the key characteristics of cloud computing systems. Existing elastic approaches focus mainly on single resource consumption such as CPU consumption, rarely considering comprehensively various features of applications. Applications deployed on a PaaS are usually heterogeneous. While sharing the same resource, these applications are usually quite different in resource consuming. How to deploy these heterogeneous applications on the smallest size of hardware thus becomes a new research topic. In this paper, we take into consideration application's CPU consumption, I/O consumption, consumption of other server resources and application's request rate, all of which are defined as application features. This paper proposes a practical and effective elasticity approach based on the analysis of application features. The evaluation experiment shows that, compared with traditional approach, our approach can save up to 32.8% VMs without significant increase of average response time and SLA violation.
Shurui Zhou, Qianxiang Wang
Internetware2