VLDB 2026 Research / reviewers in the wild / expert
Tao Wang 0006
dblp:12/5838-6
· DBLP profile ↗
74ranked-venue papers
4as first author
24since 2021 · last 2026
0000-0002-8406-8672ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 53 · 3 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 8 · 1 since 2021Systems, architecture and hardware · 3 · 1 since 2021Databases, data management, data science and information retrieval · 3Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Security and privacy · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Temporal-Enhanced Multimodal Transformer for Referring Multi-Object Tracking and SegmentationabstractReferring multi-object tracking (RMOT) is an emerging cross-modal task that aims to locate an arbitrary number of target objects and maintain their identities referred by a language expression in a video. This intricate task involves the reasoning of linguistic and visual modalities, along with the temporal association of target objects. However, the seminal work relies on loose feature fusion and neglects long-term information. In this study, we introduce a compact Transformer-based method, termed TenRMOT. We conduct feature fusion at both encoding and decoding stages to fully exploit the advantages of Transformer architecture. Specifically, we incrementally perform cross-modal fusion layer-by-layer during the encoding phase. In the decoding phase, we utilize language-guided queries to probe memory features for accurate prediction of the desired objects. Moreover, we introduce a query update module that explicitly leverages temporal prior information of the tracked objects to enhance the consistency of their trajectories. In addition, we introduce a novel task called Referring Multi-Object Tracking and Segmentation (RMOTS) and construct a new dataset named Ref-KITTI Segmentation. Our dataset consists of 18 videos with 818 expressions, and each expression averages 10.7 masks, which poses a greater challenge compared to the typical single mask in most existing referring video segmentation datasets. TenRMOT demonstrates superior performance on both the referring multi-object tracking and the segmentation tasks. Changcheng Xiao, Qiong Cao, Xiang Zhang 0008, Tao Wang 0006, Canqun Yang, Long Lan |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | An Empirical Study of Overlooked Code Review Comments in OSS ProjectsabstractOpen source software (OSS) development widely adopts modern code review to identify issues and guarantee code quality. As reported repeatedly, maintainers are under heavy workloads when reviewing code changes. Meanwhile, we notice that some code reviews were overlooked by the authors of the code changes, i.e., neither causing code modification nor being replied to. These code reviews, if requiring responses but not receiving any, might represent a significant inefficiency, risk of overlooking critical issues, and problematic social exchange. Moreover, leaving code reviews publicly unanswered may cause a negative impression on both the corresponding OSS contributors and the OSS projects. Existing literature on code review mainly focuses on the usefulness of code reviews, reviewer recommendations, factors affecting PR acceptance, and review comment generation; the nature of overlooked reviews has not been explored. To this end, we focus on a widely-used modern code review mechanism, i.e., reviewing Pull Request (PR) code before merge, and conduct the first empirical study on 80 Java OSS projects to explore the prevalence, characteristics, rationales, and possible impact of the overlooked code reviews. We find that approximately 7.5% of PRs have at least one review comment being ignored. We further show that pull requests containing no-response comments are significantly associated with longer review lifecycles and lower acceptance rates, indicating measurable negative outcomes beyond their modest prevalence. Then, we categorize these no-response comments through thematic analysis and find two main categories with seven subcategories: Review inquiry and PR management. We also extract four subcategories in Review inquiry, e.g., Give suggestions about code implementation, Point out implementation issues, and Additional task requests. PR management consists of three subcategories, i.e., PR status checks, PR merge conflict notifications, and Reject PR with uncertain reasons. To better understand the existence of no-response comments, we surveyed developers and received 45 responses. We found that the reasons for the existence of no-response comments are diverse, such as prolonged review times and a lack of consensus on opinions. Developers also hold the consensus that ignored reviews will have negative effects on software projects. These findings emphasize the need for attention from both academia and industry to the responses to review comments and optimization of the reminder mechanism. Yuxia Zhang, Qunhong Zeng, Lin Shi 0006, Xin Tan 0003, Tao Wang 0006, Yanjie Jiang, Hui Liu 0003 |
IEEE Trans. Software Eng. | 6 |
| 2026 | A Deep Dive Into Deprecation Declarations in the Rust Package EcosystemabstractUtilizing third-party open source libraries is fundamental to modern software development because it can enhance productivity and software quality. However, libraries may cease maintenance and become deprecated, negatively impacting the projects that rely on them. Promptly identifying and addressing deprecated libraries can help developers mitigate potential risks within their projects. As a programming language known for its emphasis on safety, Rust’s package manager currently does not provide a direct mechanism for deprecation. Nevertheless, Rust developers can still declare deprecation using certain methods offered by GitHub and the official Rust package registry, crates.io. However, the current usage of these deprecation mechanisms in the Rust ecosystem, as well as their effectiveness, remains underexplored. This paper addresses this gap by empirically studying the prevalence of deprecation declarations in Rust libraries, the effectiveness of different ways of declarations, and the reasons for using deprecated libraries to understand how deprecation information is disseminated and perceived in the current Rust ecosystem. We found that: 1) Among the 13,289 inactive libraries in the Rust ecosystem, only 11% of them indicate their deprecated status; 2) Among the packages that released a new version after their dependent library declared deprecation, 38.9% still chose to use the deprecated library in their new releases; 3) Despite developers being able to actively or passively discover deprecated libraries within their projects through various means, unawareness of library deprecation is a significant reason for developers using deprecated libraries. Based on these findings, we discuss practice insights to help improve the deprecation mechanism and mitigate software dependency risks. Minyu Shu, Meng Fan, Yuxia Zhang, Tao Wang 0006, Hui Liu 0003 |
IEEE Trans. Software Eng. | 4 |
| 2026 | DockerFill: Automatically Completing Dockerfile Code With Syntax-Aware Multi-Task LearningabstractAs a kind of infrastructure-as-code, Dockerfile specifies the structure and functionality of a built Docker image and thus plays an important role in the containerized software development process. Nowadays developers need to spend extra time and effort configuring their Dockerfiles in addition to their regular coding work, which requires knowledge and skills orthogonal to those entailed in other software-related experiences. Poorly written Dockerfile code often introduces errors and maintenance costs. However, little automated support is available for assisting developers in configuring Dockerfiles. In this study, we first conduct an online survey to investigate Docker developers’ perceptions of Dockerfile writing, highlighting the needs and potential benefits of Dockerfile auto-completion techniques. Then, we introduceDOCKERFILL, a pre-trained model based approach that provides completion suggestions for Dockerfile-specific code.DOCKERFILLleverages multi-layer Transformer architecture with syntax-aware multi-task learning, which includes contextual file information and three pre-training tasks, i.e., masked language modeling, syntax type identification, and masked identifier prediction. To evaluateDOCKERFILL’s effectiveness, we collect a dataset of 6,350 high-quality real-world Dockerfiles. Our empirical results show that DOCKERFILL provides up to 52.38% accuracy for token-level completion and 19.69% exact match for line-level completion, outperforming the baselines by 7.32%-37.67% and 1.97%-19.69%, respectively. Also,DOCKERFILLobtains significantly higher human evaluation scores compared to the baselines. Yiwen Wu 0001, Yang Zhang 0026, Tao Wang 0006, Bo Ding 0001, Huaimin Wang 0001 |
IEEE Trans. Software Eng. | 3 |
| 2025 | Open source oriented cross-platform survey
Simeng Yao, Xunhui Zhang, Yang Zhang 0026, Tao Wang 0006 |
Inf. Softw. Technol. | 4 |
| 2025 | What problems are MLOps practitioners talking about? A study of discussions in Stack Overflow forum and GitHub projects
Yang Zhang 0026, Yiwen Wu 0001, Tao Wang 0006, Bo Ding 0007, Huaimin Wang 0001 |
Inf. Softw. Technol. | 3 |
| 2024 | GHA-BFP: Framework for Automated Build Failure Prediction in GitHub ActionsabstractGitHub Actions (GHA), a powerful Continuous Integration and Continuous Deployment (CI/CD) service, has revolutionized the way developers automate tasks in the software development pipeline. Although GHA provides great convenience, if a GHA build fails, the time spent waiting for results and debugging is wasted, which can seriously affect development efficiency. In this study, we delve into GHA build results and introduce an automatic framework named GHA-BFP that uses ML models to predict the failure of GHA builds. Using GHA-BFP with Random Forest model, we achieved the highest performance in predicting the failure of GHA builds, with all key metrics (i.e., Accuracy, Precision, Recall, and F1 score) exceeding 75%. Furthermore, through ablation experiments, we have verified the essentiality of the four categories of input features. Lastly, we conducted an assessment of the importance of each individual input feature in relation to the model's predictive capabilities. Jiatai Li, Yang Zhang 0026, Tao Wang 0006, Yiwen Wu 0001 |
APSEC | 3 |
| 2024 | How do Developers Talk about GitHub Actions? Evidence from Online Software Development CommunityabstractContinuous integration, deployment and delivery (CI/CD) have become cornerstones of DevOps practices. In recent years, GitHub Action (GHA) has rapidly replaced the traditional CI/CD tools on GitHub, providing efficiently automated workflows for developers. With the widespread use and influence of GHA, it is critical to understand the existing problems that GHA developers face in their practices as well as the potential solutions to these problems. Unfortunately, we currently have relatively little knowledge in this area. To fill this gap, we conduct a large-scale empirical study of 6,590 Stack Overflow (SO) questions and 315 GitHub issues. Our study leads to the first comprehensive taxonomy of problems related to GHA, covering 4 categories and 16 sub-categories. Then, we analyze the popularity and difficulty of problem categories and their correlations. Further, we summarize 56 solution strategies for different GHA problems. We also distill practical implications of our findings from the perspective of different audiences. We believe that our study contributes to the research of emerging GHA practices and guides the future support of tools and technologies. Yang Zhang 0026, Yiwen Wu 0001, Tao Wang 0006, Hui Liu 0052, Huaimin Wang 0001 |
ICSE | 4 |
| 2024 | Understanding the Challenges of Data Management in the AI Application DevelopmentabstractWith the development of AI model ecology pioneered by the large language models, data plays a more important and diversified role than the traditional one. In recent years, data problems have been increasingly noticed and voiced by developers, and it is crucial to understand the existing challenges that they are facing in reality, as well as the potential features of these problems. Unfortunately, we currently have relatively little knowledge in this area. To fill this gap, we conduct an empirical study based on 2,154 posts and 785 issues by collecting model data problems from three online communities (i.e., GitHub, Stack Overflow, and Hugging Face). We present the first taxonomy of data problems around model production encountered by developers, covering 11 topics in 4 categories. Then, we analyze the popularity and the difficulty of the posts and issues in the derived topics. We distill many findings from the study and present the practical implications from the perspectives of different audiences. Junchen Li, Yang Zhang 0026, Kele Xu, Tao Wang 0006, Huaimin Wang 0001 |
JCC | 4 |
| 2023 | TLDBERT: Leveraging Further Pre-Trained Model for Issue Typed Links DetectionabstractIssue links are crucial for promoting software information flowing and development efficiency, but manually conducting typed links detection (TLD) is time-consuming and error-prone due to the numerous candidate issues. The general pre-trained NLP models, such as BERT, provide promising automated approaches for TLD task after fine-tuning. In addition, the further pre-training method, which is an intermediate pre-training process on the in-domain corpora, has been proven to further improve the performance of pre-trained models on many downstream tasks. In this paper, we apply the further pre-training method on TLD task and assess the improved performance. We further pre-train BERT by utilizing the in-domain corpora constructed by Jira issues dataset, and then finetune it on the links dataset to obtain the more applicable model for TLD task - TLDBERT. The experimental results indicate the statistically significant improvement in the performance of TLDBERT, with a range of 0.4%-6.6% for accuracy and 0.1%-9.7% for macro F1-score. Finally, based on correlation analysis, we conclude that the coverage and the ratio of issues to users are significantly correlated with the improvement, although the two directions are opposite. Huaian Zhou, Tao Wang 0006, Yang Zhang 0026 |
APSEC | 2 |
| 2023 | To Follow or Not to Follow: Understanding Issue/Pull-Request Templates on GitHubabstractFor most Open Source Software (OSS) projects, issues and Pull-requests (PR) are the primary means by which stakeholders of a project report and discuss software problems and code changes, and their descriptions are important for people to understand them. To help ensure the informational quality of issue/PR descriptions, GitHub introduced theissue/PR templatefeature, which pre-populates the description for anyone trying to open a new issue/PR. To better understand this feature, we report on a large-scale, mixed-methods empirical study of templates that explores contents, impacts, and perceptions. Our results show that templates typically contain elements to greet contributors, explain project guidelines, and collect relevant information. After template adoption, the monthly volume of incoming issues and PRs decreases, and issues have fewer monthly discussion comments and longer resolution duration. Although both contributors and maintainers positively rated the usefulness of templates from various aspects, they also reported challenges in using templates (e.g., excessive and irrelevant information request) and suggested potential improvements of the template feature (e.g., better user interaction and advanced automation). This work contributes to the informed use and targeted improvement of templates to enhance OSS practitioners’ collaboration and interaction. Yue Yu 0001, Tao Wang 0006, Yan Lei 0005, Ying Wang 0038, Huaimin Wang 0001 |
IEEE Trans. Software Eng. | 3 |
| 2022 | Who, What, Why and How? Towards the Monetary Incentive in Crowd Collaboration: A Case Study of Github's Sponsor MechanismabstractWhile many forms of financial support are currently available, there are still many complaints about inadequate financing from software maintainers. In May 2019, GitHub, the world’s most active social coding platform, launched the Sponsor mechanism as a step toward more deeply integrating open source development and financial support. This paper collects data on 8,028 maintainers, 13,555 sponsors, and 22,515 sponsorships and conducts a comprehensive analysis. We explore the relationship between the Sponsor mechanism and developers along four dimensions using a combination of qualitative and quantitative analysis, examining why developers participate, how the mechanism affects developer activity, who obtains more sponsorships, and what mechanism flaws developers have encountered in the process of using it. We find a long-tail effect in the act of sponsorship, with most maintainers’ expectations remaining unmet, and sponsorship has only a short-term, slightly positive impact on development activity but is not sustainable. While sponsors participate in this mechanism mainly as a means of thanking the developers of OSS that they use, in practice, the social status of developers is the primary influence on the number of sponsorships. We find that both the Sponsor mechanism and open source donations have certain shortcomings and need further improvements to attract more participants. Xunhui Zhang, Tao Wang 0006, Yue Yu 0001, Qiubing Zeng, Huaimin Wang 0001 |
CHI | 2 |
| 2022 | A Preliminary Study of Bots Usage in Open Source CommunityabstractBots are seen as a promising approach in software development, which help to deal with the ever-increasing complexity of modern software engineering and development. The number of bots in open source community, such as GitHub, has expanded substantially over the last three years. Due to its increasing popularity, it is essential to characterize the current usage of bots in practices. In this paper, we present an empirical study of bots usage in GitHub community. By analyzing 7,399 projects from GitHub, we find that 4,148 (56%) projects have used bots. Through automatic identification and manual detection, we collect a total of 196 bots. We then analyze and classify them into 4 categories and 14 topics. Finally, we discuss some raised implications for bots in current GitHub community. Anze Gao, Yang Zhang 0026, Tao Wang 0006 |
Internetware | 4 |
| 2022 | Understanding and Predicting Docker Build Duration: An Empirical Study of Containerized Workflow of OSS ProjectsabstractDocker building is a critical component of containerized workflow, which automates the process by which sources are packaged and transformed into container images. If not run properly, Docker builds can bring long durations (i.e., slow builds), which increases the cost in human and computing resources, and thus inevitably affect the software development. However, the current status and remedy for the duration cost in Docker builds remain unclear and need an in-depth study. To fill this gap, this paper provides the first empirical investigation on 171,439 Docker builds from 5,833 open source software (OSS) projects. Starting with an exploratory study, the Docker build durations can be characterized in real-world projects, and the developers’ perceptions of slow builds are obtained via a comprehensive survey. Driven by the results of our exploratory study, we propose a prediction modeling of Docker build duration, leveraging 27 handcrafted features from build-related context and configuration and 8 regression algorithms for the prediction task. Our results demonstrate that Random Forest model provides the superior performance with a Spearman’s correlation of 0.781, outperforming the baseline random model by 82.9% in RMSE, 90.6% in MAE, and 94.4% in MAPE, respectively. The implications of this study will facilitate research and assist practitioners in improving the Docker build process. Yiwen Wu 0001, Yang Zhang 0026, Kele Xu, Tao Wang 0006, Huaimin Wang 0001 |
ASE | 4 |
| 2022 | HAF: a hybrid annotation framework based on expert knowledge and learning technique
Yue Yu 0001, Tao Wang 0006, Gang Yin, Xinjun Mao, Huaimin Wang 0001 |
Sci. China Inf. Sci. | 3 |
| 2022 | Pull request latency explained: an empirical overview
Xunhui Zhang, Yue Yu 0001, Tao Wang 0006, Ayushi Rastogi, Huaimin Wang 0001 |
Empir. Softw. Eng. | 3 |
| 2022 | Opportunities and Challenges in Repeated Revisions to Pull-Requests: An Empirical StudyabstractBackground: The Pull-Request (PR) model is a widespread approach adopted by open source software (OSS) projects to support collaborative software development. However, it is often challenging to continuously evaluate and revise PRs in several iterations of code reviewsinvolving technical and social aspects. Aim: Our objective is twofold: identifying best practices for effective collaboration in continuous PR improvement and uncovering problems that deserve special attention to improve collaboration efficiency and productivity. Method: We conducted a mixed-methods empirical study of repeatedly revised PRs (i.e. those that have undergone a high number of revisions). Historical trace data of five long-lived popular GitHub projects were used for manual investigation of practices for requesting changes to PRs and reasons for nonacceptance of repeatedly revised PRs. Surveys of OSS practitioners were conducted to evaluate the results of manual analysis and to provide additional insights into developers' willingness regarding PR revisions and factors causing avoidable revisions in practice. Results: The main results of our research were as follows: (1) We identified 15 code review practices for requesting changes to PRs, among which practices with respect to explaining the reasoning behind requested changes and tracking the progress of PR review and revision were undervalued by reviewers; (2) While submitters can in general undergo 1-5 rounds of revisions, they are willing to offer more revisions when they are in a friendly community and receive helpful feedback; (3) We revealed 11 factors causing avoidable revisions regarding to reviewers' feedback, code review policy, pre-submission issues, and implementation of new revisions; and (4) Nonacceptance of repeatedly revised PRs was due mainly to inactivity of submitters or reviewers and being superseded for better maintenance. Finally, based on these findings, we proposed recommendations and implications for OSS practitioners and tool designers to facilitate efficient collaboration in PR revisions. Yue Yu 0001, Tao Wang 0006, Shanshan Li 0001, Huaimin Wang 0001 |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2022 | Are You Still Working on This? An Empirical Study on Pull Request AbandonmentabstractThe great success of numerous community-based open source software (OSS) is based on volunteers continuously submitting contributions, but ensuring sustainability is a persistent challenge in OSS communities. Although the motivations behind and barriers to OSS contributors’ joining and retention have been extensively studied, the impacts of, reasons for and solutions to contribution abandonment at the individual level have not been well studied, especially for pull-based development. To bridge this gap, we present an empirical study on pull request abandonment based on a sizable dataset. We manually examine 321 abandoned pull requests on GitHub and then quantify the manual observations by surveying 710 OSS developers. We find that while the lack of integrators’ responsiveness and the lack of contributors’ time and interest remain the main reasons that deter contributors from participation, limitations during the processes of patch updating and consensus reaching can also cause abandonment. We also show the significant impacts of pull request abandonment on project management and maintenance. Moreover, we elucidate the strategies used by project integrators to cope with abandoned pull requests and highlight the need for a practical handover mechanism. We discuss the actionable suggestions and implications for OSS practitioners and tool builders, which can help to upgrade the infrastructure and optimize the mechanisms of OSS communities. Yue Yu 0001, Tao Wang 0006, Gang Yin, Shanshan Li 0001, Huaimin Wang 0001 |
IEEE Trans. Software Eng. | 3 |
| 2022 | Redundancy, Context, and Preference: An Empirical Study of Duplicate Pull Requests in OSS ProjectsabstractOSS projects are being developed by globally distributed contributors, who often collaborate through the pull-based model today. While this model lowers the barrier to entry for OSS developers by synthesizing, automating and optimizing the contribution process, coordination among an increasing number of contributors remains as a challenge due to the asynchronous and self-organized nature of distributed development. In particular, duplicate contributions, where multiple different contributors unintentionally submit duplicate pull requests to achieve the same goal, are an elusive problem that may waste effort in automated testing, code review and software maintenance. While the issue of duplicate pull requests has been highlighted, to what extent duplicate pull requests affect the development in OSS communities has not been well investigated. In this paper, we conduct a mixed-approach study to bridge this gap. Based on a comprehensive dataset constructed from 26 popular GitHub projects, we obtain the following findings: (a) Duplicate pull requests result in redundant human and computing resources, exerting a significant impact on the contribution and evaluation process. (b) Contributors’ inappropriate working patterns and the drawbacks of their collaborating environment might result in duplicate pull requests. (c) Compared to non-duplicate pull requests, duplicate pull requests have significantly different features, e.g., being submitted by inexperienced contributors, being fixing bugs, touching cold files, and solving tracked issues. (d) Integrators choosing between duplicate pull requests prefer to accept those with early submission time, accurate and high-quality implementation, broad coverage, test code, high maturity, deep discussion, and active response. Finally, actionable suggestions and implications are proposed for OSS practitioners. Yue Yu 0001, Minghui Zhou 0001, Tao Wang 0006, Gang Yin, Long Lan, Huaimin Wang 0001 |
IEEE Trans. Software Eng. | 4 |
| 2022 | Motivation Under Gamification: An Empirical Study of Developers' Motivations and Contributions in Stack OverflowabstractTo encourage developers' volunteer contributions, modern programming question and answer (Q&A) sites like Stack Overflow (SO) employ gamified incentive mechanisms such as reputation and badges. Understanding developers' motivations in the presence of gamification and the relationship between their motivations and behavioral outcomes is crucial for community building and designing good incentive mechanisms. Grounded on self-determination theory, we conducted a survey with 938 developers who participate in SO to understand their participation motivations and incentive perceptions. By connecting the survey responses with the SO data, we quantitatively analyzed how the developers' motivations and satisfaction of needs relate to their effort and contribution quality. Our main findings are as follows: (1) despite the presence of gamified incentive mechanisms, developers are mainly motivated by intrinsic motivation to participate in SO; (2) developers who have strong motivations to gain gamification rewards are associated with higher intrinsic and integrated motivations, while developers with more development experiences are less motivated by the gamified incentives; (3) both extrinsic motivations (in terms of career prospects) and intrinsic motivations (regarding self-improvement and helping others) can motivate developers to make high-quantity and high-quality contributions; and (4) high-level satisfaction of needs for competency and autonomy has a positive effect on developers making high-quantity and high-quality contributions and addressing difficult problems. Based on these findings, we discuss implications for developer motivation and gamification in the crowdsourcing context and for the mechanism design of gamified crowdsourced platforms. Yao Lu 0003, Xinjun Mao, Minghui Zhou 0001, Yang Zhang 0026, Zude Li, Tao Wang 0006, Gang Yin, Huaimin Wang 0001 |
IEEE Trans. Software Eng. | 6 |
| 2021 | A Technical Capability Evaluation Model Based Concept and Prerequisite Relation in Computer Education(SEKEEO) (S)abstractEffectively assessing the results of users' online learning and enhancing social recognition has become a major development direction for online education platforms.For computer education, this article constructs a technical capability assessment model.This model integrates professional concepts in the field of computer science and extracts knowledge concepts from educational resources.The model first extracts candidate concepts, then uses a graph propagation algorithm to quantify candidate concepts and obtains concepts from them, and finally uses prerequisite relationships to further quantify the concepts mastered by students.The model combines the prerequisite relationship among concepts to quantify the skills that students have mastered.It can not only effectively evaluate the user's skill mastery but also lays a foundation for subsequent course recommendations and career recommendations for users.The model is tested in the real learning environment of 250 students.This model has been proved to own certain practicability and reliability by Kendall rank correlation coefficient, which is used as an evaluation index. Jiwen Luo, Tao Wang 0006, Junsheng Chang |
SEKE | 2 |
| 2021 | Why API documentation is insufficient for developers: an empirical study
Yue Yu 0001, Tao Wang 0006, Gang Yin, Huaimin Wang 0001 |
Sci. China Inf. Sci. | 3 |
| 2021 | Dual Channel Among Task and Contribution on OSS Communities: An Empirical StudyabstractOpen Source Software (OSS) community has attracted a large number of distributed developers to work together, e.g. reporting and discussing issues as well as submitting and reviewing code. OSS developers create links among development units (e.g. issues and pull requests in GitHub), share their opinions and promote the resolution of development units. Although previous work has examined the role of links in recommending high-priority tasks and reducing resource waste, the understanding of the actual usage of links in practice is still limited. To address the research gap, we conduct an empirical study based on the 5W1H model and data mining from five popular OSS projects on GitHub. We find that links originating from a PR are more common than the other three types of links, and links are more frequently created in Documentation. We also find that average duration between development units’ create time in a link is half a year. We observed that link behaviors are very complex and the duration of link increases with the complexity of link structure. We also observe that the reasons of link are very different, especially in P–P and I–I. Finally, future works are discussed in conclusion. Yue Yu 0001, Tao Wang 0006 |
Int. J. Softw. Eng. Knowl. Eng. | 3 |
| 2021 | Detecting Duplicate Contributions in Pull-Based Model Combining Textual and Change Similarities
Yue Yu 0001, Tao Wang 0006, Gang Yin, Xinjun Mao, Huaimin Wang 0001 |
J. Comput. Sci. Technol. | 3 |
| 2020 | Sia-RAE: A Siamese Network based on Recursive AutoEncoder for Effective Clone DetectionabstractCode clone helps improving programming productivity, while at the same time leads to many negative effects on software maintenance. Many approaches have been proposed to detect clones, but most of them fail on detecting low similarity code snippets. In this paper, we propose a Siamese network which links two recursive autoencoders (RAE) with a comparator network for clone detection. The unweighted recursive autoencoder is designed to learn code representation and then the comparator network is employed for similarity evaluation. In this Siamese network, it takes full advantages of lexical, semantic and structure information, and achieves high accuracy in revealing tiny similarity. We conduct comprehensive experiments on BigCloneBench using tagged clones as well as the whole repository respectively. The results suggest that our approach achieves good accuracy, and its recall reaches 93.02 % in WT3/T4, which outperforms state-of-the-art. Chenhui Feng, Tao Wang 0006, Yue Yu 0001, Yang Zhang 0026, Huaimin Wang 0001 |
APSEC | 2 |
| 2020 | Dockerfile Changes in Practice: A Large-Scale Empirical Study of 4, 110 Projects on GitHubabstractDocker is one of the most popular containerization tools in current DevOps practice. Particularly, Dockerfile plays an important role in the Docker-based software development process by specifying the commands and build environment of Docker containers. As a project progresses through its development stages, the content of the Dockerfile may be revised many times. Previous studies have examined Dockerfile usage in open-source projects. However, little is known about the details of Dockerfile changes in practice. In this paper, we conduct an empirical study on Dockerfile changes for 4,110 open-source projects hosted on GitHub. Based on the Dockerfile data, we measure the frequency, magnitude, and instructions of Dockerfile changes and report how Dockerfile co-changed with other files. To explore the relationship between Dockerfile changes and project outcomes, i.e., popularity, success, and productivity, we also develop regression models, by controlling for various confounds. Our findings help to characterize and understand Dockerfile changes and motivate the need for collecting more empirical evidence. Yiwen Wu 0001, Yang Zhang 0026, Tao Wang 0006, Huaimin Wang 0001 |
APSEC | 3 |
| 2020 | Using Configuration Semantic Features and Machine Learning Algorithms to Predict Build Result in Cloud-Based Container EnvironmentabstractContainer technologies are being widely used in large scale production cloud environments, of which Docker has become the de-facto industry standard. In practice, Docker builds often break, and a large amount of efforts are put into troubleshooting broken builds. Prior studies have evaluated the rate at which builds in large organizations fail. However, there is still a lack of early warning methods for predicting the Docker build result before the build starts. This paper provides a first attempt to propose an automatic method named PDBR. It aims to use the configuration semantic features extracted by AST and the machine learning algorithms to predict build result in the cloud-based container environment. The evaluation experiments based on more than 36,000 collected Docker builds show that PDBR achieves 73.45%-91.92% in F1 and 29.72%-72.16% in AUC. We also demonstrate that different ML classifiers have significant and large effects on the PDBR AUC performance. Yiwen Wu 0001, Yang Zhang 0026, Junsheng Chang, Bo Ding 0001, Tao Wang 0006, Huaimin Wang 0001 |
ICPADS | 5 |
| 2020 | Haste Makes Waste: An Empirical Study of Fast Answers in Stack OverflowabstractModern programming question & answer (Q&A) sites such as Stack Overflow (SO) employ gamified mechanisms to stimulate volunteers' contributions. To maximize the chances of winning gamification rewards such as reputation and badges, a portion of users race to post answers as quickly as possible (i.e., fast answers or FAs), which makes SO the fastest Q&A site; however, this behavior may affect the contribution quality as well. In this paper, we report on a large-scale, mixed-methods empirical study of the gamification-influenced FA phenomenon in SO. We first quantitatively investigate the popularity of the phenomenon and user behaviors regarding FAs. Then, we study the quality of FAs by using regression modeling and qualitatively analyzing 300 instances of FAs. Our main findings reveal that more than 70% and 90% of FAs are not edited by the answerers and other users, respectively, and that later incoming answers have lower chances of being voted on and accepted. Notably, we find that the answer length, code snippets length, and readability of FAs are significantly lower than those of non-fast answers. Although FAs have higher crowd assessment scores, they have no relationship with acceptance from the perspective of asker assessment, and a considerable portion of FAs solve the problem by interacting with the asker in the comments. These results help us better understand the effects of reward-based gamification on crowdsourced software engineering communitites and provide implications for designers of gamified systems. Yao Lu 0003, Xinjun Mao, Minghui Zhou 0001, Yang Zhang 0026, Tao Wang 0006, Zude Li |
ICSME | 5 |
| 2020 | An Empirical Study of Multi-discussing Pattern in Open-Source Software DevelopmentabstractGitHub enables developers to expediently contribute their comments on multiple issues and switch their discussion between issues, i.e., multi-discussing. Discussing multiple issues simultaneously is able to enhance work efficiency. However, multi-discussing also relies on developers’ rationally allocating their focus, which may result in the different influence on the resolution of issues. Therefore, investigating how multi-discussing affects the issue resolution is a meaningful research question that can help developers understand the benefits and limitations of multi-discussing. Using quantitative and qualitative methods, this paper proposes a groundbreaking study of the impact of multi-discussing on issue resolution in GitHub. First, we collect and analyze data from 624 GitHub projects to explore how multi-discussing affects the overall issue resolution of the project. Further, we investigate how multi-discussing affects the resolution of a single issue. We find that multi-discussing is a common behavior in GitHub. Also, multi-discussing is connected to a shorter average issue resolution latency of the project. However, during a single issue resolution, more multi-discussing behaviors tend to bring longer issue resolution latency. We also conduct the qualitative analysis to explore the developers’ experiences and expectations of multi-discussing. Cheng Yang 0004, Dongyang Hu, Yang Zhang 0026, Tao Wang 0006, Yue Yu 0001 |
Internetware | 4 |
| 2020 | Exploring the Dependency Network of Docker Containers: Structure, Diversity, and RelationshipabstractContainer technologies are being widely used in large scale production cloud environments, of which Docker has become the de-facto industry standard. As a key step, containers need to define their dependent base image, which makes complex dependencies exist in a large number of containers. Prior studies have shown that references between software packages could form technical dependencies, thus forming a dependency network. However, little is known about the details of docker container dependency networks. In this paper, we perform an empirical study on the dependency network of docker containers from more than 120,000 dockerfiles. We construct the container dependency network and analyze its network structure. Further, we focus on the Top-100 dominant containers and investigate their subnetworks, including diversity and relationships. Our findings help to characterize and understand the container dependencies in the docker community and motivate the need for developing container dependency management tools. Yinyuan Zhang, Yang Zhang 0026, Yiwen Wu 0001, Yao Lu 0003, Tao Wang 0006, Xinjun Mao |
Internetware | 5 |
| 2020 | An Empirical Study of Build Failures in the Docker ContextabstractDocker containers have become the de-facto industry standard. Docker builds often break, and a large amount of efforts are put into troubleshooting broken builds. Prior studies have evaluated the rate at which builds in large organizations fail. However, little is known about the frequency and fix effort of failures that occur in Docker builds of open-source projects. This paper provides a first attempt to present a preliminary study on 857,086 Docker builds from 3,828 open-source projects hosted on GitHub. Using the Docker build data, we measure the frequency of broken builds and report their fix time. Furthermore, we explore the evolution of Docker build failures across time. Our findings help to characterize and understand Docker build failures and motivate the need for collecting more empirical evidence. Yiwen Wu 0001, Yang Zhang 0026, Tao Wang 0006, Huaimin Wang 0001 |
MSR | 3 |
| 2020 | Improving students' programming quality with the continuous inspection process: a social coding perspective
Yao Lu 0003, Xinjun Mao, Tao Wang 0006, Gang Yin, Zude Li |
Frontiers Comput. Sci. | 3 |
| 2020 | GitHub's milestone tool: A mixed-methods analysis on its useabstractAbstract Social coding site GitHub provides developers with many management tools to facilitate project maintenance and developer collaboration.Milestonetool, in particular, plays an important role in organizing and tracking progress on groups of issues or pull requests in a project. However, few research has analyzed the milestone tool, even though it has been used in practice for a long time. In this paper, we want to address this literature gap and present an ongoing work aimed at investigating the use of the milestone tool in GitHub open‐source projects. We conduct a mixed‐methods analysis in a large‐scale dataset of GitHub projects, to help developers gain some insights into the milestone tool, including its usage, benefits, and limitations. We quantitatively investigate the basic adoption of milestone tool and its correlation with project properties. We also survey developers to understand the reasons for using milestone tool or not and their perceptions of the milestone tool. We find that certain types of projects use milestone tool more than others. Adopting the milestone tool is associated with more commits, more releases, and more project popularity, but the current milestone tool also has some limitations. These observations can then be forwarded to the GitHub community for follow‐up and can result in them potentially making a better milestone tool. Yang Zhang 0026, Huaimin Wang 0001, Yiwen Wu 0001, Dongyang Hu, Tao Wang 0006 |
J. Softw. Evol. Process. | 5 |
| 2020 | iLinker: a novel approach for issue knowledge acquisition in GitHub projects
Yang Zhang 0026, Yiwen Wu 0001, Tao Wang 0006, Huaimin Wang 0001 |
World Wide Web | 3 |
| 2019 | BBCPS: A Blockchain Based Open Source Contribution Protection System
Qiubing Zeng, Xunhui Zhang, Tao Wang 0006, Peichang Shi, Xiang Fu 0002, Chenhui Feng |
BlockSys | 3 |
| 2019 | A Neural-Network based Code Summarization Approach by Using Source Code and its Call DependenciesabstractCode summarization aims at generating natural language abstraction for source code, and it can be of great help for program comprehension and software maintenance. The current code summarization approaches have made progress with neural-network. However, most of these methods focus on learning the semantic and syntax of source code snippets, ignoring the dependency of codes. In this paper, we propose a novel method based on neural-network model using the knowledge of the call dependency between source code and its related codes. We extract call dependencies from the source code, transform it as a token sequence of method names, and leverage the Seq2Seq model for code summarization using the combination of source code and call dependency information. About 100,000 code data is collected from 1,000 open source Java proejects on github for experiment. The large-scale code experiment shows that by considering not only the code itself but also the codes it called, the code summarization model can be improved with the BLEU score to 33.08. Bohong Liu, Tao Wang 0006, Xunhui Zhang, Gang Yin, Jinsheng Deng |
Internetware | 2 |
| 2019 | Exploring the Relationship Between Developer Activities and Profile Images on GitHubabstractIn the GitHub platform, social media profile images are one of many visual components of developers. Besides, developer activities such as reporting issues or following other developers are regarded as important development and self-expression behaviors. However, to the best of our knowledge, no study has yet been conducted to study the relationship between GitHub developer activities and profile images. In this paper, we aim to investigate the relationship between developer activities and profile images to gain some insights into the developers' internal properties. During our experiments, we manually classify profile images into seven categories. Next, we investigate the relationship between developer's demographic information and developer activity. Further, using logistic regression analysis, when controlled for various variables, we statistically identify and quantify the relationships between developer activities and profile image categories. We find that several profile image categories significantly correlate with developer's demographic information and activities. We also provide a rich resource of research ideas for further study. Our examination and analysis provide insights into the developers' internal properties when using different profile images. Moreover, this study is the first step in understanding the relationship between developer activities and profile images on GitHub. Yiwen Wu 0001, Yang Zhang 0026, Tao Wang 0006, Huaimin Wang 0001 |
Internetware | 3 |
| 2019 | A novel approach for recommending semantically linkable issues in GitHub projects
Yang Zhang 0026, Yiwen Wu 0001, Tao Wang 0006, Huaimin Wang 0001 |
Sci. China Inf. Sci. | 3 |
| 2019 | Multi-reviewing pull-requests: An exploratory study on GitHub OSS projects
Dongyang Hu, Yang Zhang 0026, Junsheng Chang, Gang Yin, Yue Yu 0001, Tao Wang 0006 |
Inf. Softw. Technol. | 6 |
| 2019 | RepoLike: amulti-feature-based personalized recommendation approach for open-source repositoriesabstractWith the deep integration of software collaborative development and social networking, social coding represents a new style of software production and creation paradigm. Because of their good flexibility and openness, a large number of external contributors have been attracted to the open-source communities. They are playing a significant role in open-source development. However, the open-source development online is a globalized and distributed cooperative work. If left unsupervised, the contribution process may result in inefficiency. It takes contributors a lot of time to find suitable projects or tasks from thousands of open-source projects in the communities to work on. In this paper, we propose a new approach called “RepoLike,” to recommend repositories for developers based on linear combination and learning to rank. It uses the project popularity, technical dependencies among projects, and social connections among developers to measure the correlations between a developer and the given projects. Experimental results show that our approach can achieve over 25% of hit ratio when recommending 20 candidates, meaning that it can recommend closely correlated repositories to social developers. Cheng Yang 0004, Tao Wang 0006, Gang Yin, Xunhui Zhang, Yue Yu 0001, Huaimin Wang 0001 |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2018 | Recommending Similar Bug Reports: A Novel Approach Using Document Embedding ModelabstractIn the software development, it is not uncommon to find that several bug reports are related to many common code files, i.e., similar bugs. Similar bug recommendation is a meaningful task which can assist developers in bug triaging and fixing. As the state of the art, Yang et al.'s work presented an approach that combines TF-IDF method with word embedding model and achieved a good result. To further improve the performance of their approach, in this paper, we propose a novel approach using Document Embedding model. In our preliminary evaluation, we conduct the experiment on 13,090 bug reports from the Eclipse platform and the results show that our approach outperforms Yang et al.'s, with 7.89-8.96% of improvement. Dongyang Hu, Tao Wang 0006, Junsheng Chang, Gang Yin, Yue Yu 0001, Yang Zhang 0026 |
APSEC | 3 |
| 2018 | Multi-Discussing across Issues in GitHub: A Preliminary StudyabstractSocial coding sites like GitHub has enabled developers to easily contribute their comments on multiple issues and switch their discussion between issues, i.e., multi-discussing. Discussing multiple issues simultaneously may enhance the work efficiency of developers. However, multi-discussing also relies on developers' rationally allocating their time and focus, which may bring different influence to the resolution of issues. Therefore, investigating how multi-discussing affects the issue resolution is a meaningful research question which can help developers understand the benefits and limitations when they switch their discussion between issues. In this paper, we present a preliminary study of the impact of multi-discussing on issue resolution in GitHub projects, by using quantitative methods. First, we collect and analyzed data from 631 GitHub projects to explore how multi-discussing affects the average resolution latency of project issues. Further, we develop method for measuring the rate and breadth of a developers' discussionswitching behavior, and we use regression modeling to study how discussion-switching affects the single issue resolution latency. We find that multi-discussing is a common behavior of developers in GitHub projects. Also, multi-discussing is associated with shorter average issue resolution latency of project. However, during a single issue resolution, more participants' discussion-switching tend to bring longer issue resolution latency. Our study motivates the need for further research on the multi-discussing. Dongyang Hu, Tao Wang 0006, Junsheng Chang, Gang Yin, Yang Zhang 0026 |
APSEC | 2 |
| 2018 | An Insight Into the Impact of Dockerfile Evolutionary Trajectories on Quality and LatencyabstractContainerization is a software development approach aimed at packaging an application together with all its dependencies and execution environment in a light-weight, self-contained unit, of which Docker has become the de-facto industry standard. By defining the specific Docker image architecture and building orders, dockerfile plays an important role in the Docker-based containerization process. Understanding the evolution of dockerfile and which dockerfile architecture attributes enhance dockerfile quality and reduce image build latency can benefit the efficient processing of containerization. In this paper, we perform an empirical study on a large dataset of 2,840 projects to shed light on the impact of dockerfile evolutionary trajectories on quality and latency in the Docker-based containerization. Based on the six categories of dockerfile evolutionary trajectories we discovered, we build two regression models to explore the impact of dockerfile evolutionary trajectories and specific architecture attributes on dockerfile quality and image build latency, which derives a number of suggestions for practitioners. Yang Zhang 0026, Gang Yin, Tao Wang 0006, Yue Yu 0001, Huaimin Wang 0001 |
COMPSAC (1) | 3 |
| 2018 | Cross-Project Issue Classification Based on Ensemble Modeling in a Social Coding World
Yarong Zeng, Yue Yu 0001, Xunhui Zhang, Tao Wang 0006, Gang Yin, Huaimin Wang 0001 |
ICONIP (4) | 5 |
| 2018 | A Hybrid Approach for Tag Hierarchy Construction
Shangwen Wang, Tao Wang 0006, Xiaoguang Mao, Gang Yin, Yue Yu 0001 |
ICSR | 2 |
| 2018 | Who Will Become a Long-Term Contributor?: A Prediction Model based on the Early Phase BehaviorsabstractThe continuous contribution from peripheral participants is crucial for the success of open source projects. Thus, how to identify the potential Long-Term Contributors (LTC) early and retain them is of great importance. We propose a prediction model to measure the chance for an individual to become a LTC contributor through his capacity, willingness, and the opportunity to contribute at the time of joining. Using data of Rails hosted on GitHub, we find that the probability for a new joiner to become a LTC is associated with his willingness and environment. Specifically, future LTCs tend to be more active and show more community-oriented attitude than other joiners during their first month. This implies that the interaction between individual's attitude and project's climate are associated with the odds that an individual would become a valuable contributor or disengage from the project. We evaluated our prediction model by using the 10 cross-validation method. Results show that our model archives the mean AUC as 0.807, which is valuable for OSS projects to identify potential long-term contributors and adopt better strategies to retain them for continuous contribution. Tao Wang 0006, Yang Zhang 0026, Gang Yin, Yue Yu 0001, Huaimin Wang 0001 |
Internetware | 1 |
| 2018 | A dataset of duplicate pull-requests in githubabstractIn GitHub, the pull-based development model enables community contributors to collaborate in a more efficient way. However, the distributed and parallel characteristics of this model pose a potential risk for developers to submit duplicate pull-requests (PRs), which increase the extra cost of project maintenance. To facilitate the further studies to better understand and solve the issues introduced by duplicate PRs, we construct a large dataset of historical duplicate PRs extracted from 26 popular open source projects in GitHub by using a semi-automatic approach. Furthermore, we present some preliminary applications to illustrate how further researches can be conducted based on this dataset. Yue Yu 0001, Gang Yin, Tao Wang 0006, Huaimin Wang 0001 |
MSR | 4 |
| 2018 | Adaptive software search toward users' customized requirements in GitHubabstractBecause of a tremendous growth of Open Source Software (OSS) scale and the diversity of users' requirements, users now face the problem of finding OSS that meets their expectations in a huge number of OSS resources.However, current GitHub-provided search service has a shortage in adapting to user needs.When facing diverse users' requirements, it cannot always return satisfactory results.In this paper, we provide a more efficient search service for OSS on GitHub.We first design a multi-dimensional measurement model for OSS, which forms a corresponding metric system and quantitative measurement method.Then we propose a ranking algorithm based on fuzzy synthetic evaluation in order to implement an adaptive metric ranking method that is oriented to user requirements.We verify that our work is useful by setting up experiments.The experiment results show that compared with GitHub-provided search service (searching by "Best Match" & searching by "Most Stars"), the effectiveness of our method improved by 97.6% and 13.8% respectively, which means our method returns search results which meet users' expectations more, and has high self-adaptive ability. Jinze Liu, Tao Wang 0006, Yue Yu 0001, Gang Yin |
SEKE | 3 |
| 2018 | Improving code summarization by combining deep learning and empirical knowledge (S)abstractCode summaries are human-readable text that describes the functionality of code blocks.Software developers use code summaries to understand the specification of API while code retrieve system relies on code summaries for effective code search.However, code summaries are often written by software developers.Writing good code summaries usually requires great effort.It could be helpful if developers use automatic code summarization system to generate code summaries.Recently, some works have applied deep learning methods to generate code summaries for code snippets.However, those deep learning methods treat code snippets as streams of text tokens while ignoring the inherent code structure information.In this paper, we propose a novel code summarization method named the CDE-Model (Code summarization by Deep learning and Empirical knowledge) that combines inherent code structure information with deep learning models.The CDE-Model proposes several empirical strategies to transform code snippets to refined code representation and feeds them into an encoder-decoder neural network for text generation.We conduct large-scale experiments on 1500 popular Java projects on GitHub 1 with 396,184 pairs of code snippets and summaries.Experimental results show that the quality of code summaries generated by our CDE-Model is better than other two methods.To the best of our knowledge, this paper is the first to combine code structure information with deep learning. Lingbin Zeng, Xunhui Zhang, Tao Wang 0006, Xiao Li 0039, Huaimin Wang 0001 |
SEKE | 3 |
| 2018 | Correlation-based software search by leveraging software term database
Gang Yin, Tao Wang 0006, Yang Zhang 0026, Yue Yu 0001, Huaimin Wang 0001 |
Frontiers Comput. Sci. | 3 |
| 2018 | Internal quality assurance for external contributions in GitHub: An empirical investigationabstractAbstract For popular open‐source software projects, there are always a large number of worldwide developers who have been glued to making code contributions, while most of these developers play the role of casual contributors because of their very limited code commits. The frequent turnover of such a group of developers and the wide variations in their coding experiences challenge the project management on code and quality. This paper aims to investigate the status quo of internal quality assurance for external contributions in social coding sites. We first conducted a case study of 21 popular GitHub projects to estimate the code quality of the casual contributors. The quantitative results show that the casual contributors introduced greater quantity and severity of code quality issues than the main contributors; the developers who contribute to different projects as main and casual contributors did not perform significantly differently in terms of their code quality. On the basis of these findings, we further conducted a survey of 81 developers on GitHub to understand their practices on internal quality assurance. The qualitative results expose some limitations of present internal quality control for external contributions in GitHub. Finally, we discuss an alternative quality management paradigm: Continuous Inspection for industrial practices. Yao Lu 0003, Xinjun Mao, Zude Li, Yang Zhang 0026, Tao Wang 0006, Gang Yin |
J. Softw. Evol. Process. | 5 |
| 2018 | Linking Issue Tracker with Q&A Sites for Knowledge Sharing across CommunitiesabstractCollaborative development communities and knowledge sharing communities are highly correlated and mutually complementary. The knowledge sharing between these two types of open source communities can be very beneficial to both of them. However, it is a great challenge to automate this process. Current studies mainly focus on knowledge acquisition in one type of community, and few of them have tackle this problem efficiently. In this paper we take Android Issue Tracker and Stack Overflow as a case to study the mutual knowledge sharing between them. We propose an automatic approach by integrating semantic similarity with temporal locality between Android issues and Stack Overflow posts based on the internal citation-graph to reveal the potential associations between them. Our approach explores the internal citations in communities for closely related posts or issues clustering, exploits the rich semantics in fine-grained information of issues and posts for associations building, and leverages the temporal correlations between issues and posts in-depth for associations ranking. Extensive experiments show that the precision of our approach reaches 62.51 percent for top 10 recommendations when recommending Stack Overflow posts to Android issues, and 66.83 percent in reverse. Huaimin Wang 0001, Tao Wang 0006, Gang Yin, Cheng Yang 0004 |
IEEE Trans. Serv. Comput. | 2 |
| 2017 | Where Is the Road for Issue Reports Classification Based on Text Mining?abstractCurrently, open source projects receive various kinds of issues daily, because of the extreme openness of Issue Tracking System (ITS) in GitHub. ITS is a labor-intensive and time-consuming task of issue categorization for project managers. However, a contributor is only required a short textual abstract to report an issue in GitHub. Thus, most traditional classification approaches based on detailed and structured data (e.g., priority, severity, software version and so on) are difficult to adopt. In this paper, issue classification approaches on a large-scale dataset, including 80 popular projects and over 252,000 issue reports collected from GitHub, were investigated. First, four traditional text-based classification methods and their performances were discussed. Semantic perplexity (i.e., an issues description confuses bug-related sentences with nonbug-related sentences) is a crucial factor that affects the classification performances based on quantitative and qualitative study. Finally, A two-stage classifier framework based on the novel metrics of semantic perplexity of issue reports was designed. Results show that our two-stage classification can significantly improve issue classification performances. Yue Yu 0001, Gang Yin, Tao Wang 0006, Huaimin Wang 0001 |
ESEM | 4 |
| 2017 | DevRec: A Developer Recommendation System for Open Source Repositories
Xunhui Zhang, Tao Wang 0006, Gang Yin, Cheng Yang 0004, Yue Yu 0001, Huaimin Wang 0001 |
ICSR | 2 |
| 2017 | Research on 220kV GIS enclosure circulation and temporary ground potential riseabstractEnclosure circulation of Gas Insulated Switchgear (GIS) and Very Fast Transient Overvoltage (VFTO) are harmful to many respects of power system, including system's security, stability and economic. The effective inhibition of GIS enclosure circulation and VFTO has attracted wide attention. In this paper, GIS enclosure circulation and transient ground potential has been modeled and analyzed. Firstly, this paper discusses the GIS enclosure circulation VFTO, Transient Grounding Potential Rise (TPGR) and the harm, as well as its domestic and foreign researching status. Then, a simplified model has been established, according to which, the steady-state operation of GIS enclosure and connected network circulation theoretical calculation have been investigated according to the GIS shell and grounding grid circulation. The software ATP/EMTP is utilized in the simulation, which takes gas insulated equipment shell connection mode, shell grounding layout and grounding and the grounding grid structure into account. According to the results of theoretical calculation and simulation, the comprehensive optimization of the enclosure circulation and transient potential can be effectively improved and some prevention measures also can be proposed from the deeper understanding of grounding line number and distribution of circulation inhibition. Theoretical results can provide significant reference for the power grid design, operation and protection. Xuebin Lv, Xuefeng Sun, Tao Wang 0006, Shikun Wang, Yin Pei, Dongyang Hu, Hongshun Liu |
IECON | 4 |
| 2017 | Detecting Duplicate Pull-requests in GitHubabstractThe widespread use of pull-requests boosts the development and evolution for many open source software projects. However, due to the parallel and uncoordinated nature of development process in GitHub, duplicate pull-requests may be submitted by different contributors to solve the same problem. Duplicate pull-requests increase the maintenance cost of GitHub, result in the waste of time spent on the redundant effort of code review, and even frustrate developers' willing to offer continuous contribution. In this paper, we investigate using text information to automatically detect duplicate pull-requests in GitHub. For a new-arriving pull-request, we compare the textual similarity between it and other existing pull-requests, and then return a candidate list of the most similar ones. We evaluate our approach on three popular projects hosted in GitHub, namely Rails, Elasticsearch and Angular.JS. The evaluation shows that about 55.3% -- 71.0% of the duplicates can be found when we use the combination of title similarity and description similarity. Gang Yin, Yue Yu 0001, Tao Wang 0006, Huaimin Wang 0001 |
Internetware | 4 |
| 2017 | Automatic Classification of Review Comments in Pull-based Development ModelabstractThe pull-based model, widely used in distributed software development, allows any contributor to fork a public repository, package contributions as a pull-request, and then merge back to the original repository.Code review is one of the most significant stages in pull-based development.It ensures that only high-quality pull-requests are accepted, based on the in-depth discussion among reviewers.Thus, automatically identifying what reviewers are talking about in the discussions is benificial to better understand the code review process.In this paper, we conduct a case study on three popular opensource software projects hosted on GitHub and construct a finegrained taxonomy including 11 sub-categories for review comments.We then manually label over 5,600 review comments, and propose a Two-Stage Hybrid Classification (TSHC) algorithm to classify review comments automatically by combining rule-based and machine-learning techniques.Comparative experiments with a text-based method achieve a reasonable improvement on each project (9.2% in Rails, 5.3% in Elasticsearch, and 7.2% in Angular.jsrespectively) in terms of the weighted average Fmeasure. Yue Yu 0001, Gang Yin, Tao Wang 0006, Huaimin Wang 0001 |
SEKE | 4 |
| 2017 | Who Will be Interested in? A Contributor Recommendation Approach for Open Source ProjectsabstractThe crowds' continuous participation and contribution are the key factors for the success of open source projects.However, among the massive competitors, it is difficult for a project to attract enough contributors by just passively waiting for enthusiasts to join in.Instead, it should actively seek gifted developers.Most of the current studies mainly focus on recommending experts inside a repository for some specific development tasks.In this paper, we propose a novel approach ConRec to recommend potential contributors across the entire open source community for given projects.It leverages the developers' historical activities in projects to analyze their technical interests and technical connections with others.Thereafter, it combines collaborative filtering algorithm with text matching algorithm to recommend proper developers.We conducted extensive experiments on 5,995 open source projects and 2,938,620 developers in GitHub.The results show that the proposed algorithm can recommend contributors to open source projects with the best performance of 63% in accuracy, and solve the cold start problem as well. Xunhui Zhang, Tao Wang 0006, Gang Yin, Cheng Yang 0004, Huaimin Wang 0001 |
SEKE | 2 |
| 2017 | Social media in GitHub: the role of @-mention in assisting software development
Yang Zhang 0026, Huaimin Wang 0001, Gang Yin, Tao Wang 0006, Yue Yu 0001 |
Sci. China Inf. Sci. | 4 |
| 2017 | What Are They Talking About? Analyzing Code Reviews in Pull-Based Development Model
Yue Yu 0001, Gang Yin, Tao Wang 0006, Huaimin Wang 0001 |
J. Comput. Sci. Technol. | 4 |
| 2016 | Does the Role Matter? An Investigation of the Code Quality of Casual Contributors in GitHubabstractFor popular Open Source Software (OSS) projects there are always a large number of worldwide developers who have been glued to making code contributions, while most of these developers play the role of casual contributors due to their very limited code commits (for fixing defects and enhancing features, casually). The frequent turnover of such group of casual developers and the wide variations among their coding experiences challenge the project management on code and quality.This paper describes a case study which aims to estimate the quality of code made by casual contributors in 21 popular GitHub projects. The results of this case study show that: (1) casual contributors introduced greater quantity and severity of Code Quality Issues (CQIs) than main contributors; (2) developers who contribute in different projects as main and casual contributors didn't perform statistically differently in terms of code quality; (3) casual contributors who have few project stars introduced more CQIs than those who have many. Furthermore, the paper lists the CQI categories which are most frequently introduced by casual contributors in the investigated projects. These findings provide valuable insights into code quality in the OSS context, and can guide OSS developers in improving the quality of the code contributions. Yao Lu 0003, Xinjun Mao, Zude Li, Yang Zhang 0026, Tao Wang 0006, Gang Yin |
APSEC | 5 |
| 2016 | Query reformulation by leveraging crowd wisdom for scenario-based software searchabstractThe Internet-scale open source software (OSS) production in various communities are generating abundant reusable resources for software developers. However, how to retrieve and reuse the desired and mature software from huge amounts of candidates is a great challenge: there are usually big gaps between the user application contexts (that often used as queries) and the OSS key words (that often used to match the queries). In this paper, we define the scenario-based query problem for OSS retrieval, and then we propose a novel approach to reformulate the raw query by leveraging the crowd wisdom from millions of developers to improve the retrieval results. We build a software-specific domain lexical database based on the knowledge in open source communities, by which we can expand and optimize the input queries. The experiment results show that, our approach can reformulate the initial query effectively and outperforms other existing search engines significantly at finding mature software. Tao Wang 0006, Yang Zhang 0026, Yun Zhan, Gang Yin |
Internetware | 2 |
| 2016 | RepoLike: personal repositories recommendation in social coding communitiesabstractSocial coding represents a new style of software production and creation paradigm, and demands for new technologies of software reuse. Many people searching for projects package, we can provide good reuse recommendation. In this paper, we focus on an interesting research topic of recommending software repositories to social developers, which is challenging because of two points: the first is how to get the interest contexts of developers; and the second is how to rank the repository candidates for recommendation properly. We propose RepoLike, a new approach for recommending repositories to developers by predicting their interests. RepoLike explores the developers' historical development activities and the social connections with other programmers, mines the technical features of repositories and the dependencies among them, and then combines both aspects to recommend most interesting and inspiring repositories to developers. The experiment results show that our approach can surprisingly recommend closely correlated repositories to developers, and the critical test results show that the recommendation performance is strongly impacted by the interest context model. Cheng Yang 0004, Tao Wang 0006, Gang Yin, Huaimin Wang 0001 |
Internetware | 3 |
| 2016 | A Novel Open Source Software Ecosystem: From a Graphic Point of View and Its ApplicationabstractWith the rapid development of open source software, various elements such as OSS, developers, users and online posts, across different communities and their interactions constitute a novel software ecosystem.Most of the current researches about software ecosystems care the connections between software, and few of them consider the relationship across communities from an overall perspective, and fail to cover the users and their activities which should be an indispensable part of the OSS ecosystem.This paper model the OSS ecosystem as a graph, which combines different types of OSS communities as a whole.Based on this graph model, we analyze the characteristics of ecosystem, like the evolution, competition and symbiosis.In addition, we build a recommendation system as well, and the experiment results suggest the validation of our approach. Chenxi Song, Tao Wang 0006, Gang Yin, Xunhui Zhang, Cheng Yang 0004 |
SEKE | 2 |
| 2016 | Determinants of pull-based development in the context of continuous integration
Yue Yu 0001, Gang Yin, Tao Wang 0006, Cheng Yang 0004, Huaimin Wang 0001 |
Sci. China Inf. Sci. | 3 |
| 2016 | Reviewer recommendation for pull-requests in GitHub: What can we learn from code review and bug assignment?
Yue Yu 0001, Huaimin Wang 0001, Gang Yin, Tao Wang 0006 |
Inf. Softw. Technol. | 4 |
| 2015 | A supervised approach for tag hierarchy construction in open source communitiesabstractThe massive amounts of open source software provide sufficient reusable resources for software development. Most of the OSS communities adopt a kind of categorization or tagging mechanism to organize the software. However, the categorization often too coarse, while the tags are flat and fail to capture the inter-relation among them. In this paper, we propose a novel approach to reveal the latent relations between tags and build a tag hierarchy to help locate resources. We firstly build a co-occurrence network, based on which we compare the connotations of tags and construct a preliminary hierarchy. Then we leverage the domain knowledge of category in SourceForge to optimize and improve the relations between tags. At the end, we demonstrate the effectiveness of the constructed tag hierarchy with quantitative evaluation, which suggest the validation of our approach. Chongming Gu, Gang Yin, Tao Wang 0006, Cheng Yang 0004, Huaimin Wang 0001 |
Internetware | 3 |
| 2015 | Software Ranking and Analysis based on Mining Market Requirements and CharacteristicsabstractAs the rapid growth of open source software, how to choose software from many alternatives becomes a great challenge. Traditional ranking approaches mainly focus on the characteristics of the software themselves, such as qualities, security, reliable and so on. In this paper we investigate the market demands for software engineers, and propose a novel approach for ranking software by analyzing the market requirements for special software. At the same time we conclude the characteristics of software advertisements and analyze the reasons that why these situations emerge and tendency of software market requirements. As industries always need to balance several different factors for selecting software, the market demands can be a good indicator for ranking software and software evaluating. This paper provides quite a different perspective and some interesting inferences on software market requirements, and it can be a valuable supplement for traditional ranking methods, as well as software evaluating. Bingxun Liu, Gang Yin, Tao Wang 0006, Huaimin Wang 0001 |
Internetware | 3 |
| 2015 | Exploring the Use of @-mention to Assist Software Development in GitHubabstractRecently, many researches propose that social media tools can promote the collaboration among developers, which are beneficial to the software development. Nevertheless, there is little empirical evidence to confirm that using @-mention has indeed a beneficial impact on the issues in GitHub. In this paper, we analyze the data from GitHub and give some insights on how @-mention is used in the issues (general-issues and pull-requests). Our statistical results indicate that, @-mention attracts more participants and tends to be used in the difficult issues. @-mention favors the solving process of issues by enlarging the visibility of issues and facilitating the developers' collaboration. In addition to this global study, our study also build a @-network based on the @-mention database we extract. Through the @-network, we can mine the relationships and characteristics of developers in GitHub's issues. Yang Zhang 0026, Huaimin Wang 0001, Gang Yin, Tao Wang 0006, Yue Yu 0001 |
Internetware | 4 |
| 2015 | Evaluating Bug Severity Using Crowd-based Knowledge: An Exploratory StudyabstractIn bug tracking system, the high volume of incoming bug reports poses a serious challenge to project managers. Triaging these bug reports manually consumes time and resources which leads to delaying the resolution of important bugs. StackOverflow is the most popular crowdsourcing Q&A community with plenty of bug-related posts. In this paper, we explore the correlation between bug severity and the crowd attributes of linked posts. Two typical types of projects' bug repositories are studied here, e.g. Mozilla (user-centric project) and Eclipse (developer-centric project). Our results show that the bug severity is consistent with the crowd-based knowledge both in Mozilla and Eclipse, i.e. the linked posts of severe bugs have higher score etc. in StackOverflow than non-severe bugs. This interesting phenomenon inspires us that we can optimize the existing evaluation methods of bug severity by incorporating the crowd-based knowledge from a third-party in future. Yang Zhang 0026, Gang Yin, Tao Wang 0006, Yue Yu 0001, Huaimin Wang 0001 |
Internetware | 3 |
| 2014 | Linking stack overflow to issue tracker for issue resolutionabstractIssue resolution is a central task for software development. The efficiency of issue fixing largely relies on the issue report quality. Stack Overflow, which hosts rich and real-time posts about programming-specific problems, is a valuable external source of knowledge for issue resolution. In this paper, we present CrossLink, an analysis framework that automatically introduce related posts in Stack Overflow to issues in Android Issue Tracker. This helps developers to leverage the abundant and professional knowledge from massive programmers in Stack Overflow for issue resolution. CrossLink explores the semantic similarities as well as the temporal associations between the two types of repositories to recommend Stack Overflow posts to Android issues. The internal links in Stack Overflow are also employed to improve the linking accuracy. The experiments prove the effectiveness of CrossLink with precision of 62.51% for top-10 recommendations, which is significantly higher than the state-of-art method. Tao Wang 0006, Gang Yin, Huaimin Wang 0001, Cheng Yang 0004 |
Internetware | 1 |
| 2014 | Tag recommendation for open source software
Tao Wang 0006, Huaimin Wang 0001, Gang Yin, Charles Ling 0001, Xiao Li 0039 |
Frontiers Comput. Sci. | 1 |
| 2013 | Mining Software Profile across Multiple Repositories for Hierarchical CategorizationabstractThe large amounts of software repositories over the Internet are fundamentally changing the traditional paradigms of software maintenance. Efficient categorization of the massive projects for retrieving the relevant software in these repositories is of vital importance for Internet-based maintenance tasks such as solution searching, best practices learning and so on. Many previous works have been conducted on software categorization by mining source code or byte code, which are only verified on relatively small collections of projects with coarse-grained categories or clusters. However, Internet-based software maintenance requires finer-grained, more scalable and language-independent categorization approaches. In this paper, we propose a novel approach to hierarchically categorize software projects based on their online profiles across multiple repositories. We design a SVM-based categorization framework to classify the massive number of software hierarchically. To improve the categorization performance, we aggregate different types of profile attributes from multiple repositories and design a weighted combination strategy which assigns greater weights to more important attributes. Extensive experiments are carried out on more than 18,000 projects across three repositories. The results show that our approach achieves significant improvements by using weighted combination, and the overall precision, recall and F-Measure can reach 71.41%, 65.60% and 68.38% in appropriate settings. Compared to the previous work, our approach presents competitive results with 123 finer-grained and multi-layered categories. In contrast to those using source code or byte code, our approach is more effective for large-scale and language-independent software categorization. Tao Wang 0006, Huaimin Wang 0001, Gang Yin, Charles Ling 0001, Xiang Li 0012 |
ICSM | 1 |
| 2012 | Inducing Taxonomy from Tags: An Agglomerative Hierarchical Clustering Framework
Xiang Li 0012, Huaimin Wang 0001, Gang Yin, Tao Wang 0006, Cheng Yang 0004, Yue Yu 0001, Dengqing Tang |
ADMA | 4 |