VLDB 2026 Research / reviewers in the wild / expert
Gang Yin
dblp:14/3946
· DBLP profile ↗
63ranked-venue papers
5as first author
6since 2021 · last 2022
0000-0002-0839-242XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 37 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 8 · 1 first-authorDatabases, data management, data science and information retrieval · 3Systems, architecture and hardware · 2 · 1 first-authorSecurity and privacy · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | HAF: a hybrid annotation framework based on expert knowledge and learning technique
Yue Yu 0001, Tao Wang 0006, Gang Yin, Xinjun Mao, Huaimin Wang 0001 |
Sci. China Inf. Sci. | 4 |
| 2022 | Are You Still Working on This? An Empirical Study on Pull Request AbandonmentabstractThe great success of numerous community-based open source software (OSS) is based on volunteers continuously submitting contributions, but ensuring sustainability is a persistent challenge in OSS communities. Although the motivations behind and barriers to OSS contributors’ joining and retention have been extensively studied, the impacts of, reasons for and solutions to contribution abandonment at the individual level have not been well studied, especially for pull-based development. To bridge this gap, we present an empirical study on pull request abandonment based on a sizable dataset. We manually examine 321 abandoned pull requests on GitHub and then quantify the manual observations by surveying 710 OSS developers. We find that while the lack of integrators’ responsiveness and the lack of contributors’ time and interest remain the main reasons that deter contributors from participation, limitations during the processes of patch updating and consensus reaching can also cause abandonment. We also show the significant impacts of pull request abandonment on project management and maintenance. Moreover, we elucidate the strategies used by project integrators to cope with abandoned pull requests and highlight the need for a practical handover mechanism. We discuss the actionable suggestions and implications for OSS practitioners and tool builders, which can help to upgrade the infrastructure and optimize the mechanisms of OSS communities. Yue Yu 0001, Tao Wang 0006, Gang Yin, Shanshan Li 0001, Huaimin Wang 0001 |
IEEE Trans. Software Eng. | 4 |
| 2022 | Redundancy, Context, and Preference: An Empirical Study of Duplicate Pull Requests in OSS ProjectsabstractOSS projects are being developed by globally distributed contributors, who often collaborate through the pull-based model today. While this model lowers the barrier to entry for OSS developers by synthesizing, automating and optimizing the contribution process, coordination among an increasing number of contributors remains as a challenge due to the asynchronous and self-organized nature of distributed development. In particular, duplicate contributions, where multiple different contributors unintentionally submit duplicate pull requests to achieve the same goal, are an elusive problem that may waste effort in automated testing, code review and software maintenance. While the issue of duplicate pull requests has been highlighted, to what extent duplicate pull requests affect the development in OSS communities has not been well investigated. In this paper, we conduct a mixed-approach study to bridge this gap. Based on a comprehensive dataset constructed from 26 popular GitHub projects, we obtain the following findings: (a) Duplicate pull requests result in redundant human and computing resources, exerting a significant impact on the contribution and evaluation process. (b) Contributors’ inappropriate working patterns and the drawbacks of their collaborating environment might result in duplicate pull requests. (c) Compared to non-duplicate pull requests, duplicate pull requests have significantly different features, e.g., being submitted by inexperienced contributors, being fixing bugs, touching cold files, and solving tracked issues. (d) Integrators choosing between duplicate pull requests prefer to accept those with early submission time, accurate and high-quality implementation, broad coverage, test code, high maturity, deep discussion, and active response. Finally, actionable suggestions and implications are proposed for OSS practitioners. Yue Yu 0001, Minghui Zhou 0001, Tao Wang 0006, Gang Yin, Long Lan, Huaimin Wang 0001 |
IEEE Trans. Software Eng. | 5 |
| 2022 | Motivation Under Gamification: An Empirical Study of Developers' Motivations and Contributions in Stack OverflowabstractTo encourage developers' volunteer contributions, modern programming question and answer (Q&A) sites like Stack Overflow (SO) employ gamified incentive mechanisms such as reputation and badges. Understanding developers' motivations in the presence of gamification and the relationship between their motivations and behavioral outcomes is crucial for community building and designing good incentive mechanisms. Grounded on self-determination theory, we conducted a survey with 938 developers who participate in SO to understand their participation motivations and incentive perceptions. By connecting the survey responses with the SO data, we quantitatively analyzed how the developers' motivations and satisfaction of needs relate to their effort and contribution quality. Our main findings are as follows: (1) despite the presence of gamified incentive mechanisms, developers are mainly motivated by intrinsic motivation to participate in SO; (2) developers who have strong motivations to gain gamification rewards are associated with higher intrinsic and integrated motivations, while developers with more development experiences are less motivated by the gamified incentives; (3) both extrinsic motivations (in terms of career prospects) and intrinsic motivations (regarding self-improvement and helping others) can motivate developers to make high-quantity and high-quality contributions; and (4) high-level satisfaction of needs for competency and autonomy has a positive effect on developers making high-quantity and high-quality contributions and addressing difficult problems. Based on these findings, we discuss implications for developer motivation and gamification in the crowdsourcing context and for the mechanism design of gamified crowdsourced platforms. Yao Lu 0003, Xinjun Mao, Minghui Zhou 0001, Yang Zhang 0026, Zude Li, Tao Wang 0006, Gang Yin, Huaimin Wang 0001 |
IEEE Trans. Software Eng. | 7 |
| 2021 | Why API documentation is insufficient for developers: an empirical study
Yue Yu 0001, Tao Wang 0006, Gang Yin, Huaimin Wang 0001 |
Sci. China Inf. Sci. | 4 |
| 2021 | Detecting Duplicate Contributions in Pull-Based Model Combining Textual and Change Similarities
Yue Yu 0001, Tao Wang 0006, Gang Yin, Xinjun Mao, Huaimin Wang 0001 |
J. Comput. Sci. Technol. | 4 |
| 2020 | Improving students' programming quality with the continuous inspection process: a social coding perspective
Yao Lu 0003, Xinjun Mao, Tao Wang 0006, Gang Yin, Zude Li |
Frontiers Comput. Sci. | 4 |
| 2019 | A Neural-Network based Code Summarization Approach by Using Source Code and its Call DependenciesabstractCode summarization aims at generating natural language abstraction for source code, and it can be of great help for program comprehension and software maintenance. The current code summarization approaches have made progress with neural-network. However, most of these methods focus on learning the semantic and syntax of source code snippets, ignoring the dependency of codes. In this paper, we propose a novel method based on neural-network model using the knowledge of the call dependency between source code and its related codes. We extract call dependencies from the source code, transform it as a token sequence of method names, and leverage the Seq2Seq model for code summarization using the combination of source code and call dependency information. About 100,000 code data is collected from 1,000 open source Java proejects on github for experiment. The large-scale code experiment shows that by considering not only the code itself but also the codes it called, the code summarization model can be improved with the BLEU score to 33.08. Bohong Liu, Tao Wang 0006, Xunhui Zhang, Gang Yin, Jinsheng Deng |
Internetware | 5 |
| 2019 | Multi-reviewing pull-requests: An exploratory study on GitHub OSS projects
Dongyang Hu, Yang Zhang 0026, Junsheng Chang, Gang Yin, Yue Yu 0001, Tao Wang 0006 |
Inf. Softw. Technol. | 4 |
| 2019 | RepoLike: amulti-feature-based personalized recommendation approach for open-source repositoriesabstractWith the deep integration of software collaborative development and social networking, social coding represents a new style of software production and creation paradigm. Because of their good flexibility and openness, a large number of external contributors have been attracted to the open-source communities. They are playing a significant role in open-source development. However, the open-source development online is a globalized and distributed cooperative work. If left unsupervised, the contribution process may result in inefficiency. It takes contributors a lot of time to find suitable projects or tasks from thousands of open-source projects in the communities to work on. In this paper, we propose a new approach called “RepoLike,” to recommend repositories for developers based on linear combination and learning to rank. It uses the project popularity, technical dependencies among projects, and social connections among developers to measure the correlations between a developer and the given projects. Experimental results show that our approach can achieve over 25% of hit ratio when recommending 20 candidates, meaning that it can recommend closely correlated repositories to social developers. Cheng Yang 0004, Tao Wang 0006, Gang Yin, Xunhui Zhang, Yue Yu 0001, Huaimin Wang 0001 |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2018 | Recommending Similar Bug Reports: A Novel Approach Using Document Embedding ModelabstractIn the software development, it is not uncommon to find that several bug reports are related to many common code files, i.e., similar bugs. Similar bug recommendation is a meaningful task which can assist developers in bug triaging and fixing. As the state of the art, Yang et al.'s work presented an approach that combines TF-IDF method with word embedding model and achieved a good result. To further improve the performance of their approach, in this paper, we propose a novel approach using Document Embedding model. In our preliminary evaluation, we conduct the experiment on 13,090 bug reports from the Eclipse platform and the results show that our approach outperforms Yang et al.'s, with 7.89-8.96% of improvement. Dongyang Hu, Tao Wang 0006, Junsheng Chang, Gang Yin, Yue Yu 0001, Yang Zhang 0026 |
APSEC | 5 |
| 2018 | Multi-Discussing across Issues in GitHub: A Preliminary StudyabstractSocial coding sites like GitHub has enabled developers to easily contribute their comments on multiple issues and switch their discussion between issues, i.e., multi-discussing. Discussing multiple issues simultaneously may enhance the work efficiency of developers. However, multi-discussing also relies on developers' rationally allocating their time and focus, which may bring different influence to the resolution of issues. Therefore, investigating how multi-discussing affects the issue resolution is a meaningful research question which can help developers understand the benefits and limitations when they switch their discussion between issues. In this paper, we present a preliminary study of the impact of multi-discussing on issue resolution in GitHub projects, by using quantitative methods. First, we collect and analyzed data from 631 GitHub projects to explore how multi-discussing affects the average resolution latency of project issues. Further, we develop method for measuring the rate and breadth of a developers' discussionswitching behavior, and we use regression modeling to study how discussion-switching affects the single issue resolution latency. We find that multi-discussing is a common behavior of developers in GitHub projects. Also, multi-discussing is associated with shorter average issue resolution latency of project. However, during a single issue resolution, more participants' discussion-switching tend to bring longer issue resolution latency. Our study motivates the need for further research on the multi-discussing. Dongyang Hu, Tao Wang 0006, Junsheng Chang, Gang Yin, Yang Zhang 0026 |
APSEC | 4 |
| 2018 | An Insight Into the Impact of Dockerfile Evolutionary Trajectories on Quality and LatencyabstractContainerization is a software development approach aimed at packaging an application together with all its dependencies and execution environment in a light-weight, self-contained unit, of which Docker has become the de-facto industry standard. By defining the specific Docker image architecture and building orders, dockerfile plays an important role in the Docker-based containerization process. Understanding the evolution of dockerfile and which dockerfile architecture attributes enhance dockerfile quality and reduce image build latency can benefit the efficient processing of containerization. In this paper, we perform an empirical study on a large dataset of 2,840 projects to shed light on the impact of dockerfile evolutionary trajectories on quality and latency in the Docker-based containerization. Based on the six categories of dockerfile evolutionary trajectories we discovered, we build two regression models to explore the impact of dockerfile evolutionary trajectories and specific architecture attributes on dockerfile quality and image build latency, which derives a number of suggestions for practitioners. Yang Zhang 0026, Gang Yin, Tao Wang 0006, Yue Yu 0001, Huaimin Wang 0001 |
COMPSAC (1) | 2 |
| 2018 | Cross-Project Issue Classification Based on Ensemble Modeling in a Social Coding World
Yarong Zeng, Yue Yu 0001, Xunhui Zhang, Tao Wang 0006, Gang Yin, Huaimin Wang 0001 |
ICONIP (4) | 6 |
| 2018 | A Hybrid Approach for Tag Hierarchy Construction
Shangwen Wang, Tao Wang 0006, Xiaoguang Mao, Gang Yin, Yue Yu 0001 |
ICSR | 4 |
| 2018 | Who Will Become a Long-Term Contributor?: A Prediction Model based on the Early Phase BehaviorsabstractThe continuous contribution from peripheral participants is crucial for the success of open source projects. Thus, how to identify the potential Long-Term Contributors (LTC) early and retain them is of great importance. We propose a prediction model to measure the chance for an individual to become a LTC contributor through his capacity, willingness, and the opportunity to contribute at the time of joining. Using data of Rails hosted on GitHub, we find that the probability for a new joiner to become a LTC is associated with his willingness and environment. Specifically, future LTCs tend to be more active and show more community-oriented attitude than other joiners during their first month. This implies that the interaction between individual's attitude and project's climate are associated with the odds that an individual would become a valuable contributor or disengage from the project. We evaluated our prediction model by using the 10 cross-validation method. Results show that our model archives the mean AUC as 0.807, which is valuable for OSS projects to identify potential long-term contributors and adopt better strategies to retain them for continuous contribution. Tao Wang 0006, Yang Zhang 0026, Gang Yin, Yue Yu 0001, Huaimin Wang 0001 |
Internetware | 3 |
| 2018 | A dataset of duplicate pull-requests in githubabstractIn GitHub, the pull-based development model enables community contributors to collaborate in a more efficient way. However, the distributed and parallel characteristics of this model pose a potential risk for developers to submit duplicate pull-requests (PRs), which increase the extra cost of project maintenance. To facilitate the further studies to better understand and solve the issues introduced by duplicate PRs, we construct a large dataset of historical duplicate PRs extracted from 26 popular open source projects in GitHub by using a semi-automatic approach. Furthermore, we present some preliminary applications to illustrate how further researches can be conducted based on this dataset. Yue Yu 0001, Gang Yin, Tao Wang 0006, Huaimin Wang 0001 |
MSR | 3 |
| 2018 | Adaptive software search toward users' customized requirements in GitHubabstractBecause of a tremendous growth of Open Source Software (OSS) scale and the diversity of users' requirements, users now face the problem of finding OSS that meets their expectations in a huge number of OSS resources.However, current GitHub-provided search service has a shortage in adapting to user needs.When facing diverse users' requirements, it cannot always return satisfactory results.In this paper, we provide a more efficient search service for OSS on GitHub.We first design a multi-dimensional measurement model for OSS, which forms a corresponding metric system and quantitative measurement method.Then we propose a ranking algorithm based on fuzzy synthetic evaluation in order to implement an adaptive metric ranking method that is oriented to user requirements.We verify that our work is useful by setting up experiments.The experiment results show that compared with GitHub-provided search service (searching by "Best Match" & searching by "Most Stars"), the effectiveness of our method improved by 97.6% and 13.8% respectively, which means our method returns search results which meet users' expectations more, and has high self-adaptive ability. Jinze Liu, Tao Wang 0006, Yue Yu 0001, Gang Yin |
SEKE | 5 |
| 2018 | Correlation-based software search by leveraging software term database
Gang Yin, Tao Wang 0006, Yang Zhang 0026, Yue Yu 0001, Huaimin Wang 0001 |
Frontiers Comput. Sci. | 2 |
| 2018 | Internal quality assurance for external contributions in GitHub: An empirical investigationabstractAbstract For popular open‐source software projects, there are always a large number of worldwide developers who have been glued to making code contributions, while most of these developers play the role of casual contributors because of their very limited code commits. The frequent turnover of such a group of developers and the wide variations in their coding experiences challenge the project management on code and quality. This paper aims to investigate the status quo of internal quality assurance for external contributions in social coding sites. We first conducted a case study of 21 popular GitHub projects to estimate the code quality of the casual contributors. The quantitative results show that the casual contributors introduced greater quantity and severity of code quality issues than the main contributors; the developers who contribute to different projects as main and casual contributors did not perform significantly differently in terms of their code quality. On the basis of these findings, we further conducted a survey of 81 developers on GitHub to understand their practices on internal quality assurance. The qualitative results expose some limitations of present internal quality control for external contributions in GitHub. Finally, we discuss an alternative quality management paradigm: Continuous Inspection for industrial practices. Yao Lu 0003, Xinjun Mao, Zude Li, Yang Zhang 0026, Tao Wang 0006, Gang Yin |
J. Softw. Evol. Process. | 6 |
| 2018 | Linking Issue Tracker with Q&A Sites for Knowledge Sharing across CommunitiesabstractCollaborative development communities and knowledge sharing communities are highly correlated and mutually complementary. The knowledge sharing between these two types of open source communities can be very beneficial to both of them. However, it is a great challenge to automate this process. Current studies mainly focus on knowledge acquisition in one type of community, and few of them have tackle this problem efficiently. In this paper we take Android Issue Tracker and Stack Overflow as a case to study the mutual knowledge sharing between them. We propose an automatic approach by integrating semantic similarity with temporal locality between Android issues and Stack Overflow posts based on the internal citation-graph to reveal the potential associations between them. Our approach explores the internal citations in communities for closely related posts or issues clustering, exploits the rich semantics in fine-grained information of issues and posts for associations building, and leverages the temporal correlations between issues and posts in-depth for associations ranking. Extensive experiments show that the precision of our approach reaches 62.51 percent for top 10 recommendations when recommending Stack Overflow posts to Android issues, and 66.83 percent in reverse. Huaimin Wang 0001, Tao Wang 0006, Gang Yin, Cheng Yang 0004 |
IEEE Trans. Serv. Comput. | 3 |
| 2017 | Where Is the Road for Issue Reports Classification Based on Text Mining?abstractCurrently, open source projects receive various kinds of issues daily, because of the extreme openness of Issue Tracking System (ITS) in GitHub. ITS is a labor-intensive and time-consuming task of issue categorization for project managers. However, a contributor is only required a short textual abstract to report an issue in GitHub. Thus, most traditional classification approaches based on detailed and structured data (e.g., priority, severity, software version and so on) are difficult to adopt. In this paper, issue classification approaches on a large-scale dataset, including 80 popular projects and over 252,000 issue reports collected from GitHub, were investigated. First, four traditional text-based classification methods and their performances were discussed. Semantic perplexity (i.e., an issues description confuses bug-related sentences with nonbug-related sentences) is a crucial factor that affects the classification performances based on quantitative and qualitative study. Finally, A two-stage classifier framework based on the novel metrics of semantic perplexity of issue reports was designed. Results show that our two-stage classification can significantly improve issue classification performances. Yue Yu 0001, Gang Yin, Tao Wang 0006, Huaimin Wang 0001 |
ESEM | 3 |
| 2017 | DevRec: A Developer Recommendation System for Open Source Repositories
Xunhui Zhang, Tao Wang 0006, Gang Yin, Cheng Yang 0004, Yue Yu 0001, Huaimin Wang 0001 |
ICSR | 3 |
| 2017 | Detecting Duplicate Pull-requests in GitHubabstractThe widespread use of pull-requests boosts the development and evolution for many open source software projects. However, due to the parallel and uncoordinated nature of development process in GitHub, duplicate pull-requests may be submitted by different contributors to solve the same problem. Duplicate pull-requests increase the maintenance cost of GitHub, result in the waste of time spent on the redundant effort of code review, and even frustrate developers' willing to offer continuous contribution. In this paper, we investigate using text information to automatically detect duplicate pull-requests in GitHub. For a new-arriving pull-request, we compare the textual similarity between it and other existing pull-requests, and then return a candidate list of the most similar ones. We evaluate our approach on three popular projects hosted in GitHub, namely Rails, Elasticsearch and Angular.JS. The evaluation shows that about 55.3% -- 71.0% of the duplicates can be found when we use the combination of title similarity and description similarity. Gang Yin, Yue Yu 0001, Tao Wang 0006, Huaimin Wang 0001 |
Internetware | 2 |
| 2017 | An Empirical Study of Reviewer Recommendation in Pull-based Development ModelabstractCode review is an important process to reduce code defects and improve software quality. However, in social coding communities using the pull-based model, everyone can submit code changes, which increases the required code review efforts. Therefore, there is a great need of knowing the process of code review and analyzing the pre-existing reviewer recommendation algorithms. In this paper, we do an empirical study about the PRs and their reviewers in Rails project. Moreover, we reproduce a popular and effective IR-based code reviewer recommendation algorithm and validate it on our dataset which contains 16,049 PRs. We find that the inactive reviewers are very important to code reviewing process, however, the pre-existing method's recommendation result strongly depends on the activeness of reviewers. Cheng Yang 0004, Xunhui Zhang, Lingbin Zeng, Gang Yin, Huaimin Wang 0001 |
Internetware | 5 |
| 2017 | Automatic Classification of Review Comments in Pull-based Development ModelabstractThe pull-based model, widely used in distributed software development, allows any contributor to fork a public repository, package contributions as a pull-request, and then merge back to the original repository.Code review is one of the most significant stages in pull-based development.It ensures that only high-quality pull-requests are accepted, based on the in-depth discussion among reviewers.Thus, automatically identifying what reviewers are talking about in the discussions is benificial to better understand the code review process.In this paper, we conduct a case study on three popular opensource software projects hosted on GitHub and construct a finegrained taxonomy including 11 sub-categories for review comments.We then manually label over 5,600 review comments, and propose a Two-Stage Hybrid Classification (TSHC) algorithm to classify review comments automatically by combining rule-based and machine-learning techniques.Comparative experiments with a text-based method achieve a reasonable improvement on each project (9.2% in Rails, 5.3% in Elasticsearch, and 7.2% in Angular.jsrespectively) in terms of the weighted average Fmeasure. Yue Yu 0001, Gang Yin, Tao Wang 0006, Huaimin Wang 0001 |
SEKE | 3 |
| 2017 | Who Will be Interested in? A Contributor Recommendation Approach for Open Source ProjectsabstractThe crowds' continuous participation and contribution are the key factors for the success of open source projects.However, among the massive competitors, it is difficult for a project to attract enough contributors by just passively waiting for enthusiasts to join in.Instead, it should actively seek gifted developers.Most of the current studies mainly focus on recommending experts inside a repository for some specific development tasks.In this paper, we propose a novel approach ConRec to recommend potential contributors across the entire open source community for given projects.It leverages the developers' historical activities in projects to analyze their technical interests and technical connections with others.Thereafter, it combines collaborative filtering algorithm with text matching algorithm to recommend proper developers.We conducted extensive experiments on 5,995 open source projects and 2,938,620 developers in GitHub.The results show that the proposed algorithm can recommend contributors to open source projects with the best performance of 63% in accuracy, and solve the cold start problem as well. Xunhui Zhang, Tao Wang 0006, Gang Yin, Cheng Yang 0004, Huaimin Wang 0001 |
SEKE | 3 |
| 2017 | Social media in GitHub: the role of @-mention in assisting software development
Yang Zhang 0026, Huaimin Wang 0001, Gang Yin, Tao Wang 0006, Yue Yu 0001 |
Sci. China Inf. Sci. | 3 |
| 2017 | What Are They Talking About? Analyzing Code Reviews in Pull-Based Development Model
Yue Yu 0001, Gang Yin, Tao Wang 0006, Huaimin Wang 0001 |
J. Comput. Sci. Technol. | 3 |
| 2017 | Intelligent Development Environment and Software Knowledge Graph
Zeqi Lin, Yanzhen Zou, Junfeng Zhao 0001, Xuandong Li, Jun Wei 0001, Hailong Sun 0001, Gang Yin |
J. Comput. Sci. Technol. | 8 |
| 2016 | Does the Role Matter? An Investigation of the Code Quality of Casual Contributors in GitHubabstractFor popular Open Source Software (OSS) projects there are always a large number of worldwide developers who have been glued to making code contributions, while most of these developers play the role of casual contributors due to their very limited code commits (for fixing defects and enhancing features, casually). The frequent turnover of such group of casual developers and the wide variations among their coding experiences challenge the project management on code and quality.This paper describes a case study which aims to estimate the quality of code made by casual contributors in 21 popular GitHub projects. The results of this case study show that: (1) casual contributors introduced greater quantity and severity of Code Quality Issues (CQIs) than main contributors; (2) developers who contribute in different projects as main and casual contributors didn't perform statistically differently in terms of code quality; (3) casual contributors who have few project stars introduced more CQIs than those who have many. Furthermore, the paper lists the CQI categories which are most frequently introduced by casual contributors in the investigated projects. These findings provide valuable insights into code quality in the OSS context, and can guide OSS developers in improving the quality of the code contributions. Yao Lu 0003, Xinjun Mao, Zude Li, Yang Zhang 0026, Tao Wang 0006, Gang Yin |
APSEC | 6 |
| 2016 | Query reformulation by leveraging crowd wisdom for scenario-based software searchabstractThe Internet-scale open source software (OSS) production in various communities are generating abundant reusable resources for software developers. However, how to retrieve and reuse the desired and mature software from huge amounts of candidates is a great challenge: there are usually big gaps between the user application contexts (that often used as queries) and the OSS key words (that often used to match the queries). In this paper, we define the scenario-based query problem for OSS retrieval, and then we propose a novel approach to reformulate the raw query by leveraging the crowd wisdom from millions of developers to improve the retrieval results. We build a software-specific domain lexical database based on the knowledge in open source communities, by which we can expand and optimize the input queries. The experiment results show that, our approach can reformulate the initial query effectively and outperforms other existing search engines significantly at finding mature software. Tao Wang 0006, Yang Zhang 0026, Yun Zhan, Gang Yin |
Internetware | 5 |
| 2016 | RepoLike: personal repositories recommendation in social coding communitiesabstractSocial coding represents a new style of software production and creation paradigm, and demands for new technologies of software reuse. Many people searching for projects package, we can provide good reuse recommendation. In this paper, we focus on an interesting research topic of recommending software repositories to social developers, which is challenging because of two points: the first is how to get the interest contexts of developers; and the second is how to rank the repository candidates for recommendation properly. We propose RepoLike, a new approach for recommending repositories to developers by predicting their interests. RepoLike explores the developers' historical development activities and the social connections with other programmers, mines the technical features of repositories and the dependencies among them, and then combines both aspects to recommend most interesting and inspiring repositories to developers. The experiment results show that our approach can surprisingly recommend closely correlated repositories to developers, and the critical test results show that the recommendation performance is strongly impacted by the interest context model. Cheng Yang 0004, Tao Wang 0006, Gang Yin, Huaimin Wang 0001 |
Internetware | 4 |
| 2016 | A Novel Open Source Software Ecosystem: From a Graphic Point of View and Its ApplicationabstractWith the rapid development of open source software, various elements such as OSS, developers, users and online posts, across different communities and their interactions constitute a novel software ecosystem.Most of the current researches about software ecosystems care the connections between software, and few of them consider the relationship across communities from an overall perspective, and fail to cover the users and their activities which should be an indispensable part of the OSS ecosystem.This paper model the OSS ecosystem as a graph, which combines different types of OSS communities as a whole.Based on this graph model, we analyze the characteristics of ecosystem, like the evolution, competition and symbiosis.In addition, we build a recommendation system as well, and the experiment results suggest the validation of our approach. Chenxi Song, Tao Wang 0006, Gang Yin, Xunhui Zhang, Cheng Yang 0004 |
SEKE | 3 |
| 2016 | Preface
Zhi Jin 0001, Zhenjiang Hu 0002, Gang Yin |
Sci. China Inf. Sci. | 3 |
| 2016 | Determinants of pull-based development in the context of continuous integration
Yue Yu 0001, Gang Yin, Tao Wang 0006, Cheng Yang 0004, Huaimin Wang 0001 |
Sci. China Inf. Sci. | 2 |
| 2016 | Reviewer recommendation for pull-requests in GitHub: What can we learn from code review and bug assignment?
Yue Yu 0001, Huaimin Wang 0001, Gang Yin, Tao Wang 0006 |
Inf. Softw. Technol. | 3 |
| 2015 | A supervised approach for tag hierarchy construction in open source communitiesabstractThe massive amounts of open source software provide sufficient reusable resources for software development. Most of the OSS communities adopt a kind of categorization or tagging mechanism to organize the software. However, the categorization often too coarse, while the tags are flat and fail to capture the inter-relation among them. In this paper, we propose a novel approach to reveal the latent relations between tags and build a tag hierarchy to help locate resources. We firstly build a co-occurrence network, based on which we compare the connotations of tags and construct a preliminary hierarchy. Then we leverage the domain knowledge of category in SourceForge to optimize and improve the relations between tags. At the end, we demonstrate the effectiveness of the constructed tag hierarchy with quantitative evaluation, which suggest the validation of our approach. Chongming Gu, Gang Yin, Tao Wang 0006, Cheng Yang 0004, Huaimin Wang 0001 |
Internetware | 2 |
| 2015 | Software Ranking and Analysis based on Mining Market Requirements and CharacteristicsabstractAs the rapid growth of open source software, how to choose software from many alternatives becomes a great challenge. Traditional ranking approaches mainly focus on the characteristics of the software themselves, such as qualities, security, reliable and so on. In this paper we investigate the market demands for software engineers, and propose a novel approach for ranking software by analyzing the market requirements for special software. At the same time we conclude the characteristics of software advertisements and analyze the reasons that why these situations emerge and tendency of software market requirements. As industries always need to balance several different factors for selecting software, the market demands can be a good indicator for ranking software and software evaluating. This paper provides quite a different perspective and some interesting inferences on software market requirements, and it can be a valuable supplement for traditional ranking methods, as well as software evaluating. Bingxun Liu, Gang Yin, Tao Wang 0006, Huaimin Wang 0001 |
Internetware | 2 |
| 2015 | Exploring the Use of @-mention to Assist Software Development in GitHubabstractRecently, many researches propose that social media tools can promote the collaboration among developers, which are beneficial to the software development. Nevertheless, there is little empirical evidence to confirm that using @-mention has indeed a beneficial impact on the issues in GitHub. In this paper, we analyze the data from GitHub and give some insights on how @-mention is used in the issues (general-issues and pull-requests). Our statistical results indicate that, @-mention attracts more participants and tends to be used in the difficult issues. @-mention favors the solving process of issues by enlarging the visibility of issues and facilitating the developers' collaboration. In addition to this global study, our study also build a @-network based on the @-mention database we extract. Through the @-network, we can mine the relationships and characteristics of developers in GitHub's issues. Yang Zhang 0026, Huaimin Wang 0001, Gang Yin, Tao Wang 0006, Yue Yu 0001 |
Internetware | 3 |
| 2015 | A hybrid static analysis refinement approach within internetware environmentabstractIn this paper, we propose a hybrid refinement approach to improve the accuracy of static analysis. It keeps condition constraints information during forward dataflow analysis and gets the satisfiability of a warning by a constraint solver taking as input such information and path conditions; data regression analysis can remedy the capability of handling loops and library calls of abstract interpretation technique. It has been implemented in our static analysis tool, Defect Testing System (DTS) and deployed on a internetware environment TRUSTIE. Experiment on a large number of C open source projects shows the great improvement this strategy makes. Dalin Zhang 0003, Gang Yin, Dahai Jin, Yunzhan Gong, Tianshuang Wu, Hailong Zhang 0006 |
Internetware | 2 |
| 2015 | Evaluating Bug Severity Using Crowd-based Knowledge: An Exploratory StudyabstractIn bug tracking system, the high volume of incoming bug reports poses a serious challenge to project managers. Triaging these bug reports manually consumes time and resources which leads to delaying the resolution of important bugs. StackOverflow is the most popular crowdsourcing Q&A community with plenty of bug-related posts. In this paper, we explore the correlation between bug severity and the crowd attributes of linked posts. Two typical types of projects' bug repositories are studied here, e.g. Mozilla (user-centric project) and Eclipse (developer-centric project). Our results show that the bug severity is consistent with the crowd-based knowledge both in Mozilla and Eclipse, i.e. the linked posts of severe bugs have higher score etc. in StackOverflow than non-severe bugs. This interesting phenomenon inspires us that we can optimize the existing evaluation methods of bug severity by incorporating the crowd-based knowledge from a third-party in future. Yang Zhang 0026, Gang Yin, Tao Wang 0006, Yue Yu 0001, Huaimin Wang 0001 |
Internetware | 2 |
| 2014 | Who Should Review this Pull-Request: Reviewer Recommendation to Expedite Crowd CollaborationabstractGithub facilitates the pull-request mechanism as an outstanding social coding paradigm by integrating with social media. The review process of pull-requests is a typical crowd sourcing job which needs to solicit opinions of the community. Recommending appropriate reviewers can reduce the time between the submission of a pull-request and the actual review of it. In this paper, we firstly extend the traditional Machine Learning (ML) based approach of bug triaging to reviewer recommendation. Furthermore, we analyze social relations between contributors and reviewers, and propose a novel approach to recommend highly relevant reviewers by mining comment networks (CN) of given projects. Finally, we demonstrate the effectiveness of these two approaches with quantitative evaluations. The results show that CN-based approach achieves a significant improvement over the ML-based approach, and on average it reaches a precision of 78% and 67% for top-1 and top-2 recommendation respectively, and a recall of 77% for top-10 recommendation. Yue Yu 0001, Huaimin Wang 0001, Gang Yin, Charles Ling 0001 |
APSEC (1) | 3 |
| 2014 | A Exploratory Study of @-Mention in GitHub's Pull-RequestsabstractPull-request mechanism is an outstanding social development method in Git Hub. @-mention is a social media tool that deeply integrated with pull-request mechanism. Recently, many research results show that social media tools can promote the collaborative software development, but few work focuses on the impacts of @-mention. In this paper, we conduct an exploratory study of @-mention in pull-request based software development, including its current situation and benefits. We obtain some interesting findings which indicate that @-mention is beneficial to the processing of pull-request. Our work also proposes some possible research directions and problems of the @-mention. It helps the developers and researchers notice the significance of @-mention in the pull-request based software development. Yang Zhang 0026, Gang Yin, Yue Yu 0001, Huaimin Wang 0001 |
APSEC (1) | 2 |
| 2014 | Reviewer Recommender of Pull-Requests in GitHubabstractPull-Request (PR) is the primary method for code contributions from thousands of developers in GitHub. To maintain the quality of software projects, PR review is an essential part of distributed software development. Assigning new PRs to appropriate reviewers will make the review process more effective which can reduce the time between the submission of a PR and the actual review of it. However, reviewer assignment is now organized manually in GitHub. To reduce this cost, we propose a reviewer recommender to predict highly relevant reviewers of incoming PRs. Combining information retrieval with social network analyzing, our approach takes full advantage of the textual semantic of PRs and the social relations of developers. We implement an online system to show how the reviewer recommender helps project managers to find potential reviewers from crowds. Our approach can reach a precision of 74% for top-1 recommendation, and a recall of 71% for top-10 recommendation. Yue Yu 0001, Huaimin Wang 0001, Gang Yin, Charles Ling 0001 |
ICSME | 3 |
| 2014 | Linking stack overflow to issue tracker for issue resolutionabstractIssue resolution is a central task for software development. The efficiency of issue fixing largely relies on the issue report quality. Stack Overflow, which hosts rich and real-time posts about programming-specific problems, is a valuable external source of knowledge for issue resolution. In this paper, we present CrossLink, an analysis framework that automatically introduce related posts in Stack Overflow to issues in Android Issue Tracker. This helps developers to leverage the abundant and professional knowledge from massive programmers in Stack Overflow for issue resolution. CrossLink explores the semantic similarities as well as the temporal associations between the two types of repositories to recommend Stack Overflow posts to Android issues. The internal links in Stack Overflow are also employed to improve the linking accuracy. The experiments prove the effectiveness of CrossLink with precision of 62.51% for top-10 recommendations, which is significantly higher than the state-of-art method. Tao Wang 0006, Gang Yin, Huaimin Wang 0001, Cheng Yang 0004 |
Internetware | 2 |
| 2014 | Multi-dimensions of Developer Trustworthiness Assessment in OSS CommunityabstractWith the prosperity of the Open Source Software, various software communities are formed and they attract huge amounts of developers to participate in distributed software development. For such software development paradigm, how to evaluate the skills of the developers comprehensively and automatically is critical. However, most of the existing researches assess the developers based on the Implementation aspects, such as the artifacts they created or edited. They ignore the developers' contributions in Social collaboration aspects, such as answering questions, giving advices, making comments or creating social connections. In this paper, we propose a novel model which evaluate the individuals' skills from both Implementation and Social collaboration aspects. Our model defines four metrics from muti-dimensions, including collaboration index, technical skill, community influence and development contribution. We carry out experiments on a real-world online software community. The results show that our approach can make more comprehensive measurement than the previous work. Yu Bai 0010, Gang Yin, Huaimin Wang 0001 |
TrustCom | 2 |
| 2014 | Tag recommendation for open source software
Tao Wang 0006, Huaimin Wang 0001, Gang Yin, Charles Ling 0001, Xiao Li 0039 |
Frontiers Comput. Sci. | 3 |
| 2014 | Online fault diagnosis method based on Incremental Support Vector Data Description and Extreme Learning Machine with incremental output structure
Gang Yin, Ying-Tang Zhang, Zhining Li, Guo-Quan Ren, Hongbo Fan |
Neurocomputing | 1 |
| 2014 | An incentive compatible reputation mechanism for P2P systems
Junsheng Chang, Zhengbin Pang, Huaimin Wang 0001, Gang Yin |
J. Supercomput. | 5 |
| 2013 | Mining Software Profile across Multiple Repositories for Hierarchical CategorizationabstractThe large amounts of software repositories over the Internet are fundamentally changing the traditional paradigms of software maintenance. Efficient categorization of the massive projects for retrieving the relevant software in these repositories is of vital importance for Internet-based maintenance tasks such as solution searching, best practices learning and so on. Many previous works have been conducted on software categorization by mining source code or byte code, which are only verified on relatively small collections of projects with coarse-grained categories or clusters. However, Internet-based software maintenance requires finer-grained, more scalable and language-independent categorization approaches. In this paper, we propose a novel approach to hierarchically categorize software projects based on their online profiles across multiple repositories. We design a SVM-based categorization framework to classify the massive number of software hierarchically. To improve the categorization performance, we aggregate different types of profile attributes from multiple repositories and design a weighted combination strategy which assigns greater weights to more important attributes. Extensive experiments are carried out on more than 18,000 projects across three repositories. The results show that our approach achieves significant improvements by using weighted combination, and the overall precision, recall and F-Measure can reach 71.41%, 65.60% and 68.38% in appropriate settings. Compared to the previous work, our approach presents competitive results with 123 finer-grained and multi-layered categories. In contrast to those using source code or byte code, our approach is more effective for large-scale and language-independent software categorization. Tao Wang 0006, Huaimin Wang 0001, Gang Yin, Charles Ling 0001, Xiang Li 0012 |
ICSM | 3 |
| 2013 | Mining and recommending software features across multiple web repositoriesabstractThe "Internetware" paradigm is fundamentally changing the traditional way of software development. More and more software projects are developed, maintained and shared on the Internet. However, a large quantity of heterogeneous software resources have not been organized in a reasonable and efficient way. Software feature is an ideal material to characterize software resources. The effectiveness of feature-related tasks will be greatly improved, if a multi-grained feature repository is available. In this paper, we propose a novel approach for organizing, analyzing and recommending software features. Firstly, we construct a Hierarchical rEpository of Software feAture (HESA). Then, we mine the hidden affinities among the features and recommend relevant and high-quality features to stakeholders based on HESA. Finally, we conduct a user study to evaluate our approach quantitatively. The results show that HESA can organize software features in a more reasonable way compared to the traditional and the state-of-the-art approaches. The result of feature recommendation is effective and interesting. Yue Yu 0001, Huaimin Wang 0001, Gang Yin, Bo Liu 0014 |
Internetware | 3 |
| 2013 | HESA: The Construction and Evaluation of Hierarchical Software Feature Repository
Yue Yu 0001, Huaimin Wang 0001, Gang Yin, Xiang Li 0012, Cheng Yang 0004 |
SEKE | 3 |
| 2013 | An online service-oriented performance profiling tool for cloud computing systems
Haibo Mi, Huaimin Wang 0001, Yangfan Zhou 0002, Michael R. Lyu, Hua Cai, Gang Yin |
Frontiers Comput. Sci. | 6 |
| 2012 | Inducing Taxonomy from Tags: An Agglomerative Hierarchical Clustering Framework
Xiang Li 0012, Huaimin Wang 0001, Gang Yin, Tao Wang 0006, Cheng Yang 0004, Yue Yu 0001, Dengqing Tang |
ADMA | 3 |
| 2012 | Performance problems diagnosis in cloud computing systems by mining request trace logsabstractIn cloud computing systems, end-to-end request tracing approach is helpful for developers to understand the runtime behavior of user requests. Based on trace logs, we propose an approach to localize the abnormal methods that are the primary causes of performance problems. Our approach involves three steps: (1) cluster the user requests into different categories according to request call sequences and select major categories; (2) extract the principal methods that might be the causes of performance degradation; (3) pick out abnormal methods from those principal methods in each major category. We conduct four cases of performance degradations to validate our approach over a real-world enterprise-class cloud computing platform. The experimental results show that our approach can locate the prime causes of performance problems with low false-positive rate and false-negative rate. Haibo Mi, Huaimin Wang 0001, Gang Yin, Hua Cai, Tingtao Sun |
NOMS | 3 |
| 2010 | Towards Global State Identification of Nodes in DHT Based SystemsabstractPeer-to-Peer (P2P) systems receive growing acceptance, and the need of identifying the states of nodes appears increasingly in a variety of P2P based applications. In this paper, we propose Hermes, an algorithm to efficiently spread and maintain the states of all nodes in large scale systems. Hermes uses an improved push & pull style gossip process to spread the states of all nodes, and proposes a novel synopsis technique to maintain the states of all nodes locally at each node. Simulation results show that, in a N-node network with commonly accepted configurations, Hermes is 2 rounds faster than push & pull style gossip to spread the state of one node to the whole network; each node only needs to store a very small part of global state data to maintain the whole view of all nodes states in the network, which is updated dynamically with high accuracy. Dehui Liu, Gang Yin, Feng Chen 0015, Huaimin Wang 0001 |
PDP | 2 |
| 2008 | Towards Role Based Trust Management without Distributed Searching of Credentials
Gang Yin, Huaimin Wang 0001, Dian-xi Shi |
ICICS | 1 |
| 2007 | A New Reputation Mechanism Against Dishonest Recommendations in P2P Systems
Junsheng Chang, Huaimin Wang 0001, Gang Yin, Yang-Bin Tang |
WISE | 3 |
| 2006 | Trustworthiness of Internet-based software
Huaimin Wang 0001, Yang-Bin Tang, Gang Yin |
Sci. China Ser. F Inf. Sci. | 3 |
| 2004 | An Authorization Framework Based on Constrained Delegation
Gang Yin, Meng Teng, Huaimin Wang 0001, Yan Jia 0001, Dian-xi Shi |
ISPA | 1 |
| 2002 | On asymptotic properties of a constant-step-size sign-error algorithm for adaptive filtering
Gang Yin, Han-Fu Chen |
Sci. China Ser. F Inf. Sci. | 1 |
| 2002 | Hybrid singular systems of differential equationsabstractThis work develops hybrid models for large-scale singular differential system and analyzes their asymptotic properties. To take into consideration the discrete shifts in regime across which the behavior of the corresponding dynamic systems is markedly different, our goals are to develop hybrid systems in which continuous dynamics are intertwined with discrete events under random-jump disturbances and to reduce complexity of large-scale singular systems via singularly perturbed Markov chains. To reduce the complexity of large-scale hybrid singular systems, two-time scale is used in the formulation. Under general assumptions, limit behavior of the underlying system is examined. Using weak convergence methods, it is shown that the systems can be approximated by limit systems in which the coefficients are averaged out with respect to the quasi-stationary distributions. Since the limit systems have fewer states, the complexity is much reduced. Gang Yin, Ji-Feng Zhang |
Sci. China Ser. F Inf. Sci. | 1 |