Yang Zhang 0026

dblp:06/6785-26 · DBLP profile ↗
← Back
60ranked-venue papers
15as first author
26since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 35 · 9 first-author · 15 since 2021Systems, architecture and hardware · 11 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 6 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 Divergence or Convergence? A Deep Insight into the Crowd Collaboration and its Productivity in Open Source Software based on Entropy
abstract
The Fork and Pull-Request model is widely used in collaborative development of open source software (OSS), fostering innovation through independent repository copies, but it can also lead to inefficiencies and fragmentation. A key underexplored aspect is the integration effectiveness—the degree to which distributed original commits across forks are effectively integrated back into the main repository. It plays a critical role in OSS project productivity but remains poorly understood. In response, we introduce convergence entropy, a novel metric that quantifies the integration effectiveness by measuring the similarity between distributions of original and merged commits across forks, adjusted for integration ratio. This metric highlights not only the volume of contributions but also their diversity and coordination, offering a unique lens to understand forking practices. Moreover, we explore the relationship between convergence entropy and three dimensions of OSS project productivity, showing significant correlations. We also observe that other factors can alter this dynamic.
Tao Wang 0158, Xunhui Zhang, Yang Zhang 0026, Cheng Yang 0004, Bo Ding 0001, Huaimin Wang 0001
CHI4
2026 From Trigger to Impact: Knowledge-Graph Reasoning and Risk-Aware Classification for Hardware Trojan Detection
Yang Zhang 0026, Xing Hu 0012
DATE1
2026 GLRA: Graph-based leakage risk assessment via minimal transmission cost path analysis
Xing Hu 0012, Yang Zhang 0026, Shaoqing Li, Keqin Li 0001
Comput. Secur.2
2026 SubG4TJ: A collaborative subgraph classification method based on multidimensional attributes for hardware trojan detection
Xing Hu 0012, Yang Zhang 0026, Keqin Li 0001
Expert Syst. Appl.2
2026 DockerFill: Automatically Completing Dockerfile Code With Syntax-Aware Multi-Task Learning
abstract
As a kind of infrastructure-as-code, Dockerfile specifies the structure and functionality of a built Docker image and thus plays an important role in the containerized software development process. Nowadays developers need to spend extra time and effort configuring their Dockerfiles in addition to their regular coding work, which requires knowledge and skills orthogonal to those entailed in other software-related experiences. Poorly written Dockerfile code often introduces errors and maintenance costs. However, little automated support is available for assisting developers in configuring Dockerfiles. In this study, we first conduct an online survey to investigate Docker developers’ perceptions of Dockerfile writing, highlighting the needs and potential benefits of Dockerfile auto-completion techniques. Then, we introduceDOCKERFILL, a pre-trained model based approach that provides completion suggestions for Dockerfile-specific code.DOCKERFILLleverages multi-layer Transformer architecture with syntax-aware multi-task learning, which includes contextual file information and three pre-training tasks, i.e., masked language modeling, syntax type identification, and masked identifier prediction. To evaluateDOCKERFILL’s effectiveness, we collect a dataset of 6,350 high-quality real-world Dockerfiles. Our empirical results show that DOCKERFILL provides up to 52.38% accuracy for token-level completion and 19.69% exact match for line-level completion, outperforming the baselines by 7.32%-37.67% and 1.97%-19.69%, respectively. Also,DOCKERFILLobtains significantly higher human evaluation scores compared to the baselines.
Yiwen Wu 0001, Yang Zhang 0026, Tao Wang 0006, Bo Ding 0001, Huaimin Wang 0001
IEEE Trans. Software Eng.2
2025 Understanding the Faults in Serverless Computing Based Applications: An Empirical Study
abstract
Serverless computing is a novel cloud computing paradigm that enables developers to develop, deploy, and run applications in the cloud without complex and error-prone cloud resource management. However, its characteristics also introduce new types of faults (e.g., faults due to insufficient computing resource allocation) and challenges to serverless computing-based applications (abbreviated as serverless applications). While prior studies have highlighted that serverless developers encounter various challenges, no attempts have been made to understand the faults in serverless applications. These faults may cause catastrophic consequences such as application crash, thereby hindering the further spread of serverless computing. We aim in this paper to understand the symptoms, root causes, and fix patterns of faults in serverless applications. To this end, we conduct an empirical study investigating developers' issues on GitHub and posts on Stack Overflow (SO). We first identify 546 real-world serverless-related faults from GitHub and SO. Then, we manually analyze and construct taxonomies of the symptoms, root causes, and fix patterns for these faults, respectively. Our study leads to the first taxonomy for symptoms of serverlessrelated faults, covering 5 categories and 21 subcategories. The findings of our study inform that the Permission Denied error is the most common type ($\mathbf{1 0. 8 1 \%}$) of faults. Furthermore, the Incorrect Code Logic is the main cause ($\mathbf{1 7. 9 5 \%}$) behind the faults. Furthermore, we summarize 15 fix patterns that can resolve$\mathbf{7 3. 6 3 \%}$of faults in this study. Based on the results, we provide actionable implications that can potentially facilitate research and assist developers in improving the development of serverless applications. Finally, we implement a knowledge-based Q&A tool named SafHelper to help developers understand and fix faults.
Changrong Xie, Yang Zhang 0026, Xinjun Mao, Kang Yang 0001, Tanghaoran Zhang
ICSME2
2025 Decoding Serverless Security: Exploring Developer Challenges and Solutions from Stack Overflow
abstract
Serverless computing is gaining increasing attention from developers due to its simplicity of infrastructure management.With the widespread adoption of this paradigm, the security of serverless computing (abbreviated as serverless security) has become a key concern.The security responsibility model of serverless computing is different from traditional architecture, which may introduce new security-related challenges (e.g., potential attack surface increase due to eventdriven architecture).While prior studies have investigated specific problems related to serverless security, no attempts have been made to explore the serverless security challenges discussed in the developer community.In this paper, we aim to gain a better understanding of serverless security-related challenges on Stack Overflow (SO).To this end, we manually analyze 472 serverless security-related questions to construct a taxonomy of challenges and then summarize common solutions for these challenges.Moreover, we employ a series of heuristics to gauge the popularity, difficulty, and expertise status of these challenges.Our study groups serveless security-related challenges into five categories and summarize 20 common solutions for them.Our analysis informs that Authentication is the most common challenge type (32.47%) and also the most difficult to address.Based on the results, we provide actionable implications that can facilitate research and help developers enhance the security of serverless applications.
Changrong Xie, Yang Zhang 0026, Xinjun Mao
SEKE2
2025 Open source oriented cross-platform survey
Simeng Yao, Xunhui Zhang, Yang Zhang 0026, Tao Wang 0006
Inf. Softw. Technol.3
2025 What problems are MLOps practitioners talking about? A study of discussions in Stack Overflow forum and GitHub projects
Yang Zhang 0026, Yiwen Wu 0001, Tao Wang 0006, Bo Ding 0007, Huaimin Wang 0001
Inf. Softw. Technol.1
2025 Are External Contributions Important to Project Productivity in Open Source Software? A Deep Insight based on Issue Entropy
abstract
In the realm of open source software (OSS) development, the resolution of issues is not just a technical task but a pivotal activity that drives ongoing enhancement and secures a project's sustainability. Contributing to issues is the majority form for external contributors to take part in OSS projects. Although the significance of external contributors is recognized, their contributions in issue process still lacks full quantification and clarity, and the correlation between their contributions and the project productivity remains unclear. In response, we propose issue entropy, a novel metric that quantifies the complexity of event sequence in issue process applied to study the external contributions. The metric applies principles of information theory to scrutinize granular details within a project's issue-related activities, providing a unique lens to understand and assess external contributions. To explore the correlation between external contributions and project productivity, we employed issue entropy as a novel way to examine external contributions, and analyzed its correlation with new bugs, commits, and bug fix time, serving as proxies for OSS project productivity. Our findings reveal a strong positive correlation with new bugs, variable relationship with commit volume based on company sponsorship, and significant negative correlation with bug fix time. We also observe the significant interactions between external contributions and other factors, such as project age and the number of files. Moreover, issue entropy offers a new perspective on the health of the OSS project ecosystem, potentially supporting further research and practical applications.
Tao Wang 0158, Xunhui Zhang, Yang Zhang 0026, Cheng Yang 0004, Yue Yu 0001, Huaimin Wang 0001
Proc. ACM Hum. Comput. Interact.4
2024 GHA-BFP: Framework for Automated Build Failure Prediction in GitHub Actions
abstract
GitHub Actions (GHA), a powerful Continuous Integration and Continuous Deployment (CI/CD) service, has revolutionized the way developers automate tasks in the software development pipeline. Although GHA provides great convenience, if a GHA build fails, the time spent waiting for results and debugging is wasted, which can seriously affect development efficiency. In this study, we delve into GHA build results and introduce an automatic framework named GHA-BFP that uses ML models to predict the failure of GHA builds. Using GHA-BFP with Random Forest model, we achieved the highest performance in predicting the failure of GHA builds, with all key metrics (i.e., Accuracy, Precision, Recall, and F1 score) exceeding 75%. Furthermore, through ablation experiments, we have verified the essentiality of the four categories of input features. Lastly, we conducted an assessment of the importance of each individual input feature in relation to the model's predictive capabilities.
Jiatai Li, Yang Zhang 0026, Tao Wang 0006, Yiwen Wu 0001
APSEC2
2024 SFCM-HT: Hardware Trojan Detection Based on Sequence Features with a Combination Model
abstract
In the context of the globalization of the Integrated Circuit (IC) industry, the Intellectual Property Cores (IPs) assume a pivotal role, offering the potential to streamline the development process and reduce costs. However, the use of IPs from third parties introduces the potential for malicious modifications, such as the insertion of hardware Trojans (HTs), which can compromise the security and reliability of hardware designs. Despite the advent of deep learning-based HT detection methods that leverage circuit sequences and graph data, which have overcome the limitations of a lack of a golden model and scalability issues, there are still shortcomings in the utilization of global and local features. We develop an HT detection method SFCM-HT based on a combination model of the graph convolutional network (GCN) and the gated recurrent unit (GRU). We transform the Register Transfer Level (RTL) design to a graph structure and model the graph as fixed-length sequences and innovatively construct the combination model to predict sequences related to HTs. We exploit the capacity of GCN to learn features locally and utilize circuit structure information globally, as well as the aptitude of GRU for learning and predicting circuit sequence features. We evaluate the model based on benchmark circuits in Trusthub. SFCM-HT detects Trojan sequences with 98.20% Precision and 98.76% Recall and effectively reduces detection time.
Yang Zhang 0026, Xing Hu 0012, Jialong Song, Shaoqing Li
ATS2
2024 How do Developers Talk about GitHub Actions? Evidence from Online Software Development Community
abstract
Continuous integration, deployment and delivery (CI/CD) have become cornerstones of DevOps practices. In recent years, GitHub Action (GHA) has rapidly replaced the traditional CI/CD tools on GitHub, providing efficiently automated workflows for developers. With the widespread use and influence of GHA, it is critical to understand the existing problems that GHA developers face in their practices as well as the potential solutions to these problems. Unfortunately, we currently have relatively little knowledge in this area. To fill this gap, we conduct a large-scale empirical study of 6,590 Stack Overflow (SO) questions and 315 GitHub issues. Our study leads to the first comprehensive taxonomy of problems related to GHA, covering 4 categories and 16 sub-categories. Then, we analyze the popularity and difficulty of problem categories and their correlations. Further, we summarize 56 solution strategies for different GHA problems. We also distill practical implications of our findings from the perspective of different audiences. We believe that our study contributes to the research of emerging GHA practices and guides the future support of tools and technologies.
Yang Zhang 0026, Yiwen Wu 0001, Tao Wang 0006, Hui Liu 0052, Huaimin Wang 0001
ICSE1
2024 Hardware Trojan Detection Based on Circuit Sequence Features with GRU Neural Network
abstract
The globalization of the IC industry has led to hardware designers commonly adopting Third-Party Intellectual Property cores (3PIPs) to reduce design costs and time. However, over-reliance on untrustworthy 3PIPs may compromise the autonomy of ICs and provide opportunities for the Hardware Trojans (HT) implantation, thus reducing the security and trustworthiness of hardware designs. Existing methods of artificial intelligence have drawbacks of dependence on a golden HT-free model, low accuracy, and long detection time when identifying and detecting HTs in large-scale gate-level netlists (GLNs). To enhance the accuracy of HT detection in IP cores and reduce detection time, we propose a novel HT detection method using the controllability metric and the Gated Recurrent Unit (GRU) neural network to extract circuit sequence features and detect HTs. With the advantages of narrowing down the circuit detection range by the controllability metric and the simplicity and efficiency of the GRU neural network, our method achieves an average TPR detection accuracy of 95.5% and an average TNR detection accuracy of 99.6%. Compared with other existing methods, our method has high detection accuracy and effectively reduces the detection time.
Yang Zhang 0026, Xing Hu 0012, Shaoqing Li
ISCC2
2024 CA4TJ: Correlational Analysis for Always-On Information-Leakage Hardware Trojan Detecting
abstract
The presence of always-on hardware Trojans (HTs) capable of continuously leaking sensitive information poses a significant threat to security, particularly in circuits designed for sensitive information protection. These HTs are engineered to operate stealthily, making their detection a formidable challenge. While existing hardware Trojan detection methodologies often rely on the identification of rare triggering characteristics, they prove inadequate for effectively identifying always-on HTs. This paper presents a novel approach to detect such HTs, which persistently leak critical information within circuits designed for sensitive information protection. By analyzing the characteristics of always-on information-leakage HTs, we propose CA4TJ, a corresponding quantifying metric to evaluate the correlation between sensitive information and the leaked port within gate-level netlists. Distinguished from conventional methods, the proposed approach eliminates the need for a reference model, thereby enhancing its practicality in real-world applications. The efficacy of this method is validated through extensive evaluations conducted on benchmark circuits, including those with intricate designs comprising up to one hundred thousand gates.
Xing Hu 0012, Yang Zhang 0026, Shaoqing Li
ITC-Asia2
2024 ELSeM: An Efficient and Lightweight Security Mechanism for DSP
abstract
Digital Signal Processor (DSP) systems frequently manage highly sensitive data within critical sectors such as telecommunications, healthcare, and defense. However, DSP are vulnerable to attacks such as unauthorized access, reverse engineering and software vulnerabilities. To enhance the security of sensitive data without significantly increasing DSP hardware overhead, we propose ELSeM, an efficient and lightweight security mechanism specifically tailored for DSP. It utilizes Direct Memory Access(DMA) in conjunction with electronic fuse(efuse) and SM4 modules to encrypt/decrypt key data through data transmission process and obtain security permissions after successful key comparison. ELSeM offers several advantages: 1) Enables efficient encryption and decryption of data with variable pipeline stages of SM4 during DMA transfer process; 2) Lightweight, consumes minimal resources, which is specifically tailored for DSP; 3) Restrict arbitrary key access and employ redundancy backup which mitigates unauthorized chip access, thereby significantly enhancing chip security. We implement the ELSeM with RTL and conduct tape-out, verifying its security at chip level. The experimental results indicate that the hardware overhead of ELSeM is minimal, approximately 0.1%, and the activation or deactivation of SM4 functionality has virtually no effect on DMA throughout.
Yijing Peng, Yang Zhang 0026, Qijin Zhu
ITC-Asia2
2024 Understanding the Challenges of Data Management in the AI Application Development
abstract
With the development of AI model ecology pioneered by the large language models, data plays a more important and diversified role than the traditional one. In recent years, data problems have been increasingly noticed and voiced by developers, and it is crucial to understand the existing challenges that they are facing in reality, as well as the potential features of these problems. Unfortunately, we currently have relatively little knowledge in this area. To fill this gap, we conduct an empirical study based on 2,154 posts and 785 issues by collecting model data problems from three online communities (i.e., GitHub, Stack Overflow, and Hugging Face). We present the first taxonomy of data problems around model production encountered by developers, covering 11 topics in 4 categories. Then, we analyze the popularity and the difficulty of the posts and issues in the derived topics. We distill many findings from the study and present the practical implications from the perspectives of different audiences.
Junchen Li, Yang Zhang 0026, Kele Xu, Tao Wang 0006, Huaimin Wang 0001
JCC2
2024 How Do Developers Adapt Code Snippets to Their Contexts? An Empirical Study of Context-Based Code Snippet Adaptations
abstract
Reusing code snippets from online programming Q&A communities has become a common development practice, in which developers often need to adapt code snippets to their code contexts to satisfy their own programming needs. However, how developers make these code adaptations based on contexts is still unclear. To bridge this gap, we first conduct a semi-structured interview of 21 developers to investigate their adaptation practices and perceived challenges during this process. The result suggests that code snippet adaptation is a challenging and exhausting task for developers, as they should tailor the snippets to guarantee their correctness and quality with laborious work. We also note that developers all resort to their intra-file context to complete adaptations, which motivates us to further study how developers performed context-based adaptations (CAs) in real scenarios. To this end, we conduct a quantitative study on an adaptation dataset comprising 300 code snippet reuse cases with 1,384 adaptations from Stack Overflow to GitHub. For each adaptation, we manually annotate its intention and relationship with the context. Based on our annotated data, we employ frequent itemset mining to obtain four CA patterns from our dataset, includingFortification,Code Wiring,Attribute-izationandParameterization. Our main findings reveal that: (1) more than half of the code snippet reuse cases include CAs and 23.3% of the adaptations are CAs; (2) more than half of the CAs are corrective adaptations and variable is the primary adapted language construct; (3) attribute is the most frequently utilized context and 88% of the local contexts are within the nearest 10 LOCs; and (4) CAs towards different intentions are repetitive, which are useful for automatic adaptation. Overall, our study provides valuable insights into code snippet adaptation and has important implications for research, practice, and tool design.
Tanghaoran Zhang, Yao Lu 0003, Yue Yu 0001, Xinjun Mao, Yang Zhang 0026
IEEE Trans. Software Eng.5
2023 TLDBERT: Leveraging Further Pre-Trained Model for Issue Typed Links Detection
abstract
Issue links are crucial for promoting software information flowing and development efficiency, but manually conducting typed links detection (TLD) is time-consuming and error-prone due to the numerous candidate issues. The general pre-trained NLP models, such as BERT, provide promising automated approaches for TLD task after fine-tuning. In addition, the further pre-training method, which is an intermediate pre-training process on the in-domain corpora, has been proven to further improve the performance of pre-trained models on many downstream tasks. In this paper, we apply the further pre-training method on TLD task and assess the improved performance. We further pre-train BERT by utilizing the in-domain corpora constructed by Jira issues dataset, and then finetune it on the links dataset to obtain the more applicable model for TLD task - TLDBERT. The experimental results indicate the statistically significant improvement in the performance of TLDBERT, with a range of 0.4%-6.6% for accuracy and 0.1%-9.7% for macro F1-score. Finally, based on correlation analysis, we conclude that the coverage and the ratio of issues to users are significantly correlated with the improvement, although the two directions are opposite.
Huaian Zhou, Tao Wang 0006, Yang Zhang 0026
APSEC3
2022 EpiMCBN: A Kind of Epistasis Mining Approach Using MCMC Sampling Optimizing Bayesian Network
abstract
Proposing a more effective and accurate epistatic loci detection method is of great significance in improving crop quality, disease treatment, etc. Due to the characteristics of high accuracy and processing non-linear relationship, Bayesian network (BN) has been widely used in constructing the network of SNPs and phenotypes and thus to mine epistasis. However, the shortcoming of BN is that the search space is too large and unable to process large-scale SNPs. In this work, we propose a kind of epistasis mining method using Markov Chain Monte Carlo (MCMC) sampling optimizing Bayesian network (EpiMCBN). Firstly, we use the space of node order composed of SNPs and phenotype to replace the space of network structure. Then MCMC algorithm is used to do sampling to generate multiple different initial orders in linear space or partial space. We use Markov state transition matrix to transfer the initial samples along the Markov chain, thus obtaining multiple order samples. Then we use the $\alpha$-BICBN scoring function to score the Bayesian networks corresponding to these node orders. Through estimating the probability of edge occurrence in the Bayesian networks, we get an approximate Bayesian network of SNPs and phenotype, then obtain the epistatic loci affecting phenotype. Finally, we compare EpiMCBN with the current popular epistasis mining algorithms using both simulated and real age-related macular disease (AMD) datasets. Experiment results show that EpiMCBN has better epistasis detection accuracy, lower false positive rate, and higher F1-score compared to other methods. Availability and implementation: Source code and dataset are available at: http://122.205.95.139/EpiMCBN/.
Keqin Li 0001, Yang Zhang 0026, Junli Deng, Jianxiao Liu
BIBM3
2022 A Preliminary Study of Bots Usage in Open Source Community
abstract
Bots are seen as a promising approach in software development, which help to deal with the ever-increasing complexity of modern software engineering and development. The number of bots in open source community, such as GitHub, has expanded substantially over the last three years. Due to its increasing popularity, it is essential to characterize the current usage of bots in practices. In this paper, we present an empirical study of bots usage in GitHub community. By analyzing 7,399 projects from GitHub, we find that 4,148 (56%) projects have used bots. Through automatic identification and manual detection, we collect a total of 196 bots. We then analyze and classify them into 4 categories and 14 topics. Finally, we discuss some raised implications for bots in current GitHub community.
Anze Gao, Yang Zhang 0026, Tao Wang 0006
Internetware3
2022 Understanding and Predicting Docker Build Duration: An Empirical Study of Containerized Workflow of OSS Projects
abstract
Docker building is a critical component of containerized workflow, which automates the process by which sources are packaged and transformed into container images. If not run properly, Docker builds can bring long durations (i.e., slow builds), which increases the cost in human and computing resources, and thus inevitably affect the software development. However, the current status and remedy for the duration cost in Docker builds remain unclear and need an in-depth study. To fill this gap, this paper provides the first empirical investigation on 171,439 Docker builds from 5,833 open source software (OSS) projects. Starting with an exploratory study, the Docker build durations can be characterized in real-world projects, and the developers’ perceptions of slow builds are obtained via a comprehensive survey. Driven by the results of our exploratory study, we propose a prediction modeling of Docker build duration, leveraging 27 handcrafted features from build-related context and configuration and 8 regression algorithms for the prediction task. Our results demonstrate that Random Forest model provides the superior performance with a Spearman’s correlation of 0.781, outperforming the baseline random model by 82.9% in RMSE, 90.6% in MAE, and 94.4% in MAPE, respectively. The implications of this study will facilitate research and assist practitioners in improving the Docker build process.
Yiwen Wu 0001, Yang Zhang 0026, Kele Xu, Tao Wang 0006, Huaimin Wang 0001
ASE2
2022 On the Way to Microservices: Exploring Problems and Solutions from Online Q&A Community
abstract
Microservice architecture is a dominant architectural style in SaaS industry, which helps to develop a single application as a collection of independent, well-defined, and inter-communicating services. The number of microservice-related questions in Q&Awebsites, such as Stack Overflow, has expanded substantially over the last years. Due to its increasing popularity, it is essential to understand the existing problems that microservice developers face in practices as well as the potential solutions to these problems. Such an investigation of problems and solutions is vital for long-term, impactful, and qualified research and practices in microservice community. Unfortunately, we currently know relatively little about such knowledge. To fill this gap, we conduct a large-scale in-depth empirical study on 17,522 Stack Overflow microservice-related posts. Our analysis leads to the first taxonomy of microservice-related topics based on the software development process. By analyzing the characteristics of the accepted answers, we find that there are fewer experts in the microservice than other domains, and such a phenomenon is most significant with respect to the microservice design phase. Furthermore, we perform manual analysis on 6,013 answers accepted by developers and distill 47 general solution strategies for different microservice-related problems, 22 of which are proposed for the first time. For instance, several problems inherent in the delivery phase can be lessened by referring to external sources like GitHub code examples. Our findings can therefore facilitate research and development on emerging microservice systems.
Menghan Wu 0001, Yang Zhang 0026, Shangwen Wang, Zhang Zhang 0005, Xin Xia 0001, Xinjun Mao
SANER2
2022 Recommending Base Image for Docker Containers based on Deep Configuration Comprehension
abstract
Docker containers are being widely used in large-scale industrial environments. In practice, developers must manually specify the base image in the dockerfile in the process of container creation. However, finding the proper base image is a nontrivial task because manually searching is time-consuming and easily leads to the use of unsuitable base images, especially for newcomers. There is still a lack of automatic approaches for recommending related base image for developers through dockerfile configuration. To tackle this problem, this paper makes the first attempt to propose a neural network approach named DCCimagerec which is based on deep configuration comprehension. It aims to use the structural configuration features of dockerfile extracted by AST and path-attention model to recommend potentially suitable base image. The evaluation experiments based on about 83,000 dockerfiles show that DCCimagerec outperforms multiple baselines, improving Precision by 7.5%-67.5%, Recall by 6.2%-106.6%, and F1 by 7.5%-150.2%.
Yinyuan Zhang, Yang Zhang 0026, Xinjun Mao, Yiwen Wu 0001, Bo Lin 0011, Shangwen Wang
SANER2
2022 Motivation Under Gamification: An Empirical Study of Developers' Motivations and Contributions in Stack Overflow
abstract
To encourage developers' volunteer contributions, modern programming question and answer (Q&A) sites like Stack Overflow (SO) employ gamified incentive mechanisms such as reputation and badges. Understanding developers' motivations in the presence of gamification and the relationship between their motivations and behavioral outcomes is crucial for community building and designing good incentive mechanisms. Grounded on self-determination theory, we conducted a survey with 938 developers who participate in SO to understand their participation motivations and incentive perceptions. By connecting the survey responses with the SO data, we quantitatively analyzed how the developers' motivations and satisfaction of needs relate to their effort and contribution quality. Our main findings are as follows: (1) despite the presence of gamified incentive mechanisms, developers are mainly motivated by intrinsic motivation to participate in SO; (2) developers who have strong motivations to gain gamification rewards are associated with higher intrinsic and integrated motivations, while developers with more development experiences are less motivated by the gamified incentives; (3) both extrinsic motivations (in terms of career prospects) and intrinsic motivations (regarding self-improvement and helping others) can motivate developers to make high-quantity and high-quality contributions; and (4) high-level satisfaction of needs for competency and autonomy has a positive effect on developers making high-quantity and high-quality contributions and addressing difficult problems. Based on these findings, we discuss implications for developer motivation and gamification in the crowdsourcing context and for the mechanism design of gamified crowdsourced platforms.
Yao Lu 0003, Xinjun Mao, Minghui Zhou 0001, Yang Zhang 0026, Zude Li, Tao Wang 0006, Gang Yin, Huaimin Wang 0001
IEEE Trans. Software Eng.4
2021 Optimizing Data Locality by Executor Allocation in Reduce Stage for Spark Framework
Zhongming Fu, Mengsi He, Zhuo Tang, Yang Zhang 0026
PDCAT4
2020 Sia-RAE: A Siamese Network based on Recursive AutoEncoder for Effective Clone Detection
abstract
Code clone helps improving programming productivity, while at the same time leads to many negative effects on software maintenance. Many approaches have been proposed to detect clones, but most of them fail on detecting low similarity code snippets. In this paper, we propose a Siamese network which links two recursive autoencoders (RAE) with a comparator network for clone detection. The unweighted recursive autoencoder is designed to learn code representation and then the comparator network is employed for similarity evaluation. In this Siamese network, it takes full advantages of lexical, semantic and structure information, and achieves high accuracy in revealing tiny similarity. We conduct comprehensive experiments on BigCloneBench using tagged clones as well as the whole repository respectively. The results suggest that our approach achieves good accuracy, and its recall reaches 93.02 % in WT3/T4, which outperforms state-of-the-art.
Chenhui Feng, Tao Wang 0006, Yue Yu 0001, Yang Zhang 0026, Huaimin Wang 0001
APSEC4
2020 Dockerfile Changes in Practice: A Large-Scale Empirical Study of 4, 110 Projects on GitHub
abstract
Docker is one of the most popular containerization tools in current DevOps practice. Particularly, Dockerfile plays an important role in the Docker-based software development process by specifying the commands and build environment of Docker containers. As a project progresses through its development stages, the content of the Dockerfile may be revised many times. Previous studies have examined Dockerfile usage in open-source projects. However, little is known about the details of Dockerfile changes in practice. In this paper, we conduct an empirical study on Dockerfile changes for 4,110 open-source projects hosted on GitHub. Based on the Dockerfile data, we measure the frequency, magnitude, and instructions of Dockerfile changes and report how Dockerfile co-changed with other files. To explore the relationship between Dockerfile changes and project outcomes, i.e., popularity, success, and productivity, we also develop regression models, by controlling for various confounds. Our findings help to characterize and understand Dockerfile changes and motivate the need for collecting more empirical evidence.
Yiwen Wu 0001, Yang Zhang 0026, Tao Wang 0006, Huaimin Wang 0001
APSEC2
2020 A Low-Latency Successive Cancellation Hybrid Decoder for Convolutional Polar Codes
abstract
By adopting successive cancellation list decoding (SCL), polar codes demonstrate competitive error correction performance over LDPC and Turbo codes. However, SCL decoding suffers from high computational complexity and long decoding latency, especially when the list size is very large. Successive cancellation flip (SCF), as another decoding algorithm that can achieve high error correction performance, has a complexity that is close to that of successive cancellation (SC) decoding. With the observation that SCL and SCF decoding are similar at giving more chances to inspect possible codewords simultaneously or sequentially, a novel hybrid decoder is proposed in this paper, which essentially combines the ideas of SCF and SCL decoders. Moreover, in order to compensate for the degradation of performance caused by the reduction of path splitting and further reduce the decoding latency, the convolutional polar codes are adopted with a designed bit-flipping set. Simulation results demonstrate that the proposed decoder achieves the reduction of decoding latency while attaining better performance than conventional CRC-aided SCL decoder.
Yu Wang 0068, Shikai Qiu, Lirui Chen, Yang Zhang 0026, Cang Liu, Zuocheng Xing
ICASSP5
2020 Using Configuration Semantic Features and Machine Learning Algorithms to Predict Build Result in Cloud-Based Container Environment
abstract
Container technologies are being widely used in large scale production cloud environments, of which Docker has become the de-facto industry standard. In practice, Docker builds often break, and a large amount of efforts are put into troubleshooting broken builds. Prior studies have evaluated the rate at which builds in large organizations fail. However, there is still a lack of early warning methods for predicting the Docker build result before the build starts. This paper provides a first attempt to propose an automatic method named PDBR. It aims to use the configuration semantic features extracted by AST and the machine learning algorithms to predict build result in the cloud-based container environment. The evaluation experiments based on more than 36,000 collected Docker builds show that PDBR achieves 73.45%-91.92% in F1 and 29.72%-72.16% in AUC. We also demonstrate that different ML classifiers have significant and large effects on the PDBR AUC performance.
Yiwen Wu 0001, Yang Zhang 0026, Junsheng Chang, Bo Ding 0001, Tao Wang 0006, Huaimin Wang 0001
ICPADS2
2020 Haste Makes Waste: An Empirical Study of Fast Answers in Stack Overflow
abstract
Modern programming question & answer (Q&A) sites such as Stack Overflow (SO) employ gamified mechanisms to stimulate volunteers' contributions. To maximize the chances of winning gamification rewards such as reputation and badges, a portion of users race to post answers as quickly as possible (i.e., fast answers or FAs), which makes SO the fastest Q&A site; however, this behavior may affect the contribution quality as well. In this paper, we report on a large-scale, mixed-methods empirical study of the gamification-influenced FA phenomenon in SO. We first quantitatively investigate the popularity of the phenomenon and user behaviors regarding FAs. Then, we study the quality of FAs by using regression modeling and qualitatively analyzing 300 instances of FAs. Our main findings reveal that more than 70% and 90% of FAs are not edited by the answerers and other users, respectively, and that later incoming answers have lower chances of being voted on and accepted. Notably, we find that the answer length, code snippets length, and readability of FAs are significantly lower than those of non-fast answers. Although FAs have higher crowd assessment scores, they have no relationship with acceptance from the perspective of asker assessment, and a considerable portion of FAs solve the problem by interacting with the asker in the comments. These results help us better understand the effects of reward-based gamification on crowdsourced software engineering communitites and provide implications for designers of gamified systems.
Yao Lu 0003, Xinjun Mao, Minghui Zhou 0001, Yang Zhang 0026, Tao Wang 0006, Zude Li
ICSME4
2020 An Empirical Study of Multi-discussing Pattern in Open-Source Software Development
abstract
GitHub enables developers to expediently contribute their comments on multiple issues and switch their discussion between issues, i.e., multi-discussing. Discussing multiple issues simultaneously is able to enhance work efficiency. However, multi-discussing also relies on developers’ rationally allocating their focus, which may result in the different influence on the resolution of issues. Therefore, investigating how multi-discussing affects the issue resolution is a meaningful research question that can help developers understand the benefits and limitations of multi-discussing. Using quantitative and qualitative methods, this paper proposes a groundbreaking study of the impact of multi-discussing on issue resolution in GitHub. First, we collect and analyze data from 624 GitHub projects to explore how multi-discussing affects the overall issue resolution of the project. Further, we investigate how multi-discussing affects the resolution of a single issue. We find that multi-discussing is a common behavior in GitHub. Also, multi-discussing is connected to a shorter average issue resolution latency of the project. However, during a single issue resolution, more multi-discussing behaviors tend to bring longer issue resolution latency. We also conduct the qualitative analysis to explore the developers’ experiences and expectations of multi-discussing.
Cheng Yang 0004, Dongyang Hu, Yang Zhang 0026, Tao Wang 0006, Yue Yu 0001
Internetware3
2020 Exploring the Dependency Network of Docker Containers: Structure, Diversity, and Relationship
abstract
Container technologies are being widely used in large scale production cloud environments, of which Docker has become the de-facto industry standard. As a key step, containers need to define their dependent base image, which makes complex dependencies exist in a large number of containers. Prior studies have shown that references between software packages could form technical dependencies, thus forming a dependency network. However, little is known about the details of docker container dependency networks. In this paper, we perform an empirical study on the dependency network of docker containers from more than 120,000 dockerfiles. We construct the container dependency network and analyze its network structure. Further, we focus on the Top-100 dominant containers and investigate their subnetworks, including diversity and relationships. Our findings help to characterize and understand the container dependencies in the docker community and motivate the need for developing container dependency management tools.
Yinyuan Zhang, Yang Zhang 0026, Yiwen Wu 0001, Yao Lu 0003, Tao Wang 0006, Xinjun Mao
Internetware2
2020 A Paralleled Greedy LLL Algorithm for 16×16 MIMO Detection
abstract
This brief proposes a paralleled greedy Lenstra-Lenstra-Lovsz (PGLLL) algorithm for 16×16 MIMO detection. First, a paralleled constant-throughput scheme is designed for LLL algorithm. Then, greedy algorithm is adopted on this scheme to select the most urgent iterations for each stage. This selecting criterion outperforms others in that numerous iterations can be concurrently selected to reduce latency, and that the two factors of LLL potential and MIMO detection strategy are comprehensively considered by this criterion to improve bit-error-rate (BER) performance. Simulation indicates that the PGLLL can realize a comparable performance to the non-greedy algorithm and LLL algorithm with less iterations. Finally, this brief is the first to propose a hardware architecture with greedy LLL algorithm. This architecture is implemented with 65-nm 1P9M CMOS technology, which can work at a maximum frequency of 625 MHz to process 16×16 complex-valued matrices every 16 clocks. The latency is 362 ns. Comparison indicates that the proposed PGLLL architecture is superior to other existing works in terms of throughput and latency performance.
Lirui Chen, Yu Wang 0068, Zuocheng Xing, Shikai Qiu, Yang Zhang 0026
ISCAS6
2020 An Empirical Study of Build Failures in the Docker Context
abstract
Docker containers have become the de-facto industry standard. Docker builds often break, and a large amount of efforts are put into troubleshooting broken builds. Prior studies have evaluated the rate at which builds in large organizations fail. However, little is known about the frequency and fix effort of failures that occur in Docker builds of open-source projects. This paper provides a first attempt to present a preliminary study on 857,086 Docker builds from 3,828 open-source projects hosted on GitHub. Using the Docker build data, we measure the frequency of broken builds and report their fix time. Furthermore, we explore the evolution of Docker build failures across time. Our findings help to characterize and understand Docker build failures and motivate the need for collecting more empirical evidence.
Yiwen Wu 0001, Yang Zhang 0026, Tao Wang 0006, Huaimin Wang 0001
MSR2
2020 GitHub's milestone tool: A mixed-methods analysis on its use
abstract
Abstract Social coding site GitHub provides developers with many management tools to facilitate project maintenance and developer collaboration.Milestonetool, in particular, plays an important role in organizing and tracking progress on groups of issues or pull requests in a project. However, few research has analyzed the milestone tool, even though it has been used in practice for a long time. In this paper, we want to address this literature gap and present an ongoing work aimed at investigating the use of the milestone tool in GitHub open‐source projects. We conduct a mixed‐methods analysis in a large‐scale dataset of GitHub projects, to help developers gain some insights into the milestone tool, including its usage, benefits, and limitations. We quantitatively investigate the basic adoption of milestone tool and its correlation with project properties. We also survey developers to understand the reasons for using milestone tool or not and their perceptions of the milestone tool. We find that certain types of projects use milestone tool more than others. Adopting the milestone tool is associated with more commits, more releases, and more project popularity, but the current milestone tool also has some limitations. These observations can then be forwarded to the GitHub community for follow‐up and can result in them potentially making a better milestone tool.
Yang Zhang 0026, Huaimin Wang 0001, Yiwen Wu 0001, Dongyang Hu, Tao Wang 0006
J. Softw. Evol. Process.1
2020 iLinker: a novel approach for issue knowledge acquisition in GitHub projects
Yang Zhang 0026, Yiwen Wu 0001, Tao Wang 0006, Huaimin Wang 0001
World Wide Web1
2019 Exploring the Relationship Between Developer Activities and Profile Images on GitHub
abstract
In the GitHub platform, social media profile images are one of many visual components of developers. Besides, developer activities such as reporting issues or following other developers are regarded as important development and self-expression behaviors. However, to the best of our knowledge, no study has yet been conducted to study the relationship between GitHub developer activities and profile images. In this paper, we aim to investigate the relationship between developer activities and profile images to gain some insights into the developers' internal properties. During our experiments, we manually classify profile images into seven categories. Next, we investigate the relationship between developer's demographic information and developer activity. Further, using logistic regression analysis, when controlled for various variables, we statistically identify and quantify the relationships between developer activities and profile image categories. We find that several profile image categories significantly correlate with developer's demographic information and activities. We also provide a rich resource of research ideas for further study. Our examination and analysis provide insights into the developers' internal properties when using different profile images. Moreover, this study is the first step in understanding the relationship between developer activities and profile images on GitHub.
Yiwen Wu 0001, Yang Zhang 0026, Tao Wang 0006, Huaimin Wang 0001
Internetware2
2019 Algorithm and Architecture for Path Metric Aided Bit-Flipping Decoding of Polar Codes
abstract
Polar codes attract more and more attention of researchers in recent years, since its capacity achieving property. However, their error-correction performance under successive cancellation (SC) decoding is inferior to other modern channel codes at short or moderate blocklengths. SC-Flip (SCF) decoding algorithm shows higher performance than SC decoding by identifying possibly erroneous decisions made in initial SC decoding and flipping them in the sequential decoding attempts. However, it performs not well when there are more than one erroneous decisions in a codeword. In this paper, we propose a path metric aided bit-flipping decoding algorithm to identify and correct more errors efficiently. In this algorithm, the bit-flipping list is generated based on both log likelihood ratio (LLR) based path metric and bit-flipping metric. The path metric is used to verify the effectiveness of bit-flipping. In order to reduce the decoding latency and computational complexity, its corresponding pipeline architecture is designed. By applying these decoding algorithm and pipeline architecture, an improvement on error-correction performance can be got up to 0.25dB compared with SCF decoding at frame error rate of 10-4, with low average decoding latency.
Yu Wang 0068, Lirui Chen, Yang Zhang 0026, Zuocheng Xing
WCNC4
2019 A clustering-based approach for mining dockerfile evolutionary trajectories
Yang Zhang 0026, Huaimin Wang 0001, Vladimir Filkov
Sci. China Inf. Sci.1
2019 A novel approach for recommending semantically linkable issues in GitHub projects
Yang Zhang 0026, Yiwen Wu 0001, Tao Wang 0006, Huaimin Wang 0001
Sci. China Inf. Sci.1
2019 Multi-reviewing pull-requests: An exploratory study on GitHub OSS projects
Dongyang Hu, Yang Zhang 0026, Junsheng Chang, Gang Yin, Yue Yu 0001, Tao Wang 0006
Inf. Softw. Technol.2
2019 GARDENIA: A Graph Processing Benchmark Suite for Next-Generation Accelerators
abstract
This article presents the Graph Algorithm Repository for Designing Next-generation Accelerators (GARDENIA), a benchmark suite for studying irregular graph algorithms on massively parallel accelerators. Applications with limited control and data irregularity are the main focus of existing generic benchmarks for accelerators, while available graph processing benchmarks do not apply state-of-the-art algorithms and/or optimization techniques. GARDENIA includes emerging graph processing workloads from graph analytics, sparse linear algebra, and machine-learning domains, which mimic massively multithreaded commercial programs running on modern large-scale datacenters. Our characterization shows that GARDENIA exhibits irregular microarchitectural behavior, which is quite different from structured workloads and straightforward-implemented graph benchmarks.
Zhen Xu 0004, Xuhao Chen 0001, Jie Shen 0003, Yang Zhang 0026, Cheng Chen 0005, Canqun Yang
ACM J. Emerg. Technol. Comput. Syst.4
2018 Recommending Similar Bug Reports: A Novel Approach Using Document Embedding Model
abstract
In the software development, it is not uncommon to find that several bug reports are related to many common code files, i.e., similar bugs. Similar bug recommendation is a meaningful task which can assist developers in bug triaging and fixing. As the state of the art, Yang et al.'s work presented an approach that combines TF-IDF method with word embedding model and achieved a good result. To further improve the performance of their approach, in this paper, we propose a novel approach using Document Embedding model. In our preliminary evaluation, we conduct the experiment on 13,090 bug reports from the Eclipse platform and the results show that our approach outperforms Yang et al.'s, with 7.89-8.96% of improvement.
Dongyang Hu, Tao Wang 0006, Junsheng Chang, Gang Yin, Yue Yu 0001, Yang Zhang 0026
APSEC7
2018 Multi-Discussing across Issues in GitHub: A Preliminary Study
abstract
Social coding sites like GitHub has enabled developers to easily contribute their comments on multiple issues and switch their discussion between issues, i.e., multi-discussing. Discussing multiple issues simultaneously may enhance the work efficiency of developers. However, multi-discussing also relies on developers' rationally allocating their time and focus, which may bring different influence to the resolution of issues. Therefore, investigating how multi-discussing affects the issue resolution is a meaningful research question which can help developers understand the benefits and limitations when they switch their discussion between issues. In this paper, we present a preliminary study of the impact of multi-discussing on issue resolution in GitHub projects, by using quantitative methods. First, we collect and analyzed data from 631 GitHub projects to explore how multi-discussing affects the average resolution latency of project issues. Further, we develop method for measuring the rate and breadth of a developers' discussionswitching behavior, and we use regression modeling to study how discussion-switching affects the single issue resolution latency. We find that multi-discussing is a common behavior of developers in GitHub projects. Also, multi-discussing is associated with shorter average issue resolution latency of project. However, during a single issue resolution, more participants' discussion-switching tend to bring longer issue resolution latency. Our study motivates the need for further research on the multi-discussing.
Dongyang Hu, Tao Wang 0006, Junsheng Chang, Gang Yin, Yang Zhang 0026
APSEC5
2018 An Insight Into the Impact of Dockerfile Evolutionary Trajectories on Quality and Latency
abstract
Containerization is a software development approach aimed at packaging an application together with all its dependencies and execution environment in a light-weight, self-contained unit, of which Docker has become the de-facto industry standard. By defining the specific Docker image architecture and building orders, dockerfile plays an important role in the Docker-based containerization process. Understanding the evolution of dockerfile and which dockerfile architecture attributes enhance dockerfile quality and reduce image build latency can benefit the efficient processing of containerization. In this paper, we perform an empirical study on a large dataset of 2,840 projects to shed light on the impact of dockerfile evolutionary trajectories on quality and latency in the Docker-based containerization. Based on the six categories of dockerfile evolutionary trajectories we discovered, we build two regression models to explore the impact of dockerfile evolutionary trajectories and specific architecture attributes on dockerfile quality and image build latency, which derives a number of suggestions for practitioners.
Yang Zhang 0026, Gang Yin, Tao Wang 0006, Yue Yu 0001, Huaimin Wang 0001
COMPSAC (1)1
2018 Who Will Become a Long-Term Contributor?: A Prediction Model based on the Early Phase Behaviors
abstract
The continuous contribution from peripheral participants is crucial for the success of open source projects. Thus, how to identify the potential Long-Term Contributors (LTC) early and retain them is of great importance. We propose a prediction model to measure the chance for an individual to become a LTC contributor through his capacity, willingness, and the opportunity to contribute at the time of joining. Using data of Rails hosted on GitHub, we find that the probability for a new joiner to become a LTC is associated with his willingness and environment. Specifically, future LTCs tend to be more active and show more community-oriented attitude than other joiners during their first month. This implies that the interaction between individual's attitude and project's climate are associated with the odds that an individual would become a valuable contributor or disengage from the project. We evaluated our prediction model by using the 10 cross-validation method. Results show that our model archives the mean AUC as 0.807, which is valuable for OSS projects to identify potential long-term contributors and adopt better strategies to retain them for continuous contribution.
Tao Wang 0006, Yang Zhang 0026, Gang Yin, Yue Yu 0001, Huaimin Wang 0001
Internetware2
2018 One size does not fit all: an empirical study of containerized continuous deployment workflows
abstract
Continuous deployment (CD) is a software development practice aimed at automating delivery and deployment of a software product, following any changes to its code. If properly implemented, CD together with other automation in the development process can bring numerous benefits, including higher control and flexibility over release schedules, lower risks, fewer defects, and easier on-boarding of new developers. Here we focus on the (r)evolution in CD workflows caused by containerization, the virtualization technology that enables packaging an application together with all its dependencies and execution environment in a light-weight, self-contained unit, of which Docker has become the de-facto industry standard. There are many available choices for containerized CD workflows, some more appropriate than others for a given project. Owing to cross-listing of GitHub projects on Docker Hub, in this paper we report on a mixed-methods study to shed light on developers' experiences and expectations with containerized CD workflows. Starting from a survey, we explore the motivations, specific workflows, needs, and barriers with containerized CD. We find two prominent workflows, based on the automated builds feature on Docker Hub or continuous integration services, with different trade-offs. We then propose hypotheses and test them in a large-scale quantitative study.
Yang Zhang 0026, Bogdan Vasilescu, Huaimin Wang 0001, Vladimir Filkov
ESEC/SIGSOFT FSE1
2018 Correlation-based software search by leveraging software term database
Gang Yin, Tao Wang 0006, Yang Zhang 0026, Yue Yu 0001, Huaimin Wang 0001
Frontiers Comput. Sci.4
2018 Locality based warp scheduling in GPGPUs
Yang Zhang 0026, Zuocheng Xing, Cang Liu, Chuan Tang
Future Gener. Comput. Syst.1
2018 CWLP: coordinated warp scheduling and locality-protected cache allocation on GPUs
abstract
As we approach the exascale era in supercomputing, designing a balanced computer system with a powerful computing ability and low power requirements has becoming increasingly important. The graphics processing unit (GPU) is an accelerator used widely in most of recent supercomputers. It adopts a large number of threads to hide a long latency with a high energy efficiency. In contrast to their powerful computing ability, GPUs have only a few megabytes of fast on-chip memory storage per streaming multiprocessor (SM). The GPU cache is inefficient due to a mismatch between the throughput-oriented execution model and cache hierarchy design. At the same time, current GPUs fail to handle burst-mode long-access latency due to GPU’s poor warp scheduling method. Thus, benefits of GPU’s high computing ability are reduced dramatically by the poor cache management and warp scheduling methods, which limit the system performance and energy efficiency. In this paper, we put forward a coordinated warp scheduling and locality-protected (CWLP) cache allocation scheme to make full use of data locality and hide latency. We first present a locality-protected cache allocation method based on the instruction program counter (LPC) to promote cache performance. Specifically, we use a PC-based locality detector to collect the reuse information of each cache line and employ a prioritised cache allocation unit (PCAU) which coordinates the data reuse information with the time-stamp information to evict the lines with the least reuse possibility. Moreover, the locality information is used by the warp scheduler to create an intelligent warp reordering scheme to capture locality and hide latency. Simulation results show that CWLP provides a speedup up to 19.8% and an average improvement of 8.8% over the baseline methods.
Yang Zhang 0026, Zuocheng Xing, Cang Liu, Chuan Tang
Frontiers Inf. Technol. Electron. Eng.1
2018 Internal quality assurance for external contributions in GitHub: An empirical investigation
abstract
Abstract For popular open‐source software projects, there are always a large number of worldwide developers who have been glued to making code contributions, while most of these developers play the role of casual contributors because of their very limited code commits. The frequent turnover of such a group of developers and the wide variations in their coding experiences challenge the project management on code and quality. This paper aims to investigate the status quo of internal quality assurance for external contributions in social coding sites. We first conducted a case study of 21 popular GitHub projects to estimate the code quality of the casual contributors. The quantitative results show that the casual contributors introduced greater quantity and severity of code quality issues than the main contributors; the developers who contribute to different projects as main and casual contributors did not perform significantly differently in terms of their code quality. On the basis of these findings, we further conducted a survey of 81 developers on GitHub to understand their practices on internal quality assurance. The qualitative results expose some limitations of present internal quality control for external contributions in GitHub. Finally, we discuss an alternative quality management paradigm: Continuous Inspection for industrial practices.
Yao Lu 0003, Xinjun Mao, Zude Li, Yang Zhang 0026, Tao Wang 0006, Gang Yin
J. Softw. Evol. Process.4
2017 Social media in GitHub: the role of @-mention in assisting software development
Yang Zhang 0026, Huaimin Wang 0001, Gang Yin, Tao Wang 0006, Yue Yu 0001
Sci. China Inf. Sci.1
2017 Hardware Architecture Based on Parallel Tiled QRD Algorithm for Future MIMO Systems
abstract
QR decomposition (QRD) has been a vital component in the transceiver processor of future multiple-input multiple-output (MIMO) systems, in which antenna configuration will be more and more flexible. Therefore, the QRD hardware architecture in the future MIMO systems should be more flexible to meet various antenna configurations. Unfortunately, the existing QRD hardware architectures mainly focus on the matrix of one or several fixed sizes. This paper presents a new triangular systolic array QRD hardware architecture based on parallel tiled QRD algorithm to decompose an 8 × 8 real matrix. The designed hardware architecture is flexible and can be used in various MIMO systems, in which the number of antennas is smaller than 4. This paper also proposes a modified algorithm for the bottleneck operations of parallel tiled QRD algorithm to reduce the hardware overhead. To further reduce the hardware overhead, the Newton-Raphson algorithm is adopted in the proposed algorithm. The implementation results show that the normalized processing latency performance and the normalized processing efficiency performance of the designed QRD hardware architecture both are better than most of the existing QRD hardware architectures. To the best of our knowledge, the hardware architecture presented in this paper achieves the superior normalized QRD rate performance to the existing QRD hardware architectures.
Cang Liu, Chuan Tang, Zuocheng Xing, Luechao Yuan, Yang Zhang 0026
IEEE Trans. Very Large Scale Integr. Syst.5
2016 Does the Role Matter? An Investigation of the Code Quality of Casual Contributors in GitHub
abstract
For popular Open Source Software (OSS) projects there are always a large number of worldwide developers who have been glued to making code contributions, while most of these developers play the role of casual contributors due to their very limited code commits (for fixing defects and enhancing features, casually). The frequent turnover of such group of casual developers and the wide variations among their coding experiences challenge the project management on code and quality.This paper describes a case study which aims to estimate the quality of code made by casual contributors in 21 popular GitHub projects. The results of this case study show that: (1) casual contributors introduced greater quantity and severity of Code Quality Issues (CQIs) than main contributors; (2) developers who contribute in different projects as main and casual contributors didn't perform statistically differently in terms of code quality; (3) casual contributors who have few project stars introduced more CQIs than those who have many. Furthermore, the paper lists the CQI categories which are most frequently introduced by casual contributors in the investigated projects. These findings provide valuable insights into code quality in the OSS context, and can guide OSS developers in improving the quality of the code contributions.
Yao Lu 0003, Xinjun Mao, Zude Li, Yang Zhang 0026, Tao Wang 0006, Gang Yin
APSEC4
2016 Query reformulation by leveraging crowd wisdom for scenario-based software search
abstract
The Internet-scale open source software (OSS) production in various communities are generating abundant reusable resources for software developers. However, how to retrieve and reuse the desired and mature software from huge amounts of candidates is a great challenge: there are usually big gaps between the user application contexts (that often used as queries) and the OSS key words (that often used to match the queries). In this paper, we define the scenario-based query problem for OSS retrieval, and then we propose a novel approach to reformulate the raw query by leveraging the crowd wisdom from millions of developers to improve the retrieval results. We build a software-specific domain lexical database based on the knowledge in open source communities, by which we can expand and optimize the input queries. The experiment results show that, our approach can reformulate the initial query effectively and outperforms other existing search engines significantly at finding mature software.
Tao Wang 0006, Yang Zhang 0026, Yun Zhan, Gang Yin
Internetware3
2015 Exploring the Use of @-mention to Assist Software Development in GitHub
abstract
Recently, many researches propose that social media tools can promote the collaboration among developers, which are beneficial to the software development. Nevertheless, there is little empirical evidence to confirm that using @-mention has indeed a beneficial impact on the issues in GitHub. In this paper, we analyze the data from GitHub and give some insights on how @-mention is used in the issues (general-issues and pull-requests). Our statistical results indicate that, @-mention attracts more participants and tends to be used in the difficult issues. @-mention favors the solving process of issues by enlarging the visibility of issues and facilitating the developers' collaboration. In addition to this global study, our study also build a @-network based on the @-mention database we extract. Through the @-network, we can mine the relationships and characteristics of developers in GitHub's issues.
Yang Zhang 0026, Huaimin Wang 0001, Gang Yin, Tao Wang 0006, Yue Yu 0001
Internetware1
2015 Evaluating Bug Severity Using Crowd-based Knowledge: An Exploratory Study
abstract
In bug tracking system, the high volume of incoming bug reports poses a serious challenge to project managers. Triaging these bug reports manually consumes time and resources which leads to delaying the resolution of important bugs. StackOverflow is the most popular crowdsourcing Q&A community with plenty of bug-related posts. In this paper, we explore the correlation between bug severity and the crowd attributes of linked posts. Two typical types of projects' bug repositories are studied here, e.g. Mozilla (user-centric project) and Eclipse (developer-centric project). Our results show that the bug severity is consistent with the crowd-based knowledge both in Mozilla and Eclipse, i.e. the linked posts of severe bugs have higher score etc. in StackOverflow than non-severe bugs. This interesting phenomenon inspires us that we can optimize the existing evaluation methods of bug severity by incorporating the crowd-based knowledge from a third-party in future.
Yang Zhang 0026, Gang Yin, Tao Wang 0006, Yue Yu 0001, Huaimin Wang 0001
Internetware1
2014 A Exploratory Study of @-Mention in GitHub's Pull-Requests
abstract
Pull-request mechanism is an outstanding social development method in Git Hub. @-mention is a social media tool that deeply integrated with pull-request mechanism. Recently, many research results show that social media tools can promote the collaborative software development, but few work focuses on the impacts of @-mention. In this paper, we conduct an exploratory study of @-mention in pull-request based software development, including its current situation and benefits. We obtain some interesting findings which indicate that @-mention is beneficial to the processing of pull-request. Our work also proposes some possible research directions and problems of the @-mention. It helps the developers and researchers notice the significance of @-mention in the pull-request based software development.
Yang Zhang 0026, Gang Yin, Yue Yu 0001, Huaimin Wang 0001
APSEC (1)1
2014 Liquid: A Scalable Deduplication File System for Virtual Machine Images
abstract
A virtual machine (VM) has been serving as a crucial component in cloud computing with its rich set of convenient features. The high overhead of a VM has been well addressed by hardware support such as Intel virtualization technology (VT), and by improvement in recent hypervisor implementation such as Xen, KVM, etc. However, the high demand on VM image storage remains a challenging problem. Existing systems have made efforts to reduce VM image storage consumption by means of deduplication within a storage area network (SAN) cluster. Nevertheless, an SAN cannot satisfy the increasing demand of large-scale VM hosting for cloud computing because of its cost limitation. In this paper, we propose Liquid, a scalable deduplication file system that has been particularly designed for large-scale VM deployment. Its design provides fast VM deployment with peer-to-peer (P2P) data transfer and low storage consumption by means of deduplication on VM images. It also provides a comprehensive set of storage features including instant cloning for VM images, on-demand fetching through a network, and caching with local disks by copy-on-read techniques. Experiments show that Liquid's features perform well and introduce minor performance overhead.
Yang Zhang 0026, Yongwei Wu 0001, Kang Chen 0001, Jinlei Jiang, Keqin Li 0001
IEEE Trans. Parallel Distributed Syst.2