VLDB 2026 Research / reviewers in the wild / expert
Yiwen Wu 0001
dblp:12/4118-1
· DBLP profile ↗
14ranked-venue papers
6as first author
6since 2021 · last 2026
0000-0002-8652-116XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 11 · 5 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DockerFill: Automatically Completing Dockerfile Code With Syntax-Aware Multi-Task LearningabstractAs a kind of infrastructure-as-code, Dockerfile specifies the structure and functionality of a built Docker image and thus plays an important role in the containerized software development process. Nowadays developers need to spend extra time and effort configuring their Dockerfiles in addition to their regular coding work, which requires knowledge and skills orthogonal to those entailed in other software-related experiences. Poorly written Dockerfile code often introduces errors and maintenance costs. However, little automated support is available for assisting developers in configuring Dockerfiles. In this study, we first conduct an online survey to investigate Docker developers’ perceptions of Dockerfile writing, highlighting the needs and potential benefits of Dockerfile auto-completion techniques. Then, we introduceDOCKERFILL, a pre-trained model based approach that provides completion suggestions for Dockerfile-specific code.DOCKERFILLleverages multi-layer Transformer architecture with syntax-aware multi-task learning, which includes contextual file information and three pre-training tasks, i.e., masked language modeling, syntax type identification, and masked identifier prediction. To evaluateDOCKERFILL’s effectiveness, we collect a dataset of 6,350 high-quality real-world Dockerfiles. Our empirical results show that DOCKERFILL provides up to 52.38% accuracy for token-level completion and 19.69% exact match for line-level completion, outperforming the baselines by 7.32%-37.67% and 1.97%-19.69%, respectively. Also,DOCKERFILLobtains significantly higher human evaluation scores compared to the baselines. Yiwen Wu 0001, Yang Zhang 0026, Tao Wang 0006, Bo Ding 0001, Huaimin Wang 0001 |
IEEE Trans. Software Eng. | 1 |
| 2025 | What problems are MLOps practitioners talking about? A study of discussions in Stack Overflow forum and GitHub projects
Yang Zhang 0026, Yiwen Wu 0001, Tao Wang 0006, Bo Ding 0007, Huaimin Wang 0001 |
Inf. Softw. Technol. | 2 |
| 2024 | GHA-BFP: Framework for Automated Build Failure Prediction in GitHub ActionsabstractGitHub Actions (GHA), a powerful Continuous Integration and Continuous Deployment (CI/CD) service, has revolutionized the way developers automate tasks in the software development pipeline. Although GHA provides great convenience, if a GHA build fails, the time spent waiting for results and debugging is wasted, which can seriously affect development efficiency. In this study, we delve into GHA build results and introduce an automatic framework named GHA-BFP that uses ML models to predict the failure of GHA builds. Using GHA-BFP with Random Forest model, we achieved the highest performance in predicting the failure of GHA builds, with all key metrics (i.e., Accuracy, Precision, Recall, and F1 score) exceeding 75%. Furthermore, through ablation experiments, we have verified the essentiality of the four categories of input features. Lastly, we conducted an assessment of the importance of each individual input feature in relation to the model's predictive capabilities. Jiatai Li, Yang Zhang 0026, Tao Wang 0006, Yiwen Wu 0001 |
APSEC | 4 |
| 2024 | How do Developers Talk about GitHub Actions? Evidence from Online Software Development CommunityabstractContinuous integration, deployment and delivery (CI/CD) have become cornerstones of DevOps practices. In recent years, GitHub Action (GHA) has rapidly replaced the traditional CI/CD tools on GitHub, providing efficiently automated workflows for developers. With the widespread use and influence of GHA, it is critical to understand the existing problems that GHA developers face in their practices as well as the potential solutions to these problems. Unfortunately, we currently have relatively little knowledge in this area. To fill this gap, we conduct a large-scale empirical study of 6,590 Stack Overflow (SO) questions and 315 GitHub issues. Our study leads to the first comprehensive taxonomy of problems related to GHA, covering 4 categories and 16 sub-categories. Then, we analyze the popularity and difficulty of problem categories and their correlations. Further, we summarize 56 solution strategies for different GHA problems. We also distill practical implications of our findings from the perspective of different audiences. We believe that our study contributes to the research of emerging GHA practices and guides the future support of tools and technologies. Yang Zhang 0026, Yiwen Wu 0001, Tao Wang 0006, Hui Liu 0052, Huaimin Wang 0001 |
ICSE | 2 |
| 2022 | Understanding and Predicting Docker Build Duration: An Empirical Study of Containerized Workflow of OSS ProjectsabstractDocker building is a critical component of containerized workflow, which automates the process by which sources are packaged and transformed into container images. If not run properly, Docker builds can bring long durations (i.e., slow builds), which increases the cost in human and computing resources, and thus inevitably affect the software development. However, the current status and remedy for the duration cost in Docker builds remain unclear and need an in-depth study. To fill this gap, this paper provides the first empirical investigation on 171,439 Docker builds from 5,833 open source software (OSS) projects. Starting with an exploratory study, the Docker build durations can be characterized in real-world projects, and the developers’ perceptions of slow builds are obtained via a comprehensive survey. Driven by the results of our exploratory study, we propose a prediction modeling of Docker build duration, leveraging 27 handcrafted features from build-related context and configuration and 8 regression algorithms for the prediction task. Our results demonstrate that Random Forest model provides the superior performance with a Spearman’s correlation of 0.781, outperforming the baseline random model by 82.9% in RMSE, 90.6% in MAE, and 94.4% in MAPE, respectively. The implications of this study will facilitate research and assist practitioners in improving the Docker build process. Yiwen Wu 0001, Yang Zhang 0026, Kele Xu, Tao Wang 0006, Huaimin Wang 0001 |
ASE | 1 |
| 2022 | Recommending Base Image for Docker Containers based on Deep Configuration ComprehensionabstractDocker containers are being widely used in large-scale industrial environments. In practice, developers must manually specify the base image in the dockerfile in the process of container creation. However, finding the proper base image is a nontrivial task because manually searching is time-consuming and easily leads to the use of unsuitable base images, especially for newcomers. There is still a lack of automatic approaches for recommending related base image for developers through dockerfile configuration. To tackle this problem, this paper makes the first attempt to propose a neural network approach named DCCimagerec which is based on deep configuration comprehension. It aims to use the structural configuration features of dockerfile extracted by AST and path-attention model to recommend potentially suitable base image. The evaluation experiments based on about 83,000 dockerfiles show that DCCimagerec outperforms multiple baselines, improving Precision by 7.5%-67.5%, Recall by 6.2%-106.6%, and F1 by 7.5%-150.2%. Yinyuan Zhang, Yang Zhang 0026, Xinjun Mao, Yiwen Wu 0001, Bo Lin 0011, Shangwen Wang |
SANER | 4 |
| 2020 | Dockerfile Changes in Practice: A Large-Scale Empirical Study of 4, 110 Projects on GitHubabstractDocker is one of the most popular containerization tools in current DevOps practice. Particularly, Dockerfile plays an important role in the Docker-based software development process by specifying the commands and build environment of Docker containers. As a project progresses through its development stages, the content of the Dockerfile may be revised many times. Previous studies have examined Dockerfile usage in open-source projects. However, little is known about the details of Dockerfile changes in practice. In this paper, we conduct an empirical study on Dockerfile changes for 4,110 open-source projects hosted on GitHub. Based on the Dockerfile data, we measure the frequency, magnitude, and instructions of Dockerfile changes and report how Dockerfile co-changed with other files. To explore the relationship between Dockerfile changes and project outcomes, i.e., popularity, success, and productivity, we also develop regression models, by controlling for various confounds. Our findings help to characterize and understand Dockerfile changes and motivate the need for collecting more empirical evidence. Yiwen Wu 0001, Yang Zhang 0026, Tao Wang 0006, Huaimin Wang 0001 |
APSEC | 1 |
| 2020 | Using Configuration Semantic Features and Machine Learning Algorithms to Predict Build Result in Cloud-Based Container EnvironmentabstractContainer technologies are being widely used in large scale production cloud environments, of which Docker has become the de-facto industry standard. In practice, Docker builds often break, and a large amount of efforts are put into troubleshooting broken builds. Prior studies have evaluated the rate at which builds in large organizations fail. However, there is still a lack of early warning methods for predicting the Docker build result before the build starts. This paper provides a first attempt to propose an automatic method named PDBR. It aims to use the configuration semantic features extracted by AST and the machine learning algorithms to predict build result in the cloud-based container environment. The evaluation experiments based on more than 36,000 collected Docker builds show that PDBR achieves 73.45%-91.92% in F1 and 29.72%-72.16% in AUC. We also demonstrate that different ML classifiers have significant and large effects on the PDBR AUC performance. Yiwen Wu 0001, Yang Zhang 0026, Junsheng Chang, Bo Ding 0001, Tao Wang 0006, Huaimin Wang 0001 |
ICPADS | 1 |
| 2020 | Exploring the Dependency Network of Docker Containers: Structure, Diversity, and RelationshipabstractContainer technologies are being widely used in large scale production cloud environments, of which Docker has become the de-facto industry standard. As a key step, containers need to define their dependent base image, which makes complex dependencies exist in a large number of containers. Prior studies have shown that references between software packages could form technical dependencies, thus forming a dependency network. However, little is known about the details of docker container dependency networks. In this paper, we perform an empirical study on the dependency network of docker containers from more than 120,000 dockerfiles. We construct the container dependency network and analyze its network structure. Further, we focus on the Top-100 dominant containers and investigate their subnetworks, including diversity and relationships. Our findings help to characterize and understand the container dependencies in the docker community and motivate the need for developing container dependency management tools. Yinyuan Zhang, Yang Zhang 0026, Yiwen Wu 0001, Yao Lu 0003, Tao Wang 0006, Xinjun Mao |
Internetware | 3 |
| 2020 | An Empirical Study of Build Failures in the Docker ContextabstractDocker containers have become the de-facto industry standard. Docker builds often break, and a large amount of efforts are put into troubleshooting broken builds. Prior studies have evaluated the rate at which builds in large organizations fail. However, little is known about the frequency and fix effort of failures that occur in Docker builds of open-source projects. This paper provides a first attempt to present a preliminary study on 857,086 Docker builds from 3,828 open-source projects hosted on GitHub. Using the Docker build data, we measure the frequency of broken builds and report their fix time. Furthermore, we explore the evolution of Docker build failures across time. Our findings help to characterize and understand Docker build failures and motivate the need for collecting more empirical evidence. Yiwen Wu 0001, Yang Zhang 0026, Tao Wang 0006, Huaimin Wang 0001 |
MSR | 1 |
| 2020 | GitHub's milestone tool: A mixed-methods analysis on its useabstractAbstract Social coding site GitHub provides developers with many management tools to facilitate project maintenance and developer collaboration.Milestonetool, in particular, plays an important role in organizing and tracking progress on groups of issues or pull requests in a project. However, few research has analyzed the milestone tool, even though it has been used in practice for a long time. In this paper, we want to address this literature gap and present an ongoing work aimed at investigating the use of the milestone tool in GitHub open‐source projects. We conduct a mixed‐methods analysis in a large‐scale dataset of GitHub projects, to help developers gain some insights into the milestone tool, including its usage, benefits, and limitations. We quantitatively investigate the basic adoption of milestone tool and its correlation with project properties. We also survey developers to understand the reasons for using milestone tool or not and their perceptions of the milestone tool. We find that certain types of projects use milestone tool more than others. Adopting the milestone tool is associated with more commits, more releases, and more project popularity, but the current milestone tool also has some limitations. These observations can then be forwarded to the GitHub community for follow‐up and can result in them potentially making a better milestone tool. Yang Zhang 0026, Huaimin Wang 0001, Yiwen Wu 0001, Dongyang Hu, Tao Wang 0006 |
J. Softw. Evol. Process. | 3 |
| 2020 | iLinker: a novel approach for issue knowledge acquisition in GitHub projects
Yang Zhang 0026, Yiwen Wu 0001, Tao Wang 0006, Huaimin Wang 0001 |
World Wide Web | 2 |
| 2019 | Exploring the Relationship Between Developer Activities and Profile Images on GitHubabstractIn the GitHub platform, social media profile images are one of many visual components of developers. Besides, developer activities such as reporting issues or following other developers are regarded as important development and self-expression behaviors. However, to the best of our knowledge, no study has yet been conducted to study the relationship between GitHub developer activities and profile images. In this paper, we aim to investigate the relationship between developer activities and profile images to gain some insights into the developers' internal properties. During our experiments, we manually classify profile images into seven categories. Next, we investigate the relationship between developer's demographic information and developer activity. Further, using logistic regression analysis, when controlled for various variables, we statistically identify and quantify the relationships between developer activities and profile image categories. We find that several profile image categories significantly correlate with developer's demographic information and activities. We also provide a rich resource of research ideas for further study. Our examination and analysis provide insights into the developers' internal properties when using different profile images. Moreover, this study is the first step in understanding the relationship between developer activities and profile images on GitHub. Yiwen Wu 0001, Yang Zhang 0026, Tao Wang 0006, Huaimin Wang 0001 |
Internetware | 1 |
| 2019 | A novel approach for recommending semantically linkable issues in GitHub projects
Yang Zhang 0026, Yiwen Wu 0001, Tao Wang 0006, Huaimin Wang 0001 |
Sci. China Inf. Sci. | 2 |