VLDB 2026 Research / reviewers in the wild / expert
Bei Wang 0010
dblp:08/6391-10
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2025
0009-0001-8465-8189ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 5 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Refactoring Deep Learning Code: a Study of Practices and Unsatisfied Tool NeedsabstractWith the rapid development of deep learning, the implementation of intricate algorithms and substantial data processing has become a standard element of deep learning projects. As a result, the code has become progressively complex as the software evolves, which is difficult to maintain and understand. Existing studies have investigated the impact of refactoring on software quality within non-deep learning software. However, the insights of code refactoring in the context of deep learning are still unclear. This study endeavors to fill this knowledge gap by empirically examining the current state of code refactoring in the deep learning realm and practitioners' views on refactoring tools. We first manually analyze the commit history of five popular and well-maintained deep learning projects (e.g., PyTorch). We mine$\mathbf{4, 4 0 1}$refactoring practices in$\mathbf{2, 4 4 5}$historical commits and measure how different types and elements of refactoring operations are distributed. We then survey 159 practitioners about their views of code refactoring in deep learning projects and their expectations of current refactoring tools. The survey result shows that refactoring research and the development of related tools in the field of deep learning are crucial for improving project maintainability and code quality, and that current refactoring tools do not adequately meet the needs of practitioners. Lastly, we provide our perspective on the future advancement of refactoring tools and offer suggestions for developers' development practices. Xing Hu 0008, Bei Wang 0010, WenXin Yao, Xin Xia 0001, Xinyu Wang 0001 |
ICSME | 3 |
| 2022 | Multi-Dimension Convolutional Neural Network for Bug LocalizationabstractSoftware bugs remain frequent in the life cycle of software development and maintenance. Automatic localization of buggy source code files is critical for timely bug fixing and improving the efficiency of software quality assurance. Various bug localization techniques have been proposed using different dimensions of features. Recent studies have shown that different dimensions of features may play different roles in bug localization. Unfortunately, how to effectively merge these dimensions of features for improving bug localization has rarely been investigated. This article presents a Multi-Dimension Convolutional Neural Network (MD-CNN) model for bug localization automatically based on a bug report. Our approach has dual-novelty. First, we identify and extract five statistical dimensions of features. Second, we design a Convolutional Neural Network (CNN) model that takes our five statistical dimensions of features as the input and iteratively learns the complex and non-linear relationship between the features and the bug locations. The MD-CNN bug localization model is verified using six large-scale open source projects. The experimental results show that our MD-CNN outperforms the existing representative bug localization techniques in terms of the Mean Average Precision (MAP) and the number of bugs successfully localized in the top 1, 5, and 10 matched source code files. Bei Wang 0010, Meng Yan 0001, Chao Liu 0014, Ling Liu 0001 |
IEEE Trans. Serv. Comput. | 1 |
| 2021 | Improving Code Summarization Through Automated Quality AssuranceabstractThe code summarization task aims to generate brief descriptions of source code automatically. It is beneficial for developers to understand source code. However, almost all of current code summarization approaches may generate low-quality (BLEU4<40) summaries, which will mislead developers. Previous work has shown that it is possible to conduct quality assurance for document generation (QA4DG) and improve the practicability of document generation approaches. Code summarization can also be regarded as a document generation task. This work aims to investigate whether QA4DG approaches can be leveraged to improve code summarization. Specifically, we first investigate whether existing QA4DG approaches can be plugged in code summarization approaches. We find that an automated quality assurance framework for commit message generation named QACom performs best. In-spired by the idea behind QAcom, we propose an ensemble code summarization approach called Ensum. Precisely, given a code snippet, Ensum first uses current code summarization approaches to generate candidate summaries. Then, Ensum predicts the quality of each candidate summary using a collaborative filtering-based component and a retrieval-based component and selects the best candidate summary as the output. Experimental results on two public datasets show that Ensum outperforms three state-of-the-art single approaches and one ensemble approach for code summarization in terms of BLEU-4, METEOR, and ROUGE-L. Yuxing Hu, Meng Yan 0001, Zhongxin Liu 0002, Qiuyuan Chen, Bei Wang 0010 |
ISSRE | 5 |
| 2021 | Plot2API: Recommending Graphic API from Plot via Semantic Parsing Guided Neural NetworkabstractPlot-based Graphic API recommendation (Plot2API) is an unstudied but meaningful issue, which has several important applications in the context of software engineering and data visualization, such as the plotting guidance of the beginner, graphic API correlation analysis, and code conversion for plotting. Plot2API is a very challenging task, since each plot is often associated with multiple APIs and the appearances of the graphics drawn by the same API can be extremely varied due to the different settings of the parameters. Additionally, the samples of different APIs also suffer from extremely imbalanced.Considering the lack of technologies in Plot2API, we present a novel deep multi-task learning approach named Semantic Parsing Guided Neural Network (SPGNN) which translates the Plot2API issue as a multi-label image classification and an image semantic parsing tasks for the solution. In SPGNN, the recently advanced Convolutional Neural Network (CNN) named EfficientNet is employed as the backbone network for API recommendation. Meanwhile, a semantic parsing module is complemented to exploit the semantic relevant visual information in feature learning and eliminate the appearance-relevant visual information which may confuse the visual-information-based API recommendation. Moreover, the recent data augmentation technique named random erasing is also applied for alleviating the imbalance of API categories.We collect plots with the graphic APIs used to drawn them from Stack Overflow, and release three new Plot2API datasets corresponding to the graphic APIs of R and Python programming languages for evaluating the effectiveness of Plot2API techniques. Extensive experimental results not only demonstrate the superiority of our method over the recent deep learning baselines but also show the practicability of our method in the recommendation of graphic APIs. Zeyu Wang 0001, Sheng Huang 0001, Zhongxin Liu 0002, Meng Yan 0001, Xin Xia 0001, Bei Wang 0010, Dan Yang 0001 |
SANER | 6 |
| 2021 | Quality Assurance for Automated Commit Message GenerationabstractMany automated commit message generation (CMG) approaches have been proposed for facilitating the understanding of software changes. They are shown to be promising and can generate commit messages that are semantically relevant to the reference messages for a number of commits. However, a large proportion (over 50%) of semantically irrelevant commit messages are also generated simultaneously. Such messages may mislead developers, require additional efforts of developers to confirm and filter out, and hinder the application of existing CMG approaches in practice. For tackling this problem, prior work mainly focuses on proposing new methods to improve the generation accuracy. However, another promising way for bridging the gap between CMG approaches and the practice has not been well investigated, which is: can we automatically assure the semantic relevance of the generated messages?To that end, in this work, we propose an automated Quality A ssurance framework for commit message generation (QAcom). QAcom can assure the quality of generated commit messages by automatically filtering out the semantically-irrelevant generated messages and preserving the semantically-relevant ones as many as possible. In particular, QAcom consists of a Collaborative-Filtering-based (CF) component and a Retrieval-based (RE) component. Given a commit message generated by a CMG approach, QAcom estimates whether this generated message is semantically relevant to its ground truth, which is unknown when estimating, based on both the collaborative filtering algorithm and the similarity between this commit and historical commits. We evaluate the effectiveness of QAcom by "plugging" it in three state-of-the-art CMG approaches. Experimental results on three public datasets show that QAcom can effectively filter out semantically-irrelevant generated messages and preserve semantically-relevant ones. Bei Wang 0010, Meng Yan 0001, Zhongxin Liu 0002, Xin Xia 0001, Xiaohong Zhang 0002, Dan Yang 0001 |
SANER | 1 |