VLDB 2026 Research / reviewers in the wild / expert
Lanxin Yang
dblp:264/3015
· DBLP profile ↗
24ranked-venue papers
7as first author
23since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 23 · 7 first-author · 22 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | One Size Does Not Fit All: Investigating Efficacy of Perplexity in Detecting LLM-Generated CodeabstractLarge Language Model-Generated Code (LLMgCode) has become increasingly common in software development. So far LLMgCode has more quality issues than Human-Authored Code (HaCode). It is common for LLMgCode to mix with HaCode in a code change, while the change is signed by only human developers, without being carefully examined. Many automated methods have been proposed to detect LLMgCode from HaCode, in which the perplexity-based method ( Perplexity for short) is the state-of-the-art method. However, the efficacy evaluation of Perplexity has focused on detection accuracy. Yet it is unclear whether Perplexity is good enough in a wider range of realistic evaluation settings. To this end, we carry out a family of experiments to compare Perplexity against feature- and pre-training-based methods from three perspectives: detection accuracy , detection speed , and generalization capability . The experimental results show that Perplexity has the best generalization capability while having limited detection accuracy and detection speed. Based on that, we discuss the strengths and limitations of Perplexity , e.g., Perplexity is unsuitable for high-level programming languages. Finally, we provide recommendations to improve Perplexity and apply it in practice. As the first large-scale investigation on detecting LLMgCode from HaCode, this article provides a wide range of findings for future improvement. Jinwei Xu, He Zhang 0001, Yanjing Yang, Lanxin Yang, Zeru Cheng, Bohan Liu 0003, Xin Zhou 0016, Alberto Bacchelli, Yin Kia Chiam, Thiam Kian Chiew |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2026 | UntCC: Untangling Composite Commits Using Structural and Semantic InformationabstractSmall and focused commits are highly valued in modern software development. However, developers sometimes submit a commit with more than one concern, represented by several lines of code changes for a specific purpose, e.g., adding new features or fixing bugs. Such composite commits confuse developers during code reviews as well as other software activities, resulting in various issues. Existing studies predominantly leverage code structure to untangle composite commits, but without considering code semantics that have been demonstrated to be important in many related studies. In this article, we propose UNTCC, a new approach that uses structural and semantic information forUNTanglingCompositeCommits. To achieve structural information, we propose the code change graph, a fine-grained, text-attributed graph representation of a commit, incorporating before-change and after-change code dependencies; and UNTCC employs the graph autoencoder to learn its structural representation. To achieve semantic information, UNTCC leverages a large language model (Llama-3.2-3B) to learn joint embeddings of the raw commit and its aligned graph representation, which guide the division of different concerns within a composite commit. The experimental evaluation using 27,853 composite commits from 9 C# and 10 Java projects shows that in terms of Accuracya/Accuracyc, UNTCC achieves 94%/74% in C# and 77%/54% in Java, outperforming state-of-the-art approaches by 2%—623%/32%—573% in C# and 22%—285%/35%—286% in Java. The results indicate that UNTCC can effectively untangle composite commits. Yuzhe Jin, Lanxin Yang, He Zhang 0001, Gongyuan Li, Bohan Liu 0003, Xin Zhou 0016, Hongyu Kuang, Liming Dong 0001 |
IEEE Trans. Software Eng. | 2 |
| 2026 | Automated Localization of Affected Libraries and Versions from Vulnerability Reports
Jinwei Xu, He Zhang 0001, Xin Zhou 0016, Yanjing Yang, Jinghao Hu 0001, Lanxin Yang, Bohan Liu 0003 |
IEEE Trans. Software Eng. | 7 |
| 2025 | Automatic Fixing of Missing Dependency ErrorsabstractMany build systems, such as Make, rely on build scripts that are written by users to specify dependencies. As a serious dependency error in Makefiles, Missing Dependencies (MDs) can result in compiling and linking outdated artifacts in incremental builds, preventing software project updates from being applied correctly. Many studies have explored the detection of MDs. Automatically fixing those missing build dependency errors has become an apparent but challenging task. The challenges mainly result from Makefiles having complex semantics and project maintainers declaring dependencies in a variety of ways. To address these challenges, we propose a new approach to fixing MDs called MDfixer. The core idea of MDfixer is to identify the dependency declaration style in a Makefile and generate patches for the same declaration style based on declaration graphs and automatic prompt generation. Specifically, MDfixer locates dependency declarations for targets that have errors in the Makefile based on error reports, and then builds a declaration graph for each build target with errors and identifies the target’s declaration style based on a distance metric between the target and the dependencies. Based on the declaration graph and automatic prompt generation, MDfixer generates patches with the same style for the dependencies that need to be added. We evaluated the effectiveness and efficiency of MDfixer with 35 well-known projects. The evaluation results show that MDfixer can fix all MDs. We submitted fixes for 2,786 individual dependency issues across 17 projects, with 11 of them merging our pull requests, resulting in a total of 2,099 errors being fixed. MDfixer consumes an average time of 3.31 min for fixing a project, with a median of 62.999s. It can assist practitioners in the effective and efficient fixing of MDs. He Zhang 0001, Lanxin Yang, Yue Li 0047, Chenxing Zhong, Manuel Rigger |
ASE | 3 |
| 2025 | Securing Self-Managed Third-Party LibrariesabstractModern software development reuses third-party libraries to cut costs but may introduce vulnerabilities. A critical practice is to verify the security of third-party libraries against public vulnerability reports. Many automated methods have been proposed to identify vulnerable libraries from vulnerability reports. Existing methods are designed for the generic identification of vulnerable libraries, considering the security of all software libraries. Generic identification is inherently challenging, resulting in limited accuracy. However, organizations only consider the security of libraries they trust and use, by self-managing a library whitelist. Therefore, we propose LibGuard, a framework to adapt existing methods to help organizations secure the libraries they use. LibGuard supplies a library whitelist for existing methods and filters the results according to a threshold, facilitating the discovery of risks overlooked by organizations while controlling false alarms. LibGuard is implemented in two ways. The first attaches the whitelist after existing methods. The second integrates the whitelist into existing methods. We evaluated LibGuard using 5,107 vulnerability reports and the library whitelist built from 79 Google projects and 29 Huawei projects. The results show that the two implementations of LibGuard increase the average F1 score by 10.25% and 11.77%, respectively. Moreover, LibGuard performs stably during the extension of whitelists. To our knowledge, this paper is the first study dedicated to securing self-managed third-party libraries, offering insights into adapting generic software security management to self-managed contexts. Xin Zhou 0016, Jinwei Xu, He Zhang 0001, Yanjing Yang, Lanxin Yang, Bohan Liu 0003, Hongshan Tang |
ASE | 5 |
| 2025 | Automated detection of affected libraries from vulnerability reports
Jinwei Xu, He Zhang 0001, Xin Zhou 0016, Yanjing Yang, Runfeng Mao, Lanxin Yang, Haifeng Shen |
Autom. Softw. Eng. | 7 |
| 2025 | Prioritizing code review requests to improve review efficiency: a simulation study
Lanxin Yang, Bohan Liu 0003, Junyu Jia, Jinwei Xu, Junming Xue, He Zhang 0001, Alberto Bacchelli |
Empir. Softw. Eng. | 1 |
| 2025 | Correction to: A preliminary investigation on using multi-task learning to predict change performance in code reviews
Lanxin Yang, He Zhang 0001, Jinwei Xu, Xin Zhou 0016, Dong Shao, Shan Gao 0009, Alberto Bacchelli |
Empir. Softw. Eng. | 1 |
| 2025 | DLAP: A Deep Learning Augmented Large Language Model Prompting framework for software vulnerability detection
Yanjing Yang, Xin Zhou 0016, Runfeng Mao, Jinwei Xu, Lanxin Yang, Haifeng Shen, He Zhang 0001 |
J. Syst. Softw. | 5 |
| 2025 | Measuring software engineer's contribution in practice: An industrial experience reportabstractAbstract Software engineers play a centric role throughout the software development lifecycle. Their activities directly impact the quality, performance, and successful delivery of software products, in particular for enterprises with an emphasis on high levels of quality assurance and timely delivery. Proper incentives that motivate software engineers are vital to secure and continuously improve development productivity and software quality. However, most existing research ignores the positive incentives for software engineers, especially industry‐oriented research. In addition, existing research largely relies on peer assessment and lacks objectivity and transparency. To this end, this study investigates the process of contribution measurement for software engineers in a global Information and Communications Technology (ICT) enterprise, to explore the practical experiences and significance of contribution measurement. We investigated the practices of contribution measurement through multiple methods, including archival analysis, interviews, and survey. A total of 22 software engineers were interviewed to understand the practical implementation process of measuring contributions and its impact on software processes as well as engineers. In addition, 74 responses to our questionnaire were collected and used for a comprehensive impact analysis on software engineers. The analysis results reveal five benefits for software development processes and four benefits for practitioners of contribution measurement in the studied enterprise. In addition, this study reports on the best practices of contribution measurement, such as team‐specific measurements, and provides a practical reference for researchers and organizations interested in studying or performing contribution measurement. Yue Li 0047, He Zhang 0001, Lanxin Yang, Liming Dong 0001, Juzheng Zhang, Bohan Liu 0003 |
J. Softw. Evol. Process. | 3 |
| 2025 | Refactoring Microservices to Microservices in Support of Evolutionary DesignabstractEvolutionary designis a widely accepted practice for defining microservice boundaries. It is performed through a sequence of incremental refactoring tasks (we call it“microservice refactoring”), each restructuring only part of a microservice system (a.k.a., refactoring part) into well-defined services for improving the architecture in a controlled manner. Despite its popularity in practice, microservice refactoring suffers from insufficient methodological support. While there are numerous studies addressing similar software design tasks,i.e., software remodularization and microservitization, their approaches prove inadequate when applied to microservice refactoring. Our analysis reveals that their approaches may even degrade the entire architecture in microservice refactoring, as they only optimize the refactoring part in such applications, but neglect the relationships between the refactoring part and the remaining system. As the first response to the need,Micro2Microis proposed to re-partition the refactoring part while optimizing three quality objectives including the interdependence between the refactoring and non-refactoring parts. In addition, it allows architects to intervene in the decision-making process by interactively incorporating their knowledge into the iterative search for optimal refactoring solutions. An empirical study on 13 open-source projects of different sizes shows that the solutions fromMicro2Microperform well and exhibit quality improvement with an average up to 45% to the original architecture. Users ofMicro2Microfound the suggested solutions highly satisfactory. They acknowledge the advantages in terms of infusing human intelligence into decisions, providing immediate quality feedback, and quick exploration capability. Chenxing Zhong, Shanshan Li 0002, He Zhang 0001, Lanxin Yang, Yuanfang Cai |
IEEE Trans. Software Eng. | 5 |
| 2024 | Fine-SE: Integrating Semantic Features and Expert Features for Software Effort EstimationabstractReliable effort estimation is of paramount importance to software planning and management, especially in industry that requires effective and on-time delivery. Although various estimation approaches have been proposed (e.g., planning poker and analogy), they may be manual and/or subjective, which are difficult to apply to other projects. In recent years, deep learning approaches for effort estimation that rely on learning expert features or semantic features respectively have been extensively studied and have been found to be promising. Semantic features and expert features describe software tasks from different perspectives, however, in the literature, the best combination of these two features has not been explored to enhance effort estimation. Additionally, there are a few studies that discuss which expert features are useful for estimating effort in the industry. To this end, we investigate the potential 13 expert features that can be used to estimate effort by interviewing 26 enterprise employees. Based on that, we propose a novel model, called Fine-SE, that leverages semantic features and expert features for effort estimation. To validate our model, a series of evaluations are conducted on more than 30,000 software tasks from 17 industrial projects of a global ICT enterprise and four open-source software (OSS) projects. The evaluation results indicate that Fine-SE provides higher performance than the baselines on evaluation measures (i.e., mean absolute error, mean magnitude of relative error, and performance indicator), particularly in industrial projects with large amounts of software tasks, which implies a significant improvement in effort estimation. In comparison with expert estimation, Fine-SE improves the performance of evaluation measures by 32.0%-45.2% in within-project estimation. In comparison with the state-of-the-art models, Deep-SE and GPT2SP, it also achieves an improvement of 8.9%-91.4% in industrial projects. The experimental results reveal the value of integrating expert features with semantic features in effort estimation. Yue Li 0047, Lanxin Yang, Liming Dong 0001, Chenxing Zhong, He Zhang 0001 |
ICSE | 4 |
| 2024 | An Experience Report on Modeling Software Process in Industrial Context: Challenges and SolutionsabstractSoftware Process Model (SPM) is an abstraction of the software development process over time to assist in managing the process. SPM has attracted significant attention from researchers and practitioners in the past decades. Due to the complexity of SPM, building a practical process model often requires collaboration between academia and industry. Unfortunately, there are few empirical studies on SPM conducted in collaboration with enterprises. In this paper, we report on the challenges and solutions encountered while modeling software processes based on our collaboration with a global enterprise. These experiences are valuable to both researchers and practitioners. We presented the modeling process in detail and collected all the interview records during collaboration. As a result of building an SPM in the enterprise, we identify seven challenges and discussed solutions for each of them. The fundamental issue with SPM remains the quality and availability of data, even within industry settings. To enhance the value and applicability of models, we propose a checklist for building simulation models. The checklist can be used by modelers and practitioners to verify details that are easily overlooked during the modeling process. Our experience report provides a practical reference with researchers and practitioners who are interested in modeling software process. Yue Li 0047, He Zhang 0001, Liming Dong 0001, Bohan Liu 0003, Lanxin Yang |
ICSSP | 5 |
| 2024 | An Explainable Automated Model for Measuring Software Engineer ContributionabstractSoftware engineers play an important role throughout the software development life-cycle, particularly in industry emphasizing quality assurance and timely delivery. Contribution measurement provides proper incentives to software engineers that motivate them to continuously improve the quality and efficiency of their work. However, existing research tends to ignore contribution measurement for software engineers in practice, relying heavily on peer review and lacking objectivity and transparency. Specifically, these studies still have two weaknesses. First, a few studies explore which metrics can be useful for contribution measurement in practice. Second, managers measure the contribution of software engineers based on their experience and lack of explainable automated tools to assist them. Yue Li 0047, He Zhang 0001, Yuzhe Jin, Liming Dong 0001, Lanxin Yang, David Lo 0001, Dong Shao |
ASE | 7 |
| 2024 | GPP: A Graph-Powered Prioritizer for Code Review RequestsabstractPeer code review has become a must-have in modern software development. However, many code review requests (CRRs) could be a backlog for large-scale and active projects, blocking continuous integration and continuous delivery (CI/CD). Prioritizing CRRs to make the relevant ones to be reviewed first is a critical method for addressing this issue. Early studies have shown that many factors affect the review priority of a CRR, including its properties and relationships with other CRRs. However, the relationships, e.g., modifying the same files and sharing the same authors, are rarely considered when developing CRR prioritizers. In this paper, we propose a Graph-Powered Prioritizer (namely GPP) to make full use of the properties and relationships of CRRs. GPP uses the multi-graph structure to develop an initial representation of a collection of CRRs and uses the graph neural network algorithm to learn the prioritization-adapted representation, and eventually, outputs an ordered list of CRRs based on it. With experimental evaluation, we define relevant CRRs in the context of CI/CD as those that are likely to achieve three objectives, i.e., being merged while undergoing a few iterations in a short duration. We compare GPP against two rule-based and six learning-based prioritizers on 15 open-source software projects with more than 420K CRRs. The experimental results indicate that GPP outperforms the baselines on three basic ranking-aware evaluation metrics, including NDCG (82.94%), MRR (36.52%), and MAP (63.80%); while providing benefits in recommending the most relevant CRRs and balancing multiple objectives. Data&materials: https://figshare.com/s/133f23da558b7b254041 Lanxin Yang, Jinwei Xu, He Zhang 0001, Fanghao Wu, Yue Li 0047, Alberto Bacchelli |
ASE | 1 |
| 2024 | A preliminary investigation on using multi-task learning to predict change performance in code reviews
Lanxin Yang, He Zhang 0001, Jinwei Xu, Xin Zhou 0016, Dong Shao, Shan Gao 0009, Alberto Bacchelli |
Empir. Softw. Eng. | 1 |
| 2023 | An Experience Report on Assessing Software Engineer's Outputs in PracticeabstractThe success of a software organization relies heavily on the quality of its products and services, which in turn are influenced by the knowledge, capability, and experience of the software engineers involved in development processes. It is popular to apply quantitative assessments of software engineers for quality assurance. However, the extent to which it benefits software organizations and how it can be effectively implemented in industrial settings remains unclear. One global Information and Communications Technology (ICT) enterprise has implemented a quantitative assessment practice of software engineer’s outputs to improve its engineering capability and product and service quality. To investigate the benefits and experiences of adopting this practice in industrial settings, we conducted an empirical study using a mixed-method approach (i.e., archive analysis, interviews, and surveys). The results indicate that this practice can benefit the ICT enterprise in terms of standardizing development processes, optimizing team structures, and offering suggestions for training and management, etc. Meanwhile, this paper reports on the best practices to tackle the challenges during the adoption of the practice in the ICT enterprise, e.g., customization for teams and synergy of quantitative and qualitative assessment. In addition, we discuss the implications and recommendations of institutionalizing quantitative engineer assessment in software organizations. For organizations intending to improve software quality from the human aspect, this study provides empirical references on how to implement quantitative engineer assessment meanwhile mitigate potential risks. Juzheng Zhang, He Zhang 0001, Lanxin Yang, Liming Dong 0001, Yue Li 0047 |
ICSSP | 3 |
| 2023 | EvaCRC: Evaluating Code Review CommentsabstractIn code reviews, developers examine code changes authored by peers and provide feedback through comments. Despite the importance of these comments, no accepted approach currently exists for assessing their quality. Therefore, this study has two main objectives: (1) to devise a conceptual model for an explainable evaluation of review comment quality, and (2) to develop models for the automated evaluation of comments according to the conceptual model. To do so, we conduct mixed-method studies and propose a new approach: EvaCRC (Evaluating Code Review Comments). To achieve the first goal, we collect and synthesize quality attributes of review comments, by triangulating data from both authoritative documentation on code review standards and academic literature. We then validate these attributes using real-world instances. Finally, we establish mappings between quality attributes and grades by inquiring domain experts, thus defining our final explainable conceptual model. To achieve the second goal, EvaCRC leverages multi-label learning. To evaluate and refine EvaCRC, we conduct an industrial case study with a global ICT enterprise. The results indicate that EvaCRC can effectively evaluate review comments while offering reasons for the grades. Data and materials: https://doi.org/10.5281/zenodo.8297481 Lanxin Yang, Jinwei Xu, He Zhang 0001, Alberto Bacchelli |
ESEC/SIGSOFT FSE | 1 |
| 2023 | Evaluating Learning-to-Rank Models for Prioritizing Code Review Requests using Process SimulationabstractIn large-scale, active software projects, one of the main challenges with code review is prioritizing the many Code Review Requests (CRRs) these projects receive. Prior studies have developed many Learning-to-Rank (LtR) models in support of prioritizing CRRs and adopted rich evaluation metrics to compare their performances. However, the evaluation was performed before observing the complex interactions between CRRs and reviewers, activities and activities in real-world code reviews. Such a pre-review evaluation provides few indications about how effective LtR models contribute to code reviews. This study aims to perform a post-review evaluation on LtR models for prioritizing CRRs. To establish the evaluation environment, we employ Discrete-Event Simulation (DES) paradigm-based Software Process Simulation Modeling (SPSM) to simulate real-world code review processes, together with three customized evaluation metrics. We develop seven LtR models and use the historical review orders of CRRs as baselines for evaluation. The results indicate that employing LtR can effectively help to accelerate the completion of reviewing CRRs and the delivery of qualified code changes. Among the seven LtR models, LambdaMART and AdaRank are particularly beneficial for accelerating completion and delivery, respectively. This study empirically demonstrates the effectiveness of using DES-based SPSM for simulating code review processes, the benefits of using LtR for prioritizing CRRs, and the specific advantages of several LtR models. This study provides new ideas for software organizations that seek to evaluate LtR models and other artificial intelligence-powered software techniques.Data&materials: https://figshare.com/s/a033e99cd2a61e64c8bc. Lanxin Yang, Bohan Liu 0003, Junyu Jia, Junming Xue, Jinwei Xu, Alberto Bacchelli, He Zhang 0001 |
SANER | 1 |
| 2022 | Modeling Review History for Reviewer Recommendation: A Hypergraph ApproachabstractModern code review is a critical and indispensable practice in a pull-request development paradigm that prevails in Open Source Software (OSS) development. Finding a suitable reviewer in projects with massive participants thus becomes an increasingly challenging task. Many reviewer recommendation approaches (recommenders) have been developed to support this task which apply a similar strategy, i.e. modeling the review history first then followed by predicting/recommending a reviewer based on the model. Apparently, the better the model reflects the reality in review history, the higher recommender's performance we may expect. However, one typical scenario in a pull-request development paradigm, i.e. one Pull-Request (PR) (such as a revision or addition submitted by a contributor) may have multiple reviewers and they may impact each other through publicly posted comments, has not been modeled well in existing recommenders. We adopted the hypergraph technique to model this high-order relationship (i.e. one PR with multiple reviewers herein) and developed a new recommender, namely HGRec, which is evaluated by 12 OSS projects with more than 87K PRs, 680K comments in terms of accuracy and recommendation distribution. The results indicate that HGRec outperforms the state-of-the-art recommenders on recommendation accuracy. Besides, among the top three accurate recommenders, HGRec is more likely to recommend a diversity of reviewers, which can help to relieve the core reviewers' workload congestion issue. Moreover, since HGRec is based on hypergraph, which is a natural and interpretable representation to model review history, it is easy to accommodate more types of entities and realistic relationships in modern code review scenarios. As the first attempt, this study reveals the potentials of hypergraph on advancing the pragmatic solutions for code reviewer recommendation. Guoping Rong, Lanxin Yang, Fuli Zhang, Hongyu Kuang, He Zhang 0001 |
ICSE | 3 |
| 2022 | Memristive Recurrent Neural Network Circuit for Fast Solving Equality-Constrained Quadratic Programming With Parallel OperationabstractEquality-constrained quadratic programming (QP) has been one of the most basic and typical problems in the Internet of Things domain. In big data scenarios, how to quickly and accurately solve the problem in hardware has not been realized. Therefore, in this article, a memristive recurrent neural circuit that can parallel solve the QP problem in real time is proposed. First, a new memristive synaptic array is designed that can simultaneously implement parallel reading and writing. On the basis of this structure, a new neural network circuit based on memristor is designed that can perform large-scale recursive operations by parallel methods. This circuit can solve the equality-constrained QP problem in different situations by using such real-time programmable memristor arrays processing in memory. The PSpice simulation results show that the problem can be solved with 99.8% precision. Based on practical verification, the neural circuit experiment on PCB is presented with 97.34% precision. Moreover, the circuit has good robustness under the interference of weight value. And, it has an advantage in processing time compared with FPGA. Qinghui Hong, Lanxin Yang, Sichun Du, Ya Li 0008 |
IEEE Internet Things J. | 2 |
| 2021 | Survey on Pains and Best Practices of Code ReviewabstractDespite widespread agreement on the benefits of code review, its outcomes may not be as expected. The complications can undermine the purpose of the development process and even destroy the entire development cycle. Both academia and the industrial communities have invested a great deal of time and effort into code reviews. When a project team adheres to the best practices and creates a conducive environment, it is likely that code reviews could be conducted effectively and efficiently. By reviewing peer-reviewed scientific publications and gray literature on code review best practices, we summarized 57 practices as well as 19 code review pains that they address. Our review has shown that following best practices can ease the process of code review considerably. Multiple actionable practices are needed to support code review pains at the same time. To enable the adoption of best practices, OSS and industrial communities alike invest in integrating automatic techniques with code review tools. We hope that this review will provide researchers and practitioners with a comprehensive understanding of code review practices, aiding them in conducting code reviews more successfully. Liming Dong 0001, He Zhang 0001, Lanxin Yang, Zhiluo Weng, Xin Zhou 0016, Zifan Pan |
APSEC | 3 |
| 2021 | Quality Assessment in Systematic Literature Reviews: A Software Engineering Perspective
Lanxin Yang, He Zhang 0001, Haifeng Shen, Xin Huang 0019, Xin Zhou 0016, Guoping Rong, Dong Shao |
Inf. Softw. Technol. | 1 |
| 2020 | An Experimental Evaluation of Imbalanced Learning and Time-Series Validation in the Context of CI/CD PredictionabstractBackground: Machine Learning (ML) has been widely used as a powerful tool to support Software Engineering (SE). The fundamental assumptions of data characteristics required for specific ML methods have to be carefully considered prior to their applications in SE. Within the context of Continuous Integration (CI) and Continuous Deployment (CD) practices, there are two vital characteristics of data prone to be violated in SE research. First, the logs generated during CI/CD for training are imbalanced data, which is contrary to the principles of common balanced classifiers; second, these logs are also time-series data, which violates the assumption of cross-validation. Objective: We aim to systematically study the two data characteristics and further provide a comprehensive evaluation for predictive CI/CD with the data from real projects. Method: We conduct an experimental study that evaluates 67 CI/CD predictive models using both cross-validation and time-series-validation. Results: Our evaluation shows that cross-validation makes the evaluation of the models optimistic in most cases, there are a few counter-examples as well. The performance of the top 10 imbalanced models are better than the balanced models in the predictions of failed builds, even for balanced data. The degree of data imbalance has a negative impact on prediction performance. Conclusion: In research and practice, the assumptions of the various ML methods should be seriously considered for the validity of research. Even if it is used to compare the relative performance of models, cross-validation may not be applicable to the problems with time-series features. The research community need to revisit the evaluation results reported in some existing research. Bohan Liu 0003, He Zhang 0001, Lanxin Yang, Liming Dong 0001, Haifeng Shen, Kaiwen Song |
EASE | 3 |