EDBT 2026 Demo / reviewers in the wild / expert
Bohan Liu 0003
dblp:32/3191-3
· DBLP profile ↗
24ranked-venue papers
6as first author
15since 2021 · last 2026
0000-0002-0146-5411ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 24 · 6 first-author · 15 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | One Size Does Not Fit All: Investigating Efficacy of Perplexity in Detecting LLM-Generated CodeabstractLarge Language Model-Generated Code (LLMgCode) has become increasingly common in software development. So far LLMgCode has more quality issues than Human-Authored Code (HaCode). It is common for LLMgCode to mix with HaCode in a code change, while the change is signed by only human developers, without being carefully examined. Many automated methods have been proposed to detect LLMgCode from HaCode, in which the perplexity-based method ( Perplexity for short) is the state-of-the-art method. However, the efficacy evaluation of Perplexity has focused on detection accuracy. Yet it is unclear whether Perplexity is good enough in a wider range of realistic evaluation settings. To this end, we carry out a family of experiments to compare Perplexity against feature- and pre-training-based methods from three perspectives: detection accuracy , detection speed , and generalization capability . The experimental results show that Perplexity has the best generalization capability while having limited detection accuracy and detection speed. Based on that, we discuss the strengths and limitations of Perplexity , e.g., Perplexity is unsuitable for high-level programming languages. Finally, we provide recommendations to improve Perplexity and apply it in practice. As the first large-scale investigation on detecting LLMgCode from HaCode, this article provides a wide range of findings for future improvement. Jinwei Xu, He Zhang 0001, Yanjing Yang, Lanxin Yang, Zeru Cheng, Bohan Liu 0003, Xin Zhou 0016, Alberto Bacchelli, Yin Kia Chiam, Thiam Kian Chiew |
ACM Trans. Softw. Eng. Methodol. | 7 |
| 2026 | UntCC: Untangling Composite Commits Using Structural and Semantic InformationabstractSmall and focused commits are highly valued in modern software development. However, developers sometimes submit a commit with more than one concern, represented by several lines of code changes for a specific purpose, e.g., adding new features or fixing bugs. Such composite commits confuse developers during code reviews as well as other software activities, resulting in various issues. Existing studies predominantly leverage code structure to untangle composite commits, but without considering code semantics that have been demonstrated to be important in many related studies. In this article, we propose UNTCC, a new approach that uses structural and semantic information forUNTanglingCompositeCommits. To achieve structural information, we propose the code change graph, a fine-grained, text-attributed graph representation of a commit, incorporating before-change and after-change code dependencies; and UNTCC employs the graph autoencoder to learn its structural representation. To achieve semantic information, UNTCC leverages a large language model (Llama-3.2-3B) to learn joint embeddings of the raw commit and its aligned graph representation, which guide the division of different concerns within a composite commit. The experimental evaluation using 27,853 composite commits from 9 C# and 10 Java projects shows that in terms of Accuracya/Accuracyc, UNTCC achieves 94%/74% in C# and 77%/54% in Java, outperforming state-of-the-art approaches by 2%—623%/32%—573% in C# and 22%—285%/35%—286% in Java. The results indicate that UNTCC can effectively untangle composite commits. Yuzhe Jin, Lanxin Yang, He Zhang 0001, Gongyuan Li, Bohan Liu 0003, Xin Zhou 0016, Hongyu Kuang, Liming Dong 0001 |
IEEE Trans. Software Eng. | 6 |
| 2026 | Automated Localization of Affected Libraries and Versions from Vulnerability Reports
Jinwei Xu, He Zhang 0001, Xin Zhou 0016, Yanjing Yang, Jinghao Hu 0001, Lanxin Yang, Bohan Liu 0003 |
IEEE Trans. Software Eng. | 8 |
| 2025 | Securing Self-Managed Third-Party LibrariesabstractModern software development reuses third-party libraries to cut costs but may introduce vulnerabilities. A critical practice is to verify the security of third-party libraries against public vulnerability reports. Many automated methods have been proposed to identify vulnerable libraries from vulnerability reports. Existing methods are designed for the generic identification of vulnerable libraries, considering the security of all software libraries. Generic identification is inherently challenging, resulting in limited accuracy. However, organizations only consider the security of libraries they trust and use, by self-managing a library whitelist. Therefore, we propose LibGuard, a framework to adapt existing methods to help organizations secure the libraries they use. LibGuard supplies a library whitelist for existing methods and filters the results according to a threshold, facilitating the discovery of risks overlooked by organizations while controlling false alarms. LibGuard is implemented in two ways. The first attaches the whitelist after existing methods. The second integrates the whitelist into existing methods. We evaluated LibGuard using 5,107 vulnerability reports and the library whitelist built from 79 Google projects and 29 Huawei projects. The results show that the two implementations of LibGuard increase the average F1 score by 10.25% and 11.77%, respectively. Moreover, LibGuard performs stably during the extension of whitelists. To our knowledge, this paper is the first study dedicated to securing self-managed third-party libraries, offering insights into adapting generic software security management to self-managed contexts. Xin Zhou 0016, Jinwei Xu, He Zhang 0001, Yanjing Yang, Lanxin Yang, Bohan Liu 0003, Hongshan Tang |
ASE | 6 |
| 2025 | Prioritizing code review requests to improve review efficiency: a simulation study
Lanxin Yang, Bohan Liu 0003, Junyu Jia, Jinwei Xu, Junming Xue, He Zhang 0001, Alberto Bacchelli |
Empir. Softw. Eng. | 2 |
| 2025 | Measuring software engineer's contribution in practice: An industrial experience reportabstractAbstract Software engineers play a centric role throughout the software development lifecycle. Their activities directly impact the quality, performance, and successful delivery of software products, in particular for enterprises with an emphasis on high levels of quality assurance and timely delivery. Proper incentives that motivate software engineers are vital to secure and continuously improve development productivity and software quality. However, most existing research ignores the positive incentives for software engineers, especially industry‐oriented research. In addition, existing research largely relies on peer assessment and lacks objectivity and transparency. To this end, this study investigates the process of contribution measurement for software engineers in a global Information and Communications Technology (ICT) enterprise, to explore the practical experiences and significance of contribution measurement. We investigated the practices of contribution measurement through multiple methods, including archival analysis, interviews, and survey. A total of 22 software engineers were interviewed to understand the practical implementation process of measuring contributions and its impact on software processes as well as engineers. In addition, 74 responses to our questionnaire were collected and used for a comprehensive impact analysis on software engineers. The analysis results reveal five benefits for software development processes and four benefits for practitioners of contribution measurement in the studied enterprise. In addition, this study reports on the best practices of contribution measurement, such as team‐specific measurements, and provides a practical reference for researchers and organizations interested in studying or performing contribution measurement. Yue Li 0047, He Zhang 0001, Lanxin Yang, Liming Dong 0001, Juzheng Zhang, Bohan Liu 0003 |
J. Softw. Evol. Process. | 6 |
| 2025 | Detecting Build Dependency Errors by Dynamic Analysis of Build Execution Against DeclarationabstractIncompletely declared build dependencies in MAKE-based build scripts can result in incorrect or inefficient incremental builds and parallel builds for C/C++ projects. In this sense, developing MAKE-based build scripts (e.g., Makefile) is a nontrivial task, since practitioners need to manually enumerate the dependencies between the parts involved in one build, which may result in serious dependency errors such as missing dependencies or redundant dependencies. To tackle this challenge, the software engineering community has invested considerable effort in dependency error detection. However, due to issues such as incomplete or even missing static dependencies (i.e., dependencies by users declared in Makefile), existing solutions either miss certain critical dependency errors or consume significant time when parsing build dependencies, posing a major challenge to ensure both detection effectiveness and efficiency. We propose a novel approach called BuildChecker to detect the above two critical types of dependency errors in MAKE dependencies that leverages a dynamically generated build execution-declaration model to improve error detection performance and reduce detection time. We evaluate BuildChecker with state-of-the-art tools (Mkcheck, Buildfs, VeriBuild, and VirtualBuild) on 30 projects. The experimental results show that BuildChecker is able to detect a total of 13,579 dependency errors with only 29 false positives, fewer than all the state-of-the-art tools. In terms of detection efficiency, BuildChecker outperforms Buildfs by 1.38 times and Mkcheck by 66.24 times. All dependency errors had been submitted to the practitioners and maintainers of these projects. At the time of writing this article, we received responses from the maintainers of four projects, who confirmed our error reports and fixes. BuildChecker demonstrates a great potential to support practitioners effectively detect build dependency errors. Shanshan Li 0002, Bohan Liu 0003, He Zhang 0001, Guoping Rong, Chenxing Zhong |
IEEE Trans. Software Eng. | 3 |
| 2025 | Decision Support for Selecting Blockchain-Based Application Design Patterns With Layered Taxonomy and Quality AttributesabstractBackground:Along with the rapid development and widespread adoption of blockchain technology, many common practices have been summarized into blockchain-based design patterns for application development. However, the numerous and scattered patterns may cause confusion among practitioners. Therefore, adopting appropriate patterns to meet various requirements has become a major challenge, as it requires deep development experience and blockchain technology knowledge.Objective:To address this problem, this paper proposes a decision-support solution to assist with the selection of design patterns during the blockchain-based application development, including a layered taxonomy of design patterns, mappings of quality attributes with the patterns, and a decision model incorporating the taxonomy and mappings.Method:We collected 72 distinct and state-of-the-art design patterns via a Systematic Literature Review (SLR) to establish a layered taxonomy, and 18 unified quality attribute metrics were proposed for blockchain-based pattern assessment and mapping establishment. Based on the pattern taxonomy and quality attribute mappings, we developed a decision model that can provide intuitive guidance for pattern selection.Results:The proposed solution was evaluated through a case study in a seafood supply chain, in which we examined how well the decision model could help identify design flaws and provide reasonable solutions. Additionally, interviews and a questionnaire-based survey were conducted to measure the completeness, correctness, and usefulness of the proposed decision model. The evaluation results indicate that the proposed decision-support solution provides developers with comprehensive guidance, facilitates targeted decision making, and supports intuitive understanding.Conclusions:Our decision-support solution can improve the development efficiency of blockchain-based applications, especially in addressing potential design flaws, achieving targeted quality attributes, and reducing development costs. Jingyue Li, Shanshan Li 0002, He Zhang 0001, Chenxing Zhong, Bohan Liu 0003, Yue Liu 0010, Qinghua Lu 0001, Xin Zhou 0016 |
IEEE Trans. Software Eng. | 9 |
| 2024 | Mining Pull Requests to Detect Process Anomalies in Open Source Software DevelopmentabstractTrustworthy Open Source Software (OSS) development processes are the basis that secures the long-term trustworthiness of software projects and products. With the aim to investigate the trustworthiness of the Pull Request (PR) process, the common model of collaborative development in OSS community, we exploit process mining to identify and analyze the normal and anomalous patterns of PR processes, and propose our approach to identifying anomalies from both control-flow and semantic aspects, and then to analyze and synthesize the root causes of the identified anomalies. We analyze 17531 PRs of 18 OSS projects on GitHub, extracting 26 root causes of control-flow anomalies and 19 root causes of semantic anomalies. We find that most PRs can hardly contain both semantic anomalies and control-flow anomalies, and the internal custom rules in projects may be the key causes for the identified anomalous PRs. We further discover and analyze the patterns of normal PR processes. We find that PRs in the non-fork model (42%) are far more likely than the fork model (5%) to bypass the review process, indicating a higher potential risk. Besides, we analyzed nine poisoned projects whose PR practices were indeed worse. Given the complex and diverse PR processes in OSS community, the proposed approach can help identify and understand not only anomalous PRs but also normal PRs, which offers early risk indications of suspicious incidents (such as poisoning) to OSS supply chain. Bohan Liu 0003, He Zhang 0001, Weigang Ma, Hongyu Kuang, Jinwei Xu, Shan Gao 0009 |
ICSE | 1 |
| 2024 | An Experience Report on Modeling Software Process in Industrial Context: Challenges and SolutionsabstractSoftware Process Model (SPM) is an abstraction of the software development process over time to assist in managing the process. SPM has attracted significant attention from researchers and practitioners in the past decades. Due to the complexity of SPM, building a practical process model often requires collaboration between academia and industry. Unfortunately, there are few empirical studies on SPM conducted in collaboration with enterprises. In this paper, we report on the challenges and solutions encountered while modeling software processes based on our collaboration with a global enterprise. These experiences are valuable to both researchers and practitioners. We presented the modeling process in detail and collected all the interview records during collaboration. As a result of building an SPM in the enterprise, we identify seven challenges and discussed solutions for each of them. The fundamental issue with SPM remains the quality and availability of data, even within industry settings. To enhance the value and applicability of models, we propose a checklist for building simulation models. The checklist can be used by modelers and practitioners to verify details that are easily overlooked during the modeling process. Our experience report provides a practical reference with researchers and practitioners who are interested in modeling software process. Yue Li 0047, He Zhang 0001, Liming Dong 0001, Bohan Liu 0003, Lanxin Yang |
ICSSP | 4 |
| 2024 | Verification and validation of software process simulation models: A systematic mapping studyabstractAbstract Software process simulation models (SPSMs) that are based on descriptive process models offer the executability that can demonstrate dynamic changes of software processes over time. Verification and validation (V&V) is critical in SPSMs for guaranteeing the quality and reliability of models. V&V of dynamic software process models is more complex and challenging than for static software process models. This work systematically summarizes and maps V&V studies in SPSM to provide guidelines for future research and practice. Specifically, this study aims at identifying the focus of research on V&V, the methods used for V&V, and how to implement V&V of SPSMs in software engineering research. We conducted a systematic mapping study on studies of SPSMs that report on their V&V activities. Under the guidance of a V&V meta‐model for SPSMs, we study four research questions about V&V process. We identified 107 primary studies from a pool of 313 papers on SPSMs until 2021. There are two main results of our study. The first one presents the relationship between quality aspects of SPSMs and the V&V methods to assure them. The second result reveals the relationships among the modeling process, three modeling steps, five quality aspects, and 10 V&V methods. Generally, researchers do not pay sufficient attention to V&V, as 65.8% ( ) failed to mention or elaborate on their V&V process. We systematically summarize and map the state‐of‐the‐art V&V research in software process modeling field to support modelers' practice and improve their V&V process. Yue Li 0047, He Zhang 0001, Bohan Liu 0003, Liming Dong 0001, Haojie Gong, Guoping Rong |
J. Softw. Evol. Process. | 3 |
| 2024 | Metrics for software process simulation modelingabstractAbstract Software process simulation (SPS) has become an effective tool for software process management and improvement. However, its adoption in industry is less than what the research community expected due to the burden of measurement cost and the high demand for domain knowledge. The difficulty of extracting appropriate metrics with real data from process enactment is one of the great challenges. We aim to provide evidence‐based support of the process metrics for software process (simulation) modeling. A systematic literature review was performed by extending our previous review series to draw a comprehensive understanding of the metrics for process modeling following our proposed ontology of metrics in SPS. We identify 131 process modeling studies that collectively involve 1975 raw metrics and classified them into 21 categories using the coding technique. We found product and process external metrics are not used frequently in SPS modeling while resource external metrics are widely used. We analyze the causal relationships between metrics. We find that the models exhibit significant diversity, as no pairwise relationship between metrics accounts for more than 10% SPS models. We identify 17 data issues may encounter in measurement and 10 coping strategies. The results of this study provide process modelers with an evidence‐based reference of the identification and the use of metrics in SPS modeling and further contribute to the development of the body of knowledge on software metrics in the context of process modeling. Furthermore, this study is not limited to process simulation but can be extended to software process modeling, in general. Taking simulation metrics as standards and references can further motivate and guide software developers to improve the collection, governance, and application of process data in practice. Bohan Liu 0003, He Zhang 0001, Liming Dong 0001, Shanshan Li 0002 |
J. Softw. Evol. Process. | 1 |
| 2023 | On Preparing and Assessing Data for Process Simulation Modeling: An Industrial ReportabstractThe rapid growth of software industry has led to a significant increase in the production of a variety of data during software development process, highlighting the apparent need for improved data quality management. As an effective means of software process research and practice, Software Process Simulation Modeling (SPSM) requires large amount and high quality data that precisely depicts what happens during the development process. Accordingly, process simulation models can be used as a reference framework for assessing the issues in data management and data governance from a process perspective. The objective of the work reported in this paper is to provide insights into the data issues in real-world industrial settings and the corresponding coping strategies for software process modelers in particular in order to assist them in preparing and assessing data for their simulation models when conducting effective SPSM in the real-world settings. This paper reports on an empirical investigation that applies software process simulation practices to study the data issues and the data governance strategies based on an industrial case from one global ICT enterprise. As the outcome, a refined process for data preparation is presented, along with a taxonomy of the data issues and the corresponding coping strategies. This paper also explores traceability recovery approaches to mine more accurate process state information from software artifacts and analyzes the impact of the recovered data traceability information by evaluating the improved fidelity of the process simulation model. Liming Dong 0001, He Zhang 0001, Yue Li 0047, Bohan Liu 0003, Zhiluo Weng |
ICSSP | 4 |
| 2023 | Evaluating Learning-to-Rank Models for Prioritizing Code Review Requests using Process SimulationabstractIn large-scale, active software projects, one of the main challenges with code review is prioritizing the many Code Review Requests (CRRs) these projects receive. Prior studies have developed many Learning-to-Rank (LtR) models in support of prioritizing CRRs and adopted rich evaluation metrics to compare their performances. However, the evaluation was performed before observing the complex interactions between CRRs and reviewers, activities and activities in real-world code reviews. Such a pre-review evaluation provides few indications about how effective LtR models contribute to code reviews. This study aims to perform a post-review evaluation on LtR models for prioritizing CRRs. To establish the evaluation environment, we employ Discrete-Event Simulation (DES) paradigm-based Software Process Simulation Modeling (SPSM) to simulate real-world code review processes, together with three customized evaluation metrics. We develop seven LtR models and use the historical review orders of CRRs as baselines for evaluation. The results indicate that employing LtR can effectively help to accelerate the completion of reviewing CRRs and the delivery of qualified code changes. Among the seven LtR models, LambdaMART and AdaRank are particularly beneficial for accelerating completion and delivery, respectively. This study empirically demonstrates the effectiveness of using DES-based SPSM for simulating code review processes, the benefits of using LtR for prioritizing CRRs, and the specific advantages of several LtR models. This study provides new ideas for software organizations that seek to evaluate LtR models and other artificial intelligence-powered software techniques.Data&materials: https://figshare.com/s/a033e99cd2a61e64c8bc. Lanxin Yang, Bohan Liu 0003, Junyu Jia, Junming Xue, Jinwei Xu, Alberto Bacchelli, He Zhang 0001 |
SANER | 2 |
| 2023 | The Why, When, What, and How About Predictive Continuous Integration: A Simulation-Based InvestigationabstractContinuous Integration (CI) enables developers to detect defects early and thus reduce lead time. However, the high frequency and long duration of executing CI have a detrimental effect on this practice. Existing studies have focused on using CI outcome predictors to reduce frequency. Since there is no reported project using predictive CI, it is difficult to evaluate its economic impact. This research aims to investigate predictive CI from a process perspective, including why and when to adopt predictors, what predictors to be used, and how to practice predictive CI in real projects. We innovatively employ Software Process Simulation to simulate a predictive CI process with a Discrete-Event Simulation (DES) model and conduct simulation-based experiments. We develop the Rollback-based Identification of Defective Commits (RIDEC) method to account for the negative effects of false predictions in simulations. Experimental results show that: 1) using predictive CI generally improves the effectiveness of CI, reducing time costs by up to 36.8% and the average waiting time before executing CI by 90.5%; 2) the time-saving varies across projects, with higher commit frequency projects benefiting more; and 3) predictor performance does not strongly correlate with time savings, but the precision of both failed and passed predictions should be paid more attention. Simulation-based evaluation helps identify overlooked aspects in existing research. Predictive CI saves time and resources, but improved prediction performance has limited cost-saving benefits. The primary value of predictive CI lies in providing accurate and quick feedback to developers, aligning with the goal of CI. Bohan Liu 0003, He Zhang 0001, Weigang Ma, Gongyuan Li, Shanshan Li 0002, Haifeng Shen |
IEEE Trans. Software Eng. | 1 |
| 2020 | An Experimental Evaluation of Imbalanced Learning and Time-Series Validation in the Context of CI/CD PredictionabstractBackground: Machine Learning (ML) has been widely used as a powerful tool to support Software Engineering (SE). The fundamental assumptions of data characteristics required for specific ML methods have to be carefully considered prior to their applications in SE. Within the context of Continuous Integration (CI) and Continuous Deployment (CD) practices, there are two vital characteristics of data prone to be violated in SE research. First, the logs generated during CI/CD for training are imbalanced data, which is contrary to the principles of common balanced classifiers; second, these logs are also time-series data, which violates the assumption of cross-validation. Objective: We aim to systematically study the two data characteristics and further provide a comprehensive evaluation for predictive CI/CD with the data from real projects. Method: We conduct an experimental study that evaluates 67 CI/CD predictive models using both cross-validation and time-series-validation. Results: Our evaluation shows that cross-validation makes the evaluation of the models optimistic in most cases, there are a few counter-examples as well. The performance of the top 10 imbalanced models are better than the balanced models in the predictions of failed builds, even for balanced data. The degree of data imbalance has a negative impact on prediction performance. Conclusion: In research and practice, the assumptions of the various ML methods should be seriously considered for the validity of research. Even if it is used to compare the relative performance of models, cross-validation may not be applicable to the problems with time-series features. The research community need to revisit the evaluation results reported in some existing research. Bohan Liu 0003, He Zhang 0001, Lanxin Yang, Liming Dong 0001, Haifeng Shen, Kaiwen Song |
EASE | 1 |
| 2020 | Constructing a Hybrid Software Process Simulation Model in Practice: An Exemplar from IndustryabstractBackground: Software Process Simulation Modeling (SPSM) is of paramount importance to support quantitative management of software development process. Hybrid process simulation combines multiple simulation paradigms to reflect complex changes in realistic software processes, which brings inherent challenges to process management. Constructing a hybrid model requires more modeling expertise and experience than modeling by solo-paradigm. However, a few studies explicitly discuss the challenges they encountered as a topic, which may discourage practitioners. Objective: Our aim in this study is to present an industrial process modeling project as an exemplar to demonstrate and discuss the technical issues and challenges associated with hybrid process simulation in practice. Method: Based on the collaboration with a global software enterprise, we constructed a hybrid process simulation model that combines System Dynamics (SD) and Discrete Event Simulation (DES) to predict the project duration and release date for management. Results: Several challenges around hybrid process simulation of software development process are identified and discussed with the proposal of sets of solutions from different perspectives. The model is validated by comparing the simulation result with the actual enactment of the process in industry. In addition, the result confirms the rationality and efficacy of the suggested solutions to some extent. Conclusions: In the collaboration with the enterprise, five-step modeling procedure was adopted for constructing the hybrid process model. The experience reported about the detailed steps of hybrid modeling may offer reference value to the SPSM community. Yue Li 0047, He Zhang 0001, Liming Dong 0001, Bohan Liu 0003, Jinyu Ma |
ICSSP | 4 |
| 2019 | What are the factors affecting the handover process in open source development?
Bohan Liu 0003, Guoping Rong, Liming Dong 0001, He Zhang 0001, Danni Chen, Tiange Chen, Yuyan Chen |
J. Syst. Softw. | 1 |
| 2018 | A replicated experiment for evaluating the effectiveness of pairing practice in PSP education
Guoping Rong, He Zhang 0001, Bohan Liu 0003, Qi Shan, Dong Shao |
J. Syst. Softw. | 3 |
| 2017 | A Mapping Study on Mining Software ProcessabstractBackground: Mining Software Process (MSP) helps distill important information about software process enactment from software data repositories. An increasing amount of research effort is being dedicated to MSP. These studies differ in various aspects (e.g., topics, data, and techniques) of MSP. Objective: We aim to study the state of the art on MSP from following aspects, i.e., research topics, data sources, data types, mining techniques, and mining tools. Method: We conducted a systematic mapping study on the research relevant to MSP at both microprocess and macroprocess levels. Results: Our mapping study identified 40 relevant studies that can be grouped into microprocess and macroprocess levels. The identified mining techniques have been mapped onto the associated mining tools that fall into four types. Driven by the three research questions which represented in a meta-model, the findings revealed the correlations among the research topics, data sources, data types, mining techniques, and mining tools. Conclusion: It is observed that in order to discover the software process model or map, the main data source is from industrial project. Current mining techniques for microprocess research are mostly business process mining or sequence mining techniques used to recover descriptive software process. In addition, various machine learning algorithms and novel proposed methods are used to improve the accuracy of macroprocess level factors (e.g., software effort estimation). Liming Dong 0001, Bohan Liu 0003, Zheng Li 0001, Muhammad Ali Babar 0001, Bingbing Xue 0002 |
APSEC | 2 |
| 2017 | Mining Handover Process in Open Source Development: An Exploratory StudyabstractBackground: Handover is a common process in all software development projects. It is one of the most complex and diverse processes in software life cycle which could have a negative impact on software quality and progress. In open source software (OSS) development, handover is a more critical task due to poor planning. Objective: The goal of this work is to investigate whether we can automatically identify the handover process in OSS development. Furthermore, we aim to mine the process of handover and identify the factors and their influences on the duration of handover process. Method: We propose an ADC metric and an HDI algorithm to automatically identify the handover process and conduct a brief survey to evaluate it. We apply the Heuristic mining algorithm to discover the process maps of handover by mining Github repositories. To identify the factors from a large set of variables, we employ the Stepwise regression method. Results: We identified 63 pairs of handover within 44 projects from 314 most popular projects using our proposed method. Our survey received 21 responses. Conclusion: This study confirms that handover can be identified automatically. Although handover processes vary, developers follow a common work-flow during handover. The number of lines of code is positively correlated to the duration of handover process. Liming Dong 0001, Bohan Liu 0003, Zheng Li 0001, Bingbing Xue 0002, Danni Chen, Tiange Chen |
APSEC | 2 |
| 2017 | Towards Confidence with Capture-recapture Estimation: An Exploratory Study of Dependence within InspectionsabstractBackground: Capture-ReCapture (CRC), as a technique for post-inspection defect estimation, has been studied in Software Engineering (SE) community since 1990s. While most studies focused on the performance evaluation of various CRC models and estimators, few have been done on the assessment of the credibility of estimation results, rendering the difficulty of decision-making for quality management when applying CRC for defect estimation. Objective: This research aims to explore and investigate a reliable and practical approach to assess the credibility of CRC based defect estimation. Method: One fundamental assumption of applying CRC method is the statistical independence of samples that can be measured by 'Coefficient of CoVariation' (CCV). We applied CCV as an indicator of the statistical dependence between the observations (i.e., the defects detected by inspectors), and assessed the estimation results of CRC with the published datasets in SE literature by examining the correlation between Relative Error (RE) and CCV. Based on the observed correlation, we further propose CĈV, which replaces the unknown N (the actual number of defects) with the estimated number (N), to assess the credibility of CRC estimates. Results: We found that most datasets are with non-zero CCVs and the R2 (Coefficient of Determination) of non-linear curve-fitting for their CCVs and REs is higher than 0.8. Conclusions: Our study shows the evidence that the statistical dependence among inspectors is ubiquitous in the existing CRC-related studies. Besides, the significant correlation between CCV (by CĈV in practice) and RE may enable the possibility of the assessment of CRC-based estimation in support of quality management. Guoping Rong, Bohan Liu 0003, He Zhang 0001, Qiuping Zhang, Dong Shao |
EASE | 2 |
| 2017 | A systematic map on verifying and validating software process simulation modelsabstractVerification and Validation (V&V) is a critical step in software process modelling to secure the model's quality and credibility. Software Process Simulation Models (SPSMs) that are based on descriptive process models offer the executability that is able to demonstrate the dynamic changes of software process over time. The V&V of process simulation models go beyond static process models and turn to be more complex and challenging to software modelers. This study aims to identify what aspects of process simulation models are verified and validated by using which V&V methods in what conditions in software engineering research. We conducted a systematic literature review (mapping study) on the studies of software process simulation that report of their V&V activities. We identified 72 relevant studies from a pool of 331 papers on SPSM until 2015. These studies can be mapped to ten V&V methods applied for five aspects of process models to be verified and validated, i.e., syntactic quality, semantic quality, pragmatic quality, performance, and value. A systematic map is presented to illustrate the relationships between the identified V&V methods and their supporting aspects of process models. This mapping will provide the community reference value when developing, verifying, and validating software process (simulation) models. Haojie Gong, He Zhang 0001, Dexian Yu, Bohan Liu 0003 |
ICSSP | 4 |
| 2016 | An Incremental V-Model Process for Automotive DevelopmentabstractV-model and its variants have become the most common process models adopted in automotive industry guiding the development of systems on a variety of refinement levels. Along with the exponentially growing complexity of modern vehicle systems, however, the late verification and validation in the conventional V-model expand in uncontrollable ways that result in higher cost of development and higher risk of failure than ever. This paper describes an inc-V development process for automotive industry that improves the conventional V-model and variants by introducing and institutionalizing early and continuous integrated verification enabled by simulation-based development. We developed a continuous simulation model of the inc-V process, and the initial version is used to investigate the characteristics of the inc-V compared to V. The preliminary finding from the simulations of an example project is that the inc-V process is able to improve the traditional V process by saving effort, shortening duration, and increasing product quality. The finding also show how the advance of development technology impacts the systems engineering processes. Bohan Liu 0003, He Zhang 0001, Saichun Zhu |
APSEC | 1 |