VLDB 2026 Research / reviewers in the wild / expert
Liming Dong 0001
dblp:204/5556-1
· DBLP profile ↗
16ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0002-4020-3473ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 16 · 5 first-author · 11 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UntCC: Untangling Composite Commits Using Structural and Semantic InformationabstractSmall and focused commits are highly valued in modern software development. However, developers sometimes submit a commit with more than one concern, represented by several lines of code changes for a specific purpose, e.g., adding new features or fixing bugs. Such composite commits confuse developers during code reviews as well as other software activities, resulting in various issues. Existing studies predominantly leverage code structure to untangle composite commits, but without considering code semantics that have been demonstrated to be important in many related studies. In this article, we propose UNTCC, a new approach that uses structural and semantic information forUNTanglingCompositeCommits. To achieve structural information, we propose the code change graph, a fine-grained, text-attributed graph representation of a commit, incorporating before-change and after-change code dependencies; and UNTCC employs the graph autoencoder to learn its structural representation. To achieve semantic information, UNTCC leverages a large language model (Llama-3.2-3B) to learn joint embeddings of the raw commit and its aligned graph representation, which guide the division of different concerns within a composite commit. The experimental evaluation using 27,853 composite commits from 9 C# and 10 Java projects shows that in terms of Accuracya/Accuracyc, UNTCC achieves 94%/74% in C# and 77%/54% in Java, outperforming state-of-the-art approaches by 2%—623%/32%—573% in C# and 22%—285%/35%—286% in Java. The results indicate that UNTCC can effectively untangle composite commits. Yuzhe Jin, Lanxin Yang, He Zhang 0001, Gongyuan Li, Bohan Liu 0003, Xin Zhou 0016, Hongyu Kuang, Liming Dong 0001 |
IEEE Trans. Software Eng. | 9 |
| 2025 | Measuring software engineer's contribution in practice: An industrial experience reportabstractAbstract Software engineers play a centric role throughout the software development lifecycle. Their activities directly impact the quality, performance, and successful delivery of software products, in particular for enterprises with an emphasis on high levels of quality assurance and timely delivery. Proper incentives that motivate software engineers are vital to secure and continuously improve development productivity and software quality. However, most existing research ignores the positive incentives for software engineers, especially industry‐oriented research. In addition, existing research largely relies on peer assessment and lacks objectivity and transparency. To this end, this study investigates the process of contribution measurement for software engineers in a global Information and Communications Technology (ICT) enterprise, to explore the practical experiences and significance of contribution measurement. We investigated the practices of contribution measurement through multiple methods, including archival analysis, interviews, and survey. A total of 22 software engineers were interviewed to understand the practical implementation process of measuring contributions and its impact on software processes as well as engineers. In addition, 74 responses to our questionnaire were collected and used for a comprehensive impact analysis on software engineers. The analysis results reveal five benefits for software development processes and four benefits for practitioners of contribution measurement in the studied enterprise. In addition, this study reports on the best practices of contribution measurement, such as team‐specific measurements, and provides a practical reference for researchers and organizations interested in studying or performing contribution measurement. Yue Li 0047, He Zhang 0001, Lanxin Yang, Liming Dong 0001, Juzheng Zhang, Bohan Liu 0003 |
J. Softw. Evol. Process. | 4 |
| 2024 | Fine-SE: Integrating Semantic Features and Expert Features for Software Effort EstimationabstractReliable effort estimation is of paramount importance to software planning and management, especially in industry that requires effective and on-time delivery. Although various estimation approaches have been proposed (e.g., planning poker and analogy), they may be manual and/or subjective, which are difficult to apply to other projects. In recent years, deep learning approaches for effort estimation that rely on learning expert features or semantic features respectively have been extensively studied and have been found to be promising. Semantic features and expert features describe software tasks from different perspectives, however, in the literature, the best combination of these two features has not been explored to enhance effort estimation. Additionally, there are a few studies that discuss which expert features are useful for estimating effort in the industry. To this end, we investigate the potential 13 expert features that can be used to estimate effort by interviewing 26 enterprise employees. Based on that, we propose a novel model, called Fine-SE, that leverages semantic features and expert features for effort estimation. To validate our model, a series of evaluations are conducted on more than 30,000 software tasks from 17 industrial projects of a global ICT enterprise and four open-source software (OSS) projects. The evaluation results indicate that Fine-SE provides higher performance than the baselines on evaluation measures (i.e., mean absolute error, mean magnitude of relative error, and performance indicator), particularly in industrial projects with large amounts of software tasks, which implies a significant improvement in effort estimation. In comparison with expert estimation, Fine-SE improves the performance of evaluation measures by 32.0%-45.2% in within-project estimation. In comparison with the state-of-the-art models, Deep-SE and GPT2SP, it also achieves an improvement of 8.9%-91.4% in industrial projects. The experimental results reveal the value of integrating expert features with semantic features in effort estimation. Yue Li 0047, Lanxin Yang, Liming Dong 0001, Chenxing Zhong, He Zhang 0001 |
ICSE | 5 |
| 2024 | An Experience Report on Modeling Software Process in Industrial Context: Challenges and SolutionsabstractSoftware Process Model (SPM) is an abstraction of the software development process over time to assist in managing the process. SPM has attracted significant attention from researchers and practitioners in the past decades. Due to the complexity of SPM, building a practical process model often requires collaboration between academia and industry. Unfortunately, there are few empirical studies on SPM conducted in collaboration with enterprises. In this paper, we report on the challenges and solutions encountered while modeling software processes based on our collaboration with a global enterprise. These experiences are valuable to both researchers and practitioners. We presented the modeling process in detail and collected all the interview records during collaboration. As a result of building an SPM in the enterprise, we identify seven challenges and discussed solutions for each of them. The fundamental issue with SPM remains the quality and availability of data, even within industry settings. To enhance the value and applicability of models, we propose a checklist for building simulation models. The checklist can be used by modelers and practitioners to verify details that are easily overlooked during the modeling process. Our experience report provides a practical reference with researchers and practitioners who are interested in modeling software process. Yue Li 0047, He Zhang 0001, Liming Dong 0001, Bohan Liu 0003, Lanxin Yang |
ICSSP | 3 |
| 2024 | An Explainable Automated Model for Measuring Software Engineer ContributionabstractSoftware engineers play an important role throughout the software development life-cycle, particularly in industry emphasizing quality assurance and timely delivery. Contribution measurement provides proper incentives to software engineers that motivate them to continuously improve the quality and efficiency of their work. However, existing research tends to ignore contribution measurement for software engineers in practice, relying heavily on peer review and lacking objectivity and transparency. Specifically, these studies still have two weaknesses. First, a few studies explore which metrics can be useful for contribution measurement in practice. Second, managers measure the contribution of software engineers based on their experience and lack of explainable automated tools to assist them. Yue Li 0047, He Zhang 0001, Yuzhe Jin, Liming Dong 0001, Lanxin Yang, David Lo 0001, Dong Shao |
ASE | 5 |
| 2024 | Verification and validation of software process simulation models: A systematic mapping studyabstractAbstract Software process simulation models (SPSMs) that are based on descriptive process models offer the executability that can demonstrate dynamic changes of software processes over time. Verification and validation (V&V) is critical in SPSMs for guaranteeing the quality and reliability of models. V&V of dynamic software process models is more complex and challenging than for static software process models. This work systematically summarizes and maps V&V studies in SPSM to provide guidelines for future research and practice. Specifically, this study aims at identifying the focus of research on V&V, the methods used for V&V, and how to implement V&V of SPSMs in software engineering research. We conducted a systematic mapping study on studies of SPSMs that report on their V&V activities. Under the guidance of a V&V meta‐model for SPSMs, we study four research questions about V&V process. We identified 107 primary studies from a pool of 313 papers on SPSMs until 2021. There are two main results of our study. The first one presents the relationship between quality aspects of SPSMs and the V&V methods to assure them. The second result reveals the relationships among the modeling process, three modeling steps, five quality aspects, and 10 V&V methods. Generally, researchers do not pay sufficient attention to V&V, as 65.8% ( ) failed to mention or elaborate on their V&V process. We systematically summarize and map the state‐of‐the‐art V&V research in software process modeling field to support modelers' practice and improve their V&V process. Yue Li 0047, He Zhang 0001, Bohan Liu 0003, Liming Dong 0001, Haojie Gong, Guoping Rong |
J. Softw. Evol. Process. | 4 |
| 2024 | Metrics for software process simulation modelingabstractAbstract Software process simulation (SPS) has become an effective tool for software process management and improvement. However, its adoption in industry is less than what the research community expected due to the burden of measurement cost and the high demand for domain knowledge. The difficulty of extracting appropriate metrics with real data from process enactment is one of the great challenges. We aim to provide evidence‐based support of the process metrics for software process (simulation) modeling. A systematic literature review was performed by extending our previous review series to draw a comprehensive understanding of the metrics for process modeling following our proposed ontology of metrics in SPS. We identify 131 process modeling studies that collectively involve 1975 raw metrics and classified them into 21 categories using the coding technique. We found product and process external metrics are not used frequently in SPS modeling while resource external metrics are widely used. We analyze the causal relationships between metrics. We find that the models exhibit significant diversity, as no pairwise relationship between metrics accounts for more than 10% SPS models. We identify 17 data issues may encounter in measurement and 10 coping strategies. The results of this study provide process modelers with an evidence‐based reference of the identification and the use of metrics in SPS modeling and further contribute to the development of the body of knowledge on software metrics in the context of process modeling. Furthermore, this study is not limited to process simulation but can be extended to software process modeling, in general. Taking simulation metrics as standards and references can further motivate and guide software developers to improve the collection, governance, and application of process data in practice. Bohan Liu 0003, He Zhang 0001, Liming Dong 0001, Shanshan Li 0002 |
J. Softw. Evol. Process. | 3 |
| 2023 | On Preparing and Assessing Data for Process Simulation Modeling: An Industrial ReportabstractThe rapid growth of software industry has led to a significant increase in the production of a variety of data during software development process, highlighting the apparent need for improved data quality management. As an effective means of software process research and practice, Software Process Simulation Modeling (SPSM) requires large amount and high quality data that precisely depicts what happens during the development process. Accordingly, process simulation models can be used as a reference framework for assessing the issues in data management and data governance from a process perspective. The objective of the work reported in this paper is to provide insights into the data issues in real-world industrial settings and the corresponding coping strategies for software process modelers in particular in order to assist them in preparing and assessing data for their simulation models when conducting effective SPSM in the real-world settings. This paper reports on an empirical investigation that applies software process simulation practices to study the data issues and the data governance strategies based on an industrial case from one global ICT enterprise. As the outcome, a refined process for data preparation is presented, along with a taxonomy of the data issues and the corresponding coping strategies. This paper also explores traceability recovery approaches to mine more accurate process state information from software artifacts and analyzes the impact of the recovered data traceability information by evaluating the improved fidelity of the process simulation model. Liming Dong 0001, He Zhang 0001, Yue Li 0047, Bohan Liu 0003, Zhiluo Weng |
ICSSP | 1 |
| 2023 | An Experience Report on Assessing Software Engineer's Outputs in PracticeabstractThe success of a software organization relies heavily on the quality of its products and services, which in turn are influenced by the knowledge, capability, and experience of the software engineers involved in development processes. It is popular to apply quantitative assessments of software engineers for quality assurance. However, the extent to which it benefits software organizations and how it can be effectively implemented in industrial settings remains unclear. One global Information and Communications Technology (ICT) enterprise has implemented a quantitative assessment practice of software engineer’s outputs to improve its engineering capability and product and service quality. To investigate the benefits and experiences of adopting this practice in industrial settings, we conducted an empirical study using a mixed-method approach (i.e., archive analysis, interviews, and surveys). The results indicate that this practice can benefit the ICT enterprise in terms of standardizing development processes, optimizing team structures, and offering suggestions for training and management, etc. Meanwhile, this paper reports on the best practices to tackle the challenges during the adoption of the practice in the ICT enterprise, e.g., customization for teams and synergy of quantitative and qualitative assessment. In addition, we discuss the implications and recommendations of institutionalizing quantitative engineer assessment in software organizations. For organizations intending to improve software quality from the human aspect, this study provides empirical references on how to implement quantitative engineer assessment meanwhile mitigate potential risks. Juzheng Zhang, He Zhang 0001, Lanxin Yang, Liming Dong 0001, Yue Li 0047 |
ICSSP | 4 |
| 2022 | Semi-supervised pre-processing for learning-based traceability framework on real-world software projectsabstractThe traceability of software artifacts has been recognized as an important factor to support various activities in software development processes. However, traceability can be difficult and time-consuming to create and maintain manually, thereby automated approaches have gained much attention. Unfortunately, existing automated approaches for traceability suffer from practical issues. This paper aims to gain an understanding of the potential challenges for the underperforming of the state-of-the-art, ML-based trace link classifiers applied in real-world projects. By investigating different industrial datasets, we found that two critical (and classic) challenges, i.e. data imbalance and sparse problems, lie in real-world projects’ traceability automation. To overcome these challenges, we developed a framework called SPLINT to incorporate hybrid textual similarity measures and semi-supervised learning strategies as enhancements to the learning-based traceability approaches. We carried out experiments with six open-source platforms and ten industry datasets. The results confirm that SPLINT is able to operate at higher performance on two communities’ datasets. Specifically, the industrial datasets, which significantly suffer from data imbalance and sparsity problems, show an increase in F2-score over 14% and AUC over 8% on average. The adjusted class-balancing and self-training policies used in SPLINT (CBST-Adjust) also work effectively for the selection of pseudo-labels on minor classes from unlabeled trace sets, demonstrating SPLINT’s practicability. Liming Dong 0001, He Zhang 0001, Zhiluo Weng, Hongyu Kuang |
ESEC/SIGSOFT FSE | 1 |
| 2021 | Survey on Pains and Best Practices of Code ReviewabstractDespite widespread agreement on the benefits of code review, its outcomes may not be as expected. The complications can undermine the purpose of the development process and even destroy the entire development cycle. Both academia and the industrial communities have invested a great deal of time and effort into code reviews. When a project team adheres to the best practices and creates a conducive environment, it is likely that code reviews could be conducted effectively and efficiently. By reviewing peer-reviewed scientific publications and gray literature on code review best practices, we summarized 57 practices as well as 19 code review pains that they address. Our review has shown that following best practices can ease the process of code review considerably. Multiple actionable practices are needed to support code review pains at the same time. To enable the adoption of best practices, OSS and industrial communities alike invest in integrating automatic techniques with code review tools. We hope that this review will provide researchers and practitioners with a comprehensive understanding of code review practices, aiding them in conducting code reviews more successfully. Liming Dong 0001, He Zhang 0001, Lanxin Yang, Zhiluo Weng, Xin Zhou 0016, Zifan Pan |
APSEC | 1 |
| 2020 | An Experimental Evaluation of Imbalanced Learning and Time-Series Validation in the Context of CI/CD PredictionabstractBackground: Machine Learning (ML) has been widely used as a powerful tool to support Software Engineering (SE). The fundamental assumptions of data characteristics required for specific ML methods have to be carefully considered prior to their applications in SE. Within the context of Continuous Integration (CI) and Continuous Deployment (CD) practices, there are two vital characteristics of data prone to be violated in SE research. First, the logs generated during CI/CD for training are imbalanced data, which is contrary to the principles of common balanced classifiers; second, these logs are also time-series data, which violates the assumption of cross-validation. Objective: We aim to systematically study the two data characteristics and further provide a comprehensive evaluation for predictive CI/CD with the data from real projects. Method: We conduct an experimental study that evaluates 67 CI/CD predictive models using both cross-validation and time-series-validation. Results: Our evaluation shows that cross-validation makes the evaluation of the models optimistic in most cases, there are a few counter-examples as well. The performance of the top 10 imbalanced models are better than the balanced models in the predictions of failed builds, even for balanced data. The degree of data imbalance has a negative impact on prediction performance. Conclusion: In research and practice, the assumptions of the various ML methods should be seriously considered for the validity of research. Even if it is used to compare the relative performance of models, cross-validation may not be applicable to the problems with time-series features. The research community need to revisit the evaluation results reported in some existing research. Bohan Liu 0003, He Zhang 0001, Lanxin Yang, Liming Dong 0001, Haifeng Shen, Kaiwen Song |
EASE | 4 |
| 2020 | Constructing a Hybrid Software Process Simulation Model in Practice: An Exemplar from IndustryabstractBackground: Software Process Simulation Modeling (SPSM) is of paramount importance to support quantitative management of software development process. Hybrid process simulation combines multiple simulation paradigms to reflect complex changes in realistic software processes, which brings inherent challenges to process management. Constructing a hybrid model requires more modeling expertise and experience than modeling by solo-paradigm. However, a few studies explicitly discuss the challenges they encountered as a topic, which may discourage practitioners. Objective: Our aim in this study is to present an industrial process modeling project as an exemplar to demonstrate and discuss the technical issues and challenges associated with hybrid process simulation in practice. Method: Based on the collaboration with a global software enterprise, we constructed a hybrid process simulation model that combines System Dynamics (SD) and Discrete Event Simulation (DES) to predict the project duration and release date for management. Results: Several challenges around hybrid process simulation of software development process are identified and discussed with the proposal of sets of solutions from different perspectives. The model is validated by comparing the simulation result with the actual enactment of the process in industry. In addition, the result confirms the rationality and efficacy of the suggested solutions to some extent. Conclusions: In the collaboration with the enterprise, five-step modeling procedure was adopted for constructing the hybrid process model. The experience reported about the detailed steps of hybrid modeling may offer reference value to the SPSM community. Yue Li 0047, He Zhang 0001, Liming Dong 0001, Bohan Liu 0003, Jinyu Ma |
ICSSP | 3 |
| 2019 | What are the factors affecting the handover process in open source development?
Bohan Liu 0003, Guoping Rong, Liming Dong 0001, He Zhang 0001, Danni Chen, Tiange Chen, Yuyan Chen |
J. Syst. Softw. | 3 |
| 2017 | A Mapping Study on Mining Software ProcessabstractBackground: Mining Software Process (MSP) helps distill important information about software process enactment from software data repositories. An increasing amount of research effort is being dedicated to MSP. These studies differ in various aspects (e.g., topics, data, and techniques) of MSP. Objective: We aim to study the state of the art on MSP from following aspects, i.e., research topics, data sources, data types, mining techniques, and mining tools. Method: We conducted a systematic mapping study on the research relevant to MSP at both microprocess and macroprocess levels. Results: Our mapping study identified 40 relevant studies that can be grouped into microprocess and macroprocess levels. The identified mining techniques have been mapped onto the associated mining tools that fall into four types. Driven by the three research questions which represented in a meta-model, the findings revealed the correlations among the research topics, data sources, data types, mining techniques, and mining tools. Conclusion: It is observed that in order to discover the software process model or map, the main data source is from industrial project. Current mining techniques for microprocess research are mostly business process mining or sequence mining techniques used to recover descriptive software process. In addition, various machine learning algorithms and novel proposed methods are used to improve the accuracy of macroprocess level factors (e.g., software effort estimation). Liming Dong 0001, Bohan Liu 0003, Zheng Li 0001, Muhammad Ali Babar 0001, Bingbing Xue 0002 |
APSEC | 1 |
| 2017 | Mining Handover Process in Open Source Development: An Exploratory StudyabstractBackground: Handover is a common process in all software development projects. It is one of the most complex and diverse processes in software life cycle which could have a negative impact on software quality and progress. In open source software (OSS) development, handover is a more critical task due to poor planning. Objective: The goal of this work is to investigate whether we can automatically identify the handover process in OSS development. Furthermore, we aim to mine the process of handover and identify the factors and their influences on the duration of handover process. Method: We propose an ADC metric and an HDI algorithm to automatically identify the handover process and conduct a brief survey to evaluate it. We apply the Heuristic mining algorithm to discover the process maps of handover by mining Github repositories. To identify the factors from a large set of variables, we employ the Stepwise regression method. Results: We identified 63 pairs of handover within 44 projects from 314 most popular projects using our proposed method. Our survey received 21 responses. Conclusion: This study confirms that handover can be identified automatically. Although handover processes vary, developers follow a common work-flow during handover. The number of lines of code is positively correlated to the duration of handover process. Liming Dong 0001, Bohan Liu 0003, Zheng Li 0001, Bingbing Xue 0002, Danni Chen, Tiange Chen |
APSEC | 1 |