Yue Li 0047

dblp:61/500-47 · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
10since 2021 · last 2025
0000-0002-1856-7182ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 11 · 6 first-author · 10 since 2021
YearPublicationVenuePosition
2025 PairSmell: A Novel Perspective Inspecting Software Modular Structure
abstract
Enhancing the modular structure of existing systems has attracted substantial research interest, focusing on two main methods: (1) software modularization and (2) identifying design issues (e.g., smells) as refactoring opportunities. However, remodularization solutions often require extensive modifications to the original modules, and the design issues identified are generally too coarse to guide refactoring strategies. Combining the above two methods, this paper introduces a novel concept, PairSmell, which exploits modularization to pinpoint design issues necessitating refactoring. We concentrate on a granular but fundamental aspect of modularity principles-modular relation (MR), i.e., whether a pair of entities are separated or collocated. The main assumption is that, if the actual MR of a pair violates its ‘apt MR’, i.e., an MR agreed on by multiple modularization tools (as raters), it can be deemed likely a flawed architectural decision that necessitates further examination. To quantify and evaluate PairSmell, we conduct an empirical study on 20 C/C++ and Java projects, using 4 established modularization tools to identify two forms of PairSmell: inapt separated pairs$InSep$and inapt collocated pairs$InCol$. Our study on 260,003 instances reveals that their architectural impacts are substantial: (1) on average, 14.60 % and 20.44 % of software entities are involved in$InSep$and$InCol$MRs respectively; (2)$InSep$pairs are associated with 190 % more co-changes than properly separated pairs, while$InCol$pairs are associated with 35% fewer co-changes than properly collocated pairs, both indicating a successful identification of modular structures detrimental to software quality; and (3) both forms of PairSmell persist across software evolution. This evidence strongly suggests that PairSmell can provide meaningful insights for inspecting modular structure, with the identified issues being both granular and fundamental, making the enhancement of modular design more efficient.
Chenxing Zhong, Daniel Feitosa, Paris Avgeriou, Yue Li 0047, He Zhang 0001
ICSE5
2025 Automatic Fixing of Missing Dependency Errors
abstract
Many build systems, such as Make, rely on build scripts that are written by users to specify dependencies. As a serious dependency error in Makefiles, Missing Dependencies (MDs) can result in compiling and linking outdated artifacts in incremental builds, preventing software project updates from being applied correctly. Many studies have explored the detection of MDs. Automatically fixing those missing build dependency errors has become an apparent but challenging task. The challenges mainly result from Makefiles having complex semantics and project maintainers declaring dependencies in a variety of ways. To address these challenges, we propose a new approach to fixing MDs called MDfixer. The core idea of MDfixer is to identify the dependency declaration style in a Makefile and generate patches for the same declaration style based on declaration graphs and automatic prompt generation. Specifically, MDfixer locates dependency declarations for targets that have errors in the Makefile based on error reports, and then builds a declaration graph for each build target with errors and identifies the target’s declaration style based on a distance metric between the target and the dependencies. Based on the declaration graph and automatic prompt generation, MDfixer generates patches with the same style for the dependencies that need to be added. We evaluated the effectiveness and efficiency of MDfixer with 35 well-known projects. The evaluation results show that MDfixer can fix all MDs. We submitted fixes for 2,786 individual dependency issues across 17 projects, with 11 of them merging our pull requests, resulting in a total of 2,099 errors being fixed. MDfixer consumes an average time of 3.31 min for fixing a project, with a median of 62.999s. It can assist practitioners in the effective and efficient fixing of MDs.
He Zhang 0001, Lanxin Yang, Yue Li 0047, Chenxing Zhong, Manuel Rigger
ASE4
2025 Measuring software engineer's contribution in practice: An industrial experience report
abstract
Abstract Software engineers play a centric role throughout the software development lifecycle. Their activities directly impact the quality, performance, and successful delivery of software products, in particular for enterprises with an emphasis on high levels of quality assurance and timely delivery. Proper incentives that motivate software engineers are vital to secure and continuously improve development productivity and software quality. However, most existing research ignores the positive incentives for software engineers, especially industry‐oriented research. In addition, existing research largely relies on peer assessment and lacks objectivity and transparency. To this end, this study investigates the process of contribution measurement for software engineers in a global Information and Communications Technology (ICT) enterprise, to explore the practical experiences and significance of contribution measurement. We investigated the practices of contribution measurement through multiple methods, including archival analysis, interviews, and survey. A total of 22 software engineers were interviewed to understand the practical implementation process of measuring contributions and its impact on software processes as well as engineers. In addition, 74 responses to our questionnaire were collected and used for a comprehensive impact analysis on software engineers. The analysis results reveal five benefits for software development processes and four benefits for practitioners of contribution measurement in the studied enterprise. In addition, this study reports on the best practices of contribution measurement, such as team‐specific measurements, and provides a practical reference for researchers and organizations interested in studying or performing contribution measurement.
Yue Li 0047, He Zhang 0001, Lanxin Yang, Liming Dong 0001, Juzheng Zhang, Bohan Liu 0003
J. Softw. Evol. Process.1
2024 Fine-SE: Integrating Semantic Features and Expert Features for Software Effort Estimation
abstract
Reliable effort estimation is of paramount importance to software planning and management, especially in industry that requires effective and on-time delivery. Although various estimation approaches have been proposed (e.g., planning poker and analogy), they may be manual and/or subjective, which are difficult to apply to other projects. In recent years, deep learning approaches for effort estimation that rely on learning expert features or semantic features respectively have been extensively studied and have been found to be promising. Semantic features and expert features describe software tasks from different perspectives, however, in the literature, the best combination of these two features has not been explored to enhance effort estimation. Additionally, there are a few studies that discuss which expert features are useful for estimating effort in the industry. To this end, we investigate the potential 13 expert features that can be used to estimate effort by interviewing 26 enterprise employees. Based on that, we propose a novel model, called Fine-SE, that leverages semantic features and expert features for effort estimation. To validate our model, a series of evaluations are conducted on more than 30,000 software tasks from 17 industrial projects of a global ICT enterprise and four open-source software (OSS) projects. The evaluation results indicate that Fine-SE provides higher performance than the baselines on evaluation measures (i.e., mean absolute error, mean magnitude of relative error, and performance indicator), particularly in industrial projects with large amounts of software tasks, which implies a significant improvement in effort estimation. In comparison with expert estimation, Fine-SE improves the performance of evaluation measures by 32.0%-45.2% in within-project estimation. In comparison with the state-of-the-art models, Deep-SE and GPT2SP, it also achieves an improvement of 8.9%-91.4% in industrial projects. The experimental results reveal the value of integrating expert features with semantic features in effort estimation.
Yue Li 0047, Lanxin Yang, Liming Dong 0001, Chenxing Zhong, He Zhang 0001
ICSE1
2024 An Experience Report on Modeling Software Process in Industrial Context: Challenges and Solutions
abstract
Software Process Model (SPM) is an abstraction of the software development process over time to assist in managing the process. SPM has attracted significant attention from researchers and practitioners in the past decades. Due to the complexity of SPM, building a practical process model often requires collaboration between academia and industry. Unfortunately, there are few empirical studies on SPM conducted in collaboration with enterprises. In this paper, we report on the challenges and solutions encountered while modeling software processes based on our collaboration with a global enterprise. These experiences are valuable to both researchers and practitioners. We presented the modeling process in detail and collected all the interview records during collaboration. As a result of building an SPM in the enterprise, we identify seven challenges and discussed solutions for each of them. The fundamental issue with SPM remains the quality and availability of data, even within industry settings. To enhance the value and applicability of models, we propose a checklist for building simulation models. The checklist can be used by modelers and practitioners to verify details that are easily overlooked during the modeling process. Our experience report provides a practical reference with researchers and practitioners who are interested in modeling software process.
Yue Li 0047, He Zhang 0001, Liming Dong 0001, Bohan Liu 0003, Lanxin Yang
ICSSP1
2024 An Explainable Automated Model for Measuring Software Engineer Contribution
abstract
Software engineers play an important role throughout the software development life-cycle, particularly in industry emphasizing quality assurance and timely delivery. Contribution measurement provides proper incentives to software engineers that motivate them to continuously improve the quality and efficiency of their work. However, existing research tends to ignore contribution measurement for software engineers in practice, relying heavily on peer review and lacking objectivity and transparency. Specifically, these studies still have two weaknesses. First, a few studies explore which metrics can be useful for contribution measurement in practice. Second, managers measure the contribution of software engineers based on their experience and lack of explainable automated tools to assist them.
Yue Li 0047, He Zhang 0001, Yuzhe Jin, Liming Dong 0001, Lanxin Yang, David Lo 0001, Dong Shao
ASE1
2024 GPP: A Graph-Powered Prioritizer for Code Review Requests
abstract
Peer code review has become a must-have in modern software development. However, many code review requests (CRRs) could be a backlog for large-scale and active projects, blocking continuous integration and continuous delivery (CI/CD). Prioritizing CRRs to make the relevant ones to be reviewed first is a critical method for addressing this issue. Early studies have shown that many factors affect the review priority of a CRR, including its properties and relationships with other CRRs. However, the relationships, e.g., modifying the same files and sharing the same authors, are rarely considered when developing CRR prioritizers. In this paper, we propose a Graph-Powered Prioritizer (namely GPP) to make full use of the properties and relationships of CRRs. GPP uses the multi-graph structure to develop an initial representation of a collection of CRRs and uses the graph neural network algorithm to learn the prioritization-adapted representation, and eventually, outputs an ordered list of CRRs based on it. With experimental evaluation, we define relevant CRRs in the context of CI/CD as those that are likely to achieve three objectives, i.e., being merged while undergoing a few iterations in a short duration. We compare GPP against two rule-based and six learning-based prioritizers on 15 open-source software projects with more than 420K CRRs. The experimental results indicate that GPP outperforms the baselines on three basic ranking-aware evaluation metrics, including NDCG (82.94%), MRR (36.52%), and MAP (63.80%); while providing benefits in recommending the most relevant CRRs and balancing multiple objectives. Data&materials: https://figshare.com/s/133f23da558b7b254041
Lanxin Yang, Jinwei Xu, He Zhang 0001, Fanghao Wu, Yue Li 0047, Alberto Bacchelli
ASE6
2024 Verification and validation of software process simulation models: A systematic mapping study
abstract
Abstract Software process simulation models (SPSMs) that are based on descriptive process models offer the executability that can demonstrate dynamic changes of software processes over time. Verification and validation (V&V) is critical in SPSMs for guaranteeing the quality and reliability of models. V&V of dynamic software process models is more complex and challenging than for static software process models. This work systematically summarizes and maps V&V studies in SPSM to provide guidelines for future research and practice. Specifically, this study aims at identifying the focus of research on V&V, the methods used for V&V, and how to implement V&V of SPSMs in software engineering research. We conducted a systematic mapping study on studies of SPSMs that report on their V&V activities. Under the guidance of a V&V meta‐model for SPSMs, we study four research questions about V&V process. We identified 107 primary studies from a pool of 313 papers on SPSMs until 2021. There are two main results of our study. The first one presents the relationship between quality aspects of SPSMs and the V&V methods to assure them. The second result reveals the relationships among the modeling process, three modeling steps, five quality aspects, and 10 V&V methods. Generally, researchers do not pay sufficient attention to V&V, as 65.8% ( ) failed to mention or elaborate on their V&V process. We systematically summarize and map the state‐of‐the‐art V&V research in software process modeling field to support modelers' practice and improve their V&V process.
Yue Li 0047, He Zhang 0001, Bohan Liu 0003, Liming Dong 0001, Haojie Gong, Guoping Rong
J. Softw. Evol. Process.1
2023 On Preparing and Assessing Data for Process Simulation Modeling: An Industrial Report
abstract
The rapid growth of software industry has led to a significant increase in the production of a variety of data during software development process, highlighting the apparent need for improved data quality management. As an effective means of software process research and practice, Software Process Simulation Modeling (SPSM) requires large amount and high quality data that precisely depicts what happens during the development process. Accordingly, process simulation models can be used as a reference framework for assessing the issues in data management and data governance from a process perspective. The objective of the work reported in this paper is to provide insights into the data issues in real-world industrial settings and the corresponding coping strategies for software process modelers in particular in order to assist them in preparing and assessing data for their simulation models when conducting effective SPSM in the real-world settings. This paper reports on an empirical investigation that applies software process simulation practices to study the data issues and the data governance strategies based on an industrial case from one global ICT enterprise. As the outcome, a refined process for data preparation is presented, along with a taxonomy of the data issues and the corresponding coping strategies. This paper also explores traceability recovery approaches to mine more accurate process state information from software artifacts and analyzes the impact of the recovered data traceability information by evaluating the improved fidelity of the process simulation model.
Liming Dong 0001, He Zhang 0001, Yue Li 0047, Bohan Liu 0003, Zhiluo Weng
ICSSP3
2023 An Experience Report on Assessing Software Engineer's Outputs in Practice
abstract
The success of a software organization relies heavily on the quality of its products and services, which in turn are influenced by the knowledge, capability, and experience of the software engineers involved in development processes. It is popular to apply quantitative assessments of software engineers for quality assurance. However, the extent to which it benefits software organizations and how it can be effectively implemented in industrial settings remains unclear. One global Information and Communications Technology (ICT) enterprise has implemented a quantitative assessment practice of software engineer’s outputs to improve its engineering capability and product and service quality. To investigate the benefits and experiences of adopting this practice in industrial settings, we conducted an empirical study using a mixed-method approach (i.e., archive analysis, interviews, and surveys). The results indicate that this practice can benefit the ICT enterprise in terms of standardizing development processes, optimizing team structures, and offering suggestions for training and management, etc. Meanwhile, this paper reports on the best practices to tackle the challenges during the adoption of the practice in the ICT enterprise, e.g., customization for teams and synergy of quantitative and qualitative assessment. In addition, we discuss the implications and recommendations of institutionalizing quantitative engineer assessment in software organizations. For organizations intending to improve software quality from the human aspect, this study provides empirical references on how to implement quantitative engineer assessment meanwhile mitigate potential risks.
Juzheng Zhang, He Zhang 0001, Lanxin Yang, Liming Dong 0001, Yue Li 0047
ICSSP5
2020 Constructing a Hybrid Software Process Simulation Model in Practice: An Exemplar from Industry
abstract
Background: Software Process Simulation Modeling (SPSM) is of paramount importance to support quantitative management of software development process. Hybrid process simulation combines multiple simulation paradigms to reflect complex changes in realistic software processes, which brings inherent challenges to process management. Constructing a hybrid model requires more modeling expertise and experience than modeling by solo-paradigm. However, a few studies explicitly discuss the challenges they encountered as a topic, which may discourage practitioners. Objective: Our aim in this study is to present an industrial process modeling project as an exemplar to demonstrate and discuss the technical issues and challenges associated with hybrid process simulation in practice. Method: Based on the collaboration with a global software enterprise, we constructed a hybrid process simulation model that combines System Dynamics (SD) and Discrete Event Simulation (DES) to predict the project duration and release date for management. Results: Several challenges around hybrid process simulation of software development process are identified and discussed with the proposal of sets of solutions from different perspectives. The model is validated by comparing the simulation result with the actual enactment of the process in industry. In addition, the result confirms the rationality and efficacy of the suggested solutions to some extent. Conclusions: In the collaboration with the enterprise, five-step modeling procedure was adopted for constructing the hybrid process model. The experience reported about the detailed steps of hybrid modeling may offer reference value to the SPSM community.
Yue Li 0047, He Zhang 0001, Liming Dong 0001, Bohan Liu 0003, Jinyu Ma
ICSSP1