VLDB 2026 Research / reviewers in the wild / expert
Hongmin Lu
dblp:89/4782
· DBLP profile ↗
26ranked-venue papers
1as first author
8since 2021 · last 2026
0000-0002-3558-9754ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 23 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Line-level bug-finding power of static analysis rules: a case study of Teamscale
Liwei Ye, Yuge Nie, Yibiao Yang, Hongmin Lu, Junyan Qian, Yuming Zhou |
Empir. Softw. Eng. | 5 |
| 2026 | Random test generators demystified: Differences and potential for compiler reliability
Yang Wang 0165, Beining Wu, Yibiao Yang, Hongmin Lu, Yuming Zhou |
Sci. Comput. Program. | 5 |
| 2024 | Towards a framework for reliable performance evaluation in defect prediction
Xutong Liu 0003, Shiran Liu, Zhaoqiang Guo, Peng Zhang 0083, Yibiao Yang, Hongmin Lu, Yanhui Li 0001, Lin Chen 0015, Yuming Zhou |
Sci. Comput. Program. | 7 |
| 2023 | Deriving Thresholds of Object-Oriented Metrics to Predict Defect-Proneness of Classes: A Large-Scale Meta-AnalysisabstractMany studies have explored the methods of deriving thresholds of object-oriented (i.e. OO) metrics. Unsupervised methods are mainly based on the distributions of metric values, while supervised methods principally rest on the relationships between metric values and defect-proneness of classes. The objective of this study is to empirically examine whether there are effective threshold values of OO metrics by analyzing existing threshold derivation methods with a large-scale meta-analysis. Based on five representative threshold derivation methods (i.e. VARL, ROC, BPP, MFM, and MGM) and 3268 releases from 65 Java projects, we first employ statistical meta-analysis and sensitivity analysis techniques to derive thresholds for 62 OO metrics on the training data. Then, we investigate the predictive performance of five candidate thresholds for each metric on the validation data to explore which of these candidate thresholds can be served as the threshold. Finally, we evaluate their predictive performance on the test data. The experimental results show that 26 of 62 metrics have the threshold effect and the derived thresholds by meta-analysis achieve promising results of GM values and significantly outperform almost all five representative (baseline) thresholds. Yuanqing Mei, Shiran Liu, Zhaoqiang Guo, Yibiao Yang, Hongmin Lu, Yutian Tang, Yuming Zhou |
Int. J. Softw. Eng. Knowl. Eng. | 6 |
| 2022 | A Target Detection and Tracking Method for Multiple Radar SystemsabstractMultiple radar systems represent an attractive option for target tracking because they can significantly enlarge the area coverage and improve both the probability of trajectory detection and the localization accuracy. The presence of multiple extended targets or weak targets is a challenge for multiple radar systems. Moreover, their performance may be severely deteriorated by regions characterized by a high clutter density. In this paper, an algorithm for detection and tracking of multiple targets, extended or weak, based on measurements provided by multiple radars in an environment with heavily cluttered regions, is proposed. The proposed method features three stages. In the first stage, past measurements are exploited to build a spatio-temporal clutter map in each radar; a weight is then assigned to each measurement to assess its significance. In the second stage, a track-before-detect algorithm, based on a weighted three-dimensional Hough transform, is applied to obtain target tracklets. In the third stage, a low-complexity tracklet association method, exploiting a lion reproduction model, is applied to associate tracklets of the same target. Three experiments are presented to illustrate the effectiveness of the proposed approach. The first experiment is based on synthetic data, the second one on actual data from a radar network with two homogeneous air surveillance radars, and the third one on actual data from a radar network with four different marine surveillance radars. The results reveal that the proposed method can outperform competing approaches. Bo Yan 0006, Enrico Paolini, Hongmin Lu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | CBUA: A Probabilistic, Predictive, and Practical Approach for Evaluating Test Suite EffectivenessabstractKnowing the effectiveness of a test suite is essential for many activities such as assessing the test adequacy of code and guiding the generation of new test cases. Mutation testing is a commonly used defect injection technique for evaluating the effectiveness of a test suite. However, it is usually computationally expensive, as a large number of mutants (buggy versions) are needed to be generated from a production code under test and executed against the test suite. In order to reduce the expensive testing cost, recent studies proposed to use supervised models to predict the effectiveness of a test suite without executing the test suite against the mutants. Nonetheless, the training of such a supervised model requires labeled data, which still depends on the costly mutant execution. Furthermore, existing models are based on traditional supervised learning techniques, which assume that the training and testing data come from the same distribution. But, in practice, software systems are subject to considerable concept drifts, i.e., the same distribution assumption usually does not hold. This can lead to inaccurate predictions of a learned supervised model on the target code as time progresses. To tackle these problems, in this paper, we propose a Coverage-Based Unsupervised Approach (CBUA) for evaluating the effectiveness of a test suite. Given a production code under test, the corresponding mutants, and a test suite, CBUA first collects the coverage information of the mutated statements in the target production code under the execution of the test suite. Then, CBUA employs coverage to estimate the probability of each mutant being alive. As such, a mutation score is computed to evaluate the test suite effectiveness and the predicted labels (i.e., killed or alive) are obtained. The whole process only requires a one-time execution of the test suite against the target production code, without involving any mutant execution and any training data. CBUA can ensure the score monotonicity property (i.e., adding test cases to a test suite does not decrease its mutation score), which may be violated by a supervised approach. The experimental results show that CBUA is very competitive with the state-of-the-art supervised approaches in prediction accuracy. In particular, CBUA is shown to be more effective in finding mutants that are covered but not killed by a test suite, which is helpful in identifying the weaknesses in the current test suite and generating new test cases accordingly. Since CBUA is an easy-to-implement approach with a low cost, we suggest that it should be used as a baseline approach for comparison when any novel prediction approach is proposed in future studies. Peng Zhang 0083, Yanhui Li 0001, Wanwangying Ma, Yibiao Yang, Lin Chen 0015, Hongmin Lu, Yuming Zhou, Baowen Xu |
IEEE Trans. Software Eng. | 6 |
| 2021 | Prioritizing code documentation effort: Can we do it simpler but better?
Shiran Liu, Zhaoqiang Guo, Yanhui Li 0001, Hongmin Lu, Lin Chen 0015, Lei Xu 0003, Yuming Zhou, Baowen Xu |
Inf. Softw. Technol. | 4 |
| 2021 | How Far Have We Progressed in Identifying Self-admitted Technical Debts? A Comprehensive Empirical StudyabstractBackground. Self-admitted technical debt (SATD) is a special kind of technical debt that is intentionally introduced and remarked by code comments. Those technical debts reduce the quality of software and increase the cost of subsequent software maintenance. Therefore, it is necessary to find out and resolve these debts in time. Recently, many automatic approaches have been proposed to identify SATD. Problem. Popular IDEs support a number of predefined task annotation tags for indicating SATD in comments, which have been used in many projects. However, such clear prior knowledge is neglected by existing SATD identification approaches when identifying SATD. Objective. We aim to investigate how far we have really progressed in the field of SATD identification by comparing existing approaches with a simple approach that leverages the predefined task tags to identify SATD. Method. We first propose a simple heuristic approach that fuzzily Matches task Annotation Tags ( MAT ) in comments to identify SATD. In nature, MAT is an unsupervised approach, which does not need any data to train a prediction model and has a good understandability. Then, we examine the real progress in SATD identification by comparing MAT against existing approaches. Result. The experimental results reveal that: (1) MAT has a similar or even superior performance for SATD identification compared with existing approaches, regardless of whether non-effort-aware or effort-aware evaluation indicators are considered; (2) the SATDs (or non-SATDs) correctly identified by existing approaches are highly overlapped with those identified by MAT ; and (3) supervised approaches misclassify many SATDs marked with task tags as non-SATDs, which can be easily corrected by their combinations with MAT . Conclusion. It appears that the problem of SATD identification has been (unintentionally) complicated by our community, i.e., the real progress in SATD comments identification is not being achieved as it might have been envisaged. We hence suggest that, when many task tags are used in the comments of a target project, future SATD identification studies should use MAT as an easy-to-implement baseline to demonstrate the usefulness of any newly proposed approach. Zhaoqiang Guo, Shiran Liu, Yanhui Li 0001, Lin Chen 0015, Hongmin Lu, Yuming Zhou |
ACM Trans. Softw. Eng. Methodol. | 6 |
| 2020 | Boosting crash-inducing change localization with rank-performance-based feature subset selection
Zhaoqiang Guo, Yanhui Li 0001, Wanwangying Ma, Yuming Zhou, Hongmin Lu, Lin Chen 0015, Baowen Xu |
Empir. Softw. Eng. | 5 |
| 2019 | Automatic Self-Validation for Code Coverage ProfilersabstractCode coverage as the primitive dynamic program behavior information, is widely adopted to facilitate a rich spectrum of software engineering tasks, such as testing, fuzzing, debugging, fault detection, reverse engineering, and program understanding. Thanks to the widespread applications, it is crucial to ensure the reliability of the code coverage profilers. Unfortunately, due to the lack of research attention and the existence of testing oracle problem, coverage profilers are far away from being tested sufficiently. Bugs are still regularly seen in the widely deployed profilers, like gcov and llvm-cov, along with gcc and llvm, respectively. This paper proposes Cod, an automated self-validator for effectively uncovering bugs in the coverage profilers. Starting from a test program (either from a compiler's test suite or generated randomly), Cod detects profiler bugs with zero false positive using a metamorphic relation in which the coverage statistics of that program and a mutated variant are bridged. We evaluated Cod over two of the most well-known code coverage profilers, namely gcov and llvm-cov. Within a four-month testing period, a total of 196 potential bugs (123 for gcov, 73 for llvm-cov) are found, among which 23 are confirmed by the developers. Yibiao Yang, Yanyan Jiang 0001, Zhiqiang Zuo 0002, Yang Wang 0165, Hao Sun 0021, Hongmin Lu, Yuming Zhou, Baowen Xu |
ASE | 6 |
| 2018 | How Far We Have Progressed in the Journey? An Examination of Cross-Project Defect PredictionabstractBackground. Recent years have seen an increasing interest in cross-project defect prediction (CPDP), which aims to apply defect prediction models built on source projects to a target project. Currently, a variety of (complex) CPDP models have been proposed with a promising prediction performance. Problem. Most, if not all, of the existing CPDP models are not compared against those simple module size models that are easy to implement and have shown a good performance in defect prediction in the literature. Objective. We aim to investigate how far we have really progressed in the journey by comparing the performance in defect prediction between the existing CPDP models and simple module size models. Method. We first use module size in the target project to build two simple defect prediction models, ManualDown and ManualUp, which do not require any training data from source projects. ManualDown considers a larger module as more defect-prone, while ManualUp considers a smaller module as more defect-prone. Then, we take the following measures to ensure a fair comparison on the performance in defect prediction between the existing CPDP models and the simple module size models: using the same publicly available data sets, using the same performance indicators, and using the prediction performance reported in the original cross-project defect prediction studies. Result. The simple module size models have a prediction performance comparable or even superior to most of the existing CPDP models in the literature, including many newly proposed models. Conclusion. The results caution us that, if the prediction performance is the goal, the real progress in CPDP is not being achieved as it might have been envisaged. We hence recommend that future studies should include ManualDown/ManualUp as the baseline models for comparison when developing new CPDP models to predict defects in a complete target project. Yuming Zhou, Yibiao Yang, Hongmin Lu, Lin Chen 0015, Yanhui Li 0001, Junyan Qian, Baowen Xu |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2017 | Training Data Selection for Cross-Project Defection Prediction: Which Approach Is Better?abstractBackground: Many relevancy filters have been proposed to select training data for building cross-project defect prediction (CPDP) models. However, up to now, there is no consensus about which relevancy filter is better for CPDP. Goal: In this paper, we conduct a thorough experiment to compare nine relevancy filters proposed in the recent literature. Method: Based on 33 publicly available data sets, we compare not only the retaining ratio of the original training data and the overlapping degree among the retained data but also the prediction performance of the resulting CPDP models under the ranking and classification scenarios. Results: In terms of retaining ratio and overlapping degree, there are important differences among these filters. According to the defect prediction performance, global filter always stays in the first level. Conclusions: For practitioners, it appears that there is no need to filter source project data, as this may lead to better defect prediction results. Yi Bin, Hongmin Lu, Yuming Zhou, Baowen Xu |
ESEM | 3 |
| 2017 | Code Churn: A Neglected Metric in Effort-Aware Just-in-Time Defect PredictionabstractBackground: An increasing research effort has devoted to just-in-time (JIT) defect prediction. A recent study by Yang et al. at FSE'16 leveraged individual change metrics to build unsupervised JIT defect prediction model. They found that many unsupervised models performed similarly to or better than the state-of-the-art supervised models in effort-aware JIT defect prediction. Goal: In Yang et al.'s study, code churn (i.e. the change size of a code change) was neglected when building unsupervised defect prediction models. In this study, we aim to investigate the effectiveness of code churn based unsupervised defect prediction model in effort-aware JIT defect prediction. Methods: Consistent with Yang et al.'s work, we first use code churn to build a code churn based unsupervised model (CCUM). Then, we evaluate the prediction performance of CCUM against the state-of-the-art supervised and unsupervised models under the following three prediction settings: cross-validation, time-wise cross-validation, and cross-project prediction. Results: In our experiment, we compare CCUM against the state-of-the-art supervised and unsupervised JIT defect prediction models. Based on six open-source projects, our experimental results show that CCUM performs better than all the prior supervised and unsupervised models. Conclusions: The result suggests that future JIT defect prediction studies should use CCUM as a baseline model for comparison when a novel model is proposed. Yuming Zhou, Yibiao Yang, Hongmin Lu, Baowen Xu |
ESEM | 4 |
| 2017 | An empirical investigation into the cost-effectiveness of test effort allocation strategies for finding faultsabstractIn recent years, it has been shown that fault prediction models could effectively guide test effort allocation in finding faults if they have a high enough fault prediction accuracy (Norm(Popt) > 0.78). However, it is often difficult to achieve such a high fault prediction accuracy in practice. As a result, fault-prediction-model-guided allocation (FPA) methods may be not applicable in real development environments. To attack this problem, in this paper, we propose a new type of test effort allocation strategy: reliability-growth-model-guided allocation (RGA) method. For a given project release V, RGA attempts to predict the optimal test effort allocation for V by learning the fault distribution information from the previous releases. Based on three open-source projects, we empirically investigate the cost-effectiveness of three test effort allocation strategies for finding faults: RGA, FPA, and structural-complexity-guided allocation (SCA) method. The experimental results show that RGA shows a promising performance in finding faults when compared with SCA and FPA. Yiyang Feng, Wanwangying Ma, Yibiao Yang, Hongmin Lu, Yuming Zhou, Baowen Xu |
SANER | 4 |
| 2017 | Understanding the value of considering client usage context in package cohesion for fault-proneness prediction
Yibiao Yang, Hongmin Lu, Hareton K. N. Leung, Yansong Wu, Yuming Zhou, Baowen Xu |
Autom. Softw. Eng. | 3 |
| 2016 | Effort-aware just-in-time defect prediction: simple unsupervised models could be better than supervised modelsabstractUnsupervised models do not require the defect data to build the prediction models and hence incur a low building cost and gain a wide application range. Consequently, it would be more desirable for practitioners to apply unsupervised models in effort-aware just-in-time (JIT) defect prediction if they can predict defect-inducing changes well. However, little is currently known on their prediction effectiveness in this context. We aim to investigate the predictive power of simple unsupervised models in effort-aware JIT defect prediction, especially compared with the state-of-the-art supervised models in the recent literature. We first use the most commonly used change metrics to build simple unsupervised models. Then, we compare these unsupervised models with the state-of-the-art supervised models under cross-validation, time-wise-cross-validation, and across-project prediction settings to determine whether they are of practical value. The experimental results, from open-source software systems, show that many simple unsupervised models perform better than the state-of-the-art supervised models in effort-aware JIT defect prediction. Yibiao Yang, Yuming Zhou, Hongmin Lu, Lei Xu 0003, Baowen Xu, Hareton K. N. Leung |
SIGSOFT FSE | 5 |
| 2016 | An empirical investigation into the effect of slice types on slice-based cohesion metrics
Yibiao Yang, Changsong Liu, Hongmin Lu, Yuming Zhou, Baowen Xu |
Inf. Softw. Technol. | 4 |
| 2015 | Predicting Vulnerable Components via Text Mining or Software Metrics? An Effort-Aware PerspectiveabstractIn order to identify vulnerable software components, developers can take software metrics as predictors or use text mining techniques to build vulnerability prediction models. A recent study reported that text mining based models have higher recall than software metrics based models. However, this conclusion was drawn without considering the sizes of individual components which affects the code inspection effort to determine whether a component is vulnerable. In this paper, we investigate the predictive power of these two kinds of prediction models in the context of effort-aware vulnerability prediction. To this end, we use the same data sets, containing 223 vulnerabilities found in three web applications, to build vulnerability prediction models. The experimental results show that: (1) in the context of effort-aware ranking scenario, text mining based models only slightly outperform software metrics based models, (2) in the context of effort-aware classification scenario, text mining based models perform similarly to software metrics based models in most cases, and (3) most of the effect sizes (i.e. the magnitude of the differences) between these two kinds of models are trivial. These results suggest that, from the viewpoint of practical application, software metrics based models are comparable to text mining based models. Therefore, for developers, software metrics based models are practical choices for vulnerability prediction, as the cost to build and apply these models is much lower. Yaming Tang, Yibiao Yang, Hongmin Lu, Yuming Zhou, Baowen Xu |
QRS | 4 |
| 2015 | Is Learning-to-Rank Cost-Effective in Recommending Relevant Files for Bug Localization?abstractSoftware bug localization aiming to determine the locations needed to be fixed for a bug report is one of the most tedious and effort consuming activities in software debugging. Learning-to-rank (LR) is the state-of-the-art approach proposed by Ye et al. to recommending relevant files for bug localization. Ye et al.'s experimental results show that the LR approach significantly outperforms previous bug localization approaches in terms of "precision" and "accuracy". However, this evaluation does not take into account the influence of the size of the recommended files on the efficiency in detecting bugs. In practice, developers will generally spend more code inspection effort to detect bugs if larger files are recommended. In this paper, we use six large-scale open-source Java projects to evaluate the LR approach in the context of effort-aware bug localization. Our results, surprisingly, show that, when taking into account the code inspection effort to detect bugs, the LR approach is similar to or even worse than the standard VSM (Vector Space Model), a naïve IR-based bug localization approach. Yaming Tang, Yibiao Yang, Hongmin Lu, Yuming Zhou, Baowen Xu |
QRS | 4 |
| 2015 | An empirical analysis of package-modularization metrics: Implications for software fault-proneness
Yibiao Yang, Hongmin Lu, Yuming Zhou, Qinbao Song, Baowen Xu |
Inf. Softw. Technol. | 3 |
| 2015 | Are Slice-Based Cohesion Metrics Actually Useful in Effort-Aware Post-Release Fault-Proneness Prediction? An Empirical StudyabstractBackground. Slice-based cohesion metrics leverage program slices with respect to the output variables of a module to quantify the strength of functional relatedness of the elements within the module. Although slice-based cohesion metrics have been proposed for many years, few empirical studies have been conducted to examine their actual usefulness in predicting fault-proneness. Objective. We aim to provide an in-depth understanding of the ability of slice-based cohesion metrics in effort-aware post-release fault-proneness prediction, i.e. their effectiveness in helping practitioners find post-release faults when taking into account the effort needed to test or inspect the code. Method. We use the most commonly used code and process metrics, including size, structural complexity, Halstead's software science, and code churn metrics, as the baseline metrics. First, we employ principal component analysis to analyze the relationships between slice-based cohesion metrics and the baseline metrics. Then, we use univariate prediction models to investigate the correlations between slice-based cohesion metrics and post-release fault-proneness. Finally, we build multivariate prediction models to examine the effectiveness of slice-based cohesion metrics in effort-aware post-release fault-proneness prediction when used alone or used together with the baseline code and process metrics. Results. Based on open-source software systems, our results show that: 1) slice-based cohesion metrics are not redundant with respect to the baseline code and process metrics; 2) most slice-based cohesion metrics are significantly negatively related to post-release fault-proneness; 3) slice-based cohesion metrics in general do not outperform the baseline metrics when predicting post-release fault-proneness; and 4) when used with the baseline metrics together, however, slice-based cohesion metrics can produce a statistically significant and practically important improvement of the effectiveness in effort-aware post-release fault-proneness prediction. Conclusion. Slice-based cohesion metrics are complementary to the most commonly used code and process metrics and are of practical value in the context of effort-aware post-release fault-proneness prediction. Yibiao Yang, Yuming Zhou, Hongmin Lu, Lin Chen 0015, Zhenyu Chen 0001, Baowen Xu, Hareton K. N. Leung, Zhenyu Zhang 0004 |
IEEE Trans. Software Eng. | 3 |
| 2012 | An in-depth investigation into the relationships between structural metrics and unit testability in object-oriented systems
Yuming Zhou, Hareton K. N. Leung, Qinbao Song, Jianjun Zhao 0001, Hongmin Lu, Lin Chen 0015, Baowen Xu |
Sci. China Inf. Sci. | 5 |
| 2012 | The ability of object-oriented metrics to predict change-proneness: a meta-analysis
Hongmin Lu, Yuming Zhou, Baowen Xu, Hareton K. N. Leung, Lin Chen 0015 |
Empir. Softw. Eng. | 1 |
| 2005 | DMC: a more precise cohesion measure for classes
Jianmin Wang 0001, Yuming Zhou, Lijie Wen 0001, Yujian Chen, Hongmin Lu, Baowen Xu |
Inf. Softw. Technol. | 5 |
| 2005 | An improved accuracy measure for rough sets
Baowen Xu, Yuming Zhou, Hongmin Lu |
J. Comput. Syst. Sci. | 3 |
| 2003 | DRC: A Dependence Relationships Based Cohesion Measure for ClassesabstractA large number of cohesion measures based on method-attribute references have been proposed. However, virtually no attention has been paid to the abstract representation that objectively depicts the relationships among the members of a class. Specially, the flow dependence relationship among attributes, the indirect and potential dependence relationships among class members, and the direction of method-attribute references are ignored. To address this problem, we first identifies four types of basic dependence relationships and uses a class member dependence graph to represent all dependences among the members of a class. Then, a dependence relationships based measure for measuring the class cohesiveness is proposed. Finally, we compare our class cohesion measure with typical cohesion measures. Yuming Zhou, Lijie Wen 0001, Jianmin Wang 0001, Yujian Chen, Hongmin Lu, Baowen Xu |
APSEC | 5 |