VLDB 2026 Research / reviewers in the wild / expert
Heng Dai
dblp:14/7036
· DBLP profile ↗
4ranked-venue papers
0as first author
2since 2021 · last 2023
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Finding the best learning to rank algorithms for effort-aware defect predictionabstractContext: Effort-Aware Defect Prediction (EADP) ranks software modules or changes based on their predicted number of defects (i.e., considering modules or changes as effort) or defect density (i.e., considering LOC as effort) by using learning to rank algorithms . Ranking instability refers to the inconsistent conclusions produced by existing empirical studies of EADP. The major reason is the poor experimental design , such as comparison of few learning to rank algorithms, the use of small number of datasets or datasets without indicating numbers of defects, and evaluation with inappropriate or few metrics. Objective: To find a stable ranking of learning to rank algorithms to investigate the best ones for EADP, Method: We examine the practical effects of 34 algorithms on 49 datasets for EADP. We measure the performance of these algorithms using 7 module-based and 7 LOC-based metrics and run experiments under cross-release and cross-project settings, respectively. Finally, we obtain the ranking of these algorithms by performing the Scott-Knott ESD test. Results: When module is used as effort, random forest regression performs the best under cross-release setting, and linear regression performs the best under cross-project setting among the learning to rank algorithms; (2) when LOC is used as effort, LTR-linear (Learning-to-Rank with the linear model) performs the best under cross-release setting, and Ranking SVM performs the best under cross-project setting. Conclusion: This comprehensive experimental procedure allows us to discover a stable ranking of the studied algorithms to select the best ones according to the requirement of software projects. Xiao Yu 0008, Heng Dai, Li Li 0029, Xiaodong Gu 0002, Jacky W. Keung, Kwabena Ebo Bennin, Jin Liu 0016 |
Inf. Softw. Technol. | 2 |
| 2022 | Predicting the precise number of software defects: Are we there yet?abstractContext: Defect Number Prediction (DNP) models can offer more benefits than classification-based defect prediction . Recently, many researchers proposed to employ regression algorithms for DNP, and found that the algorithms achieve low Average Absolute Error (AAE) and high Pred(0.3) values. However, since the defect datasets generally contain many non-defective modules, even if a DNP model predicts the number of defects in all modules as zero, the AAE value of the model will be low and Pred(0.3) value will be high. Therefore, the good performance of the regression algorithms in terms of AAE and Pred(0.3) may be questioned due to the imbalanced distribution of the number of defects. Objective: To revisit the impact of regression algorithms for predicting the precise number of defects. Method: We examine the practical effects of 12 widely-used regression algorithms, two data resampling algorithm (SmoteR and ROS), and three ensemble learning algorithms (gradient boosting regression, AdaBoost .R2, and Bagging), one feature selection method (information gain) and one parameter optimization method (grid search) for predicting the precise number of defects on the 18 PROMISE datasets. We propose to evaluate the AAE and Pred(0.3) values for the modules with different numbers of defects separately. Results: The AAE values for defective modules are very high and the Pred(0.3) values are very low, i.e., the regression algorithms are very inaccurate for predicting the precise number of defects in defective modules. Conclusion: The problem of predicting the precise number of defects via regression algorithms is far from being solved. We recommend that software testers use regression algorithms to rank modules for testing resource allocation , rather than predict the precise number of defects to evaluate the software reliability and maintenance effort. In addition, most existing DNP studies employing the whole AAE and Pred(0.3) values of all modules as the evaluation metrics for the proposed DNP algorithms should be revisited. Xiao Yu 0008, Jacky W. Keung, Yan Xiao 0002, Shuo Feng 0003, Heng Dai |
Inf. Softw. Technol. | 6 |
| 2019 | BreakID: genomics breakpoints identification to detect gene fusion events using discordant pairs and split readsabstractSUMMARY: Here we developed a tool called Breakpoint Identification (BreakID) to identity fusion events from targeted sequencing data. Taking discordant read pairs and split reads as supporting evidences, BreakID can identify gene fusion breakpoints at single nucleotide resolution. After validation with confirmed fusion events in cancer cell lines, we have proved that BreakID can achieve high sensitivity of 90.63% along with PPV of 100% at sequencing depth of 500× and perform better than other available fusion detection tools. We anticipate that BreakID will have an extensive popularity in the detection and analysis of fusions involved in clinical and research sequencing scenarios. AVAILABILITY AND IMPLEMENTATION: Source code is freely available at https://github.com/SinOncology/BreakID. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Linfang Jin, Jinhuo Lai, Yang Zhang 0056, Ying Fu 0002, Shuhang Wang, Heng Dai, Bingding Huang |
Bioinform. | 6 |
| 2010 | An effective scheduling scheme for multi-hop multicast in wireless mesh networks
Heng Dai, Farouk Y. M. Alkadhi, Jufeng Dai |
Frontiers Comput. Sci. China | 2 |