VLDB 2026 Research / reviewers in the wild / expert
Maojin Li
dblp:254/8872
· DBLP profile ↗
7ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0003-4117-8430ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 7 · 1 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HiFlaky: Hierarchy-aware flakiness classification
Yan Lei 0005, Huan Xie 0002, Maojin Li |
J. Syst. Softw. | 5 |
| 2026 | Mutants Will Tell: Statistical Mutation-Based Multiple Fault Localization for Deep Learning ProgramsabstractAs deep learning (DL) systems are increasingly deployed in safety-critical domains, e.g., intelligent planning and autonomous driving, localizing faults that occur in such systems becomes indispensable. Inevitably, DL systems also suffer from faults like traditional software. Although single fault localization for DL programs has been studied, the multiple-fault localization for DL programs remains underexplored. We notice that mutation analysis is a powerful technique for locating multiple faults since it can simulate the faulty behaviors of a DL program by generating multiple mutants simultaneously. Thus, we propose MuMuFL: StatisticalMutation-basedMultipleFaultLocalization approach to locate the multiple faulty statements residing in a faulty DL program. The insight of MuMuFL is that the different behaviors of mutants provide valuable information for pinpointing the faulty statements of a DL fault. MuMuFL defines and leverages DL mutation operators on a DL program to simulate the faulty DL behavior. Then, MuMuFL evaluates the difference in the accuracy between the original DL model and the mutated DL model to quantify the suspiciousness of each statement being faulty. Finally, the large-scale experiments show that MuMuFL effectively localizes DL faults, e.g., localizing 36% of multiple-fault DL programs, whereas the best-performing baseline can only localize 14% of them. Huan Xie 0002, Zhengxiong Deng, Yan Lei 0005, Maojin Li, Meng Yan 0001, David Lo 0001 |
IEEE Trans. Software Eng. | 4 |
| 2024 | Combining Coverage and Expert Features with Semantic Representation for Coincidental Correctness DetectionabstractCoincidental correctness (CC) can be misleading for developers because it gives the impression that the code is functioning correctly when there are hidden faults. To mitigate the negative impacts of CC test cases, extensive research has been conducted on their detection, employing either coverage-based or expert-based features. These studies have yielded promising results. Coverage and expert features each provide unique insights into program execution, yet the literature has not fully explored the combined potential of these two feature sets to enhance the detection of CC. Additionally, the rich semantics of the test code and focal method have not been fully utilized. Therefore, we propose to build a unified model, CORE, that integrates coverage and expert features with semantic representations of test and focal methods to improve the detection of CC test cases. We make a comprehensive evaluation with six state-of-the-art baselines on the widely-used Defects4J benchmark. The experimental results show that CORE outperforms the baselines in terms of CC detection accuracy, with a substantial improvement (i.e., 40% improvement on average in terms of F1 score). Then, we conduct the ablation experiment to show that the coverage, expert, and semantics contribute to CORE. CORE can also improve the effectiveness of spectrum-based and mutation-based fault localization performance (e.g., 50% improvements for spectrum-based formula Dstar and 44% improvements for mutation-based method MUSE under relabeling strategy). Huan Xie 0002, Yan Lei 0005, Maojin Li, Meng Yan 0001 |
ASE | 3 |
| 2024 | Flakyrank: Predicting Flaky Tests Using Augmented Learning to RankabstractThe ideal principle of software testing is that test results ought to be deterministic: a test failure indicates the presence of a software bug, while a test success suggests the absence of a bug. Nevertheless, flaky tests break the principle. Flaky tests yield inconsistent results when executed repeatedly under the same conditions. The most straightforward approach runs the tests multiple times to predict flaky tests whereas it is highly time-consuming. Many researchers have proposed efficient approaches to reduce the cost, e.g., recent approaches leverage machine learning techniques for the prediction of flaky tests. However, traditional machine learning primarily focuses on predicting specific instances, which is not conducive to identify flaky tests across an entire project. Therefore, we propose Flakyrank, a ranking framework based on augmented learning to rank to predict flaky tests. The insight is that learning to rank, as compared with traditional machine learning, not only concentrates on individual samples but also optimizes the overall ranking. Since flaky tests constitute a small proportion of the dataset (i.e., approximately 3.6% of the total tests), we utilize generative adversarial networks to generate some synthetic flaky tests to augment the dataset. Based on the augmented dataset, FLAKYRANK treats predicting flaky tests as an information retrieval task, where newly detected flaky tests and test cases serve as queries and documents, respectively. For each newly detected flaky test (i.e., query), FLAKYRANK combines multiple relevant features into a learning to rank model to predict flaky tests candidate tests. We conduct large-scale experiments on different learning to rank models, and the results show that FLAKYRANK with the LambdaMART algorithm yields the best performance. In addition, the experimental results on 23 benchmark projects show that FLAKYRANK outperforms the state-of-the-art predictors. Jiaguo Wang, Yan Lei 0005, Maojin Li, Guanyu Ren, Huan Xie 0002, Shifeng Jin |
SANER | 3 |
| 2024 | Labelrepair: Sequence Labelling for Compilation Errors RepairabstractManual fixing of compilation errors could be a tedious and time-consuming task for novice programmers, and even for experienced ones. In recent years, an increasing number of automated repair techniques have been proposed to guide novice programmers and improve the efficiency of software development. Among them, learning-based automated repair techniques have achieved promising results in terms of repair accuracy. However, existing approaches neglect the time efficiency of patch generation, and often treat the compilation errors repair as a neural machine translation task. The end-to-end repair model decoding cannot be parallelized during the inference stage and suffers from redundant decoding search space. Furthermore, the large search space brought by the model poses a potential risk of semantic tampering. To this end, we propose Labelrepair, a novel repair technique that treats the repair of compilation errors as a sequence labelling task. Labelrepair discards the decoding model and converts the search for patches to a search for the error mapping actions between broken code and patch code pairs. In this way, the search space for patch tokens is plum-meted from the entire vocabulary to the size of edit action labels, and the logical semantics of the original program are preserved to some extent. The time complexity of inference is reduced from$O(n)$to$O$(1) owing to the parallel generation of edit action labels. Through a comprehensive evaluation of Labelrepair on two datasets, we demonstrate that Labelrepair is able to generate patches instantly (0.71ms on average), which is 28 times faster than existing end-to-end repair models. Compared with existing edit-based repair models, Labelrepair achieves the state-of-art repair accuracy. Deheng Yang, Yan Lei 0005, Huan Xie 0002, Minghua Tang, Maojin Li |
SANER | 6 |
| 2023 | On the Reliability of Coverage Data for Fault LocalizationabstractThe high quality of input data serves as the foundation for various tasks. Inaccurate data may decrease the effectiveness of elaborate algorithms and significantly impact the output. This also applies to fault localization, as accurate and reliable data is crucial for effective fault localization techniques. Many fault localization techniques analyze the coverage information for detecting bug positions. However, the source coverage data suffers from various problems, such as the imbalanced data and the coincidental correctness. These problems make the source coverage data unreliable for fault localization. To mitigate the potential adverse effect of these unreliable factors, we propose Orlando, a cOveRage-based decoupLing And recoNstructingData apprOach for fault localization. Or-landooptimizes the coverage data by synthesizing passing coverage with less coincidental correctness and failing coverage with more balanced data. The reconstructed data can provide more reliable source data for fault localization. We evaluate Orlando using the widely used Defects4J benchmark and demonstrate its effectiveness in improving two spectrum-based and two deep learning-based methods. Furthermore, Orlando outperforms state-of-the-art data optimization approaches in fault localization. Huan Xie 0002, Maojin Li, Yan Lei 0005, Shanshan Li 0001, Xiaoguang Mao, Yue Yu 0001 |
APSEC | 2 |
| 2023 | Contrastive Coincidental Correctness Representation LearningabstractA test suite is indispensable for fault localization by providing useful execution information of its test cases for locating suspicious statements of being faulty. There exists a type of test cases known as coincidental correctness (CC) test cases, which executes the faulty statement whereas produces the anticipated output. The existing studies have shown CC test cases harmfully impact fault localization effectiveness. Therefore, it is crucial to detect CC test cases to mitigate the adverse impact of CC test cases on fault localization.To address this issue, we propose ContraCC: a CC test cases detection method using contrastive learning. The insight of ContraCC is that the internal structural information of source test case execution data should be beneficial for CC detection whereas there is a lack of suitable representation methods. Inspired by the insight, ContraCC uses contrastive learning to learn new differentiated representations as test case vectors, which differentiate between similar and dissimilar pairs of test cases by maximizing their similarity within the same class and minimizing it between different classes. Based on the contrastive learning representations (i.e., test case vectors), ContraCC adopts multi-layer perceptron for binary classification to detect CC in downstream tasks. To evaluate the effectiveness of ContraCC, we conduct large-scale experiments on widely-used benchmarks by comparing ContraCC with five state-of-the-art CC test cases detection methods and applying ContraCC for fault localization. The experimental results show that ContraCC outperforms four state-of-the-art methods (e.g., from 10% to 84% improvement in Top-N on the best-performing baseline NeuralCCD) and significantly improves fault localization effectiveness (e.g., 24% improvement on the best-performing baseline Dstar). Maojin Li, Yan Lei 0005, Huan Xie 0002, Jiaguo Wang, Zhengxiong Deng |
ISSRE | 1 |