VLDB 2026 Research / reviewers in the wild / expert
Yonghao Wu
dblp:240/3883
· DBLP profile ↗
18ranked-venue papers
6as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 10 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploring the potential and limitations of large language models for novice program fault localization
Hexiang Xu, Hengyuan Liu, Yonghao Wu, Xiaolan Kang, Xiang Chen 0005, Yong Liu 0030 |
J. Syst. Softw. | 3 |
| 2026 | A visual-neural network for specific objects-of-interest inpainting
Yonghao Wu, Ruirong Wang, Kangwen Wu, Chang Liu 0057, Vladimir F. Filaretov, Dmitry Yukhimets |
Mach. Vis. Appl. | 1 |
| 2026 | MFHS: Mutual consistency learning-based foundation model integrates hypergraph for semi-supervised medical image segmentation
Zhaichao Tang, Yonghao Wu, Ruixiang Zhai, Xuanhe Dong, Zikang Du, Shujun Cao |
Pattern Recognit. | 3 |
| 2025 | Leveraging Retrieval Augmented Generation to Enhance LLM-Based Fault Localization for Novice ProgramsabstractFault localization (FL) in novice programs is critical for computer science education, yet it remains insufficiently explored compared to industrial programs. Traditional FL methods such as Spectrum-Based Fault Localization (SBFL) and Mutation-Based Fault Localization (MBFL) primarily rely on test coverage or mutation analysis, exhibiting significant limitations when handling novice programs. In recent years, Large Language Models (LLMs) have enabled more effective fault localization through semantic code analysis. However, existing LLMs primarily rely on internal knowledge acquired during training and lack access to specialized external knowledge bases, particularly limiting their effectiveness when addressing novice programs with distinctive fault patterns. To address these limitations, we propose NPFL-RAG, a LLM-based fault localization framework that combines Retrieval Augmented Generation (RAG). This framework regards the FL task as a three-step process: Knowledge Base Construction, which builds a comprehensive repository containing multidimensional fault-fixing information; Relevant Fix Cases Identification, which employs multidimensional retrieval and Retrieval Results Optimization (RRO) mechanism to identify the most relevant fix cases; and Fault Localization Result Generation, which integrates the retrieved knowledge with the LLM's inherent understanding to localize potential faults. We conduct a comprehensive experimental evaluation on the TutorCode dataset. The results demonstrate that NPFL-RAG significantly outperforms baselines across multiple evaluation metrics. In particular, when incorporating the RAG mechanism, DeepSeek-v3 achieves a 7.1 % improvement in the TOP-1 metric compared to the without RAG technique and performs approximately 7 times better than traditional MBFL techniques. Furthermore, we confirm the indispensability of the components in NPFL-RAG with the ablation study and demonstrate the usability of RRO mechanism in fault localization and explanation tasks. Xiaolan Kang, Hexiang Xu, Yonghao Wu, Yong Liu 0030 |
QRS | 4 |
| 2024 | Multi-objective optimization-based and fault localization-oriented test case generation for novice programsabstractSummary Online judgment (OJ) systems are capable of evaluating program results by automatically executing test cases, significantly improving the efficiency of traditional guidance approaches. Moreover, existing studies attempt to assist novices through automated fault localization techniques to provide feedback to novices, which can help them quickly find the location of faulty statements. Among them, spectrum‐based fault localization (SBFL) techniques have been widely used for their lightweight and efficiency, which only requires coverage information and test results of test cases to conduct fault localization. However, manually constructing high‐quality test cases for a large number of OJ questions is tough work to complete. To solve this problem, we propose the novice program‐orientedMulti‐Objective Optimization‐BasedFault Localization‐OrientedTestCaseGeneration (MFTCG) for automatically generating test inputs. Specifically, we use multi‐objective optimization algorithms to evolve the test case in terms of both fault localization and faulty code detection capability. We conduct experiments with 8911 programs from the well‐known public OJ platform AtCoder. The results show that our proposed approach MFTCG can achieve the best fault localization performance compared with existing automated test case generation approaches in most cases and can achieve the similar faulty code detection capability compared to manually designed test cases. Yong Liu 0030, Zezhong Yang, Luxi Fan, Yonghao Wu, Xiang Chen 0005, Xiaotang Zhou |
J. Softw. Evol. Process. | 4 |
| 2023 | Identifying Coincidental Correct Test Cases with Multiple Features Extraction for Fault LocalizationabstractSpectrum-Based Fault Localization (SBFL) technique is widely applied for fault localization, identifying faulty statements potentially resulting in unexpected faulty programs’ behavior. However, researchers have approved that Coincidental Correct (CC) test cases contained in test suites can negatively affect the accuracy of SBFL. Previous researchers sought to identify CC test cases through machine learning algorithms, but the feature representation is insufficient, leading to limited accuracy. To address this challenge, we propose the Machine Learning-based CC test cases Identification approach (MLCCI), which leverages multiple features extracted from the program under test to identify CC test cases and map the CC identification task to a learning problem. To evaluate the performance of MLCCI, we conduct experiments in the well-known dataset Defects4J. The experimental results compared with state-of-the-art baselines indicate that: (1) MLCCI achieves higher CC identifying accuracy, with the average Recall, P recision, and F -measure values of MLCCI are 65.93%, 71.69%, and 53.74%, respectively; (2) The fault localization accuracy of MLCCI with the Jaccard formula outperforms baselines, where the values of Accuracy@ 1, 3, and 5 are 347, 369, and 393, achieving the maximum 137.67%, 67.73%, and 47.74% improvement against baselines, respectively. Besides, we perform ablation analysis to reveal the effectiveness of features utilized in this study. Yonghao Wu, Shuaihua Tian, Zezhong Yang, Zheng Li 0002, Yong Liu 0030, Xiang Chen 0005 |
COMPSAC | 1 |
| 2023 | Extended Abstract of SeCNN: A semantic CNN parser for code comment generationabstractCode comments are essential for software development and maintenance, as they provide natural language descriptions of the code that help developers understand the program and reduce the time spent on comprehension. However, writing code comments can be tedious and time-consuming, and many software projects lack comprehensive and up-to-date comments, which can impair the readability and maintainability of programs. Zheng Li 0002, Yonghao Wu, Xiang Chen 0005, Zeyu Sun 0004, Yong Liu 0030, Deli Yu |
SANER | 2 |
| 2023 | Urban ride-hailing demand prediction with multi-view information fusion deep learning framework
Yonghao Wu, Huyin Zhang, Shiming Tao |
Appl. Intell. | 1 |
| 2023 | VsusFL: Variable-suspiciousness-based Fault Localization for novice programsabstractAutomatically localizing faulty statements is a desired feature for effective learning programming. Most of the existing automated fault localization techniques are developed and evaluated on commercial or well-known open-source projects, which performed poorly on novice programs. In this paper, we propose a novel fault localization technique VsusFL (Variable-suspiciousness-based Fault Localization) for novice programs. VsusFL is inspired by simulating the manual program debugging process and takes advantage of variable value sequences. VsusFL can trace variable value changes, determine whether the intermediate state of the variables is correct, and report the potential faulty statements for novice programs. This paper presents the implementation of VsusFL and conducts empirical studies on 422 real faulty novice programs. Experimental results show that VsusFL performs much better than Grace, ANGELINA, VSBFL, Spectrum-Based Fault Localization (SBFL), and Variable-based Fault Localization (VFL) in terms of T O P -1, T O P -3, and T O P -5 metrics. Specifically, VsusFL can localize 90%, 35% and 9% more faulty statements than the best-performing baseline Grace. Moreover, We analyze the correlation between VsusFL and other techniques and find a weak correlation since they perform well on different programs, indicating the potential to further enhance fault localization performance through strategic integration of VsusFL with other methods. Zheng Li 0002, Shumei Wu, Yong Liu 0030, Jitao Shen, Yonghao Wu, Zhanwen Zhang, Xiang Chen 0005 |
J. Syst. Softw. | 5 |
| 2023 | SeTransformer: A Transformer-Based Code Semantic Parser for Code Comment GenerationabstractAutomated code comment generation technologies can help developers understand code intent, which can significantly reduce the cost of software maintenance and revision. The latest studies in this field mainly depend on deep neural networks, such as convolutional neural networks and recurrent neural network. However, these methods may not generate high-quality and readable code comments due to the long-term dependence problem, which means that the code blocks used to summarize information are far from each other. Owing to the long-term dependence problem, these methods forget the previous input data’s feature information during the training process. In this article, to solve the long-term dependence problem and extract both the text and structure information from the program code, we propose a novel improved-Transformer-based comment generation method, named SeTransformer. Specifically, the SeTransformer utilizes the code tokens and an abstract syntax tree (AST) of programs to extract information as the inputs, and then, it leverages the self-attention mechanism to analyze the text and structural features of code simultaneously. Experimental results based on public corpus gathered from large-scale open-source projects show that our method can significantly outperform five state-of-the-art baselines (such as Hybrid-DeepCom and AST-attendgru). Furthermore, we also conduct a questionnaire survey for developers, and the results show that the SeTransformer can generate higher quality comments than those of other baselines. Zheng Li 0002, Yonghao Wu, Xiang Chen 0005, Zeyu Sun 0004, Yong Liu 0030, Paul Doyle |
IEEE Trans. Reliab. | 2 |
| 2023 | A target behavior pattern mining and abnormal behavior monitoring based on multidimensional similarity metric
Chang Liu 0057, Yonghao Wu, Ruslan Antypenko |
Wirel. Networks | 3 |
| 2022 | MVDLSTM: MultiView deep LSTM framework for online ride-hailing order prediction
Yonghao Wu, Huyin Zhang, Shiming Tao |
J. Supercomput. | 1 |
| 2022 | Theoretical Analysis and Empirical Study on the Impact of Coincidental Correct Test Cases in Multiple Fault LocalizationabstractTo improve the efficiency of the fault localization process, different automatic fault localization approaches have been proposed. Among these approaches, the spectrum-based fault localization (SBFL) approach has been widely used and studied due to its lightweight and high effectiveness. However, while the existence of coincidental correct (CC) test cases can influence the usefulness of SBFL in single-fault programs, their influence on multiple fault programs has not been thoroughly investigated. Therefore, in this article, we conduct a theoretical analysis and an empirical study to investigate the effect of CC test cases on multiple fault localization. The theoretical analysis is based on a suspiciousness calculation formula of SBFL, which divides CC test cases into three categories (specific, irrelevant, and unspecific) according to their association with a specific faulty statement. Following this analysis, we conduct an empirical study on two well-known open-source repositories (SIR and Defects4J), and the experimental results verify the correctness of our theoretical analysis. Specifically, reducing the number of specific CC test cases for a faulty statement can improve or maintain fault localization accuracy, while eliminating irrelevant CC test cases can have a negative effect. Finally, we design a CC test case identification solution based on the isolation-based multiple fault localization approach and demonstrate its effectiveness via a simulation experiment. Yonghao Wu, Yong Liu 0030, Weibo Wang 0007, Zheng Li 0002, Xiang Chen 0005, Paul Doyle |
IEEE Trans. Reliab. | 1 |
| 2021 | SeCNN: A semantic CNN parser for code comment generation
Zheng Li 0002, Yonghao Wu, Xiang Chen 0005, Zeyu Sun 0004, Yong Liu 0030, Deli Yu |
J. Syst. Softw. | 2 |
| 2020 | Using Fine-Grained Test Cases for Improving Novice Program Fault LocalizationabstractOnline Judge (OJ) system, which can automatically evaluate the results (right or wrong) of programs by executing them on standard test cases, is widely used in programming education. While an OJ system with personalized feedback can not only give execution results, but also provide information to assist students in locating their problems quickly. Automatically fault localization techniques are designed to find the exact faults in programs automatically, experimental results showed their effect on locating artificial faults, but their effectiveness on novice programs needs to be investigated. In this paper, we first evaluate the effectiveness of several widely-studied fault localization techniques on novice programs, and then we use fine-grained test cases to improve the fault localization accuracy. Empirical studies are conducted on 77 real student programs and the results show that, compared with original test cases in OJ system, the fault localization accuracy can be improved obviously when using fine-grained test cases. More specifically, in terms of TOP-1, TOP-3 and TOP-5 metrics, the maximum results can be improved from 5, 22, 37 to 9, 24, 48, respectively. The results indicate that more faults can be located when checking the top 1, 3 or 5 statements, so the fault localization accuracy is enhanced. Furthermore, a Test Case Granularity (TCG) concept is introduced to describe fine-grained test cases, and empirically studies demonstrate that there is a strong correlation between TCG and fault localization accuracy. Zheng Li 0002, Deli Yu, Yonghao Wu, Yong Liu 0030 |
COMPSAC | 3 |
| 2020 | FATOC: Bug Isolation Based Multi-Fault Localization by Using OPTICS Clustering
Yonghao Wu, Zheng Li 0002, Yong Liu 0030, Xiang Chen 0005 |
J. Comput. Sci. Technol. | 1 |
| 2019 | An Empirical Study of Bug Isolation on the Effectiveness of Multiple Fault LocalizationabstractBug isolation is the main approach to multi-fault localization, where failed test cases are divided into groups, and each group failed test cases are used to localize a single fault combined with all passed test cases. Ideally, all failed test cases within a single group execute the same faulty statements. However, misgrouping usually occurs due to the clustering algorithms may not able to divide failed test cases accurately. This paper focuses on the impact of fault localization by the accuracy of the clustering algorithm. A large quantitative empirical study is conducted on 12786 version programs with multiple faults, in which the misgrouping are simulated with different accuracy by a controlled experiment. The results indicate that the effect of fault localization will become worse as the accuracy of clustering decreases. Zheng Li 0002, Yonghao Wu, Yong Liu 0030 |
QRS | 2 |
| 2019 | A weighted fuzzy classification approach to identify and manipulate coincidental correct test cases for fault localization
Yong Liu 0030, Meiying Li, Yonghao Wu, Zheng Li 0002 |
J. Syst. Softw. | 3 |