VLDB 2026 Research / reviewers in the wild / expert
Honglin Shu
dblp:340/7859
· DBLP profile ↗
8ranked-venue papers
4as first author
8since 2021 · last 2026
0009-0005-7311-7060ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 6 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evaluating large language models for multilingual vulnerability detection at dual granularities
Honglin Shu, Junji Yu, Dong Wang 0044, Chakkrit Tantithamthavorn, Junjie Chen 0003, Yasutaka Kamei |
Empir. Softw. Eng. | 1 |
| 2026 | An Empirical Study on Language Models for Generating Log Statements in Test CodeabstractLog statements play a critical role in modern software development, capturing essential run-time information necessary for software maintenance. Recently, new techniques have been developed to automate logging activities, allowing log statements to be injected into code by identifying specific code locations, selecting the appropriate log level, and generating meaningful log messages that describe the behavior being logged. Although automated logging in production code has attracted significant attention, little focus has been given to the injection of logs in test code. To fill this gap, we conduct an empirical study on 5,206,759 Java test methods collected from 6,405 GitHub projects to explore and disclose the effectiveness and limitations of Pre-Trained Language Models (PLMs) and Large Language Models (LLMs) for generating and injecting test log statements. Our findings demonstrate that general-purpose LLMs like GPT-3.5-Turbo, when properly instructed to inject logging statements in test methods, performs comparably to the best-performing PLMs on predicting log level. Additionally, GPT-3.5-Turbo substantially outperforms the best in PLMs on predicting log position, with a 33.97% improvement while also achieving superior performance in predicting log messages in terms of BLEU and ROUGE . This work takes the first step toward evaluating the capability of PLMs and LLMs to generate test log statement. To facilitate future research, we have open sourced all data and source code used in this work. Honglin Shu, Dong Wang 0044, Antonio Mastropaolo, Gabriele Bavota, Yasutaka Kamei |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2025 | How Small is Enough? Empirical Evidence of Quantized Small Language Models for Automated Program RepairabstractBackground: Large language models (LLMs) have greatly improved the accuracy of automated program repair (APR) methods. However, LLMs are constrained by high computational resource requirements. Aims: We focus on small language models (SLMs), which perform well even with limited computational resources compared to LLMs. We aim to evaluate whether SLMs can achieve competitive performance in APR tasks. Method: We conducted experiments on the QuixBugs benchmark to compare the bug-fixing accuracy of SLMs and LLMs. We also analyzed the impact of int8 quantization on APR performance. Results: The latest SLMs can fix bugs as accurately as-or even more accurately than-LLMs. Also, int8 quantization had minimal effect on APR accuracy while significantly reducing memory requirements. Conclusions: SLMs present a viable alternative to LLMs for APR, offering competitive accuracy with lower computational costs, and quantization can further enhance their efficiency without compromising effectiveness. Kazuki Kusama, Honglin Shu, Masanari Kondo, Yasutaka Kamei |
ESEM | 2 |
| 2025 | A Preliminary Study on Large Language Models Self-Negotiation in Software EngineeringabstractLarge Language Models (LLMs) have shown great potential in code-related software engineering tasks, including code generation, classification, and understanding. While current research primarily focuses on direct inference with single LLMs, this approach may fall short for complex tasks due to ambiguous instructions. To address this limitation, we propose LLM selfnegotiation, where multiple LLMs collaborate and debate to reach consensus on code-related tasks. This approach aims to better handle unclear instructions and improve overall effectiveness. We evaluated LLM self-negotiation in three key software engineering domains: Equivalent Mutant Detection (EMD), Automated Vulnerability Detection (AVD), and Automated Program Repair (APR). These domains represent distinct aspects of coderelated tasks: functionality understanding, code classification, and code generation, respectively. Our experimental results revealed varying effectiveness across domains. In EMD, LLM self-negotiation demonstrated remarkable improvements, with most models showing performance gains between 114.72 % and 351.01% (though CodeLlama experienced a minor 4.5% decrease in F1-score). For APR tasks, self-negotiation performed comparably to single LLM implementations. However, in AVD, the results were mixed - while Vicuna showed improved F1-scores, most models exhibited lower recall rates. These findings indicate that LLM self-negotiation is particularly promising for functionality understanding tasks, while its application to code classification and generation requires further research and refinement. Chunrun Tao, Honglin Shu, Masanari Kondo, Yasutaka Kamei |
ICSME | 2 |
| 2025 | My Fuzzers Won't Build: An Empirical Study of Fuzzing Build FailuresabstractFuzzing is an automated software testing technique used to find software vulnerabilities that works by sending large amounts of inputs to a software system to trigger bad behaviors. In recent years, the open source software ecosystem has seen a significant increase in the adoption of fuzzing to avoid spreading vulnerabilities throughout the ecosystem. While fuzzing can uncover vulnerabilities, there is currently a lack of knowledge regarding the challenges of conducting fuzzing activities over time. Specifically, fuzzers are very complex tools to set up and build before they can be used. We set out to empirically find out how challenging is build maintenance in the context of fuzzing. We mine over 1.2 million build logs from Google’s OSS-Fuzz service to investigate fuzzing build failures. We first conduct a quantitative analysis to quantify the prevalence of fuzzing build failures. We then manually investigate 677 failing fuzzing builds logs and establish a taxonomy of 25 root causes of build failures. We finally train a machine learning model to recognize common failure patterns in failing build logs. Our taxonomy can serve as a reference for practitioners conducting fuzzing build maintenance. Our modeling experiment shows the potential of using automation to simplify the process of fuzzing. Olivier Nourry, Yutaro Kashiwa, Weiyi Shang, Honglin Shu, Yasutaka Kamei |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2024 | Large Language Models for Equivalent Mutant Detection: How Far Are We?abstractMutation testing is vital for ensuring software quality. However, the presence of equivalent mutants is known to introduce redundant cost and bias issues, hindering the effectiveness of mutation testing in practical use. Although numerous equivalent mutant detection (EMD) techniques have been proposed, they exhibit limitations due to the scarcity of training data and challenges in generalizing to unseen mutants. Recently, large language models (LLMs) have been extensively adopted in various code-related tasks and have shown superior performance by more accurately capturing program semantics. Yet the performance of LLMs in equivalent mutant detection remains largely unclear. In this paper, we conduct an empirical study on 3,302 method-level Java mutant pairs to comprehensively investigate the effectiveness and efficiency of LLMs for equivalent mutant detection. Specifically, we assess the performance of LLMs compared to existing EMD techniques, examine the various strategies of LLMs, evaluate the orthogonality between EMD techniques, and measure the time overhead of training and inference. Our findings demonstrate that LLM-based techniques significantly outperform existing techniques (i.e., the average improvement of 35.69% in terms of F1-score), with the fine-tuned code embedding strategy being the most effective. Moreover, LLM-based techniques offer an excellent balance between cost (relatively low training and inference time) and effectiveness. Based on our findings, we further discuss the impact of model size and embedding quality, and provide several promising directions for future research. This work is the first to examine LLMs in equivalent mutant detection, affirming their effectiveness and efficiency. Zhao Tian 0002, Honglin Shu, Dong Wang 0044, Xuejie Cao, Yasutaka Kamei, Junjie Chen 0003 |
ISSTA | 2 |
| 2023 | Drugs Resistance Analysis from Scarce Health Records via Multi-task Graph Representation
Honglin Shu, Pei Gao, Lingwei Zhu, Zheng Chen 0012, Yasuko Matsubara, Yasushi Sakurai |
ADMA (3) | 1 |
| 2023 | MetaGC-MC: A graph-based meta-learning approach to cold-start recommendation with/without auxiliary information
Honglin Shu, Korris Fu-Lai Chung |
Inf. Sci. | 1 |