EDBT 2026 Demo / reviewers in the wild / expert
Yichen Li 0003
dblp:27/2248-3
· DBLP profile ↗
15ranked-venue papers
3as first author
15since 2021 · last 2026
0009-0009-8370-644XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 13 · 3 first-author · 13 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Proactive Change Risk Detection in Production Cloud Systems: ByteDance's ExperienceabstractModern cloud services rely on a high volume of changes for rapid innovation, yet these changes are a primary cause of production incidents. To manage this risk, industry practice employs tiered change management pipelines that concentrate static rules and human reviews. However, our analysis of ByteDance's cloud platform (Volcano Engine) reveals a critical long tail problem: while high-risk changes undergo extensive scrutiny, 78.1% of change-induced incidents originate from the sheer volume of low-risk changes that receive minimal review. At this scale, exhaustive manual review is fundamentally infeasible. To address this, we present Aegis, a novel knowledge-driven system that provides interpretable change risk assessment. Aegis automatically constructs a knowledge base from historical operational data, distilling it into generalized risk heuristics and identifying relevant precedent cases. When a new change is proposed, Aegis generates human-readable warnings grounded in this historical evidence, explaining why a change is risky. We have deployed Aegis in Volcano Engine's production environment for five months, where it processed tens of thousands of change requests. Its risk escalations achieved a 75% acceptance rate from production engineers and successfully prevented multiple potential incidents. Jinyang Liu 0002, Yichen Li 0003, Tieying Zhang, Binbin Chen 0005, Xiao He 0008, Yi Li 0098 |
EuroSys | 2 |
| 2026 | AutoLogger: A Multi-Agent Framework for the End-to-End Automated LoggingabstractSoftware logging is critical for system observability, yet developers face a dual crisis of costly overlogging and risky underlogging. Existing automated logging tools often overlook the fundamental whether-to-log decision and struggle with the composite nature of logging. In this paper, we propose AutoLogger, a novel hybrid framework that addresses the complete the end-to-end logging pipeline. AutoLogger first employs a fine-tuned classifier, the Judger, to accurately determine if a method requires new logging statements. If logging is needed, a multi-agent system is activated. The system includes specialized agents: a Locator dedicated to determining where to log, and a Generator focused on what to log. These agents work together, utilizing our designed program analysis and retrieval tools. We evaluate AutoLogger on a large corpus from three mature open-source projects against state-of-the-art baselines. Our results show that AutoLogger achieves 96.63% F1-score on the crucial whether-to-log decision. In an end-to-end setting, AutoLogger improves the overall quality of generated logging statements by 16.13% over the strongest baseline, as measured by an LLM-as-a-judge score. We also demonstrate that our framework is generalizable, consistently boosting the performance of various backbone LLMs. Renyi Zhong, Yintong Huo, Wenwei Gu, Yichen Li 0003, Michael R. Lyu |
ICPC | 4 |
| 2026 | LogUpdater: Automated Detection and Repair of Specific Defects in Logging StatementsabstractDevelopers write logging statements to monitor software runtime behaviors and system state. However, poorly constructed or misleading log messages can inadvertently obfuscate actual program execution patterns, thereby impeding effective software maintenance. Existing research on analyzing issues within logging statements is limited, primarily focusing on detecting a singular type of defect and relying on manual intervention for fixes rather than automated solutions. To address the limitation, we initiate a systematic study that pinpoints four specific types of defects in logging statements (i.e., statement code inconsistency, static dynamic inconsistency, temporal relation inconsistency, and readability issues) through the analysis of real-world log-centric changes. We then propose LogUpdater , a two-stage framework for automatically detecting and updating logging statements for these specific defects. In the offline stage, LogUpdater constructs a similarity-based classifier on a set of synthetic defective logging statements to identify specific defect types. During the online testing phase, this classifier first evaluates logging statements in a given code snippet to determine the necessity and type of improvements required. Then, LogUpdater constructs type-aware prompts from historical logging update changes for an LLM-based recommendation framework to suggest updates addressing these specific defects. We evaluate the effectiveness of LogUpdater on a dataset containing real-world logging changes, a synthetic dataset, and a new real-world project dataset. The results indicate that our approach is highly effective in detecting logging defects, achieving an F1-score of 0.625. Additionally, it exhibits significant improvements in suggesting precise static text and dynamic variables, with enhancements of 48.12% and 24.90%, respectively. Furthermore, LogUpdater achieves a 61.49% success rate in recommending correct updates on new real-world projects. We reported 40 problematic logging statements and their fixes to GitHub via pull requests, resulting in 25 changes confirmed and merged across 11 different projects. Renyi Zhong, Yichen Li 0003, Jinxi Kuang, Wenwei Gu, Yintong Huo, Michael R. Lyu |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2025 | COCA: Generative Root Cause Analysis for Distributed Systems with Code KnowledgeabstractRuntime failures are commonplace in modern distributed systems. When such issues arise, users often turn to platforms such as Github or JIRA to report them and request assistance. Automatically identifying the root cause of these failures is critical for ensuring high reliability and availability. However, prevailing automatic root cause analysis (RCA) approaches rely significantly on comprehensive runtime monitoring data, which is often not fully available in issue platforms. Recent methods leverage large language models (LLMs) to analyze issue reports, but their effectiveness is limited by incomplete or ambiguous user-provided information. To obtain more accurate and comprehensive RCA results, the core idea of this work is to extract additional diagnostic clues from code to supplement data-limited issue reports. Specifically, we propose COCA, a code knowledge enhanced root cause analysis approach for issue reports. Based on the data within issue reports, COCA intelligently extracts relevant code snippets and reconstructs execution paths, providing a comprehensive execution context for further RCA. Subsequently, COCA constructs a prompt combining historical issue reports along with profiled code knowledge, enabling the LLMs to generate detailed root cause summaries and localize responsible components. Our evaluation on datasets from five real-world distributed systems demonstrates that COCA significantly outperforms existing methods, achieving a$\mathbf{2 8. 3 \%}$improvement in root cause localization and a 22.0 % improvement in root cause summarization. Furthermore, COCA's performance consistency across various LLMs underscores its robust generalizability. Yichen Li 0003, Jinyang Liu 0002, Zhuangbin Chen, Guangba Yu, Michael R. Lyu |
ICSE | 1 |
| 2025 | LogPilot: Intent-aware and Scalable Alert Diagnosis for Large-scale Online Service SystemsabstractEffective alert diagnosis is essential for ensuring the reliability of large-scale online service systems. However, on-call engineers are often burdened with manually inspecting massive volumes of logs to identify root causes. While various automated tools have been proposed, they struggle in practice due to alert-agnostic log scoping and the inability to organize complex data effectively for reasoning. To overcome these limitations, we introduce LogPilot, an intent-aware and scalable framework powered by Large Language Models (LLMs) for automated log-based alert diagnosis. LogPilot introduces an intent-aware approach, interpreting the logic in alert definitions (e.g., PromQL) to precisely identify causally related logs and requests. To achieve scalability, it reconstructs each request’s execution into a spatiotemporal log chain, clusters similar chains to identify recurring execution patterns, and provides representative samples to the LLMs for diagnosis. This clustering-based approach ensures the input is both rich in diagnostic detail and compact enough to fit within the LLM’s context window. Evaluated on real-world alerts from Volcano Engine Cloud, LogPilot improves the usefulness of root cause summarization by 50.34% and exact localization accuracy by 54.79% over state-of-the-art methods. With a diagnosis time under one minute and a cost of only $0.074 per alert, LogPilot has been successfully deployed in production, offering an automated and practical solution for service alert diagnosis. Jinyang Liu 0002, Yichen Li 0003, Haiyu Huang 0008, Xiao He 0008, Tieying Zhang, Jianjun Chen 0001, Yi Li 0098, Michael R. Lyu |
ASE | 3 |
| 2025 | Automated Proactive Logging Quality Improvement for Large-Scale CodebasesabstractHigh-quality logging is critical for the reliability of cloud services, yet the industrial process for improving it is typically manual, reactive, and unscalable. Existing automated tools inherit this reactive nature, failing to answer the crucial whether-to-log question and are constrained to simple logging statement insertion, thus addressing only a fraction of the real-world logging improvement.To address these gaps and cope with logging debt in large-scale codebases, we propose LogImprover, a framework powered by Large Language Models (LLMs) that automates proactive logging quality improvement. LogImprover introduces two paradigm shifts: from reactive generation to proactive discovery, and from simple insertion to holistic logging patch generation. First, it identifies potential logging gaps based on principles distilled from industrial best practices. Then, it grounds each candidate through a cascading, structure-aware RAG module. Next, it prunes false positives by analyzing call-stack logging responsibilities and implicit logger inheritance. Finally, it generates holistic and explainable logging patches that reflect real-world development practices.Our evaluation provides dual confirmation of its effectiveness: LogImprover significantly outperforms state-of-the-art base-lines in closed-world experiments and achieves 68.12% developer acceptance rate in its real-world deployment. This success demonstrates the practical value of automating the entire logging quality improvement lifecycle, from discovery to recommendation. Yichen Li 0003, Jinyang Liu 0002, Junsong Pu, Zhuangbin Chen, Xiao He 0008, Tieying Zhang, Jianjun Chen 0001, Yi Li 0098, Michael R. Lyu |
ASE | 1 |
| 2025 | Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing TasksabstractAutonomous agent systems powered by Large Language Models (LLMs) have demonstrated promising capabilities in automating complex tasks. However, current evaluations largely rely on success rates without systematically analyzing the interactions, communication mechanisms, and failure causes within these systems. To bridge this gap, we present a benchmark of 34 representative programmable tasks designed to rigorously assess autonomous agents. Using this benchmark, we evaluate three popular open-source agent frameworks combined with two LLM backbones, observing a task completion rate of approximately 50%. Through in-depth failure analysis, we develop a three-tier taxonomy of failure causes aligned with task phases, highlighting planning errors, task execution issues, and incorrect response generation. Based on these insights, we propose actionable improvements to enhance agent planning and self-diagnosis capabilities. Our failure taxonomy, together with mitigation advice, provides an empirical foundation for developing more robust and effective autonomous agent systems in the future. Ruofan Lu, Yichen Li 0003, Yintong Huo |
ASE | 2 |
| 2025 | ErrorPrism: Reconstructing Error Propagation Paths in Cloud Service SystemsabstractReliability management in cloud service systems is challenging due to the cascading effect of failures. Error wrapping, a practice prevalent in modern microservice development, enriches errors with context at each layer of the function call stack, constructing an error chain that describes a failure from its technical origin to its business impact. However, this also presents a significant traceability problem when recovering the complete error propagation path from the final log message back to its source. Existing approaches are ineffective at addressing this problem. To fill this gap, we present ErrorPrism in this work for automated reconstruction of error propagation paths in production microservice systems. ErrorPrism first performs static analysis on service code repositories to build a function call graph and map log strings to relevant candidate functions. This significantly reduces the path search space for subsequent analysis. Then, ErrorPrism employs an LLM agent to perform an iterative backward search to accurately reconstruct the complete, multi-hop error path. Evaluated on 67 production microservices at ByteDance, ErrorPrism achieves 97.0% accuracy in reconstructing paths for 102 real-world errors, outperforming existing static analysis and LLM-based approaches. ErrorPrism provides an effective and practical tool for root cause analysis in industrial microservice systems. Junsong Pu, Yichen Li 0003, Zhuangbin Chen, Jinyang Liu 0002, Jianjun Chen 0001, Zibin Zheng, Tieying Zhang |
ASE | 2 |
| 2024 | A Large-Scale Evaluation for Log Parsing Techniques: How Far Are We?abstractLog data have facilitated various tasks of software development and maintenance, such as testing, debugging and diagnosing. Due to the unstructured nature of logs, log parsing is typically required to transform log messages into structured data for automated log analysis. Given the abundance of log parsers that employ various techniques, evaluating these tools to comprehend their characteristics and performance becomes imperative. Loghub serves as a commonly used dataset for benchmarking log parsers, but it suffers from limited scale and representativeness, posing significant challenges for studies to comprehensively evaluate existing log parsers or develop new methods. This limitation is particularly pronounced when assessing these log parsers for production use. To address these limitations, we provide a new collection of annotated log datasets, denoted Loghub-2.0, which can better reflect the characteristics of log data in real-world software systems. Loghub-2.0 comprises 14 datasets with an average of 3.6 million log lines in each dataset. Based on Loghub-2.0, we conduct a thorough re-evaluation of 15 state-of-the-art log parsers in a more rigorous and practical setting. Particularly, we introduce a new evaluation metric to mitigate the sensitivity of existing metrics to imbalanced data distributions. We are also the first to investigate the granular performance of log parsers on logs that represent rare system events, offering in-depth details for software diagnosis. Accurately parsing such logs is essential, yet it remains a challenge. We believe this work could shed light on the evaluation and design of log parsers in practical settings, thereby facilitating their deployment in production systems. Jinyang Liu 0002, Junjie Huang 0008, Yichen Li 0003, Yintong Huo, Jiazhen Gu, Zhuangbin Chen, Jieming Zhu, Michael R. Lyu |
ISSTA | 4 |
| 2024 | Face It Yourselves: An LLM-Based Two-Stage Strategy to Localize Configuration Errors via LogsabstractConfigurable software systems are prone to configuration errors, resulting in significant losses to companies. However, diagnosing these errors is challenging due to the vast and complex configuration space. These errors pose significant challenges for both experienced maintainers and new end-users, particularly those without access to the source code of the software systems. Given that logs are easily accessible to most end-users, we conduct a preliminary study to outline the challenges and opportunities of utilizing logs in localizing configuration errors. Based on the insights gained from the preliminary study, we propose an LLM-based two-stage strategy for end-users to localize the root-cause configuration properties based on logs. We further implement a tool, LogConfigLocalizer, aligned with the design of the aforementioned strategy, hoping to assist end-users in coping with configuration errors through log analysis. Shiwen Shan, Yintong Huo, Yuxin Su 0001, Yichen Li 0003, Dan Li 0016, Zibin Zheng |
ISSTA | 4 |
| 2024 | Exploring the Effectiveness of LLMs in Automated Logging Statement Generation: An Empirical StudyabstractAutomated logging statement generation supports developers in documenting critical software runtime behavior. While substantial recent research has focused on retrieval-based and learning-based methods, results suggest they fail to provide appropriate logging statements in real-world complex software. Given the great success in natural language generation and programming language comprehension, large language models (LLMs) might help developers generate logging statements, but this has not yet been investigated. To fill the gap, this paper performs the first study on exploring LLMs for logging statement generation. We first build a logging statement generation dataset,LogBench, with two parts: (1)LogBench-O:3,870methods with6,849logging statements collected from GitHub repositories, and (2)LogBench-T: the transformed unseen code from LogBench-O. Then, we leverage LogBench to evaluate theeffectivenessandgeneralization capabilities(usingLogBench-T) of 13 top-performing LLMs, from 60M to 405B parameters. In addition, we examine the performance of these LLMs against classical retrieval-based and machine learning-based logging methods from the era preceding LLMs. Specifically, we evaluate the logging effectiveness of LLMs by studying their ability to determine logging ingredients and the impact of prompts and external program information. We further evaluate LLM's logging generalization capabilities using unseen data (LogBench-T) derived from code transformation techniques. While existing LLMs deliver decent predictions on logging levels and logging variables, our study indicates that they only achieve a maximum BLEU score of0.249, thus calling for improvements. The paper also highlights the importance of prompt constructions and external factors (e.g., programming contexts and code comments) for LLMs’ logging performance. In addition, we observed that existing LLMs show a significant performance drop (8.2%-16.2%decrease) when dealing with logging unseen code, revealing their unsatisfactory generalization capabilities. Based on these findings, we identify five implications and provide practical advice for future logging research. Our empirical analysis discloses the limitations of current logging approaches while showcasing the potential of LLM-based logging tools, and provides actionable guidance for building more practical models. Yichen Li 0003, Yintong Huo, Renyi Zhong, Pinjia He, Yuxin Su 0001, Lionel C. Briand, Michael R. Lyu |
IEEE Trans. Software Eng. | 1 |
| 2024 | Meta-Path Based Attentional Graph Learning Model for Vulnerability DetectionabstractIn recent years, deep learning (DL)-based methods have been widely used in code vulnerability detection. The DL-based methods typically extract structural information from source code, e.g., code structure graph, and adopt neural networks such as Graph Neural Networks (GNNs) to learn the graph representations. However, these methods fail to consider the heterogeneous relations in the code structure graph, i.e., the heterogeneous relations mean that the different types of edges connect different types of nodes in the graph, which may obstruct the graph representation learning. Besides, these methods are limited in capturing long-range dependencies due to the deep levels in the code structure graph. In this paper, we propose aMeta-path basedAttentionalGraph learning model for code vulNErability deTection, calledMAGNET. MAGNET constructs a multi-granularity meta-path graph for each code snippet, in which the heterogeneous relations are denoted as meta-paths to represent the structural information. A meta-path based hierarchical attentional graph neural network is also proposed to capture the relations between distant nodes in the graph. We evaluate MAGNET on three public datasets and the results show that MAGNET outperforms the best baseline method in terms of F1 score by 6.32%, 21.50%, and 25.40%, respectively. MAGNET also achieves the best performance among all the baseline methods in detecting Top-25 most dangerous Common Weakness Enumerations (CWEs), further demonstrating its effectiveness in vulnerability detection. Xin-Cheng Wen, Cuiyun Gao 0001, Jiaxin Ye, Yichen Li 0003, Zhihong Tian 0001, Yan Jia 0001, Xuan Wang 0002 |
IEEE Trans. Software Eng. | 4 |
| 2023 | Improving the Transferability of Adversarial Samples by Path-Augmented MethodabstractDeep neural networks have achieved unprecedented success on diverse vision tasks. However, they are vulnerable to adversarial noise that is imperceptible to humans. This phenomenon negatively affects their deployment in real-world scenarios, especially security-related ones. To evaluate the robustness of a target model in practice, transfer-based attacks craft adversarial samples with a local model and have attracted increasing attention from researchers due to their high efficiency. The state-of-the-art transfer-based attacks are generally based on data augmentation, which typically augments multiple training images from a linear path when learning adversarial samples. However, such methods selected the image augmentation path heuristically and may augment images that are semantics-inconsistent with the target images, which harms the transferability of the generated adversarial samples. To overcome the pitfall, we propose the Path-Augmented Method (PAM). Specifically, PAM first constructs a candidate augmentation path pool. It then settles the employed augmentation paths during adversarial sample generation with greedy search. Furthermore, to avoid augmenting semantics-inconsistent images, we train a Semantics Predictor (SP) to constrain the length of the augmentation path. Extensive experiments confirm that PAM can achieve an improvement of over 4.8% on average compared with the state-of-the-art baselines in terms of the attack success rates. Jianping Zhang 0002, Jen-tse Huang 0001, Wenxuan Wang 0001, Yichen Li 0003, Weibin Wu 0002, Xiaosen Wang, Yuxin Su 0001, Michael R. Lyu |
CVPR | 4 |
| 2023 | AutoLog: A Log Sequence Synthesis Framework for Anomaly DetectionabstractThe rapid progress of modern computing systems has led to a growing interest in informative run-time logs. Various log-based anomaly detection techniques have been proposed to ensure software reliability. However, their implementation in the industry has been limited due to the lack of high-quality public log resources as training datasets. While some log datasets are available for anomaly detection, they suffer from limitations in (1) comprehensiveness of log events; (2) scalability over diverse systems; and (3) flexibility of log utility. To address these limitations, we propose AUTOLOG, the first automated log generation methodology for anomaly detection. AUTOLOG uses program analysis to generate runtime log sequences without actually running the system. AUTOLOG starts with probing comprehensive logging statements associated with the call graphs of an application. Then, it constructs execution graphs for each method after pruning the call graphs to find log-related execution paths in a scalable manner. Finally, AUTOLOG propagates the anomaly label to each acquired execution path based on human knowledge. It generates flexible log sequences by walking along the log execution paths with controllable parameters. Experiments on 50 popular Java projects show that AUTOLOG acquires significantly more (9x-58x) log events than existing log datasets from the same system, and generates log messages much faster (15x) with a single machine than existing passive data collection approaches. AUTOLOG also provides hyper-parameters to adjust the data size, anomaly rate, and component indicator for simulating different real-world scenarios. We further demonstrate AUTOLOG's practicality by showing that AUTOLOG enables log-based anomaly detectors to achieve better performance (1.93%) compared to existing log datasets. We hope AUTOLOG can facilitate the benchmarking and adoption of automated log analysis techniques. Yintong Huo, Yichen Li 0003, Yuxin Su 0001, Pinjia He, Zifan Xie, Michael R. Lyu |
ASE | 2 |
| 2023 | Revisiting, Benchmarking and Exploring API Recommendation: How Far Are We?abstractApplication Programming Interfaces (APIs), which encapsulate the implementation of specific functions as interfaces, greatly improve the efficiency of modern software development. As the number of APIs grows up fast nowadays, developers can hardly be familiar with all the APIs and usually need to search for appropriate APIs for usage. So lots of efforts have been devoted to improving the API recommendation task. However, it has been increasingly difficult to gauge the performance of new models due to the lack of a uniform definition of the task and a standardized benchmark. For example, some studies regard the task as a code completion problem, while others recommend relative APIs given natural language queries. To reduce the challenges and better facilitate future research, in this paper, we revisit the API recommendation task and aim at benchmarking the approaches. Specifically, the paper groups the approaches into two categories according to the task definition, i.e., query-based API recommendation and code-based API recommendation. We study 11 recently-proposed approaches along with 4 widely-used IDEs. One benchmark named APIBench is then built for the two respective categories of approaches. Based on APIBench, we distill some actionable insights and challenges for API recommendation. We also achieve some implications and directions for improving the performance of recommending APIs, including appropriate query reformulation, data source selection, low resource setting, user-defined APIs, and query-based API recommendation with usage patterns. Yun Peng 0003, Shuqing Li 0001, Wenwei Gu, Yichen Li 0003, Wenxuan Wang 0001, Cuiyun Gao 0001, Michael R. Lyu |
IEEE Trans. Software Eng. | 4 |