Siyu Yu

dblp:135/8812 · DBLP profile ↗
← Back
15ranked-venue papers
4as first author
15since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 10 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 An Empirical Study on Commit Message Generation Using LLMs via In-Context Learning
abstract
Commit messages concisely describe code changes in natural language and are important for software maintenance. Several approaches have been proposed to automatically generate commit messages, but they still suffer from critical limitations, such as time-consuming training and poor generalization ability. To tackle these limitations, we propose to borrow the weapon of large language models (LLMs) and in-context learning (ICL). Our intuition is based on the fact that the training corpora of LLMs contain extensive code changes and their pairwise commit messages, which makes LLMs capture the knowledge about commits, while ICL can exploit the knowledge hidden in the LLMs and enable them to perform downstream tasks without model tuning. However, it remains unclear how well LLMs perform on commit message generation via ICL. In this paper, we conduct an empirical study to investigate the capability of LLMs to generate commit messages via ICL. Specifically, we first explore the impact of different settings on the performance of ICL-based commit message generation. We then compare ICL-based commit message generation with state-of-the-art approaches on a popular multilingual dataset and a new dataset we created to mitigate potential data leakage. The results show that ICL-based commit message generation significantly outperforms state-of-the-art approaches on subjective evaluation and achieves better generalization ability. We further analyze the root causes for LLM's underperformance and propose several implications, which shed light on future research directions for using LLMs to generate commit messages.
Yifan Wu 0002, Ying Li 0012, Siyu Yu, Wei Jiang 0041
ICSE5
2025 Log Parsing Using LLMs with Self-Generated In-Context Learning and Self-Correction
abstract
Log parsing transforms log messages into structured formats, serving as a crucial step for log analysis. Despite a variety of log parsers that have been proposed, their performance on evolving log data remains unsatisfactory due to reliance on human-crafted rules or learning-based models with limited training data. The recent emergence of large language models (LLMs) has demonstrated strong abilities in understanding natural language and code, making it promising to apply LLMs for log parsing. Consequently, several studies have proposed LLM-based log parsers. However, LLMs may produce inaccurate templates, and existing LLM-based log parsers directly use the template generated by the LLM as the parsing result, hindering the accuracy of log parsing. Furthermore, these log parsers depend heavily on historical log data as demonstrations, which poses challenges in maintaining accuracy when dealing with scarce historical log data or evolving log data. To address these challenges, we propose AdaParser, an effective and adaptive log parsing framework using LLMs with self-generated in-context learning (SG-ICL) and self-correction. To facilitate accurate log parsing, AdaParser incorporates a novel component, a template corrector, which utilizes the LLM to correct potential parsing errors in the templates it generates. In addition, AdaParser maintains a dynamic candidate set composed of previously generated templates as demonstrations to adapt evolving log data. Extensive experiments on public large-scale datasets indicate that AdaParser outperforms state-of-the-art methods across all metrics, even in zero-shot scenarios. Moreover, when integrated with different LLMs, AdaParser consistently enhances the performance of the utilized LLMs by a large margin.
Yifan Wu 0002, Siyu Yu, Ying Li 0012
ICPC2
2025 LogAction: Consistent Cross-system Anomaly Detection through Logs via Active Domain Adaptation
abstract
Log-based anomaly detection is a essential task for ensuring the reliability and performance of software systems. However, the performance of existing anomaly detection methods heavily relies on labeling, while labeling a large volume of logs is highly challenging. To address this issue, many approaches based on transfer learning and active learning have been proposed. Nevertheless, their effectiveness is hindered by issues such as the gap between source and target system data distributions and cold-start problems. In this paper, we propose LogAction, a novel log-based anomaly detection model based on active domain adaptation. LogAction integrates transfer learning and active learning techniques. On one hand, it uses labeled data from a mature system to train a base model, mitigating the cold-start issue in active learning. On the other hand, LogAction utilize free energy-based sampling and uncertainty-based sampling to select logs located at the distribution boundaries for manual labeling, thus addresses the data distribution gap in transfer learning with minimal human labeling efforts. Experimental results on six different combinations of datasets demonstrate that LogAction achieves an average 93.01% F1 score with only 2% of manual labels, outperforming some state-of-the-art methods by 26.28%. Website: https://logaction.github.io
Chiming Duan, Minghua He, Pei Xiao 0005, Zhewei Zhong, Yan Niu, Lingzhe Zhang, Siyu Yu, Yifan Wu 0002, Weijie Hong, Ying Li 0012, Gang Huang 0001
ASE10
2025 United We Stand: Towards End-to-End Log-based Fault Diagnosis via Interactive Multi-Task Learning
abstract
Log-based fault diagnosis is essential for maintaining software system availability. However, existing fault diagnosis methods are built using a task-independent manner, which fails to bridge the gap between anomaly detection and root cause localization in terms of data form and diagnostic objectives, resulting in three major issues: 1) Diagnostic bias accumulates in the system; 2) System deployment relies on expensive monitoring data; 3) The collaborative relationship between diagnostic tasks is overlooked. Facing this problems, we propose a novel end-to-end log-based fault diagnosis method, Chimera, whose key idea is to achieve end-to-end fault diagnosis through bidirectional interaction and knowledge transfer between anomaly detection and root cause localization. Chimera is based on interactive multitask learning, carefully designing interaction strategies between anomaly detection and root cause localization at the data, feature, and diagnostic result levels, thereby achieving both sub-tasks interactively within a unified end-to-end framework. Evaluation on two public datasets and one industrial dataset shows that Chimera outperforms existing methods in both anomaly detection and root cause localization, achieving improvements of over 2.92%~5.00% and 19.01% ~ 37.09%, respectively. It has been successfully deployed in production, serving an industrial cloud platform.
Minghua He, Chiming Duan, Pei Xiao 0005, Siyu Yu, Lingzhe Zhang, Weijie Hong, Yifan Wu 0002, Ying Li 0012, Gang Huang 0001
ASE5
2024 Unlocking the Power of Numbers: Log Compression via Numeric Token Parsing
abstract
Parser-based log compressors have been widely explored in recent years because the explosive growth of log volumes makes the compression performance of general-purpose compressors unsatisfactory. These parser-based compressors preprocess logs by grouping the logs based on the parsing result and then feed the preprocessed files into a general-purpose compressor. However, parser-based compressors have their limitations. First, the goals of parsing and compression are misaligned, so the inherent characteristics of logs were not fully utilized. In addition, the performance of parser-based compressors depends on the sample logs and thus it is very unstable. Moreover, parser-based compressors often incur a long processing time. To address these limitations, we propose Denum, a simple, general log compressor with high compression ratio and speed. The core insight is that a majority of the tokens in logs are numeric tokens (i.e. pure numbers, tokens with only numbers and special characters, and numeric variables) and effective compression of them is critical for log compression. Specifically, Denum contains a Numeric Token Parsing module, which extracts all numeric tokens and applies tailored processing methods (e.g. store the differences of incremental numbers like timestamps), and a String Processing module, which processes the remaining log content without numbers. The processed files of the two modules are then fed as input to a general-purpose compressor and it outputs the final compression results. Denum has been evaluated on 16 log datasets and it achieves an 8.7% -- 434.7% higher average compression ratio and 2.6× -- 37.7× faster average compression speed (i.e. 26.2 MB/S) compared to the baselines. Moreover, integrating Denum's Numeric Token Parsing module into existing log compressors can provide a 11.8% improvement in their average compression ratio and achieve 37% faster average compression speed.
Siyu Yu, Yifan Wu 0002, Ying Li 0012, Pinjia He
ASE1
2024 Prospective Learning: Learning for a Dynamic Future
abstract
In real-world applications, the distribution of the data, and our goals, evolve over time. The prevailing theoretical framework for studying machine learning, namely probably approximately correct (PAC) learning, largely ignores time. As a consequence, existing strategies to address the dynamic nature of data and goals exhibit poor real-world performance. This paper develops a theoretical framework called "Prospective Learning" that is tailored for situations when the optimal hypothesis changes over time. In PAC learning, empirical risk minimization (ERM) is known to be consistent. We develop a learner called Prospective ERM, which returns a sequence of predictors that make predictions on future data. We prove that the risk of prospective ERM converges to the Bayes risk under certain assumptions on the stochastic process generating the data. Prospective ERM, roughly speaking, incorporates time as an input in addition to the data. We show that standard ERM as done in PAC learning, without incorporating time, can result in failure to learn when distributions are dynamic. Numerical experiments illustrate that prospective ERM can learn synthetic and visual recognition problems constructed from MNIST and CIFAR-10. Code at https://github.com/neurodata/prolearn.
Ashwin De Silva, Rahul Ramesh, Rubing Yang, Siyu Yu, Joshua T. Vogelstein, Pratik Chaudhari
NeurIPS4
2024 XDrain: Effective log parsing in log streams using fixed-depth forest
Yang Tian 0008, Siyu Yu, Donghui Gao, Yifan Wu 0002, Suqun Huang, Xiaochun Hu, Ningjiang Chen
Inf. Softw. Technol.3
2024 Combining Image Editing and SinGAN for Conditional Sedimentary Facies Modeling
abstract
Traditional sedimentary facies modeling using generative adversarial networks (GANs) usually requires extensive datasets for network training. However, obtaining large datasets that align with reservoir depositional characteristics is often complex and costly. This letter introduces a conditional generative adversarial network (CSinGAN) based on a single training image. CSinGAN does not use conditional data in the training phase. In the model generation stage, conditional facies simulation is achieved by adjusting the intermediate model using image editing techniques. The conditional realizations of the three sets of training images successfully matched the well data. The variogram function and connectivity function indicate that CSinGAN can generate heterogeneous structures that conform to the statistical characteristics of the training image. We used multiscale sliced Wasserstein distance to verify that the realizations of the CSinGAN outperform the classical multipoint geostatistical algorithms. This study demonstrates the viability of using a single training image in GANs for conditional sedimentary facies modeling.
Changsheng Lu, Xixin Wang, Siyu Yu
IEEE Geosci. Remote. Sens. Lett.4
2024 Reliable and adaptive computation offload strategy with load and cost coordination for edge computing
Weicheng Tang, Donghui Gao, Siyu Yu, Zhanrong Li, Ningjiang Chen
Pervasive Mob. Comput.3
2023 A Monolithic GaN Power Stage with Low Propagation Delay and High Reliability Level Shifting for High Frequency Power Converter
abstract
In this article, a monolithic gallium nitride (GaN)-based half-bridge power stage is proposed for high frequency power converter. A level shifting technique with buffering device and shielding capacitance is designed to perform communications in half-bridge topologies with high reliability and low propagation delay. To prevent the missing pulse or output latch during deadtime, a negative voltage pull-down circuit is designed. Implemented with a$0.25\mu \mathrm{m}$15-V enhanced mode GaN (eGaN) process, this work consists of optimized delay matching, monolithic gate drivers and power high electron mobility transistors(HEMTs) as well. The implementation results show that this work can achieve a propagation delay of no more than 6.5ns, a delay mismatch within 2.5ns and a CMTI capability above 150V/ns at 30Mhz switching frequency.
Rongxing Lai, Ze-kun Zhou, Jinyang He, Siyu Yu, William Li, Bo Zhang 0027
ISCAS4
2023 Transient Data Caching Based on Maximum Entropy Actor-Critic in Internet-of-Things Networks
abstract
With the rapid development of the Internet of Things (IoT), a massive amount of transient data is transmitted in edge networks.Transient data is highly time-sensitive, such as monitoring data generated by industrial devices.Due to their inefficiency, traditional caching strategies in edge networks are inadequate for handling transient data.Thus, to improve the efficiency of transient data caching, we construct a freshness model of transient data and propose a maximum entropy Actor-Critic based caching strategy, TD-MEAC-which can improve the freshness of cached data and reduce the long-term caching cost.Simulation results show that the proposed TD-MEAC achieves a higher cache hit rate and maintains a higher average freshness of cached transient data compared with the existing DRL and baseline caching strategies.
Ningjiang Chen, Siyu Yu
SEKE3
2023 Log Parsing with Generalization Ability under New Log Types
abstract
Log parsing, which converts semi-structured logs into structured logs, is the first step for automated log analysis. Existing parsers are still unsatisfactory in real-world systems due to new log types in new-coming logs. In practice, available logs collected during system runtime often do not contain all the possible log types of a system because log types related to infrequently activated system states are unlikely to be recorded and new log types are frequently introduced with system updates. Meanwhile, most existing parsers require preprocessing to extract variables in advance, but preprocessing is based on the operator’s prior knowledge of available logs and therefore may not work well on new log types. In addition, parser parameters set based on available logs are difficult to generalize to new log types. To support new log types, we propose a variable generation imitation strategy to craft a novel log parsing approach with generalization ability, called Log3T. Log3T employs a pre-trained transformer encoder-based model to extract log templates and can update parameters at parsing time to adapt to new log types by a modified test-time training. Experimental results on 16 benchmark datasets show that Log3T outperforms the state-of-the-art parsers in terms of parsing accuracy. In addition, Log3T can automatically adapt to new log types in new-coming logs.
Siyu Yu, Yifan Wu 0002, Zhijing Li 0007, Pinjia He, Ningjiang Chen
ESEC/SIGSOFT FSE1
2023 Self-supervised log parsing using semantic contribution difference
Siyu Yu, Ningjiang Chen, Yifan Wu 0002, Wensheng Dou
J. Syst. Softw.1
2023 Brain: Log Parsing With Bidirectional Parallel Tree
abstract
Automated log analysis can facilitate failure diagnosis for developers and operators using a large volume of logs. Log parsing is a prerequisite step for automated log analysis, which parses semi-structured logs into structured logs. However, existing parsers are difficult to apply to software-intensive systems, due to their unstable parsing accuracy on various software. Although neural network-based approaches are stable, their inefficiency makes it challenging to keep up with the speed of log production. In this work, we found that the longest common pattern among logs is likely to be part of the log template. Inspired by this key insight, we propose a new stable log parsing approach, called Brain, which creates initial groups according to the longest common pattern. Then a bidirectional tree is used to hierarchically complement the constant words to the longest common pattern to form the complete log template efficiently. Experimental results on 16 benchmark datasets show that our approach outperforms the state-of-the-art parsers on two widely-used parsing accuracy metrics, and it only takes around 46 seconds to process one million lines of logs.
Siyu Yu, Pinjia He, Ningjiang Chen, Yifan Wu 0002
IEEE Trans. Serv. Comput.1
2021 Crowdsourcing the perceived urban built environment via social media: The case of underutilized land
Yan Wang 0066, Shangde Gao, Nan Li 0014, Siyu Yu
Adv. Eng. Informatics4