Hetong Dai

dblp:256/1720 · DBLP profile ↗
← Back
2ranked-venue papers
2as first author
2since 2021 · last 2023
0000-0001-8465-3380ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2023 PILAR: Studying and Mitigating the Influence of Configurations on Log Parsing
abstract
The significance of logs has been widely acknowledged with the adoption of various log analysis techniques that assist in software engineering tasks. Many log analysis techniques require structured logs as input while raw logs are typically unstructured. Automated log parsing is proposed to convert unstructured raw logs into structured log templates. Some log parsers achieve promising accuracy, yet they rely on significant efforts from the users to tune the parameters to achieve optimal results. In this paper, we first conduct an empirical study to understand the influence of the configurable parameters of six state-of-the-art log parsers on their parsing results on three aspects: 1) varying the parameters while using the same dataset, 2) keeping the same parameters while using different datasets, and 3) using different samples from the same dataset. Our results show that all these parsers are sensitive to the parameters, posing challenges to their adoption in practice. To mitigate such challenges, we propose PILAR (Parameter Insensitive Log Parser), an entropy-based log parsing approach. We compare PILAR with the existing log parsers on the same three aspects and find that PILAR is the most parameter-insensitive one. In addition, PILAR achieves the second highest parsing accuracy and efficiency among all the state-of-the-art log parsers. This paper paves the road for easing the adoption of log analysis in software engineer practices.
Hetong Dai, Yiming Tang 0002, Heng Li 0007, Weiyi Shang
ICSE1
2022 Logram: Efficient Log Parsing Using $n$n-Gram Dictionaries
abstract
Software systems usually record important runtime information in their logs. Logs help practitioners understand system runtime behaviors and diagnose field failures. As logs are usually very large in size, automated log analysis is needed to assist practitioners in their software operation and maintenance efforts. Typically, the first step of automated log analysis is log parsing, i.e., converting unstructured raw logs into structured data. However, log parsing is challenging, because logs are produced by static templates in the source code (i.e., logging statements) yet the templates are usually inaccessible when parsing logs. Prior work proposed automated log parsing approaches that have achieved high accuracy. However, as the volume of logs grows rapidly in the era of cloud computing, efficiency becomes a major concern in log parsing. In this work, we propose an automated log parsing approach,Logram, which leverages$n$-gram dictionaries to achieve efficient log parsing. We evaluatedLogramon 16 public log datasets and comparedLogramwith five state-of-the-art log parsing approaches. We found thatLogramachieves a higher parsing accuracy than the best existing approaches (i.e., at least 10 percent higher, on average) and also outperforms these approaches in efficiency (i.e., 1.8 to 5.1 times faster than the second-fastest approaches in terms of end-to-end parsing time). Furthermore, we deployedLogramonSparkand we found thatLogramscales out efficiently with the number ofSparknodes (e.g., with near-linear scalability for some logs) without sacrificing parsing accuracy. In addition, we demonstrated thatLogramcan support effective online parsing of logs, achieving similar parsing results and efficiency to the offline mode.
Hetong Dai, Heng Li 0007, Che-Shao Chen, Weiyi Shang, Tse-Hsun (Peter) Chen
IEEE Trans. Software Eng.1