VLDB 2026 Research / reviewers in the wild / expert
Shayan Hashemi
dblp:284/8497
· DBLP profile ↗
4ranked-venue papers
4as first author
4since 2021 · last 2026
0000-0001-6031-1765ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 4 · 4 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Token interdependency parsing (Tipping) - A map-reduced based statistical log parserabstractContext: Over the last decade, impressive growth in software adaptations has led to a surge in log data production, making manual log analysis impractical and underscoring the need for automated methods. On the other hand, most automated analysis tools benefit from a component that separates log templates from their parameters, commonly referred to as a “log parser”. Effectiveness and efficiency are inherently competing attributes, such that enhancing one typically requires compromising the other, making the simultaneous maximization of both challenging. Objective: This paper aims to introduce a new log parser capable of processing logs faster than the statistics-based methods while maintaining a respectable effectiveness, named “Tipping”. Method: Tipping combines rule-based tokenizers, interdependency token graphs, strongly connected components, and various techniques to ensure rapid, scalable, and precise log parsing. Furthermore, Tipping was designed with map-reduce principles to be parallelizable, enabling it to run on multiple processing cores to speed up the process. We evaluated Tipping against other statistics-based log parsers in terms of effectiveness, efficiency, and the downstream task of anomaly detection. Results: Accordingly, our evaluations demonstrate that Tipping outperforms most statistics-based parsing methods in both effectiveness and efficiency. More in-depth, Tipping can parse 11 million lines of logs in less than 20 s on a development machine. Furthermore, Tipping ranks among the top performers in terms of effectiveness and efficiency across the Loghub-2k, Loghub-2.0, LogPM, and LogLead benchmarks. Conclusion: Tipping’s robustness, versatility, efficiency, and scalability make it a viable tool for modern automated log analysis. However, Tipping currently cannot operate in online or streaming mode, which limits its applicability in environments where logs arrive continuously and must be processed incrementally in real time. Shayan Hashemi, Mika Mäntylä |
Inf. Softw. Technol. | 1 |
| 2024 | LogPM: Character-Based Log Parser BenchmarkabstractLog parsers transform free-form textual log messages into categorical data and are important tools in automated log analysis pipelines. However, selecting a suitable log parsing algorithm poses a formidable obstacle, thereby underscoring the importance of having a comprehensive benchmark to facilitate decision-making. This paper introduces a novel log parsing benchmark, focusing on predicted template precision at the character level rather than accurate grouping used in the past. We present a new metric called Parameter Mask Agreement that measures template accuracy at the character level alongside a dataset tailored for the task. We identified several challenges that parsers encounter for each dataset, which can aid in developing new log parsers. Moreover, a small empirical study was conducted using the proposed benchmark, evaluating the performance of three renowned parsers: Drain, Spell, and Lenma. The findings revealed that Lenma demonstrated the highest parsing accuracy, whereas Drain exhibited superior parsing speed. Finally, we propose that our benchmark is more appropriate than previous approaches in scenarios where accurate template detection is essential and computational efficiency needs to be assessed. Shayan Hashemi, Jesse Nyyssölä, Mika Mäntylä |
SANER | 1 |
| 2024 | OneLog: towards end-to-end software log anomaly detectionabstractAbstract With the growth of online services, IoT devices, and DevOps-oriented software development, software log anomaly detection is becoming increasingly important. Prior works mainly follow a traditional four-staged architecture (Preprocessor, Parser, Vectorizer, and Classifier). This paper proposes OneLog, which utilizes a single deep neural network instead of multiple separate components. OneLog harnesses convolutional neural network (CNN) at the character level to take digits, numbers, and punctuations, which were removed in prior works, into account alongside the main natural language text. We evaluate our approach in six message- and sequence-based data sets: HDFS, Hadoop, BGL, Thunderbird, Spirit, and Liberty. We experiment with Onelog with single-, multi-, and cross-project setups. Onelog offers state-of-the-art performance in our datasets. Onelog can utilize multi-project datasets simultaneously during training, which suggests our model can generalize between datasets. Multi-project training also improves Onelog performance making it ideal when limited training data is available for an individual project. We also found that cross-project anomaly detection is possible with a single project pair (Liberty and Spirit). Analysis of model internals shows that one log has multiple modes of detecting anomalies and that the model learns manually validated parsing rules for the log messages. We conclude that character-based CNNs are a promising approach toward end-to-end learning in log anomaly detection. They offer good performance and generalization over multiple datasets. We will make our scripts publicly available upon the acceptance of this paper. Shayan Hashemi, Mika Mäntylä |
Autom. Softw. Eng. | 1 |
| 2022 | SiaLog: detecting anomalies in software execution logs using the siamese networkabstractAbstract Detecting anomalies in software logs has become a notable concern for software engineers and maintainers as they represent anomalies in software execution paths and states. This paper propose a novel anomaly detection approach based on the Siamese network on top of Recurrent Neural Networks(RNN). Accordingly, we introduce a novel training pair generation algorithm to train the Siamese network which reduces generated training significantly while maintaining the $$F_1$$ F 1 score. Additionally, we propose a hybrid model by combining the Siamese network with a traditional feedforward neural network to make end-to-end training possible, reducing engineering effort in setting up a deep-learning-based log anomaly detector. Furthermore, we provides validations of the approach on the Hadoop Distributed File System (HDFS), Blue Gene/L (BGL), and Hadoop map-reduce task log datasets. To the best of our knowledge, the proposed approach outperforms other methods on the same dataset at the $$F_1$$ F 1 scores of respectively 0.99, 0.99, and 0.94 on HDFS, BGL, and Hadoop datasets, resulting in a new state-of-the-art performance.To further evaluate the proposed method, we examine our method’s robustness to log evolutions by evaluating the model on synthetically evolved log sequences; we got the $$F_1$$ F 1 score of 0.95 on the HDFS dataset at the noise ratio of $$20\%$$ 20 % . Finally, we dive deep into some of the side benefits of the Siamese network. Accordingly, we introduce an unsupervised log evolution monitoring method alongside a visualization technique that facilitates model interpretability. Shayan Hashemi, Mika Mäntylä |
Autom. Softw. Eng. | 1 |