VLDB 2026 Research / reviewers in the wild / expert
Jesse Nyyssölä
dblp:326/0245
· DBLP profile ↗
8ranked-venue papers
4as first author
8since 2021 · last 2026
0009-0006-7276-5696ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 8 · 4 first-author · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VisualLogAnalyzer: An Interactive Web Application for Multi-Level Log Analysis
Jesse Nyyssölä, Simo Sipilä, Mika Mäntylä |
SANER | 1 |
| 2025 | Cross-System Software Log-based Anomaly Detection Using Meta-LearningabstractModern software systems produce vast amounts of logs, serving as an essential resource for anomaly detection. Artificial Intelligence for IT Operations (AIOps) tools have been developed to automate the process of log-based anomaly detection for software systems. Three practical challenges are widely recognized in this field: data labeling costs, evolving logs in dynamic systems, and adaptability across different systems. In this paper, we propose CroSysLog, an AIOps tool for log-event level anomaly detection, considering these challenges. Following prior approaches, CroSysLog uses a neural representation approach to gain a nuanced understanding of logs and generate representations for individual log events accordingly. CroSysLog can be trained on source systems with sufficient labeled logs from open datasets to achieve robustness, and then efficiently adapt to target systems with a few labeled log events for effective anomaly detection. We evaluate CroSysLog using open datasets of four large-scale distributed supercomputing systems: BGL, Thunderbird, Liberty, and Spirit. We used random log splits, maintaining the chronological order of consecutive log events, from these systems to train and evaluate CroSysLog. These splits were widely distributed across a one/two-year span of each system's log collection duration, capturing the evolving nature of the logs in each system. Our results show that, after training CroSysLog on Liberty and BGL as source systems, CroSysLog can efficiently adapt to target systems Thunderbird and Spirit using a few labeled log events from each target system, effectively performing anomaly detection for these target systems. The results demonstrate that CroSysLog is a practical, scalable, and adaptable tool for log-event level anomaly detection in operational and maintenance contexts of software systems. Yuqing Wang 0002, Mika Mäntylä, Jesse Nyyssölä, Ke Ping |
SANER | 3 |
| 2024 | A Dataset of Microservices-based Open-Source ProjectsabstractResearchers in the microservices community often resort to demonstrating the impact of their proposed advancements on custom-made microservices projects. This is a possible source of bias that can reduce the trustworthiness of the results. Moreover, it is hard to compare advances in small projects, often developed due to lack of time. It is common across disciplines to recognize benchmarks that mitigate bias and unify the advancements' impact. To facilitate the identification of available open-source microservice projects (OSS-MS), we performed a comprehensive study to identify, curate, and catalog OSS-MS. We started with 389559 projects and filtered them down to 3804 projects that we manually labeled. After manual labeling, our dataset contains 378 projects with three or more microservices and with over 100 commits. We document the projects from many perspectives, including project size, platform, number of contributors, project purpose, and foundation support. This dataset can serve researchers as a roadmap to identify benchmarks, as our dataset can be used to answer questions such as whether the number of services impacts the issue count. Dario Amoroso d'Aragona, Alexander Bakhtin, Xiaozhou Li 0002, Ruoyu Su, Lauren Adams, Ernesto Aponte, Francis Boyle, Patrick Boyle, Rachel Koerner, Joseph Lee, Fangchao Tian, Yuqing Wang 0002, Jesse Nyyssölä, Ernesto Quevedo Caballero, Md Shahidur Rahaman, Amr S. Abdelfattah, Mika Mäntylä, Tomás Cerný, Davide Taibi 0001 |
MSR | 13 |
| 2024 | Event-level Anomaly Detection on Software logs: Role of Algorithm, Threshold, and Window SizeabstractAnomaly detection on software logs has proven to be an efficient way to identify root causes of software related issues. Rather than identifying anomalies in sequences, this study aims to utilize an event-based anomaly detection approach to perform a ground-truth evaluation on the BlueGene/L (BGL) dataset. In the study we determined two approaches for adjusting the threshold of event-based anomaly detection: one that accounts for the absolute number of predictions (top-k) and another that accounts for cumulative probability of those predictions (top-p). These approaches were assessed on precision, recall and F1-score, and we found that top-k generally yields better results. We found that on deep learning (DL) models adjusting the threshold has the largest effect on F1-score among all our test configurations. If possible, we recommend using the top-k approach with a kvalue tailored to the dataset. Adjusting the threshold increased the F1-score from 0.846 of the baseline to 0.957 with optimal k-value. Regarding window size, we found that the optimal size is dependent on the model. N-Gram model benefits from as short sequences as possible while the DL models experience improvement up to the window size of 5. Adjusting the mask position of the prediction within the window can also improve the F1-score. Regarding the algorithm selection, with the best configuration the results show similar F1-scores on LSTM (0.9727) and CNN (0.9731) while the N-Gram (0.970) and Transformer (0.9675) model performed somewhat worse. As the best configurations allow us to consistently reach F1-scores of over 0.970 on BGL, we reached state of the art results on event-based anomaly detection. Jesse Nyyssölä, Mika Mäntylä |
QRS | 1 |
| 2024 | Speed and Performance of Parserless and Unsupervised Anomaly Detection Methods on Software LogsabstractSoftware log analysis can be laborious and time consuming. Time and labeled data are usually lacking in industrial settings. This paper studies unsupervised and time efficient methods for anomaly detection. We study two custom and two established models. The custom models are: an OOV (Out-Of-Vocabulary) detector, which counts the terms in the test data that are not present in the training data, and the Rarity Model (RM), which calculates a rarity score for terms based on their infrequency. The established models are KMeans and Isolation Forest. The models are evaluated on four public datasets (BGL, Thunderbird, Hadoop, HDFS) with three different representation techniques for the log messages (Words, character Trigrams, Parsed events). For training, we used both normal-only data, which is free of all anomalies, and unfiltered data, which contains both normal and anomalous instances. We used primarily the AUC-ROC metric for evaluation due to challenges in setting a threshold but we also include F1-scores for further insight. Different configurations are advised based on specific requirements. When training data is unfiltered, includes both normal and anomalous instances, the most effective combination is the Isolation Forest with event representation, achieving an AUC-ROC of 0.829. If it’s possible to create a normal-only training dataset, combining the Out-Of-Vocabulary (OOV) detector with trigram representation yields the highest AUC-ROC of 0.846. For speed considerations, the OOV detector is optimal for filtered data, while the Rarity Model is the best choice for unfiltered data. Jesse Nyyssölä, Mika Mäntylä |
QRS | 1 |
| 2024 | LogPM: Character-Based Log Parser BenchmarkabstractLog parsers transform free-form textual log messages into categorical data and are important tools in automated log analysis pipelines. However, selecting a suitable log parsing algorithm poses a formidable obstacle, thereby underscoring the importance of having a comprehensive benchmark to facilitate decision-making. This paper introduces a novel log parsing benchmark, focusing on predicted template precision at the character level rather than accurate grouping used in the past. We present a new metric called Parameter Mask Agreement that measures template accuracy at the character level alongside a dataset tailored for the task. We identified several challenges that parsers encounter for each dataset, which can aid in developing new log parsers. Moreover, a small empirical study was conducted using the proposed benchmark, evaluating the performance of three renowned parsers: Drain, Spell, and Lenma. The findings revealed that Lenma demonstrated the highest parsing accuracy, whereas Drain exhibited superior parsing speed. Finally, we propose that our benchmark is more appropriate than previous approaches in scenarios where accurate template detection is essential and computational efficiency needs to be assessed. Shayan Hashemi, Jesse Nyyssölä, Mika Mäntylä |
SANER | 2 |
| 2024 | LogLead - Fast and Integrated Log Loader, Enhancer, and Anomaly DetectorabstractThis paper introduces LogLead, a tool designed for efficient log analysis benchmarking. LogLead combines three essential steps in log processing: loading, enhancing, and anomaly detection. The tool leverages Polars, a high-speed DataFrame library. We currently have Loaders for eight systems that are publicly available (HDFS, Hadoop, BGL, Thunderbird, Spirit, Liberty, TrainTicket, and GC Webshop). We have multiple enhancers with three parsers (Drain, Spell, LenMa), Bert embedding creation and other log representation techniques like bag-of-words. LogLead integrates to five supervised and four unsupervised machine learning algorithms for anomaly detection from SKLearn. By integrating diverse datasets, log representation methods and anomaly detectors, LogLead facilitates comprehensive benchmarking in log analysis research. We show that log loading from raw file to dataframe is over 10x faster with LogLead compared to past solutions. We demonstrate roughly 2x improvement in Drain parsing speed by off-loading log message normalization to LogLead. Our brief benchmarking on HDFS indicates that log representations extending beyond the bag-of-words approach offer limited additional benefits. Tool URL: https://github.com/EvoTestOps/LogLead. Mika Mäntylä, Yuqing Wang 0002, Jesse Nyyssölä |
SANER | 3 |
| 2022 | How to Configure Masked Event Anomaly Detection on Software Logs?abstractSoftware Log anomaly event detection with masked event prediction has various technical approaches with countless configurations and parameters. Our objective is to provide a baseline of settings for similar studies in the future. The models we use are the N-Gram model, which is a classic approach in the field of natural language processing (NLP), and two deep learning (DL) models long short-term memory (LSTM) and convolutional neural network (CNN). For datasets we used four datasets Profilence, BlueGene/L (BGL), Hadoop Distributed File System (HDFS) and Hadoop. Other settings are the size of the sliding window which determines how many surrounding events we are using to predict a given event, mask position (the position within the window we are predicting), the usage of only unique sequences, and the portion of data that is used for training. The results show clear indications of settings that can be generalized across datasets. The performance of the DL models does not deteriorate as the window size increases while the N-Gram model shows worse performance with large window sizes on the BGL and Profilence datasets. Despite the popularity of Next Event Prediction, the results show that in this context it is better not to predict events at the edges of the subsequence, i.e., first or last event, with the best result coming from predicting the fourth event when the window size is five. Regarding the amount of data used for training, the results show differences across datasets and models. For example, the N-Gram model appears to be more sensitive toward the lack of data than the DL models. Overall, for similar experimental setups we suggest the following general baseline: Window size 10, mask position second to last, do not filter out non-unique sequences, and use a half of the total data for training. Jesse Nyyssölä, Mika Mäntylä, Martín Varela 0001 |
ICSME | 1 |