VLDB 2026 Research / reviewers in the wild / expert
Alexander Acker
dblp:213/1424
· DBLP profile ↗
15ranked-venue papers
3as first author
4since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 2 first-authorSoftware engineering, systems software and programming languages · 4 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-authorComputer networks · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | QuLog: data-driven approach for log instruction quality assessmentabstractIn the current IT world, developers write code while system operators run the code mostly as a black box. The connection between both worlds is typically established with log messages: the developer provides hints to the (unknown) operator, where the cause of an occurred issue is, and vice versa, the operator can report bugs during operation. To fulfil this purpose, developers write log instructions that are structured text commonly composed of a log level (e.g., "info", "error"), static text ("IP {} cannot be reached"), and dynamic variables (e.g. IP {}). However, opposed to well-adopted coding practices, there are no widely adopted guidelines on how to write log instructions with good quality properties. For example, a developer may assign a high log level (e.g., "error") for a trivial event that can confuse the operator and increase maintenance costs. Or the static text can be insufficient to hint at a specific issue. In this paper, we address the problem of log quality assessment and provide the first step towards its automation. We start with an in-depth analysis of quality log instruction properties in nine software systems and identify two quality properties: 1) correct log level assignment assessing the correctness of the log level, and 2) sufficient linguistic structure assessing the minimal richness of the static text necessary for verbose event description. Based on these findings, we developed a data-driven approach that adapts deep learning methods for each of the two properties. An extensive evaluation on large-scale open-source systems shows that our approach correctly assesses log level assignments with an accuracy of 0.88, and the sufficient linguistic structure with an F1 score of 0.99, outperforming the baselines. Our study highlights the potential of the data-driven methods in assessing log instructions quality. Jasmin Bogatinovski, Sasho Nedelkoski, Alexander Acker, Jorge Cardoso 0001, Odej Kao |
ICPC | 3 |
| 2021 | Bellamy: Reusing Performance Models for Distributed Dataflow Jobs Across ContextsabstractDistributed dataflow systems enable the use of clusters for scalable data analytics. However, selecting appropriate cluster resources for a processing job is often not straightforward. Performance models trained on historical executions of a concrete job are helpful in such situations, yet they are usually bound to a specific job execution context (e.g. node type, software versions, job parameters) due to the few considered input parameters. Even in case of slight context changes, such supportive models need to be retrained and cannot benefit from historical execution data from related contexts.This paper presents Bellamy, a novel modeling approach that combines scale-outs, dataset sizes, and runtimes with additional descriptive properties of a dataflow job. It is thereby able to capture the context of a job execution. Moreover, Bellamy is realizing a two-step modeling approach. First, a general model is trained on all the available data for a specific scalable analytics algorithm, hereby incorporating data from different contexts. Subsequently, the general model is optimized for the specific situation at hand, based on the available data for the concrete context. We evaluate our approach on two publicly available datasets consisting of execution data from various dataflow jobs carried out in different environments, showing that Bellamy outperforms state-of-the-art methods. Dominik Scheinert, Lauritz Thamsen, Houkun Zhu, Jonathan Will, Alexander Acker, Thorsten Wittkopp, Odej Kao |
CLUSTER | 5 |
| 2021 | LogLAB: Attention-Based Labeling of Log Data Anomalies via Weak Supervision
Thorsten Wittkopp, Philipp Wiesner, Dominik Scheinert, Alexander Acker |
ICSOC | 4 |
| 2021 | Enel: Context-Aware Dynamic Scaling of Distributed Dataflow Jobs using Graph PropagationabstractDistributed dataflow systems like Spark and Flink enable the use of clusters for scalable data analytics. While runtime prediction models can be used to initially select appropriate cluster resources given target runtimes, the actual runtime performance of dataflow jobs depends on several factors and varies over time. Yet, in many situations, dynamic scaling can be used to meet formulated runtime targets despite significant performance variance.This paper presents Enel, a novel dynamic scaling approach that uses message propagation on an attributed graph to model dataflow jobs and, thus, allows for deriving effective rescaling decisions. For this, Enel incorporates descriptive properties that capture the respective execution context, considers statistics from individual dataflow tasks, and propagates predictions through the job graph to eventually find an optimized new scale-out. Our evaluation of Enel with four iterative Spark jobs shows that our approach is able to identify effective rescaling actions, reacting for instance to node failures, and can be reused across different execution contexts. Dominik Scheinert, Houkun Zhu, Lauritz Thamsen, Morgan Geldenhuys, Jonathan Will, Alexander Acker, Odej Kao |
IPCCC | 6 |
| 2020 | Towards AIOps in Edge Computing EnvironmentsabstractEdge computing was introduced as a technical enabler for the demanding requirements of new network technologies like 5G. It aims to overcome challenges related to centralized cloud computing environments by distributing computational resources to the edge of the network towards the customers. The complexity of the emerging infrastructures increases significantly, together with the ramifications of outages on critical use cases such as self-driving cars or health care. Artificial Intelligence for IT Operations (AIOps) aims to support human operators in managing complex infrastructures by using machine learning methods. This paper describes the system design of an AIOps platform which is applicable in heterogeneous, distributed environments. The overhead of a high-frequency monitoring solution on edge devices is evaluated and performance experiments regarding the applicability of three anomaly detection algorithms on edge devices are conducted. The results show, that it is feasible to collect metrics with a high frequency and simultaneously run specific anomaly detection algorithms directly on edge devices with a reasonable overhead on the resource utilization. Sören Becker 0001, Florian Schmidt 0006, Anton Gulenko, Alexander Acker, Odej Kao |
IEEE BigData | 4 |
| 2020 | Superiority of Simplicity: A Lightweight Model for Network Device Workload PredictionabstractThe rapid growth and distribution of IT systems increases their complexity and aggravates operation and maintenance.To sustain control over large sets of hosts and the connecting networks, monitoring solutions are employed and constantly enhanced.They collect diverse key performance indicators (KPIs) (e.g.CPU utilization, allocated memory, etc.) and provide detailed information about the system state.Predicting the future progress of those KPIs allows ahead of time optimizations like anomaly detection or predictive maintenance and can be defined as a time series forecasting problem.Although, a variety of time series forecasting methods exist, forecasting the progress of IT system KPIs is very hard.First, KPI types like CPU utilization or allocated memory are very different and hard to be modelled by the same model.Second, system components are interconnected and constantly changing due to soft-or firmware updates and hardware modernization.Thus a frequent model retraining or fine-tuning must be expected.Therefore, we propose a lightweight solution for KPI series forecasting.It consists of a weighted heterogeneous ensemble method composed of two models -a neural network and a mean predictor.As ensemble method a weighted summation is used, whereby a heuristic is employed to set the weights.The modelling approach is evaluated on the available FedCSIS 2020 challenge dataset and achieves an overall R 2 score of 0.10 on the preliminary 10% test data and 0.15 on the complete test data.We publish our code on the following github repository: https://github.com/citlab/fed_challenge Alexander Acker, Thorsten Wittkopp, Sasho Nedelkoski, Jasmin Bogatinovski, Odej Kao |
FedCSIS | 1 |
| 2020 | AI-Governance and Levels of Automation for AIOps-supported System AdministrationabstractArtificial Intelligence for IT Operations (AIOps) describes the process of maintaining and operating large IT infrastructures in data centers using AI-supported methods and tools, e.g. for automated anomaly detection, root cause analysis, for remediation, optimization, and for automated initiation of self-stabilizing activities. Initial results and products show that AIOps platforms can help to reach the required level of availability, reliability, dependability, and serviceability for future settings, where latency and response times are of crucial importance. The human operators see the benefits, but also the risks of losing a control over the system while still being accountable for the AIOps-managed infrastructure. While automation is mandatory due to the system complexity and the criticality of a QoS-bounded response, the measures compiled and deployed by the AI-controlled administration are not easily understood or reproducible. Therefore, explainable actions taken by the automated system is becoming a regulatory requirement for future IT infrastructures. In this paper we address several important sub-aspects of the AI-Governance with focus on IT service and infrastructure management and provide a set of rules and levels of automation that precisely describe the shared responsibility between human operators and the AIOps-controlled administration. We aim at providing guidance, decision-support, and explainable processes for AIOps. Anton Gulenko, Alexander Acker, Odej Kao |
ICCCN | 2 |
| 2020 | Self-Attentive Classification-Based Anomaly Detection in Unstructured LogsabstractThe detection of anomalies is an essential data mining task for achieving security and reliability in computer systems. Logs are a common and major data source for anomaly detection methods in almost every computer system. Recent studies have focused predominantly on one-class deep learning methods on manually specified log representations. The main limitation is that these models are not able to learn log representations describing the semantic differences between normal and anomaly logs, leading to a poor generalization on unseen logs. We propose Logsy, a classification-based method to learn log representations that allow to distinguish between normal system log data and anomaly samples from auxiliary log datasets, easily accessible via the internet. The idea behind such an approach to anomaly detection is that the auxiliary dataset is sufficiently informative to enhance the representation of the normal data, yet diverse to regularize against overfitting and improve generalization. We perform several experiments on publicly available datasets to evaluate the performance and properties, where we show improvement of 0.25 in F1 compared to previous methods. Sasho Nedelkoski, Jasmin Bogatinovski, Alexander Acker, Jorge Cardoso 0001, Odej Kao |
ICDM | 3 |
| 2020 | Self-supervised Log Parsing
Sasho Nedelkoski, Jasmin Bogatinovski, Alexander Acker, Jorge Cardoso 0001, Odej Kao |
ECML/PKDD (4) | 3 |
| 2019 | Silent Consensus: Probabilistic Packet Sampling for Lightweight Network Monitoring
Marcel Wallschläger, Alexander Acker, Odej Kao |
ICCSA (1) | 2 |
| 2018 | Detecting Anomalous Behavior of Black-Box Services Modeled with Distance-Based Online ClusteringabstractReliable deployment of services is especially challenging in virtualized infrastructures, where the deep tech-nological stack and the multitude of components necessitate automatic anomaly detection and remediation mechanisms. Traditional monitoring solutions observe the system and generate alarms when the collected metrics exceed predefined thresholds. The fixed thresholds rely on expert knowledge and can lead to numerous false alarms, while abnormal behavior that spans over multiple metrics, components, or system layers, may not be detected. We propose to use an unsupervised online clustering algorithm to create a model of the normal behavior of each monitored component with minimal human interaction and no impact on the monitored system. When an anomaly is detected, a human administrator or automatic remediation system can subsequently revert the component into a normal state. An experimental evaluation resulted in a high accuracy of our approach, indicating that it is suitable for anomaly detection in productive systems. Anton Gulenko, Florian Schmidt 0006, Alexander Acker, Marcel Wallschläger, Odej Kao |
IEEE CLOUD | 3 |
| 2018 | Online Density Grid Pattern Analysis to Classify Anomalies in Cloud and NFV SystemsabstractTechnologies like machine-to-machine communication, autonomous driving or virtual reality applications form an increasingly diverse service landscape. This entails individual and dynamic requirements regarding scalability, availability, latency or throughput from the underlying IT infrastructure. To meet those, telecommunication and network providers started a transformation process towards virtualized technologies like network function virtualization (NFV). However, this drastically increases the infrastructure complexity to a point where more autonomous management is required. In order to meet the reliability of dedicated hardware, virtualized solutions are in demand of autonomous recovery and remediation systems. For critical network systems, actions must be selected very cautiously to not disrupt the operational process. To enable a precise handling, anomaly situations need to be accurately identified based on monitoring data streams. Therefore, we present a supervised machine learning method for an online classification of anomaly states based on similarities between anomaly type-specific density grid patterns. For evaluation, we created an extensive NFV testbed running a virtual implementation of the IP multimedia subsystem. Applying our method to classify various synthetically injected anomaly situations, the results reveal an average overall accuracy of 0.94. Further results also show that the classification model is applicable for identifying previously unknown anomaly situations. Thus, our approach provides a valuable step towards autonomous maintenance of virtualized IT infrastructures. Alexander Acker, Florian Schmidt 0006, Anton Gulenko, Odej Kao |
CloudCom | 1 |
| 2018 | Unsupervised Anomaly Event Detection for VNF Service Monitoring Using Multivariate Online ArimaabstractCloud computing provides companies large scale access to virtual resources, offering cost efficient and flexible usage of digital resources at any time. Thus, companies digitalize their dedicated hardware solutions to virtualized services, which can run in a cloud environment. For example, telecommunication providers move their IP multimedia subsystems, which currently run on dedicated hardware, into the cloud. As the dedicated hardware solutions provided a reliability of 99.999% in the past, the same high reliability is demanded for the virtualized services. But these come with higher complexity due to the fragile computation stack and cannot provide such high requirements. Future zero touch administration systems can help to detect automatically anomalies, find root causes and execute automated remediation actions, providing, providing the opportunity to increase the reliability of the overall system. This work focusses on the detection of degraded state anomalies. We propose an unsupervised detection approach using a multivariate version of the Online Arima forecasting algorithm consuming real-time monitoring data. This approach is evaluated on a testbed running an open source implementation of the IP multimedia subsystem (Clearwater) executed on a replicated Openstack cloud. Results show the applicability of the Online Arima detection approach with high detection rates and low number of false alarms. Florian Schmidt 0006, Florian Suri-Payer, Anton Gulenko, Marcel Wallschläger, Alexander Acker, Odej Kao |
CloudCom | 5 |
| 2018 | IFTM - Unsupervised Anomaly Detection for Virtualized Network Function ServicesabstractTelecommunication system providers move their IP multimedia subsystems to virtualized services in the cloud. For such systems, dedicated hardware solutions provided a reliability of 99.999% in the past. Although virtualization offers more cost efficient usage of such services, it comes with higher complexity for providing reliable running software components due to the fragile computation stack. In order to hide the impact of such problematic behaviors, automatic mechanisms may help to detect degraded state anomalies in order to execute remediation actions. This work introduces IFTM as a framework for unsupervised anomaly detection in a distributed environment based on real-time monitoring data. The proposed approach consists of two key concepts using an automatic identity function and threshold learning to distinguish between normal and abnormal system behaviors. The evaluation is performed on a testbed running an open source implementation of the IP multimedia subsystem (Clearwater) executed on a replicated Openstack cloud environment. Results show the applicability of IFTM with high detection rates (98%) and low number of false alarms. Florian Schmidt 0006, Anton Gulenko, Marcel Wallschläger, Alexander Acker, Vincent Hennig, Odej Kao |
ICWS | 4 |
| 2017 | Patient-individual morphological anomaly detection in multi-lead electrocardiography data streamsabstractCardiac diseases like myocardial infarction, which possibly result in cardiac death, are still a relevant topic. To achieve recognitions in early stages, long term ECG monitoring devices are used. Such devices produce large amounts of data, either directly streamed or stored in databases. Manually analysing this data by experts is inefficient. Thus, automated preprocessing methods are needed to minimize the temporal effort dedicated to the inspection. The proposed method helps to identify morphological anomalies within the ECG data stream. It determines a set of meaningful time series features based on a Kolmogorov-Smirnov test (KST) and after that, applies the BICO online clustering algorithm. Thereby, the system learns the patient-individual PQRST-complex segment morphologies and after that, uses the learned models for detecting anomalies within the ECG data stream. For evaluation, real world patient data was used, which was previously tagged by electrophysiologists. As a result, the KST selected set of features was revealed to be especially suitable for analysing ECG data streams, resulting in average sensitivity rates of 98.82% and average specificity rates of 98.13%. Alexander Acker, Florian Schmidt 0006, Anton Gulenko, Reinhard Kietzmann, Odej Kao |
IEEE BigData | 1 |