EDBT 2026 Demo / reviewers in the wild / expert
Thorsten Wittkopp
dblp:270/0285
· DBLP profile ↗
8ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0001-5154-7813ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Adaptive LLM Routing for Scientific Workflows: Predicting Query Complexity to Optimize CostabstractScientific workflows increasingly incorporate large-language-model (LLM) services for tasks such as metadata curation, code generation, and interactive analysis. However, inference cost in terms of resource usage varies widely with query complexity, both for service-based or self-hosted setups. Current “always-XL” deployments tend to over-provision: they safeguard response quality but squander compute usage. We introduce a routing framework based on query complexity estimations that trains a lightweight classifier on past multi-model performance data to predict the smallest model likely to satisfy a target quality threshold and then routes queries accordingly. Evaluated on$\text{4 0, 0 0 4}$Mix-Instruct prompts, our method matches an XL-only baseline ($+0.27 \%$quality) while reducing floating-point operations by 12.9 %, achieving competitive performance at a fraction of the cost. The method thus enables scalable, complexity-aware routing for LLM services within data- and compute-intensive scientific workflows. Jonas Thierfeldt, Dominik Scheinert, Thorsten Wittkopp, Odej Kao |
IPCCC | 3 |
| 2024 | LogRCA: Log-Based Root Cause Analysis for Distributed Services
Thorsten Wittkopp, Philipp Wiesner, Odej Kao |
Euro-Par (2) | 1 |
| 2023 | Karasu: A Collaborative Approach to Efficient Cluster Configuration for Big Data AnalyticsabstractSelecting the right resources for big data analytics jobs is hard because of the wide variety of configuration options like machine type and cluster size. As poor choices can have a significant impact on resource efficiency, cost, and energy usage, automated approaches are gaining popularity. Most existing methods rely on profiling recurring workloads to find near-optimal solutions over time. Due to the cold-start problem, this often leads to lengthy and costly profiling phases. However, big data analytics jobs across users can share many common properties: they often operate on similar infrastructure, using similar algorithms implemented in similar frameworks. The potential in sharing aggregated profiling runs to collaboratively address the cold start problem is largely unexplored. We present Karasu, an approach to more efficient resource configuration profiling that promotes data sharing among users working with similar infrastructures, frameworks, algorithms, or datasets. Karasu trains lightweight performance models using aggregated runtime information of collaborators and combines them into an ensemble method to exploit inherent knowledge of the configuration search space. Moreover, Karasu allows the optimization of multiple objectives simultaneously. Our evaluation is based on performance data from diverse workload executions in a public cloud environment. We show that Karasu is able to significantly boost existing methods in terms of performance, search time, and cost, even when few comparable profiling runs are available that share only partial common characteristics with the target job. Dominik Scheinert, Philipp Wiesner, Thorsten Wittkopp, Lauritz Thamsen, Jonathan Will, Odej Kao |
IPCCC | 3 |
| 2022 | Cucumber: Renewable-Aware Admission Control for Delay-Tolerant Cloud and Edge Workloads
Philipp Wiesner, Dominik Scheinert, Thorsten Wittkopp, Lauritz Thamsen, Odej Kao |
Euro-Par | 3 |
| 2021 | On the Potential of Execution Traces for Batch Processing Workload Optimization in Public CloudsabstractWith the growing amount of data, data processing workloads and the management of their resource usage becomes increasingly important. Since managing a dedicated infrastructure is in many situations infeasible or uneconomical, users progressively execute their respective workloads in the cloud. As the configuration of workloads and resources is often challenging, various methods have been proposed that either quickly profile towards a good configuration or determine one based on data from previous runs. Still, performance data to train such methods is often lacking and must be costly collected.In this paper, we propose a collaborative approach for sharing anonymized workload execution traces among users, mining them for general patterns, and exploiting clusters of historical workloads for future optimizations. We evaluate our prototype implementation for mining workload execution graphs on a publicly available trace dataset and demonstrate the predictive value of workload clusters determined through traces only. Dominik Scheinert, Alireza Alamgiralem, Jonathan Bader, Jonathan Will, Thorsten Wittkopp, Lauritz Thamsen |
IEEE BigData | 5 |
| 2021 | Bellamy: Reusing Performance Models for Distributed Dataflow Jobs Across ContextsabstractDistributed dataflow systems enable the use of clusters for scalable data analytics. However, selecting appropriate cluster resources for a processing job is often not straightforward. Performance models trained on historical executions of a concrete job are helpful in such situations, yet they are usually bound to a specific job execution context (e.g. node type, software versions, job parameters) due to the few considered input parameters. Even in case of slight context changes, such supportive models need to be retrained and cannot benefit from historical execution data from related contexts.This paper presents Bellamy, a novel modeling approach that combines scale-outs, dataset sizes, and runtimes with additional descriptive properties of a dataflow job. It is thereby able to capture the context of a job execution. Moreover, Bellamy is realizing a two-step modeling approach. First, a general model is trained on all the available data for a specific scalable analytics algorithm, hereby incorporating data from different contexts. Subsequently, the general model is optimized for the specific situation at hand, based on the available data for the concrete context. We evaluate our approach on two publicly available datasets consisting of execution data from various dataflow jobs carried out in different environments, showing that Bellamy outperforms state-of-the-art methods. Dominik Scheinert, Lauritz Thamsen, Houkun Zhu, Jonathan Will, Alexander Acker, Thorsten Wittkopp, Odej Kao |
CLUSTER | 6 |
| 2021 | LogLAB: Attention-Based Labeling of Log Data Anomalies via Weak Supervision
Thorsten Wittkopp, Philipp Wiesner, Dominik Scheinert, Alexander Acker |
ICSOC | 1 |
| 2020 | Superiority of Simplicity: A Lightweight Model for Network Device Workload PredictionabstractThe rapid growth and distribution of IT systems increases their complexity and aggravates operation and maintenance.To sustain control over large sets of hosts and the connecting networks, monitoring solutions are employed and constantly enhanced.They collect diverse key performance indicators (KPIs) (e.g.CPU utilization, allocated memory, etc.) and provide detailed information about the system state.Predicting the future progress of those KPIs allows ahead of time optimizations like anomaly detection or predictive maintenance and can be defined as a time series forecasting problem.Although, a variety of time series forecasting methods exist, forecasting the progress of IT system KPIs is very hard.First, KPI types like CPU utilization or allocated memory are very different and hard to be modelled by the same model.Second, system components are interconnected and constantly changing due to soft-or firmware updates and hardware modernization.Thus a frequent model retraining or fine-tuning must be expected.Therefore, we propose a lightweight solution for KPI series forecasting.It consists of a weighted heterogeneous ensemble method composed of two models -a neural network and a mean predictor.As ensemble method a weighted summation is used, whereby a heuristic is employed to set the weights.The modelling approach is evaluated on the available FedCSIS 2020 challenge dataset and achieves an overall R 2 score of 0.10 on the preliminary 10% test data and 0.15 on the complete test data.We publish our code on the following github repository: https://github.com/citlab/fed_challenge Alexander Acker, Thorsten Wittkopp, Sasho Nedelkoski, Jasmin Bogatinovski, Odej Kao |
FedCSIS | 2 |