EDBT 2026 Demo / reviewers in the wild / expert
Dorothea Wiesmann
dblp:72/3640 · also Dorothea Wiesmann Rothuizen
· DBLP profile ↗
15ranked-venue papers
0as first author
2since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 1 since 2021Databases, data management, data science and information retrieval · 5Systems, architecture and hardware · 3Computer networks · 2 · 1 since 2021Security and privacy · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Trustworthy machine learning · 33% Knowledge representation and reasoning · 33% Language models and text generation · 33% | |
| Databases, data mining, and information retrieval
2 papers |
Data mining · 95% Machine learning and data management · 5% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Storage systems · 98% Emerging computing paradigms · 2% | |
| Software engineering, system software, and programming languages
1 paper |
Services computing and microservices · 100% |
Topics — the 12 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
interpretability |
0.9 | 1 | 2025 | Large Language Models are Interpretable Learners · ICLR 2025 |
Natural language and speech › Language models and text generation
large language model |
0.9 | 1 | 2025 | Large Language Models are Interpretable Learners · ICLR 2025 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
neuro-symbolic reasoning |
0.9 | 1 | 2025 | Large Language Models are Interpretable Learners · ICLR 2025 |
Data mining
clustering |
0.4 | 2 | 2015 | Multi-View Incident Ticket Clustering for Optimal Ticket Dispatching · KDD 2015 Hierarchical Incident Ticket Classification with Minimal Supervision · ICDM 2014 |
Storage systems › storage reliability
disk failure prediction |
0.2 | 1 | 2016 | Predicting Disk Replacement towards Reliable Data Centers · KDD 2016 |
Storage systems
storage reliability |
0.2 | 1 | 2016 | Predicting Disk Replacement towards Reliable Data Centers · KDD 2016 |
Services computing and microservices › service management
IT service management |
0.2 | 1 | 2015 | Multi-View Incident Ticket Clustering for Optimal Ticket Dispatching · KDD 2015 |
Data mining › predictive modeling
classification |
0.2 | 1 | 2014 | Hierarchical Incident Ticket Classification with Minimal Supervision · ICDM 2014 |
Data mining › clustering
hierarchical clustering |
0.2 | 1 | 2014 | Hierarchical Incident Ticket Classification with Minimal Supervision · ICDM 2014 |
Data mining › text mining
text classification |
0.2 | 1 | 2014 | Hierarchical Incident Ticket Classification with Minimal Supervision · ICDM 2014 |
Machine learning and data management
active learning |
0.1 | 1 | 2014 | Hierarchical Incident Ticket Classification with Minimal Supervision · ICDM 2014 |
Data mining › clustering
graph clustering |
0.1 | 1 | 2014 | Hierarchical Incident Ticket Classification with Minimal Supervision · ICDM 2014 |
Methods — techniques the papers use, named apart from their topics
symbolic programs · 0.9prompt tuning · 0.9statistical analysis · 0.2topic modelling · 0.2community finding · 0.2active learning · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Large Language Models are Interpretable LearnersabstractThe trade-off between expressiveness and interpretability remains a core challenge when building human-centric models for classification and decision-making. While symbolic rules offer interpretability, they often lack expressiveness, whereas neural networks excel in performance but are known for being black boxes. This paper shows a combination of Large Language Models (LLMs) and symbolic programs can bridge this gap. In the proposed LLM-based Symbolic Programs (LSPs), the pretrained LLM with natural language prompts provides a massive set of interpretable modules that can transform raw input into natural language concepts. Symbolic programs then integrate these modules into interpretable decision rules. To train LSPs, we develop a divide-and-conquer approach to incrementally build the program from scratch, where the learning process of each step is guided by LLMs. To evaluate the effectiveness of LSPs in extracting interpretable and accurate knowledge from data, we introduce IL-Bench, a collection of diverse tasks, including both synthetic and real-world scenarios across different modalities. Empirical results demonstrate LSP's superior performance compared to traditional neurosymbolic programs and vanilla automatic prompt tuning methods. Moreover, as the knowledge learned by LSP is a combination of natural language descriptions and symbolic rules, it is easily transferable to humans (interpretable), and other LLMs, and generalizes well to out-of-distribution samples. Our code and benchmark will be released for future research. Si Si, Felix X. Yu, Dorothea Wiesmann, Cho-Jui Hsieh, Inderjit S. Dhillon |
ICLR | 4 |
| 2021 | ezNL2SQL: A System for Network Devices Management with a Natural Language Interface for Databases
Jasmina Bogojeska, David Lanyi, Mirela Botezatu, Dorothea Wiesmann |
IM | 4 |
| 2018 | Transfer learning for server behavior classification in small IT environmentsabstractTechnology refresh is an important component in data-center management that needs to be properly justified because of its high cost and associated migration risk. The goal of this paper is to support the technology refresh decision process for small target IT environments with a statistical learning method that automatically identifies and ranks their servers with problematic behavior based on incident ticket and server attribute data. Since the IT environments are heterogeneous, in practice, a separate model is trained for each of them. To address the small sample sizes available for many IT environments, we develop a random forest transfer learning solution that leverages information from large IT environments in a selective manner. It trains a model for each target IT environment that uses properly derived resampling weights such that the distribution of the pool of all examples from the large accounts is matched to the target distribution of the small target IT environment. In this way, a tailored predictive model that uses the information available from many large IT environments provides good quality predictions for small IT environments. We demonstrate the superior prediction quality of our model on a large set of real data. Jasmina Bogojeska, Dorothea Wiesmann |
NOMS | 2 |
| 2017 | "Memory loss" in commodity hardware?: predicting DIMM failures with machine learningabstractFailures of memory modules have been a concern for a long time, as they are costly both in terms of hardware replacement and service disruption. These failures can be preceded by correctable (soft) and then uncorrectable (hard) errors, which accumulate over time. Valuable large scale studies of DIMM errors in the wild [2, 1] analyze in depth hard and soft errors and their correlations with specific sensors. However, little has been reported on how these findings could be used to automatically predict future DIMM failures. We show that by understanding which factors drive such failures, we can build intelligent predictive models with off-the-shelf machine learning techniques to predict DIMM failures ahead of time with high accuracy. Such models not only provide early signs of failures, but also allow administrators to proactively replace DIMMs at risk weeks in advance, thus avoiding "memory loss" of their commodity hardware. Ioana Giurgiu, Dorothea Wiesmann, John Bird |
SYSTOR | 2 |
| 2016 | Predicting Disk Replacement towards Reliable Data CentersabstractDisks are among the most frequently failing components in today's IT environments. Despite a set of defense mechanisms such as RAID, the availability and reliability of the system are still often impacted severely. In this paper, we present a highly accurate SMART-based analysis pipeline that can correctly predict the necessity of a disk replacement even 10-15 days in advance. Our method has been built and evaluated on more than 30000 disks from two major manufacturers, monitored over 17 months. Our approach employs statistical techniques to automatically detect which SMART parameters correlate with disk replacement and uses them to predict the replacement of a disk with even 98% accuracy. Mirela Botezatu, Ioana Giurgiu, Jasmina Bogojeska, Dorothea Wiesmann |
KDD | 4 |
| 2015 | Comprehensible Models for Reconfiguring Enterprise Relational Databases to Avoid IncidentsabstractConfiguring enterprise database management systems is a notoriously hard problem. The combinatorial parameter space makes it intractable to run and observe the DBMS behavior in all scenarios. Thus, the database administrator has the difficult task of choosing DBMS configurations that potentially lead to critical incidents, thus hindering its availability or performance. We propose using machine learning to understand how configuring a DBMS can lead to such high risk incidents. We collect historical data from three IT environments that run both IBM DB2 and Oracle DBMS. Then, we implement several linear and non-linear multivariate models to identify and learn from high risk configurations. We analyze their performance, in terms of accuracy, cost, generalization and interpretability. Results show that high risk configurations can be identified with extremely high accuracy and that the database administrator can potentially benefit from the rules extracted to reconfigure in order to prevent incidents. Ioana Giurgiu, Mirela Botezatu, Dorothea Wiesmann |
CIKM | 3 |
| 2015 | Do you know how to configure your enterprise relational database to reduce incidents?abstractWith the advancement of relational databases, the number of configuration parameters that control memory allocation, concurrency, cost of query plans, I/O optimization, logging, recovery or transaction consistency, increases. Users and even expert database administrators struggle to tune these parameters in order to ensure high availability and performance, and in many cases rely on their experience and some rules of thumb. Research on improving database manageability has shown that this is a critical, but hard problem. In this paper, we propose a highly accurate multivariate statistical model that identifies databases which are bound to raise high volumes of incidents over time. Moreover, we show that by adding detailed configuration parameters to the model, we can better link the problems reported in incident tickets to specific poor database configurations. Finally, we analyze trends of top-ranked parameters and compare their values between problematic and non-problematic databases, in order to suggest better configurations. Ioana Giurgiu, Adela-Diana Almasi, Dorothea Wiesmann |
IM | 3 |
| 2015 | Multi-View Incident Ticket Clustering for Optimal Ticket DispatchingabstractWe present a novel technique that optimizes the dispatching of incident tickets to the agents in an IT Service Support Environment. Unlike the common skill-based dispatching, our approach also takes empirical evidence on the agent's speed from historical data into account. Our solution consists of two parts. First, a novel technique clusters historic tickets into incident categories that are discriminative in terms of agent's performance. Second, a dispatching policy selects, for an incoming ticket, the fastest available agent according to the target cluster. We show that, for ticket data collected from several Service Delivery Units, our new dispatching technique can reduce service time between $35\%$ and $44\%$. Mirela Botezatu, Jasmina Bogojeska, Ioana Giurgiu, Hagen Völzer, Dorothea Wiesmann |
KDD | 5 |
| 2014 | Analysis of Labor Efforts and their Impact Factors to Solve Server Incidents in DatacentersabstractA company's IT infrastructure delivers the basic hardware, networking, operating system, and middleware support to the business' applications. IT service providers perform incident and problem resolution, as well as user administration and change implementation required to maintain the availability and service provided for the business. As a result, they become increasingly challenged with delivering better, faster, and cheaper services to their customers. With the variety of incident tickets reported on a daily basis, understanding where and how much effort is spent to resolve them is critical. Moreover, analyzing the effort data identifies opportunities for self-service and automation, as well as what modernization strategies businesses should implement to reduce incident volumes and, by association, labor effort. In this paper, we conduct a large scale study on the incident and server factors that affect technician effort and quantify their impact. We show that the nature of the incidents and their complexity, the assigned support groups, as well as the underlying OS type play a major role in how much labor effort is spent towards resolving such tickets. Ioana Giurgiu, Jasmina Bogojeska, Sergii Nikolaiev, George Stark, Dorothea Wiesmann |
CCGRID | 5 |
| 2014 | Failure Analysis of Virtual and Physical Machines: Patterns, Causes and CharacteristicsabstractIn today's commercial data centers, the computation density grows continuously as the number of hardware components and workloads in units of virtual machines increase. The service availability guaranteed by data centers heavily depends on the reliability of the physical and virtual servers. In this study, we conduct an analysis on 10K virtual and physical machines hosted on five commercial data centers over an observation period of one year. Our objective is to establish a sound understanding of the differences and similarities between failures of physical and virtual machines. We first capture their failure patterns, i.e., the failure rates, the distributions of times between failures and of repair times, as well as, the time and space dependency of failures. Moreover, we correlate failures with the resource capacity and run-time usage to identify the characteristics of failing servers. Finally, we discuss how virtual machine management actions, i.e., consolidation and on/off frequency, impact virtual machine failures. Robert Birke, Ioana Giurgiu, Lydia Y. Chen, Dorothea Wiesmann, Antonius P. J. Engbersen |
DSN | 4 |
| 2014 | Hierarchical Incident Ticket Classification with Minimal SupervisionabstractIn this paper, we introduce a novel approach for incident ticket classification that aims at minimizing the manual labelling effort while achieving good-quality predictions. To accomplish this, we devise a two-stage technique that employs hierarchical clustering using a combination of graph clustering (community finding) and topic modelling as first stage, followed by either another round of hierarchical clustering or an active learning approach as second stage. We evaluate the performance of our method in terms of manual labelling effort, prediction quality and efficiency on three real-world datasets and demonstrate that classical approaches to text classification are not well suited for incident ticket texts. Andrii Maksai, Jasmina Bogojeska, Dorothea Wiesmann |
ICDM | 3 |
| 2014 | Impact of HW and OS type and currency on server availability derived from problem ticket analysisabstractTechnology refresh is an important component in data center management. The goal of this paper is to assess the impact of HW and OS currency on server availability based on a large set of incident tickets and server attributes data collected from several different IT environments. In order to achieve this we first identify the server failure incidents using a machine learning method for automatic ticket classification. Then we conduct the data analysis to inspect the impact of HW and OS type along with their currency on the rates of server failures. This can further be used to derive guidelines to support the technology refresh decisions in the data centers. Jasmina Bogojeska, Ioana Giurgiu, David Lanyi, George Stark, Dorothea Wiesmann |
NOMS | 5 |
| 2013 | Classifying server behavior and predicting impact of modernization actionsabstractToday the decision of when to modernize which elements of the server HW/SW stack is often done manually based on simple business rules. In this paper we alleviate this problem by supporting the decision process with an automated approach based on incident tickets and server attributes data. As a first step we identify and rank servers with problematic behavior as candidates for modernization using a random forest classifier. Second, this predictive model is used to evaluate the impact of different modernization actions and suggest the most effective ones. We show that our chosen model yields high quality predictions and outperforms traditional linear regression models on a large set of real data. Jasmina Bogojeska, David Lanyi, Ioana Giurgiu, George Stark, Dorothea Wiesmann |
CNSM | 5 |
| 2012 | Change Risk Expert: Leveraging advanced classification and risk management techniques for systematic change failure reductionabstractThis application track paper describes the Change Risk Expert (CRE) tool, which is designed to help reduce change failure rates. CRE assists Change Requesters to adequately plan changes by semi- automatically classifying change tickets, by informing about past failure rates and reasons, by systematically managing change risks, and by providing standard change implementation plans. Sinem Güven, Catalin-Mihai Barbu, Dirk Husemann, Dorothea Wiesmann |
NOMS | 4 |
| 2003 | A Nanotechnology-based Approach to Data Storage
Evangelos Eleftheriou, Peter Bächtold, Giovanni Cherubini, Ajay Dholakia, Christoph Hagleitner, Teddy Loeliger, Angeliki Pantazi, Haralampos Pozidis, T. R. Albrecht, Gerd Karl Binnig, Michel Despont, Ute Drechsler, Urs Dürig, Bernd Gotsmann, Daniel Jubin, Walter Häberle, Mark A. Lantz, Hugo E. Rothuizen, Richard Stutz, Peter Vettiger, Dorothea Wiesmann |
VLDB | 21 |