Georg Krempl

dblp:03/8005 · DBLP profile ↗
← Back
20ranked-venue papers
8as first author
8since 2021 · last 2026
0000-0002-4153-2594ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 7 first-author · 8 since 2021Databases, data management, data science and information retrieval · 9 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 The Window Dilemma: Why Concept Drift Detection is Ill-Posed
Brandon Gower-Winter, Misja Groen, Georg Krempl
IDA3
2025 Identifying Predictions That Influence the Future: Detecting Performative Concept Drift in Data Streams
abstract
Concept Drift has been extensively studied within the context of Stream Learning. However, it is often assumed that the deployed model's predictions play no role in the concept drift the system experiences. Closer inspection reveals that this is not always the case. Automated trading might be prone to self-fulfilling feedback loops. Likewise, malicious entities might adapt to evade detectors in the adversarial setting resulting in a self-negating feedback loop that requires the deployed models to constantly retrain. Such settings where a model may induce concept drift are called performative. In this work, we investigate this phenomenon. Our contributions are as follows: First, we define performative drift within a stream learning setting and distinguish it from other causes of drift. We introduce a novel type of drift detection task, aimed at identifying potential performative concept drift in data streams. We propose a first such performative drift detection approach, called CheckerBoard Performative Drift Detection (CB-PDD). We apply CB-PDD to both synthetic and semi-synthetic datasets that exhibit varying degrees of self-fulfilling feedback loops. Results are positive with CB-PDD showing high efficacy, low false detection rates, resilience to intrinsic drift, comparability to other drift detection techniques, and an ability to effectively detect performative drift in semi-synthetic datasets. Secondly, we highlight the role intrinsic (traditional) drift plays in obfuscating performative drift and discuss the implications of these findings as well as the limitations of CB-PDD.
Brandon Gower-Winter, Georg Krempl, Sergey Dragomiretskiy, Tineke Jelsma, Arno Siebes
AAAI2
2025 Performative Drift Resistant Classification Using Generative Domain Adversarial Networks
Maciej Makowski, Brandon Gower-Winter, Georg Krempl
IDA3
2022 A Stopping Criterion for Transductive Active Learning
abstract
Abstract In transductive active learning, the goal is to determine the correct labels for an unlabeled, known dataset. Therefore, we can either ask an oracle to provide the right label at some cost or use the prediction of a classifier which we train on the labels acquired so far. In contrast, the commonly used (inductive) active learning aims to select instances for labeling out of the unlabeled set to create a generalized classifier, which will be deployed on unknown data. This article formally defines the transductive setting and shows that it requires new solutions. Additionally, we formalize the theoretically cost-optimal stopping point for the transductive scenario. Building upon the probabilistic active learning framework, we propose a new transductive selection strategy that includes a stopping criterion and show its superiority.
Daniel Kottke, Christoph Sandrock, Georg Krempl, Bernhard Sick
ECML/PKDD (4)3
2022 Stream-based active learning for sliding windows under the influence of verification latency
abstract
Abstract Stream-based active learning (AL) strategies minimize the labeling effort by querying labels that improve the classifier’s performance the most. So far, these strategies neglect the fact that an oracle or expert requires time to provide a queried label. We show that existing AL methods deteriorate or even fail under the influence of such verification latency. The problem with these methods is that they estimate a label’s utility on the currently available labeled data. However, when this label would arrive, some of the current data may have gotten outdated and new labels have arrived. In this article, we propose to simulate the available data at the time when the label would arrive. Therefore, our method Forgetting and Simulating (FS) forgets outdated information and simulates the delayed labels to get more realistic utility estimates. We assume to know the label’s arrival date a priori and the classifier’s training data to be bounded by a sliding window. Our extensive experiments show that FS improves stream-based AL strategies in settings with both, constant and variable verification latency.
Daniel Kottke, Georg Krempl, Bernhard Sick
Mach. Learn.3
2021 Statistical Analysis of Pairwise Connectivity
Georg Krempl, Daniel Kottke
DS1
2021 Active Selection of Classification Features
Thomas T. Kok, Rachel M. Brouwer, René C. W. Mandl, Hugo G. Schnack, Georg Krempl
IDA5
2021 Toward optimal probabilistic active learning using a Bayesian approach
abstract
Abstract Gathering labeled data to train well-performing machine learning models is one of the critical challenges in many applications. Active learning aims at reducing the labeling costs by an efficient and effective allocation of costly labeling resources. In this article, we propose a decision-theoretic selection strategy that (1) directly optimizes the gain in misclassification error, and (2) uses a Bayesian approach by introducing a conjugate prior distribution to determine the class posterior to deal with uncertainties. By reformulating existing selection strategies within our proposed model, we can explain which aspects are not covered in current state-of-the-art and why this leads to the superior performance of our approach. Extensive experiments on a large variety of datasets and different kernels validate our claims.
Daniel Kottke, Marek Herde, Christoph Sandrock, Denis Huseljic, Georg Krempl, Bernhard Sick
Mach. Learn.5
2020 Constructing and predicting school advice for academic achievement: a comparison of item response theory and machine learning techniques
abstract
Educational tests can be used to estimate pupils' abilities and thereby give an indication of whether their school type is suitable for them. However, tests in education are usually conducted for each content area separately which makes it difficult to combine these results into one single school advice. To help with school advice, we provide a comparison between both domain-specific and domain-agnostic methods for predicting school types. Both use data from a pupil monitoring system in the Netherlands, a system that keeps track of pupils' educational progress over several years by a series of tests measuring multiple skills.
Koen Niemeijer, Remco Feskens, Georg Krempl, Jesse Koops, Matthieu J. S. Brinkhuis
LAK3
2019 Temporal density extrapolation using a dynamic basis approach
abstract
Density estimation is a versatile technique underlying many data mining tasks and techniques, ranging from exploration and presentation of static data, to probabilistic classification, or identifying changes or irregularities in streaming data. With the pervasiveness of embedded systems and digitisation, this latter type of streaming and evolving data becomes more important. Nevertheless, research in density estimation has so far focused on stationary data, leaving the task of of extrapolating and predicting density at time points outside a training window an open problem. For this task, temporal density extrapolation (TDX) is proposed. This novel method models and predicts gradual monotonous changes in a distribution. It is based on the expansion of basis functions, whose weights are modelled as functions of compositional data over time by using an isometric log-ratio transformation. Extrapolated density estimates are then obtained by extrapolating the weights to the requested time point, and querying the density from the basis functions with back-transformed weights. Our approach aims for broad applicability by neither being restricted to a specific parametric distribution, nor relying on cluster structure in the data. It requires only two additional extrapolation-specific parameters, for which reasonable defaults exist. Experimental evaluation on various data streams, synthetic as well as from the real-world domains of credit scoring and environmental health, shows that the model manages to capture monotonous drift patterns accurately and better than existing methods. Thereby, it requires not more than 1.5 times the run time of a corresponding static density estimation approach.
Georg Krempl, Dominik Lang, Vera Hofer
Data Min. Knowl. Discov.1
2016 Multi-Class Probabilistic Active Learning
abstract
This work addresses active learning for multi-class classification. Active learning algorithms optimize classifier performance by successively selecting the most beneficial instances from a pool of unlabeled instances to be labeled by an oracle. In this work, we study the influence of the following factors for active learning: (1) an instance's impact, (2) its posterior, and (3) the reliability of this posterior. To do so, we propose a new decision-theoretic approach, called multi-class probabilistic active learning (McPAL). Building on a probabilistic active learning framework, our approach is non-myopic, fast, and optimizes a performance measure (like accuracy) directly. Considering all influence factors, McPAL determines the expected gain in performance to compare the usefulness of instances. For this purpose, it calculates the density weighted expectation over the true posterior and over all possible labeling combinations in a closed-form solution. Thus, in contrast to other multi-class algorithms, it considers the posterior's reliability which improved the performance. In our experimental evaluation, we show that the combination of the selected influence factors works best and that McPAL is superior in comparison to various other multi-class active learning algorithms on six datasets.
Daniel Kottke, Georg Krempl, Dominik Lang, Johannes Teschner, Myra Spiliopoulou
ECAI2
2015 Clustering-Based Optimised Probabilistic Active Learning (COPAL)
Georg Krempl, Tuan Cuong Ha, Myra Spiliopoulou
Discovery Science1
2015 Probabilistic Active Learning in Datastreams
Daniel Kottke, Georg Krempl, Myra Spiliopoulou
IDA2
2015 Optimised probabilistic active learning (OPAL) - For fast, non-myopic, cost-sensitive active classification
Georg Krempl, Daniel Kottke, Vincent Lemaire 0001
Mach. Learn.1
2014 Probabilistic Active Learning: Towards Combining Versatility, Optimality and Efficiency
Georg Krempl, Daniel Kottke, Myra Spiliopoulou
Discovery Science1
2014 Probabilistic Active Learning: A Short Proposition
abstract
Active Mining of Big Data requires fast approaches that ideally select for a user-specified performance measure and arbitrary classifier the optimal instance for improving the classification performance. Existing generic approaches are either slow, like error reduction, or heuristics, like uncertainty sampling. We propose a novel, fast yet versatile approach that directly optimises any user-specified performance measure: Probabilistic Active Learning (PAL).
Georg Krempl, Daniel Kottke, Myra Spiliopoulou
ECAI1
2013 Correcting the Usage of the Hoeffding Inequality in Stream Mining
Pawel Matuszyk, Georg Krempl, Myra Spiliopoulou
IDA2
2012 A hierarchical tree layout algorithm with an application to corporate management in a change process
Vera Hofer, Georg Krempl
Expert Syst. Appl.2
2011 The Algorithm APT to Classify in Concurrence of Latency and Drift
Georg Krempl
IDA1
2011 Online Clustering of High-Dimensional Trajectories under Concept Drift
Georg Krempl, Zaigham Faraz Siddiqui, Myra Spiliopoulou
ECML/PKDD (2)1