EDBT 2026 Demo / reviewers in the wild / expert
Wouter Duivesteijn
dblp:23/1341
· DBLP profile ↗
31ranked-venue papers in the field
9as first author
11since 2021 · last 2026
0000-0003-0412-8864ORCID · verified
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 29 (9 first)Database Systems & Data Management · 1Knowledge Engineering, Semantic Web & Information Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exceptional Model Residual Mining, and Three Richer EMM Description Languages
Aniket Mishra, Cristiana Carbunaru, Wouter Duivesteijn |
IDA | 3 |
| 2025 | Beyond Discriminant Patterns: On the Robustness of Decision Rule EnsemblesabstractLocal decision rules are highly regarded for their interpretability, offering insights into granular patterns that are critical for explainable machine learning. While existing methods emphasize the identification of discriminative patterns to achieve high predictive accuracy, they often fail to account for robustness against distributional shifts that occur during deployment. This paper addresses this gap by proposing a novel approach to learning and ensembling local decision rules that are inherently robust across diverse training and deployment environments. Our method leverages causal inference principles, viewing distributional shifts as interventions on the underlying system. We incorporate two regularization techniques: graph-based regularization, which decomposes invariant features using causal graphs, and variance-based regularization, which promotes stability by introducing artificial features to guide decision boundaries. These techniques enable the generation of decision rules that excel in predictive power while maintaining stability under changing environmental conditions. Extensive experiments on synthetic and benchmark datasets validate the effectiveness of the proposed method. The results demonstrate significant improvements in robustness, outperforming traditional boosting ensembles when subjected to diverse and challenging environments. Quantitative and qualitative analyses further highlight how the integration of causal knowledge and adaptive regularization encourages the utilization of invariant features, leading to better generalization. This work emphasizes the importance of causal reasoning in the design of machine learning models, paving the way for future research into robust, interpretable, and reliable decision-making frameworks for real-world applications. Xin Du 0006, Subramanian Ramamoorthy, Wouter Duivesteijn, Mykola Pechenizkiy |
ICDM | 3 |
| 2025 | Local Subgroup Discovery on Attributed Network Graphs
Carl Vico Heinrich, Tommie Lombarts, Jules Mallens, Luc Tortike, David Wolf, Wouter Duivesteijn |
IDA | 6 |
| 2025 | Characterizing the Risk of Atrial Fibrillation in Cardiac Patients with Exceptional Electrocardiogram PhenotypesabstractWe provide a transparent method to characterize Atrial Fibrillation (AF) caused by cardiac surgery, using Electrocardiogram (ECG) phenotypes. Current practice in the hospital is reactive rather than preventive and is based on a third party's proprietary alarms on vitals. Assistance with detection and prediction methods often lacks sufficient insights into their decisions toward the users, i.e., the hospital workers. This aspect is necessary to gain the trust of medical workers and patients in the decisions that are made. Our objective of transparently identifying risk factors for AF helps experts increase their understanding of the problem, and assists in decision-making about administering preventive medication to risk groups. With the deployment of the Exceptional Model Mining (EMM) framework on AF-related ECG phenotypes, we introduce a transparent and actionable method that assists the hospital in preventive treatment. We find several subgroups with EMM that align with known risk factors in the existing literature, confirming the ability of our method to identify risk groups of AF successfully. In addition, new hypotheses on found characteristics and combinations thereof have originated from the deployment. The hospital is advised to administer preventive medications to patients who match the descriptions of the risk groups found and perform follow-up clinical studies to validate the found hypotheses. Lieke van den Biggelaar, Rianne Margaretha Schouten, Ashley De Bie, R. Arthur Bouwman, Wouter Duivesteijn |
KDD (2) | 5 |
| 2025 | Conformalized Exceptional Model Mining: Telling Where Your Model Performs (Not) Well
Xin Du 0006, Sikun Yang, Wouter Duivesteijn, Mykola Pechenizkiy |
ECML/PKDD (3) | 3 |
| 2024 | RMI-RRG: A Soft Protocol to Postulate Monotonicity Constraints for Tabular Datasets
Iko Vloothuis, Wouter Duivesteijn |
IDA (1) | 2 |
| 2024 | Exceptional Subitizing Patterns: Exploring Mathematical Abilities of Finnish Primary School Children with Piecewise Linear Regression
Rianne Margaretha Schouten, Wouter Duivesteijn, Pekka Räsänen, Jacob M. Paul, Mykola Pechenizkiy |
ECML/PKDD (10) | 2 |
| 2022 | Efficient Subgroup Discovery Through Auto-Encoding
Joost F. van der Haar, Sander C. Nagelkerken, Igor G. Smit, Kjell van Straaten, Janneke A. Tack, Rianne Margaretha Schouten, Wouter Duivesteijn |
IDA | 7 |
| 2022 | Exceptional Model Mining for Repeated Cross-Sectional Data (EMM-RCS)abstractRepeated Cross-Sectional (RCS) data measures a phenomenon by repeatedly sampling new cases from a population at successive measurement moments. It allows for analyzing societal trends without the need to follow individuals. To gain a deeper understanding of these trends, we propose EMM-RCS, an Exceptional Model Mining instance designed to find subgroups displaying exceptional trend behavior in RCS data. We build quality measures on the standard error, finding various types of exceptionalities within trends (exceptional flattening, slope, deviation from the norm). Additionally, EMM-RCS can handle practical RCS data problems, including uneven spacing of measurements over time, fluctuating sample sizes, and missing data. Rianne Margaretha Schouten, Wouter Duivesteijn, Mykola Pechenizkiy |
SDM | 2 |
| 2022 | Mining sequences with exceptional transition behaviour of varying order using quality measures based on information-theoretic scoring functionsabstractAbstract Discrete Markov chains are frequently used to analyse transition behaviour in sequential data. Here, the transition probabilities can be estimated using varying order Markov chains, where order k specifies the length of the sequence history that is used to model these probabilities. Generally, such a model is fitted to the entire dataset, but in practice it is likely that some heterogeneity in the data exists and that some sequences would be better modelled with alternative parameter values, or with a Markov chain of a different order. We use the framework of Exceptional Model Mining (EMM) to discover these exceptionally behaving sequences. In particular, we propose an EMM model class that allows for discovering subgroups with transition behaviour of varying order. To that end, we propose three new quality measures based on information-theoretic scoring functions. Our findings from controlled experiments show that all three quality measures find exceptional transition behaviour of varying order and are reasonably sensitive. The quality measure based on Akaike’s Information Criterion is most robust for the number of observations. We furthermore add to existing work by seeking for subgroups of sequences, as opposite to subgroups of transitions. Since we use sequence-level descriptive attributes, we form subgroups of entire sequences, which is practically relevant in situations where you want to identify the originators of exceptional sequences, such as patients. We show this relevance by analysing sequences of blood glucose values of adult persons with diabetes type 2. In the experiments, we find subgroups of patients based on age and glycated haemoglobin (HbA1c), a measure known to correlate with average blood glucose values. Clinicians and domain experts confirmed the transition behaviour as estimated by the fitted Markov chain models. Rianne Margaretha Schouten, Marcos L. P. Bueno, Wouter Duivesteijn, Mykola Pechenizkiy |
Data Min. Knowl. Discov. | 3 |
| 2021 | Adversarial balancing-based representation learning for causal effect inference with observational dataabstractAbstract Learning causal effects from observational data greatly benefits a variety of domains such as health care, education, and sociology. For instance, one could estimate the impact of a new drug on specific individuals to assist clinical planning and improve the survival rate. In this paper, we focus on studying the problem of estimating the Conditional Average Treatment Effect (CATE) from observational data. The challenges for this problem are two-fold: on the one hand, we have to derive a causal estimator to estimate the causal quantity from observational data, in the presence of confounding bias; on the other hand, we have to deal with the identification of the CATE when the distributions of covariates over the treatment group units and the control units are imbalanced. To overcome these challenges, we propose a neural network framework called Adversarial Balancing-based representation learning for Causal Effect Inference (ABCEI), based on recent advances in representation learning. To ensure the identification of the CATE, ABCEI uses adversarial learning to balance the distributions of covariates in the treatment and the control group in the latent representation space, without any assumptions on the form of the treatment selection/assignment function. In addition, during the representation learning and balancing process, highly predictive information from the original covariate space might be lost. ABCEI can tackle this information loss problem by preserving useful information for predicting causal effects under the regularization of a mutual information estimator. The experimental results show that ABCEI is robust against treatment selection bias, and matches/outperforms the state-of-the-art approaches. Our experiments show promising results on several datasets, encompassing several health care (and other) domains. Xin Du 0006, Wouter Duivesteijn, Alexander G. Nikolaev, Mykola Pechenizkiy |
Data Min. Knowl. Discov. | 3 |
| 2020 | Predicting Remaining Useful Life with Similarity-Based PriorsabstractPrognostics is the area of research that is concerned with predicting the remaining useful life of machines and machine parts. The remaining useful life is the time during which a machine or part can be used, before it must be replaced or repaired. To create accurate predictions, predictive techniques must take external data into account on the operating conditions of the part and events that occurred during its lifetime. However, such data is often not available. Similarity-based techniques can help in such cases. They are based on the hypothesis that if a curve developed similarly to other curves up to a point, it will probably continue to do so. This paper presents a novel technique for similarity-based remaining useful life prediction. In particular, it combines Bayesian updating with priors that are based on similarity estimation. The paper shows that this technique outperforms other techniques on long-term predictions by a large margin, although other techniques still perform better on short-term predictions. Youri Soons, Remco M. Dijkman, Maurice Jilderda, Wouter Duivesteijn |
IDA | 4 |
| 2020 | Exceptional spatio-temporal behavior mining through Bayesian non-parametric modelingabstractAbstract Collective social media provides a vast amount of geo-tagged social posts, which contain various records on spatio-temporal behavior. Modeling spatio-temporal behavior on collective social media is an important task for applications like tourism recommendation, location prediction and urban planning. Properly accomplishing this task requires a model that allows for diverse behavioral patterns on each of the three aspects: spatial location, time, and text. In this paper, we address the following question: how to find representative subgroups of social posts, for which the spatio-temporal behavioral patterns are substantially different from the behavioral patterns in the whole dataset? Selection and evaluation are the two challenging problems for finding the exceptional subgroups. To address these problems, we propose BNPM: a Bayesian non-parametric model, to model spatio-temporal behavior and infer the exceptionality of social posts in subgroups. By training BNPM on a large amount of randomly sampled subgroups, we can get the global distribution of behavioral patterns. For each given subgroup of social posts, its posterior distribution can be inferred by BNPM. By comparing the posterior distribution with the global distribution, we can quantify the exceptionality of each given subgroup. The exceptionality scores are used to guide the search process within the exceptional model mining framework to automatically discover the exceptional subgroups. Various experiments are conducted to evaluate the effectiveness and efficiency of our method. On four real-world datasets our method discovers subgroups coinciding with events, subgroups distinguishing professionals from tourists, and subgroups whose consistent exceptionality can only be truly appreciated by combining exceptional spatio-temporal and exceptional textual behavior. Xin Du 0006, Yulong Pei, Wouter Duivesteijn, Mykola Pechenizkiy |
Data Min. Knowl. Discov. | 3 |
| 2019 | DEvIANT: Discovering Significant Exceptional (Dis-)Agreement Within Groups
Adnene Belfodil, Wouter Duivesteijn, Marc Plantevit, Sylvie Cazalens, Philippe Lamarre |
ECML/PKDD (1) | 2 |
| 2019 | k Is the Magic Number - Inferring the Number of Clusters Through Nonparametric Concentration Inequalities
Sibylle Hess, Wouter Duivesteijn |
ECML/PKDD (1) | 2 |
| 2018 | Subjectively Interesting Subgroup Discovery on Real-Valued TargetsabstractDeriving insights from high-dimensional data is one of the core problems in data mining. The difficulty mainly stems from the large number of variable combinations to potentially consider. Hence, an obvious question is whether we can automate the search for interesting patterns. Here, we consider the setting where a user wants to learn as efficiently as possible about real-valued attributes. We introduce a method to find subgroups in the data that are maximally informative (in the Information Theoretic sense) with respect to one or more real-valued target attributes. The succinct subgroup descriptions are in terms of arbitrarily-typed description attributes. The approach is based on the Subjective Interestingness framework FORSIED to use prior knowledge when mining most informative patterns. Jefrey Lijffijt, Bo Kang, Wouter Duivesteijn, Kai Puolamäki, Emilia Oikarinen, Tijl De Bie |
ICDE | 3 |
| 2017 | Have It Both Ways - From A/B Testing to A&B Testing with Exceptional Model Mining
Wouter Duivesteijn, Tara Farzami, Thijs Putman, Evertjan Peer, Hilde J. P. Weerts, Jasper N. Adegeest, Gerson Foks, Mykola Pechenizkiy |
ECML/PKDD (3) | 1 |
| 2017 | Exceptionally monotone models - the rank correlation model class for Exceptional Model Mining
Lennart Downar, Wouter Duivesteijn |
Knowl. Inf. Syst. | 2 |
| 2016 | Exceptional Model Mining - Supervised descriptive local pattern mining with complex target concepts
Wouter Duivesteijn, A. J. Feelders, Arno J. Knobbe |
Data Min. Knowl. Discov. | 1 |
| 2015 | Exceptionally Monotone Models - The Rank Correlation Model Class for Exceptional Model MiningabstractExceptional Model Mining strives to find coherent subgroups of the dataset where multiple target attributes interact in an unusual way. One instance of such an investigated form of interaction is Pearson's correlation coefficient between two targets. EMM then finds subgroups with an exceptionally linear relation between the targets. In this paper, we enrich the EMM toolbox by developing the more general rank correlation model class. We find subgroups with an exceptionally monotone relation between the targets. Apart from catering for this richer set of relations, the rank correlation model class does not necessarily require the assumption of target normality, which is implicitly invoked in the Pearson's correlation model class. Furthermore, it is less sensitive to outliers. Lennart Downar, Wouter Duivesteijn |
ICDM | 2 |
| 2015 | Understanding Where Your Classifier Does (Not) Work
Wouter Duivesteijn, Julia Thaele |
ECML/PKDD (3) | 1 |
| 2015 | Cost-based quality measures in subgroup discovery
Rob M. Konijn, Wouter Duivesteijn, Marvin Meeng, Arno J. Knobbe |
J. Intell. Inf. Syst. | 2 |
| 2014 | Understanding Where Your Classifier Does (Not) Work - The SCaPE Model Class for EMMabstractFACT, the First G-APD Cherenkov Telescope, detects air showers induced by high-energetic cosmic particles. It is desirable to classify a shower as being induced by a gamma ray or a background particle. Generally, it is nontrivial to get any feedback on the real-life training task, but we can attempt to understand how our classifier works by investigating its performance on Monte Carlo simulated data. To this end, in this paper we develop the SCaPE (Soft Classifier Performance Evaluation) model class for Exceptional Model Mining, which is a Local Pattern Mining framework devoted to highlighting unusual interplay between multiple targets. In our Monte Carlo simulated data, we take as targets the computed classifier probabilities and the binary column containing the ground truth: which kind of particle induced the corresponding shower. Using a newly developed quality measure based on ranking loss, the SCaPE model class highlights subspaces of the search space where the classifier performs particularly well or poorly. These subspaces arrive in terms of conditions on attributes of the data, hence they come in a language a domain expert understands, which should aid him in understanding where his/her classifier does (not) work. Found subgroups highlight subspaces whose difficulty for classification is corroborated by astrophysical interpretation, as well as subspaces that warrant further investigation. Wouter Duivesteijn, Julia Thaele |
ICDM | 1 |
| 2014 | RealKrimp - Finding Hyperintervals that Compress with MDL for Real-Valued Data
Jouke Witteveen, Wouter Duivesteijn, Arno J. Knobbe, Peter Grünwald |
IDA | 2 |
| 2014 | ROCsearch - An ROC-guided Search Strategy for Subgroup DiscoveryabstractSubgroup Discovery (SD) aims to find coherent, easy-to-interpret subsets of the dataset at hand, where something exceptional is going on. Since the resulting subgroups are defined in terms of conditions on attributes of the dataset, this data mining task is ideally suited to be used by non-expert analysts. The typical SD approach uses a heuristic beam search, involving parameters that strongly influence the outcome. Unfortunately, these parameters are often hard to set properly for someone who is not a data mining expert; correct settings depend on properties of the dataset, and on the resulting search landscape. To remove this potential obstacle for casual SD users, we introduce ROCsearch, a new ROC-based beam search variant for Subgroup Discovery. On each search level of the beam search, ROCsearch analyzes the intermediate results in ROC space to automatically determine a sensible search width for the next search level. Thus, beam search parameter setting is taken out of the domain expert's hands, lowering the threshold for using Subgroup Discovery. Also, ROCsearch automatically adapts its search behavior to the properties and resulting search landscape of the dataset at hand. Aside form these advantages, we also show that ROCsearch is an order of magnitude more efficient than traditional beam search, while its results are equivalent and on large datasets even better than traditional beam search results. Marvin Meeng, Wouter Duivesteijn, Arno J. Knobbe |
SDM | 2 |
| 2013 | Discovering Local Subgroups, with an Application to Fraud Detection
Rob M. Konijn, Wouter Duivesteijn, Wojtek Kowalczyk, Arno J. Knobbe |
PAKDD (1) | 2 |
| 2012 | Multi-label LeGo - Enhancing Multi-label Classifiers with Local Patterns
Wouter Duivesteijn, Eneldo Loza Mencía, Johannes Fürnkranz, Arno J. Knobbe |
IDA | 1 |
| 2012 | Different slopes for different folks: mining for exceptional regression models with cook's distanceabstractExceptional Model Mining (EMM) is an exploratory data analysis technique that can be regarded as a generalization of subgroup discovery. In EMM we look for subgroups of the data for which a model fitted to the subgroup differs substantially from the same model fitted to the entire dataset. In this paper we develop methods to mine for exceptional regression models. We propose a measure for the exceptionality of regression models (Cook's distance), and explore the possibilities to avoid having to fit the regression model to each candidate subgroup. The algorithm is evaluated on a number of real life datasets. These datasets are also used to illustrate the results of the algorithm. We find interesting subgroups with deviating models on datasets from several different domains. We also show that under certain circumstances one can forego fitting regression models on up to 40% of the subgroups, and these 40% are the relatively expensive regression models to compute. Wouter Duivesteijn, A. J. Feelders, Arno J. Knobbe |
KDD | 1 |
| 2011 | Exploiting False Discoveries - Statistical Validation of Patterns and Quality Measures in Subgroup DiscoveryabstractSubgroup discovery suffers from the multiple comparisons problem: we search through a large space, hence whenever we report a set of discoveries, this set will generally contain false discoveries. We propose a method to compare subgroups found through subgroup discovery with a statistical model we build for these false discoveries. We determine how much the subgroups we find deviate from the model, and hence statistically validate the found subgroups. Furthermore we propose to use this subgroup validation to objectively compare quality measures used in subgroup discovery, by determining how much the top subgroups we find with each measure deviate from the statistical model generated with that measure. We thus aim to determine how good individual measures are in selecting significant findings. We invoke our method to experimentally compare popular quality measures in several subgroup discovery settings. Wouter Duivesteijn, Arno J. Knobbe |
ICDM | 1 |
| 2010 | Subgroup Discovery Meets Bayesian Networks -- An Exceptional Model Mining ApproachabstractWhenever a dataset has multiple discrete target variables, we want our algorithms to consider not only the variables themselves, but also the interdependencies between them. We propose to use these interdependencies to quantify the quality of subgroups, by integrating Bayesian networks with the Exceptional Model Mining framework. Within this framework, candidate subgroups are generated. For each candidate, we fit a Bayesian network on the target variables. Then we compare the network's structure to the structure of the Bayesian network fitted on the whole dataset. To perform this comparison, we define an edit distance-based distance metric that is appropriate for Bayesian networks. We show interesting subgroups that we experimentally found with our method on datasets from music theory, semantic scene classification, biology and zoogeography. Wouter Duivesteijn, Arno J. Knobbe, A. J. Feelders, Matthijs van Leeuwen |
ICDM | 1 |
| 2008 | Nearest Neighbour Classification with Monotonicity Constraints
Wouter Duivesteijn, A. J. Feelders |
ECML/PKDD (1) | 1 |