EDBT 2026 Demo / reviewers in the wild / expert
Siegfried Nijssen
dblp:71/709
· DBLP profile ↗
38ranked-venue papers in the field
7as first author
9since 2021 · last 2025
0000-0003-2678-1266ORCID · verified
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 32 (7 first)Database Systems & Data Management · 3Information Retrieval & Web Search · 2Big Data, Cloud & Distributed Data Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TabFairGDT: A Fast Fair Tabular Data Generator Using Autoregressive Decision TreesabstractEnsuring fairness in machine learning remains a significant challenge, as models often inherit biases from their training data. Generative models have recently emerged as a promising approach to mitigate bias at the data level while preserving utility. However, many rely on deep architectures, despite evidence that simpler models can be highly effective for tabular data. In this work, we introduce TabFairGDT, a novel method for generating fair synthetic tabular data using autoregressive decision trees. To enforce fairness, we propose a soft leaf resampling technique that adjusts decision tree outputs to reduce bias while preserving predictive performance. Our approach is non-parametric, effectively capturing complex relationships between mixed feature types, without relying on assumptions about the underlying data distributions. We evaluate TabFairGDT on benchmark fairness datasets and demonstrate that it outperforms state-of-the-art (SOTA) deep generative models, achieving better fairness-utility trade-off for downstream tasks, as well as higher synthetic data quality. Moreover, our method is lightweight, highly efficient, and CPU-compatible, requiring no data preprocessing. Remarkably, TabFairGDT achieves a 72% average speedup over the fastest SOTA baseline across various dataset sizes, and can generate fair synthetic data for medium-sized datasets (10 features, 10K samples) in just one second on a standard CPU, making it an ideal solution for realworld fairness-sensitive applications. Emmanouil Panagiotou, Benoît Ronval, Arjun Roy 0001, Ludwig Bothmann, Bernd Bischl, Siegfried Nijssen, Eirini Ntoutsi |
ICDM | 6 |
| 2025 | Detection of Large Language Model Contamination with Tabular Data
Benoît Ronval, Pierre Dupont, Siegfried Nijssen |
IDA | 3 |
| 2024 | Efficient Lookahead Decision Trees
Harold Silvère Kiossou, Pierre Schaus, Siegfried Nijssen, Gaël Aglin |
IDA (2) | 3 |
| 2024 | Interpretable Quantile Regression by Optimal Decision Trees
Valentin Lemaire, Gaël Aglin, Siegfried Nijssen |
IDA (2) | 3 |
| 2023 | RL-Net: Interpretable Rule Learning with Neural Networks
Lucile Dierckx, Rosana Veroneze, Siegfried Nijssen |
PAKDD (1) | 3 |
| 2022 | Detection and Multi-label Classification of Bats
Lucile Dierckx, Mélanie Beauvois, Siegfried Nijssen |
IDA | 3 |
| 2022 | Learning Optimal Decision Trees Under Memory Constraints
Gaël Aglin, Siegfried Nijssen, Pierre Schaus |
ECML/PKDD (5) | 2 |
| 2022 | Time Constrained DL8.5 Using Limited Discrepancy Search
Harold Silvère Kiossou, Pierre Schaus, Siegfried Nijssen, Vinasétan Ratheil Houndji |
ECML/PKDD (5) | 3 |
| 2021 | Iterated Matrix Reordering
Gauthier Van Vracem, Siegfried Nijssen |
ECML/PKDD (3) | 2 |
| 2018 | ConvoMap: Using Convolution to Order Boolean Data
Thomas Bollen, Guillaume Leurquin, Siegfried Nijssen |
IDA | 3 |
| 2017 | Discovering interesting patterns in large graph cubesabstractDue to the increasing importance and volume of highly interconnected data, such as in social or information networks, a plethora of graph mining techniques have been designed to enable the analysis of such data. In this work, we focus on the mining of associations between entity features in networks. We model each entity feature as a dimension to be analyzed. Consequently we build our approach on top of the existing graph cube framework which is an extension of the concept of the data cube to networks. Our task is particularly challenging because it requires the analysis of both the initial multidimensional network and all its subsequent aggregate forms. As soon as we deal with a big data situation it is impossible for an analyst to consider manually all the possible views of the network data. The aim of this work is to design an algorithm for the discovery of interesting patterns in large graph cubes. Thus, instead of examining all the possible aggregations manually, the proposed technique leads the analyst to the interesting associations or patterns in the multidimensional network. Furthermore, we study the application of existing algorithms from the frequent itemset mining literature on graph data and propose a mapping between the two settings. Florian Demesmaeker, Amine Ghrab, Siegfried Nijssen, Sabri Skhiri |
IEEE BigData | 3 |
| 2017 | Biclustering Multivariate Time Series
Ricardo Cachucho, Siegfried Nijssen, Arno J. Knobbe |
IDA | 2 |
| 2017 | Semiring Rank Matrix FactorizationabstractRank data, in which each row is a complete or partial ranking of available items (columns), is ubiquitous. Among others, it can be used to represent preferences of users, levels of gene expression, and outcomes of sports events. It can have many types of patterns, among which consistent rankings of a subset of the items in multiple rows, and multiple rows that rank the same subset of the items highly. In this article, we show that the problems of finding such patterns can be formulated within a single generic framework that is based on the concept of semiring matrix factorization. In this framework, we employ the max-product semiring rather than the plus-product semiring common in traditional linear algebra. We apply this semiring matrix factorization framework on two tasks: sparse rank matrix factorization and rank matrix tiling. Experiments on both synthetic and real world datasets show that the framework is capable of discovering different types of structure as well as obtaining high quality solutions. Thanh Le Van, Siegfried Nijssen, Matthijs van Leeuwen, Luc De Raedt |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2016 | Bipeline: A Web-Based Visualization Tool for Biclustering of Multivariate Time Series
Ricardo Cachucho, Siegfried Nijssen, Arno J. Knobbe |
ECML/PKDD (3) | 3 |
| 2015 | Constraint-Based Querying for Bayesian Network Exploration
Behrouz Babaki, Tias Guns, Siegfried Nijssen, Luc De Raedt |
IDA | 3 |
| 2015 | Rank Matrix Factorisation
Thanh Le Van, Matthijs van Leeuwen, Siegfried Nijssen, Luc De Raedt |
PAKDD (1) | 3 |
| 2015 | Efficient algorithms for finding optimal binary features in numeric and nominal labeled data
Michael Mampaey, Siegfried Nijssen, A. J. Feelders, Rob M. Konijn, Arno J. Knobbe |
Knowl. Inf. Syst. | 2 |
| 2014 | Ranked Tiling
Thanh Le Van, Matthijs van Leeuwen, Siegfried Nijssen, Ana Carolina Fierro, Kathleen Marchal, Luc De Raedt |
ECML/PKDD (2) | 3 |
| 2013 | Mining characteristic multi-scale motifs in sensor-based time seriesabstractMore and more, physical systems are being fitted with various kinds of sensors in order to monitor their behavior, health or intensity of use. The large quantities of time series data collected from these complex systems often exhibit two important characteristics: the data is a combination of various superimposed effects operating at different time scales, and each effect shows a fair degree of repetition. Each of these effects can be described by a small collection of motifs: recurring temporal patterns in the data. We propose a method to discover characteristic and potentially overlapping motifs at multiple time scales, taking into account systemic deformations and temporal warping. Our method is based on a combination of scale-space theory and the Minimum Description Length principle. We show its effectiveness on two time series datasets from real world applications. Ugo Vespier, Siegfried Nijssen, Arno J. Knobbe |
CIKM | 2 |
| 2013 | Dominance Programming for Itemset MiningabstractFinding small sets of interesting patterns is an important challenge in pattern mining. In this paper, we argue that several well-known approaches that address this challenge are based on performing pair wise comparisons between patterns. Examples include finding closed patterns, free patterns, relevant subgroups and skyline patterns. Although progress has been made on each of these individual problems, a generic approach for solving these problems (and more) is still lacking. This paper tackles this challenge. It proposes a novel, generic approach for handling pattern mining problems that involve pair wise comparisons between patterns. Our key contributions are the following. First, we propose a novel algebra for programming pattern mining problems. This algebra extends relational algebras in a novel way towards pattern mining. It allows for the generic combination of constraints on individual patterns with dominance relations between patterns. Second, we introduce a modified generic constraint satisfaction system to evaluate these algebraic expressions. Experiments show that this generic approach can indeed effectively identify patterns expressed in the algebra. Benjamin Négrevergne, Anton Dries, Tias Guns, Siegfried Nijssen |
ICDM | 4 |
| 2013 | Guest editor's introduction: special issue of the ECML PKDD 2013 journal track
Hendrik Blockeel, Kristian Kersting, Siegfried Nijssen, Filip Zelezný |
Data Min. Knowl. Discov. | 3 |
| 2013 | k-Pattern Set Mining under ConstraintsabstractWe introduce the problem of k-pattern set mining, concerned with finding a set of k related patterns under constraints. This contrasts to regular pattern mining, where one searches for many individual patterns. The k-pattern set mining problem is a very general problem that can be instantiated to a wide variety of well-known mining tasks including concept-learning, rule-learning, redescription mining, conceptual clustering and tiling. To this end, we formulate a large number of constraints for use in k-pattern set mining, both at the local level, that is, on individual patterns, and on the global level, that is, on the overall pattern set. Building general solvers for the pattern set mining problem remains a challenge. Here, we investigate to what extent constraint programming (CP) can be used as a general solution strategy. We present a mapping of pattern set constraints to constraints currently available in CP. This allows us to investigate a large number of settings within a unified framework and to gain insight in the possibilities and limitations of these solvers. This is important as it allows us to create guidelines in how to model new problems successfully and how to model existing problems more efficiently. It also opens up the way for other solver technologies. Tias Guns, Siegfried Nijssen, Luc De Raedt |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2012 | Efficient Algorithms for Finding Richer Subgroup Descriptions in Numeric and Nominal DataabstractSubgroup discovery systems are concerned with finding interesting patterns in labeled data. How these systems deal with numeric and nominal data has a large impact on the quality of their results. In this paper, we consider two ways to extend the standard pattern language of subgroup discovery: using conditions that test for interval membership for numeric attributes, and value set membership for nominal attributes. We assume a greedy search setting, that is, iteratively refining a given subgroup, with respect to a (convex) quality measure. For numeric attributes, we propose an algorithm that finds the optimal interval in linear (rather than quadratic) time, with respect to the number of examples and split points. Similarly, for nominal attributes, we show that finding the optimal set of values can be achieved in linear (rather than exponential) time, with respect to the number of examples and the size of the domain of the attribute. These algorithms operate by only considering subgroup refinements that lie on a convex hull in ROC space, thus significantly narrowing down the search space. We further provide efficient algorithms specifically for the popular Weighted Relative Accuracy quality measure, taking advantage of some of its properties. Our algorithms are shown to perform well in practice, and furthermore provide additional expressive power leading to higher-quality results. Michael Mampaey, Siegfried Nijssen, A. J. Feelders, Arno J. Knobbe |
ICDM | 2 |
| 2012 | MDL-Based Analysis of Time Series at Multiple Time-Scales
Ugo Vespier, Arno J. Knobbe, Siegfried Nijssen, Joaquin Vanschoren |
ECML/PKDD (2) | 3 |
| 2012 | Mining Patterns in Networks using HomomorphismabstractIn recent years many algorithms have been developed for finding patterns in graphs and networks. A disadvantage of these algorithms is that they use subgraph isomorphism to determine the support of a graph pattern; subgraph isomorphism is a well-known NP complete problem. In this paper, we propose an alternative approach which mines tree patterns in networks by using subgraph homomorphism. The advantage of homomorphism is that it can be computed in polynomial time, which allows us to develop an algorithm that mines tree patterns in arbitrary graphs in incremental polynomial time. Homomorphism however entails two problems not found when using isomorphism: (1) two patterns of different size can be equivalent; (2) patterns of unbounded size can be frequent. In this paper we formalize these problems and study solutions that easily fit within our algorithm. Anton Dries, Siegfried Nijssen |
SDM | 2 |
| 2011 | Evaluating Pattern Set Mining Strategies in a Constraint Programming Framework
Tias Guns, Siegfried Nijssen, Luc De Raedt |
PAKDD (2) | 2 |
| 2010 | Integrating Constraint Programming and Itemset Mining
Siegfried Nijssen, Tias Guns |
ECML/PKDD (2) | 1 |
| 2010 | Optimal constraint-based decision tree induction from itemset lattices
Siegfried Nijssen, Élisa Fromont |
Data Min. Knowl. Discov. | 1 |
| 2010 | Mining Predictive k-CNF ExpressionsabstractWe adapt Mitchell's version space algorithm for mining k-CNF formulas. Advantages of this algorithm are that it runs in a single pass over the data, is conceptually simple, can be used for missing value prediction, and has interesting theoretical properties, while an empirical evaluation on classification tasks yields competitive predictive results. Anton Dries, Luc De Raedt, Siegfried Nijssen |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2009 | A query language for analyzing networksabstractWith more and more large networks becoming available, mining and querying such networks are increasingly important tasks which are not being supported by database models and querying languages. This paper wants to alleviate this situation by proposing a data model and a query language for facilitating the analysis of networks. Key features include support for executing external tools on the networks, flexible contexts on the network each resulting in a different graph, primitives for querying subgraphs (including paths) and transforming graphs. Anton Dries, Siegfried Nijssen, Luc De Raedt |
CIKM | 2 |
| 2009 | Correlated itemset mining in ROC space: a constraint programming approachabstractCorrelated or discriminative pattern mining is concerned with finding the highest scoring patterns w.r.t. a correlation measure (such as information gain). By reinterpreting correlation measures in ROC space and formulating correlated itemset mining as a constraint programming problem, we obtain new theoretical insights with practical benefits. More specifically, we contribute 1) an improved bound for correlated itemset miners, 2) a novel iterative pruning algorithm to exploit the bound, and 3) an adaptation of this algorithm to mine all itemsets on the convex hull in ROC space. The algorithm does not depend on a minimal frequency threshold and is shown to outperform several alternative approaches by orders of magnitude, both in runtime and in memory requirements. Siegfried Nijssen, Tias Guns, Luc De Raedt |
KDD | 1 |
| 2009 | Grammar MiningabstractWe introduce the problem of grammar mining, where patterns are context-free grammars, as a generalization of a large number of common pattern mining tasks, such as tree, sequence and itemset mining. The proposed system offers data miners the possibility to specify and explore pattern domains declaratively, in a way which is very similar to the declarative specification of regular expressions in popular scripting languages. Siegfried Nijssen, Luc De Raedt |
SDM | 1 |
| 2008 | Constraint programming for itemset miningabstractThe relationship between constraint-based mining and constraint programming is explored by showing how the typical constraints used in pattern mining can be formulated for use in constraint programming environments. The resulting framework is surprisingly flexible and allows us to combine a wide range of mining constraints in different ways. We implement this approach in off-the-shelf constraint programming systems and evaluate it empirically. The results show that the approach is not only very expressive, but also works well on complex benchmark problems. Luc De Raedt, Tias Guns, Siegfried Nijssen |
KDD | 3 |
| 2008 | What Is Frequent in a Single Graph?
Björn Bringmann, Siegfried Nijssen |
PAKDD | 2 |
| 2007 | Mining optimal decision trees from itemset latticesabstractWe present DL8, an exact algorithm for finding a decision tree that optimizes a ranking function under size, depth, accuracy and leaf constraints. Because the discovery of optimal trees has high theoretical complexity, until now few efforts have been made to compute such trees for real-world datasets. An exact algorithm is of both scientific and practical interest. From a scientific point of view, it can be used as a gold standard to evaluate the performance of heuristic constraint-based decision tree learners and to gain new insight in traditional decision tree learners. From the application point of view, it can be used to discover trees that cannot be found by heuristic decision tree learners. The key idea behind our algorithm is that there is a relation between constraints on decision trees and constraints on itemsets. We show that optimal decision trees can be extracted from lattices of itemsets in linear time. We give several strategies to efficiently build these lattices. Experiments show that under the same constraints, DL8 obtains better results than C4.5, which confirms that exhaustive search does not always imply overfitting. The results also show that DL8 is a useful and interesting tool to learn decision trees under constraints. Siegfried Nijssen, Élisa Fromont |
KDD | 1 |
| 2006 | Don't Be Afraid of Simpler Patterns
Björn Bringmann, Albrecht Zimmermann, Luc De Raedt, Siegfried Nijssen |
PKDD | 4 |
| 2004 | A quickstart in frequent structure mining can make a differenceabstractGiven a database, structure mining algorithms search for substructures that satisfy constraints such as minimum frequency, minimum confidence, minimum interest and maximum frequency. Examples of substructures include graphs, trees and paths. For these substructures many mining algorithms have been proposed. In order to make graph mining more efficient, we investigate the use of the "quickstart principle", which is based on the fact that these classes of structures are contained in each other, thus allowing for the development of structure mining algorithms that split the search into steps of increasing complexity. We introduce the GrAph/Sequence/Tree extractiON (Gaston) algorithm that implements this idea by searching first for frequent paths, then frequent free trees and finally cyclic graphs. We investigate two alternatives for computing the frequency of structures and present experimental results to relate these alternatives. Siegfried Nijssen, Joost N. Kok |
KDD | 1 |
| 2003 | Efficient Frequent Query Discovery in FARMER
Siegfried Nijssen, Joost N. Kok |
PKDD | 1 |