EDBT 2026 Demo / reviewers in the wild / expert
Filip S. Slijkhuis
dblp:336/7568
· DBLP profile ↗
3ranked-venue papers in the field
1as first author
3since 2021 · last 2025
0000-0001-5926-368XORCID · corroborated
Domains — venue-derived; a paper can count in several
Other / Interdisciplinary · 3 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Exploiting Causal Structures for Data-Efficient Neural Network DesignabstractWhen employing neural networks in the real world, we often encounter challenges related to dataset size, data imbalance and model selection. In this paper, we investigate the benefits of incorporating information obtained from causal structures or Qualitative Models for Data Generation Processes (QM-DGP) into neural network design and training, by leveraging causal reasoning principles such as d-separation and the Markov blanket to determine relevant input variables. Through empirical analysis, we compare networks which exploit causal structures and those that do not, focusing on aspects related to dataset efficiency. Our findings show that networks which exploit causal structures require fewer samples to achieve (near-)optimal performance on average, making them more suitable in frugal learning scenarios. We also propose a causally-decomposed neural network architecture based on causal structural information and show that it is more data efficient than its fully-connected counterpart. Our results highlight the practical advantage of using causal reasoning in neural network design, particularly in settings where data efficiency plays a crucial role. Filip S. Slijkhuis, Kathryn B. Laskey, Franck Mignet, Gregor Pavlin, L. Jansen |
FUSION | 1 |
| 2024 | A Qualitative Causal Approach to Determining Adequate Training Data Quantity for Machine LearningabstractThis paper proposes an improved analysis of the Qualitative Models of Data Generating Processes (QM-DGP). The approach supports (i) determination of the complexity of a Machine Learning problem and (ii) a coarse determination of the quantities of training data that are needed to train good quality models. Compared to the previously published approach to the QM-DGP analysis, this paper introduces a more thorough and theoretically sound treatment of the learning complexity. Firstly, the approach provides more rigorous determination of the complexity of the data generating processes (DGP). Secondly, the determination of the learning complexity and the required training data volumes is based on sound statistical principles for the estimation of the distributions over categorical variables. The effectiveness of the proposed method was experimentally confirmed in controlled settings. Different ground truth models were used to sample test and training data. The approach correctly predicts the size of the training data sets for which machine learning yields models supporting classification close to Bayes Error. While the majority of the experiments were carried out on probabilistic graphical models (PGM), the experiments with Neural Networks confirmed that the QM-DGP approach is not limited to PGMs. Franck Mignet, Filip S. Slijkhuis, A. Abouhafc, Gregor Pavlin, Kathryn B. Laskey |
FUSION | 2 |
| 2023 | Qualitative Models of Data Generation Processes: Facilitating Data-Intensive AI SolutionsabstractAI-based decision support solutions require life cycles that adequately address critical steps, such as (i) finding suitable machine learning (ML) methods for the problem at hand, (ii) preparing and executing adequate data acquisition processes and (iii) tractable evaluation of the overall solution. Understanding the data generating processes is key in achieving this. Training and test data can be seen as a result of a causal data generation process, a sampling process in which the data is collected from different sources that are influenced by multiple interdependent phenomena. This is represented by a Qualitative Model of Data Generation Processes (QM-DGP), a causal graphical model. QM-DGP facilitates analysis of the complexity of the underlying data generating processes that can inform the development of trustable ML-based solutions in multiple ways. Firstly, this analysis is the basis for the determination of the required complexity of the ML models. Secondly, it facilitates the determination of the quantities of training data supporting good learning results. Thirdly, it can provide guidance for a systematic simplification of the models, supporting tractable solutions without significantly reduced performance. The construction of QM-DGP and the analysis benefit from sound theoretical concepts, such as d-separation and I-Maps. Experimental results with simulated data indicate that the approach can be effective in predicting the required quantities of training data and the determination of the modelling complexity using different types of models. Gregor Pavlin, Kathryn B. Laskey, Franck Mignet, Filip S. Slijkhuis, Erik Blasch, Valentina Dragos, Johan Pieter de Villiers, Lennard Jansen |
FUSION | 4 |