VLDB 2026 Research / reviewers in the wild / expert
Tim Verdonck
dblp:95/1823
· DBLP profile ↗
17ranked-venue papers
1as first author
17since 2021 · last 2026
0000-0003-1105-2028ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 1 first-author · 12 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Inductive inference of gradient-boosted decision trees on graphs for insurance fraud detection
Félix Vandervorst, Bruno Deprez, Wouter Verbeke, Tim Verdonck |
Data Min. Knowl. Discov. | 4 |
| 2025 | Differentiable Causal Structure Learning with Identifiability by NOTIMEabstractThe introduction of the NOTEARS algorithm resulted in a wave of research on differentiable Directed Acyclic Graph (DAG) learning. Differentiable DAG learning transforms the combinatorial problem of identifying the DAG underlying a Structural Causal Model (SCM) into a constrained continuous optimization problem. Being differentiable, these problems can be solved using gradient-based tools which allow integration into other differentiable objectives. However, in contrast to classical constrained-based algorithms, the identifiability properties of differentiable algorithms are poorly understood. We illustrate that even in the well-known Linear Non-Gaussian Additive Model (LiNGAM), the current state-of-the-art methods do not identify the true underlying DAG. To address the issue, we propose NOTIME (\emph{Non-combinatorial Optimization of Trace exponential and Independence MEasures}), the first differentiable DAG learning algorithm with \emph{provable} identifiability guarantees under the LiNGAM by building on a measure of (joint) independence. With its identifiability guarantees, NOTIME remains invariant to normalization of the data on a population level, a property lacking in existing methods. NOTIME compares favourably against NOTEARS and other (scale-invariant) differentiable DAG learners, across different noise distributions and normalization procedures. Introducing the first identifiability guarantees to general LiNGAM is an important step towards practical adoption of differentiable DAG learners. Jeroen Berrevoets, Jakob Raymaekers, Mihaela van der Schaar, Tim Verdonck, Ruicong Yao |
AISTATS | 4 |
| 2025 | Causal discovery in mixed additive noise modelsabstractUncovering causal relationships in datasets that include both categorical and continuous variables is a challenging problem. The overwhelming majority of existing methods restrict their application to dealing with a single type of variable. Our contribution is a structural causal model designed to handle mixed-type data through a general function class. We present a theoretical foundation that specifies the conditions under which the directed acyclic graph underlying the causal model can be identified from observed data. In addition, we propose Mixed-type data Extension for Regression and Independence Testing (MERIT), enabling the discovery of causal connections in real-world classification settings. Our empirical studies demonstrate that MERIT outperforms its state-of-the-art competitor in causal discovery on relatively low-dimensional data. Ruicong Yao, Tim Verdonck, Jakob Raymaekers |
AISTATS | 2 |
| 2025 | AutoCATE: End-to-End, Automated Treatment Effect EstimationabstractEstimating causal effects is crucial in domains like healthcare, economics, and education. Despite advances in machine learning (ML) for estimating conditional average treatment effects (CATE), the practical adoption of these methods remains limited, due to the complexities of implementing, tuning, and validating them. To address these challenges, we formalize the search for an optimal ML pipeline for CATE estimation as a counterfactual Combined Algorithm Selection and Hyperparameter (CASH) optimization. We introduce AutoCATE, the first end-to-end, automated solution for CATE estimation. Unlike prior approaches that address only parts of this problem, AutoCATE integrates evaluation, estimation, and ensembling in a unified framework. AutoCATE enables comprehensive comparisons of different protocols, yielding novel insights into CATE estimation and a final configuration that outperforms commonly used strategies. To facilitate broad adoption and further research, we release AutoCATE as an open-source software package. Toon Vanderschueren, Tim Verdonck, Mihaela van der Schaar, Wouter Verbeke |
ICML | 2 |
| 2025 | Can causal machine learning reveal individual bid responses of bank customers? - A study on mortgage loan applications in Belgium
Christopher Bockel-Rickermann, Sam Verboven, Tim Verdonck, Wouter Verbeke |
Decis. Support Syst. | 3 |
| 2025 | Enhancing explainability in real-world scenarios: Towards a robust stability measure for local interpretability
Eduardo Sepúlveda, Félix Vandervorst, Bart Baesens, Tim Verdonck |
Expert Syst. Appl. | 4 |
| 2024 | A new perspective on classification: Optimally allocating limited resources to uncertain tasks
Toon Vanderschueren, Bart Baesens, Tim Verdonck, Wouter Verbeke |
Decis. Support Syst. | 3 |
| 2024 | Fast linear model trees by PILOTabstractAbstract Linear model trees are regression trees that incorporate linear models in the leaf nodes. This preserves the intuitive interpretation of decision trees and at the same time enables them to better capture linear relationships, which is hard for standard decision trees. But most existing methods for fitting linear model trees are time consuming and therefore not scalable to large data sets. In addition, they are more prone to overfitting and extrapolation issues than standard regression trees. In this paper we introduce PILOT, a new algorithm for linear model trees that is fast, regularized, stable and interpretable. PILOT trains in a greedy fashion like classic regression trees, but incorporates an L 2 boosting approach and a model selection rule for fitting linear models in the nodes. The abbreviation PILOT stands for PIecewise Linear Organic Tree, where ‘organic’ refers to the fact that no pruning is carried out. PILOT has the same low time and space complexity as CART without its pruning. An empirical study indicates that PILOT tends to outperform standard decision trees and other linear model trees on a variety of data sets. Moreover, we prove its consistency in an additive model setting under weak assumptions. When the data is generated by a linear model, the convergence rate is polynomial. Jakob Raymaekers, Peter J. Rousseeuw, Tim Verdonck, Ruicong Yao |
Mach. Learn. | 3 |
| 2024 | Special issue on feature engineering editorial
Tim Verdonck, Bart Baesens, María Óskarsdóttir, Seppe K. L. M. vanden Broucke |
Mach. Learn. | 1 |
| 2023 | Exploiting sensor data in professional road cycling: personalized data-driven approach for frequent fitness monitoring
Arie-Willem de Leeuw, Mathieu Heijboer, Tim Verdonck, Arno J. Knobbe, Steven Latré |
Data Min. Knowl. Discov. | 3 |
| 2023 | Interpretable cost-sensitive regression through one-step boostingabstractIn most practical prediction problems, such as regression and classification, the different types of prediction errors are not equally costly in the decision-making process. Although there exist numerous real-world cost-sensitive regression problems, ranging from loan charge-off forecasting to house price predictions, the literature on cost-sensitive learning mainly focuses on classification and only a few solutions are proposed for regression problems. These regressions are typically characterized by an asymmetric cost structure, where over- and underpredictions of a similar magnitude face vastly different costs. In this paper, we present a one-step boosting method (OSB) for cost-sensitive regression. The proposed methodology leverages a secondary learner to incorporate cost-sensitivity into an already trained cost-insensitive regression model. The secondary learner is defined as a linear function of certain variables deemed interesting for cost-sensitivity. These variables do not necessarily need to be the same as in the already trained model. An efficient optimization algorithm is achieved through iteratively reweighted least squares using the asymmetric cost function . The obtained results become interpretable through bootstrapping, enabling decision makers to distinguish important variables for cost-sensitivity as well as facilitating statistical inference . Applying different cost functions and various initial cost-insensitive learning methods on several public datasets consistently yields a significant reduction in the average misprediction cost, illustrating the excellent performance of our approach. Thomas Decorte, Jakob Raymaekers, Tim Verdonck |
Decis. Support Syst. | 3 |
| 2023 | Fraud analytics: A decade of research: Organizing challenges and solutions in the field
Christopher Bockel-Rickermann, Tim Verdonck, Wouter Verbeke |
Expert Syst. Appl. | 2 |
| 2023 | Regularization oversampling for classification tasks: To exploit what you do not know
Lennert Van der Schraelen, Kristof Stouthuysen, Seppe K. L. M. vanden Broucke, Tim Verdonck |
Inf. Sci. | 4 |
| 2023 | Smart initialisation and approximating loss function for robust regression
Thomas Servotte, Jakob Raymaekers, Tim Verdonck |
Inf. Sci. | 3 |
| 2022 | Data misrepresentation detection for insurance underwriting fraud prevention
Félix Vandervorst, Wouter Verbeke, Tim Verdonck |
Decis. Support Syst. | 3 |
| 2022 | Predict-then-optimize or predict-and-optimize? An empirical evaluation of cost-sensitive learning strategies
Toon Vanderschueren, Tim Verdonck, Bart Baesens, Wouter Verbeke |
Inf. Sci. | 2 |
| 2021 | Data engineering for fraud detection
Bart Baesens, Sebastiaan Höppner, Tim Verdonck |
Decis. Support Syst. | 3 |