EDBT 2026 Demo / reviewers in the wild / expert
Wouter Verbeke
dblp:07/7256
· DBLP profile ↗
35ranked-venue papers
1as first author
20since 2021 · last 2026
0000-0002-8438-0535ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 1 first-author · 14 since 2021Databases, data management, data science and information retrieval · 8 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Inductive inference of gradient-boosted decision trees on graphs for insurance fraud detection
Félix Vandervorst, Bruno Deprez, Wouter Verbeke, Tim Verdonck |
Data Min. Knowl. Discov. | 3 |
| 2025 | AutoCATE: End-to-End, Automated Treatment Effect EstimationabstractEstimating causal effects is crucial in domains like healthcare, economics, and education. Despite advances in machine learning (ML) for estimating conditional average treatment effects (CATE), the practical adoption of these methods remains limited, due to the complexities of implementing, tuning, and validating them. To address these challenges, we formalize the search for an optimal ML pipeline for CATE estimation as a counterfactual Combined Algorithm Selection and Hyperparameter (CASH) optimization. We introduce AutoCATE, the first end-to-end, automated solution for CATE estimation. Unlike prior approaches that address only parts of this problem, AutoCATE integrates evaluation, estimation, and ensembling in a unified framework. AutoCATE enables comprehensive comparisons of different protocols, yielding novel insights into CATE estimation and a final configuration that outperforms commonly used strategies. To facilitate broad adoption and further research, we release AutoCATE as an open-source software package. Toon Vanderschueren, Tim Verdonck, Mihaela van der Schaar, Wouter Verbeke |
ICML | 4 |
| 2025 | Can causal machine learning reveal individual bid responses of bank customers? - A study on mortgage loan applications in Belgium
Christopher Bockel-Rickermann, Sam Verboven, Tim Verdonck, Wouter Verbeke |
Decis. Support Syst. | 4 |
| 2025 | Uplift Model Evaluation with Ordinal Dominance GraphsabstractUplift modelling is a subfield of causal learning that focuses on ranking entities by individual treatment effects. Uplift models are typically evaluated using Qini curves or Qini scores. While intuitive, the theoretical grounding for Qini in the literature is limited, and the mathematical connection to the well-understood Receiver Operating Characteristic (ROC) curve is unclear. In this paper, we introduce pROCini, a novel uplift evaluation metric that improves upon Qini in two important ways. First, it explicitly incorporates more information by taking into account negative outcomes. Second, it leverages this additional information within the Ordinal Dominance Graph framework, which is the basis behind the well known ROC curve, resulting in a mathematically well-behaved metric that facilitates theoretical analysis. We derive confidence bounds for pROCini, exploiting its theoretical properties. Finally, we empirically validate the improved discriminative power of ROCini and pROCini in a simulation study as well as via experiments on real data. Brecht Verbeken, Marie-Anne Guerry, Wouter Verbeke, Sam Verboven |
J. Mach. Learn. Res. | 3 |
| 2025 | A prescriptive analytics framework for jointly optimizing retention incentives and targeting
Paolo Latorre, Armando Meza, Héctor A. López-Ospina, Wouter Verbeke, Juan Pérez |
Knowl. Based Syst. | 4 |
| 2024 | A new perspective on classification: Optimally allocating limited resources to uncertain tasks
Toon Vanderschueren, Bart Baesens, Tim Verdonck, Wouter Verbeke |
Decis. Support Syst. | 4 |
| 2024 | Evaluating text classification: A benchmark study
Manon Reusens, Alexander Stevens, Jonathan Tonglet, Johannes De Smedt, Wouter Verbeke, Seppe K. L. M. vanden Broucke, Bart Baesens |
Expert Syst. Appl. | 5 |
| 2024 | Data-driven internal mobility: Similarity regularization gets the job done
Simon De Vos, Johannes De Smedt, Marijke Verbruggen, Wouter Verbeke |
Knowl. Based Syst. | 4 |
| 2024 | Probabilistic Forecasting With Modified N-BEATS NetworksabstractIn this article, we present a modification to the state-of-the-art N-BEATS deep learning architecture for the univariate time series point forecasting problem for generating parametric probabilistic forecasts. Next, we propose an extension to this probabilistic N-BEATS architecture to allow optimizing probabilistic forecasts from both a traditional forecast accuracy perspective as well as a forecast stability perspective, where the latter is defined in terms of a change in the forecast distribution for a specific time period caused by updating the probabilistic forecast for this time period when new observations become available (i.e., as time passes). We empirically show that this extension leads to more stable forecast distributions without causing considerable losses in forecast accuracy for the M4 monthly dataset. Finally, we present a second extension to the probabilistic N-BEATS network which makes it possible to jointly optimize single-period marginal and multiperiod cumulative (i.e., aggregated over multiple time periods) probabilistic forecasts. Empirical results are reported for the M4 monthly dataset and indicate that improvements in accuracy can be obtained over basic but well-established methods to produce probabilistic cumulative forecasts. The proposed probabilistic N-BEATS network and the extensions are all useful in a supply chain planning context. Jente Van Belle, Ruben Crevits, Daan Caljon, Wouter Verbeke |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Client Recruitment for Federated Learning in ICU Length of Stay PredictionabstractMachine and deep learning methods for medical and healthcare applications have shown significant progress and performance improvement in recent years. These methods require vast amounts of training data which are available in the medical sector, albeit decentralized. Medical institutions generate vast amounts of data for which sharing and centralizing remains a challenge as the result of data and privacy regulations. Federated Learning (FL) is well-suited to tackle these challenges. However, FL comes with a new set of open problems related to communication overhead, efficient parameter aggregation, client selection strategies and more. In this work, we address the step prior to the initiation of a federated network for model training, client recruitment. By intelligently recruiting clients, communication overhead and overall cost of training can be reduced without sacrificing predictive performance. Client recruitment aims at pre-excluding potential clients from partaking in the federation based on a set of criteria indicative of their eventual contributions to the federation. In this work, we propose a client recruitment approach using only the output distribution and sample size at the client site. We show how a subset of clients can be recruited without sacrificing model performance whilst significantly improving computation time. By applying the recruitment approach to the training of federated models for accurate patient Length of Stay prediction using data from 189 intensive care units (clients), we show how the models trained in federations made up from only recruited clients significantly outperform federated models trained with the standard procedure in terms of predictive power and training time. Vincent Scheltjens, Lyse Naomi Wamba Momo, Wouter Verbeke, Bart De Moor |
e-Science | 3 |
| 2023 | Accounting For Informative Sampling When Learning to Forecast Treatment Outcomes Over TimeabstractMachine learning (ML) holds great potential for accurately forecasting treatment outcomes over time, which could ultimately enable the adoption of more individualized treatment strategies in many practical applications. However, a significant challenge that has been largely overlooked by the ML literature on this topic is the presence of informative sampling in observational data. When instances are observed irregularly over time, sampling times are typically not random, but rather informative–depending on the instance’s characteristics, past outcomes, and administered treatments. In this work, we formalize informative sampling as a covariate shift problem and show that it can prohibit accurate estimation of treatment outcomes if not properly accounted for. To overcome this challenge, we present a general framework for learning treatment outcomes in the presence of informative sampling using inverse intensity-weighting, and propose a novel method, TESAR-CDE, that instantiates this framework using Neural CDEs. Using a simulation environment based on a clinical use case, we demonstrate the effectiveness of our approach in learning under informative sampling. Toon Vanderschueren, Alicia Curth, Wouter Verbeke, Mihaela van der Schaar |
ICML | 3 |
| 2023 | HydaLearn
Sam Verboven, Muhammad Hafeez Chaudhary, Jeroen Berrevoets, Vincent Ginis, Wouter Verbeke |
Appl. Intell. | 5 |
| 2023 | Fraud analytics: A decade of research: Organizing challenges and solutions in the field
Christopher Bockel-Rickermann, Tim Verdonck, Wouter Verbeke |
Expert Syst. Appl. | 3 |
| 2022 | Cost-sensitive ensemble learning: a unifying frameworkabstractAbstract Over the years, a plethora of cost-sensitive methods have been proposed for learning on data when different types of misclassification errors incur different costs. Our contribution is a unifying framework that provides a comprehensive and insightful overview on cost-sensitive ensemble methods, pinpointing their differences and similarities via a fine-grained categorization. Our framework contains natural extensions and generalisations of ideas across methods, be it AdaBoost, Bagging or Random Forest, and as a result not only yields all methods known to date but also some not previously considered. George Petrides, Wouter Verbeke |
Data Min. Knowl. Discov. | 2 |
| 2022 | Data misrepresentation detection for insurance underwriting fraud prevention
Félix Vandervorst, Wouter Verbeke, Tim Verdonck |
Decis. Support Syst. | 2 |
| 2022 | Predict-then-optimize or predict-and-optimize? An empirical evaluation of cost-sensitive learning strategies
Toon Vanderschueren, Tim Verdonck, Bart Baesens, Wouter Verbeke |
Inf. Sci. | 4 |
| 2022 | Learning to Rank for Uplift ModelingabstractCausal classification concerns the estimation of the net effect of a treatment on an outcome of interest at the instance level, i.e., of the individual treatment effect (ITE). For binary treatment and outcome variables, causal classification models produce ITE estimates that essentially allow one to rank instances from a large positive effect to a large negative effect. Often, as in uplift modeling (UM), one is merely interested in this ranking, rather than in the ITE estimates themselves. In this regard, we investigate the potential of learning to rank (L2R) techniques to learn a ranking of the instances directly. We propose a unified formalization of different binary causal classification performance measures from the UM literature and explore how these can be integrated into the L2R framework. Additionally, we introduce a new metric for UM with L2R called thepromoted cumulative gain(PCG). We employ the L2R technique LambdaMART to optimize the ranking according to PCG and show improved results over the use of standard L2R metrics and equal to improved results when compared with state-of-the-art UM. Finally, we show how L2R techniques can be used to specifically optimize for the top-$k$fraction of the ranking in a UM context, however, these results do not generalize to the test set. Floris Devriendt, Jente Van Belle, Tias Guns, Wouter Verbeke |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2021 | Redefining profit metrics for boosting student retention in higher education
Sebastián Maldonado 0001, Jaime Miranda, Diego Olaya, Jonathan Vásquez, Wouter Verbeke |
Decis. Support Syst. | 5 |
| 2021 | Autoencoders for strategic decision support
Sam Verboven, Jeroen Berrevoets, Chris Wuytens, Bart Baesens, Wouter Verbeke |
Decis. Support Syst. | 5 |
| 2021 | Why you should stop predicting customer churn and start using uplift modelsabstractUplift modeling has received increasing interest in both the business analytics research community and the industry as an improved paradigm for predictive analytics for data-driven operational decision-making. The literature, however, does not provide conclusive empirical evidence that uplift modeling outperforms predictive modeling . Case studies that directly compare both approaches are lacking, and the performance of predictive models and uplift models as reported in various experimental studies cannot be compared indirectly since different evaluation measures are used to assess their performance. Therefore, in this paper, we introduce a novel evaluation metric called the maximum profit uplift (MPU) measure that allows assessing the performance in terms of the maximum potential profit that can be achieved by adopting an uplift model. This measure, developed for evaluating customer churn uplift models, extends the maximum profit measure for evaluating customer churn prediction models. While introducing the MPU measure, we describe the generally applicable liftup curve and liftup measure for evaluating uplift models as counterparts of the lift curve and lift measure that are broadly used to evaluate predictive models. These measures are subsequently applied to assess and compare the performance of customer churn prediction and uplift models in a case study that applies uplift modeling to customer retention in the financial industry. We observe that uplift models outperform predictive models and lead to improved profitability of retention campaigns. Floris Devriendt, Jeroen Berrevoets, Wouter Verbeke |
Inf. Sci. | 3 |
| 2020 | A survey and benchmarking study of multitreatment uplift modelingabstractAbstract Uplift modeling is an instrument used to estimate the change in outcome due to a treatment at the individual entity level. Uplift models assist decision-makers in optimally allocating scarce resources. This allows the selection of the subset of entities for which the effect of a treatment will be largest and, as such, the maximization of the overall returns. The literature on uplift modeling mostly focuses on queries concerning the effect of a single treatment and rarely considers situations where more than one treatment alternative is utilized. This article surveys the current literature on multitreatment uplift modeling and proposes two novel techniques: the naive uplift approach and the multitreatment modified outcome approach. Moreover, a benchmarking experiment is performed to contrast the performances of different multitreatment uplift modeling techniques across eight data sets from various domains. We verify and, if needed, correct the imbalance among the pretreatment characteristics of the treatment groups by means of optimal propensity score matching, which ensures a correct interpretation of the estimated uplift. Conventional and recently proposed evaluation metrics are adapted to the multitreatment scenario to assess performance. None of the evaluated techniques consistently outperforms other techniques. Hence, it is concluded that performance largely depends on the context and problem characteristics. The newly proposed techniques are found to offer similar performances compared to state-of-the-art approaches. Diego Olaya, Kristof Coussement, Wouter Verbeke |
Data Min. Knowl. Discov. | 3 |
| 2020 | Uplift Modeling for preventing student dropout in higher educationabstractUplift modeling is an approach for estimating the incremental effect of an action or treatment at the individual level. It has gained attention in the marketing and analytics communities due to its ability to adequately model the effect of direct marketing actions via predictive analytics. The main contribution of our study is the implementation of the uplift modeling framework to maximize the effectiveness of retention efforts in higher education institutions i.e., improvement of academic performance by offering tutorials. The objective is to improve the design of retention programs by tailoring them to students who are more likely to be retained if targeted. Data from three different bachelor programs from a Chilean university were collected. Students who participated in the tutorials are considered the treatment group, otherwise, they are assigned to the nontreatment group. Our results demonstrate the virtues of uplift modeling in tailoring retention efforts in higher education over conventional predictive modeling approaches. Diego Olaya, Jonathan Vásquez, Sebastián Maldonado 0001, Jaime Miranda, Wouter Verbeke |
Decis. Support Syst. | 5 |
| 2020 | A Model for Range Estimation and Energy-Efficient Routing of Electric Vehicles in Real-World ConditionsabstractThis paper presents an integrated model for energy consumption and range estimation, capable of energy-efficient routing. This integrated model predicts the energy consumption on all road segments in the road network and applies shortest path algorithms to calculate energy-efficient routes. A temperature-dependent model of the battery internal resistance based on real-world data translates the energy consumption prediction into driving range. The graph representation of the road network is transformed from a node-based graph to an edge-based graph to allow cost allocation of one edge based on the characteristics of the preceding edge. The integrated model is used to perform an analysis of energy-efficient routes for a real-life scenario. The analysis showed that the energy-optimal route is different from the time- and distance-optimal route with energy-efficiency gains up to 37% for the chosen scenario. The energy-efficient route tends to lean toward a distance-optimal route in the case of low auxiliary consumption and to lean toward a time-optimal route in the case of high auxiliary consumption. The energy-efficient route generation is tested in real life for one case of the chosen scenario. The measured results showed a good match with the predicted values for the energy consumption with a 9% mean relative error and preserved the ranking of all routes in terms of the travel distance, travel time, and energy consumption. The proposed integrated model is a functional model for energy consumption in real-life that differentiates many energy consumption influencing factors and produces energy-efficient routes with good accuracy. Cedric De Cauwer, Wouter Verbeke, Joeri Van Mierlo, Thierry Coosemans |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2019 | Reducing inferior member community participation using uplift modeling: Evidence from a field experiment
Steven Debaere, Floris Devriendt, Johanna Brunneder, Wouter Verbeke, Tom De Ruyck, Kristof Coussement |
Decis. Support Syst. | 4 |
| 2019 | Monitoring Urban-Freight Transport Based on GPS Trajectories of Heavy-Goods VehiclesabstractFor designing transport policy measures, it is crucial to base decisions on evidence-based insights regarding transport flows and behavior. This paper introduces practical indicators for urban transport, which can be derived from large collections of GPS trajectories of heavy-goods vehicles. The indicators framework enables cities, municipalities, and regions to gain insights into the urban transport activities in their region. We motivate our indicators based on the objectives and action plans described in the strategic plan for goods traffic of the Brussels-Capital Region in Belgium. We provide a case study on data collected from the on-board units of heavy-goods vehicles, which became mandatory in Belgium as part of its dynamic road-pricing scheme in 2016. This paper contributes to the exciting new capabilities that are obtained through GPS trajectories data and the possibilities this offers for smart cities at the tactical, operational, and strategic level. Sheida Hadavi, Sara Verlinde, Wouter Verbeke, Cathy Macharis, Tias Guns |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2018 | A Robust profit measure for binary classification model evaluation
Franco Garrido, Wouter Verbeke, Cristián Bravo |
Expert Syst. Appl. | 2 |
| 2017 | Social network analytics for churn prediction in telco: Model building, evaluation and network architecture
María Óskarsdóttir, Cristián Bravo, Wouter Verbeke, Carlos Sarraute, Bart Baesens, Jan Vanthienen |
Expert Syst. Appl. | 3 |
| 2016 | A comparative study of social network classifiers for predicting churn in the telecommunication industryabstractRelational learning in networked data has been shown to be effective in a number of studies. Relational learners, composed of relational classifiers and collective inference methods, enable the inference of nodes in a network given the existence and strength of links to other nodes. These methods have been adapted to predict customer churn in telecommunication companies showing that incorporating them may give more accurate predictions. In this research, the performance of a variety of relational learners is compared by applying them to a number of CDR datasets originating from the telecommunication industry, with the goal to rank them as a whole and investigate the effects of relational classifiers and collective inference methods separately. Our results show that collective inference methods do not improve the performance of relational classifiers and the best performing relational classifier is the network-only link-based classifier, which builds a logistic model using link-based measures for the nodes in the network. María Óskarsdóttir, Cristián Bravo, Wouter Verbeke, Carlos Sarraute, Bart Baesens, Jan Vanthienen |
ASONAM | 3 |
| 2014 | Predicting online channel acceptance with social network data
Thomas Verbraken, Frank G. Goethals, Wouter Verbeke, Bart Baesens |
Decis. Support Syst. | 3 |
| 2014 | Profit optimizing customer churn prediction with Bayesian network classifiersabstractCustomer churn prediction is becoming an increasingly important business analytics problem for telecom operators. In order to increase the efficiency of customer retention campaigns, churn prediction models need to be accurate as well as compact and Thomas Verbraken, Wouter Verbeke, Bart Baesens |
Intell. Data Anal. | 2 |
| 2013 | A Novel Profit Maximizing Metric for Measuring Classification Performance of Customer Churn Prediction ModelsabstractThe interest for data mining techniques has increased tremendously during the past decades, and numerous classification techniques have been applied in a wide range of business applications. Hence, the need for adequate performance measures has become more important than ever. In this paper, a cost-benefit analysis framework is formalized in order to define performance measures which are aligned with the main objectives of the end users, i.e., profit maximization. A new performance measure is defined, the expected maximum profit criterion. This general framework is then applied to the customer churn problem with its particular cost-benefit structure. The advantage of this approach is that it assists companies with selecting the classifier which maximizes the profit. Moreover, it aids with the practical implementation in the sense that it provides guidance about the fraction of the customer base to be included in the retention campaign. Thomas Verbraken, Wouter Verbeke, Bart Baesens |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2012 | Data Mining Techniques for Software Effort Estimation: A Comparative StudyabstractA predictive model is required to be accurate and comprehensible in order to inspire confidence in a business setting. Both aspects have been assessed in a software effort estimation setting by previous studies. However, no univocal conclusion as to which technique is the most suited has been reached. This study addresses this issue by reporting on the results of a large scale benchmarking study. Different types of techniques are under consideration, including techniques inducing tree/rule-based models like M5 and CART, linear models such as various types of linear regression, nonlinear models (MARS, multilayered perceptron neural networks, radial basis function networks, and least squares support vector machines), and estimation techniques that do not explicitly induce a model (e.g., a case-based reasoning approach). Furthermore, the aspect of feature subset selection by using a generic backward input selection wrapper is investigated. The results are subjected to rigorous statistical testing and indicate that ordinary least squares regression in combination with a logarithmic transformation performs best. Another key finding is that by selecting a subset of highly predictive attributes such as project size, development, and environment related attributes, typically a significant increase in estimation accuracy can be obtained. Karel Dejaeger, Wouter Verbeke, David Martens, Bart Baesens |
IEEE Trans. Software Eng. | 2 |
| 2011 | Performance of classification models from a user perspective
David Martens, Jan Vanthienen, Wouter Verbeke, Bart Baesens |
Decis. Support Syst. | 3 |
| 2011 | Building comprehensible customer churn prediction models with advanced rule induction techniques
Wouter Verbeke, David Martens, Christophe Mues, Bart Baesens |
Expert Syst. Appl. | 1 |
| 2010 | Software Effort Prediction Using Regression Rule Extraction from Neural NetworksabstractNeural networks are often selected as tool for software effort prediction because of their capability to approximate any continuous function with arbitrary accuracy. A major drawback of neural networks is the complex mapping between inputs and output, which is not easily understood by a user. This paper describes a rule extraction technique that derives a set of comprehensible IF-THEN rules from a trained neural network applied to the domain of software effort prediction. The suitability of this technique is tested on the ISBSG R11 data set by a comparison with linear regression, radial basis function networks, and CART. It is found that the most accurate results are obtained by CART, though the large number of rules limits comprehensibility. Considering comprehensible models only, the concise set of extracted rules outperform the pruned CART tree, making neural network rule extraction the most suitable technique for software effort prediction when comprehensibility is important. Rudy Setiono, Karel Dejaeger, Wouter Verbeke, David Martens, Bart Baesens |
ICTAI (2) | 3 |