Seppe K. L. M. vanden Broucke

dblp:117/1587 · also Seppe vanden Broucke · DBLP profile ↗
← Back
22ranked-venue papers in the field
3as first author
12since 2021 · last 2026
0000-0002-8781-3906ORCID · verified

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 8 (1 first)Database Systems & Data Management · 5 (1 first)Knowledge Engineering, Semantic Web & Information Systems · 5Business Process & Enterprise Data · 3 (1 first)Information Retrieval & Web Search · 1
YearPublicationVenuePosition
2026 When subgraphs outperform graphs: a scalable training strategy for churn prediction on large class-imbalanced networks
Yameng Guo, Seppe K. L. M. vanden Broucke
Data Min. Knowl. Discov.2
2025 On the Performance of LLMs for Real Estate Appraisal
Margot Geerts, Manon Reusens, Bart Baesens, Seppe K. L. M. vanden Broucke, Jochen De Weerdt
ECML/PKDD (9)4
2024 SuTraN: an Encoder-Decoder Transformer for Full-Context-Aware Suffix Prediction of Business Processes
abstract
Predictive Process Monitoring (PPM) in Process Mining (PM) uses predictive analytics to forecast business process progression. A key challenge is suffix prediction, forecasting future event sequences with activity labels, timestamps, and remaining runtime. Current techniques often focus on one-step-ahead predictions, rely on iterative feedback loops for suffix generation, and underutilize payload data. Additionally, many lag behind recent advancements in model architectures, sticking to LSTM-based models. Addressing these gaps, we propose SuTraN, a novel transformer architecture for PPM suffix prediction. SuTraN avoids iterative prediction loops and utilizes all available data, including event features, to forecast entire event suffixes in a single forward pass. Our approach integrates autoregressive suffix generation, data awareness, and seq2seq learning. Experimental results on real-life event logs demonstrate SuTraN’s superior performance in suffix prediction, highlighting its contributions often overlooked in current research.
Brecht Wuyts, Seppe K. L. M. vanden Broucke, Jochen De Weerdt
ICPM2
2024 GeoRF: a geospatial random forest
Margot Geerts, Seppe K. L. M. vanden Broucke, Jochen De Weerdt
Data Min. Knowl. Discov.2
2024 Enhancing geospatial prediction models with feature engineering from road networks: a graph-driven approach
abstract
Traditional geospatial predictive models for property valuation have naturally relied on coordinates as well as ‘hedonic’ (internal and external) features. In particular, location-centric methods such as Geographically Weighted Regression (GWR) and Kriging have focused on intrinsic target characteristics together with distances between individual targets. However, especially in the context of heavily urbanized areas, these approaches might overlook crucial aspects arising from the underlying topological structure that presents itself in such areas. Concretely, in this work, we focus on the structure arising from the road network connecting properties. We introduce a novel though straightforward technique for feature engineering based on graphs constructed on a road network. We then extract relevant features from these and utilize those as inputs for predictive models and assess their performance benefits when used together with a variety of both well-known geospatial models and state-of-art machine learning models. To this end, we present an exhaustive experiment using four different real-life data sets across various regions and exhibiting sizes outperforming many comparative works in the field. Our findings reveal that our feature engineering approach offers significant improvements in predictive performance. Finally, we apply Shapley values as an interpretability technique to confirm the reliability and effectiveness of our approach.
Yameng Guo, Seppe K. L. M. vanden Broucke
Int. J. Geogr. Inf. Sci.2
2024 LS-ICE: A Load State Intercase Encoding framework for improved predictive monitoring of business processes
Björn Rafn Gunnarsson, Seppe K. L. M. vanden Broucke, Jochen De Weerdt
Inf. Syst.2
2024 Validation set sampling strategies for predictive process monitoring
Jari Peeperkorn, Seppe K. L. M. vanden Broucke, Jochen De Weerdt
Inf. Syst.2
2023 A two-step anomaly detection based method for PU classification in imbalanced data sets
Carlos Ortega Vázquez, Seppe K. L. M. vanden Broucke, Jochen De Weerdt
Data Min. Knowl. Discov.2
2023 Regularization oversampling for classification tasks: To exploit what you do not know
Lennert Van der Schraelen, Kristof Stouthuysen, Seppe K. L. M. vanden Broucke, Tim Verdonck
Inf. Sci.3
2023 Can recurrent neural networks learn process model structure?
Jari Peeperkorn, Seppe K. L. M. vanden Broucke, Jochen De Weerdt
J. Intell. Inf. Syst.2
2022 A GAN-based hybrid sampling method for imbalanced customer classification
Bing Zhu 0005, Seppe K. L. M. vanden Broucke, Jin Xiao 0003
Inf. Sci.3
2021 Expert-driven trace clustering with instance-level constraints
Pieter De Koninck, Klaas Nelissen, Seppe K. L. M. vanden Broucke, Bart Baesens, Monique Snoeck, Jochen De Weerdt
Knowl. Inf. Syst.3
2018 Incorporating negative information to process discovery of complex systems
Hernán Ponce de León, Lucio Nardelli, Josep Carmona 0001, Seppe K. L. M. vanden Broucke
Inf. Sci.4
2018 Swipe and Tell: Using Implicit Feedback to Predict User Engagement on Tablets
abstract
When content consumers explicitly judge content positively, we consider them to be engaged. Unfortunately, explicit user evaluations are difficult to collect, as they require user effort. Therefore, we propose to use device interactions as implicit feedback to detect engagement. We assess the usefulness of swipe interactions on tablets for predicting engagement and make the comparison with using traditional features based on time spent. We gathered two unique datasets of more than 250,000 swipes, 100,000 unique article visits, and over 35,000 explicitly judged news articles by modifying two commonly used tablet apps of two newspapers. We tracked all device interactions of 407 experiment participants during one month of habitual news reading. We employed a behavioral metric as a proxy for engagement, because our analysis needed to be scalable to many users, and scanning behavior required us to allow users to indicate engagement quickly. We point out the importance of taking into account content ordering, report the most predictive features, zoom in on briefly read content and on the most frequently read articles. Our findings demonstrate that fine-grained tablet interactions are useful indicators of engagement for newsreaders on tablets. The best features successfully combine both time-based aspects and swipe interactions.
Klaas Nelissen, Monique Snoeck, Seppe K. L. M. vanden Broucke, Bart Baesens
ACM Trans. Inf. Syst.3
2017 An Approach for Incorporating Expert Knowledge in Trace Clustering
Pieter De Koninck, Klaas Nelissen, Bart Baesens, Seppe K. L. M. vanden Broucke, Monique Snoeck, Jochen De Weerdt
CAiSE4
2017 Explaining clusterings of process instances
Pieter De Koninck, Jochen De Weerdt, Seppe K. L. M. vanden Broucke
Data Min. Knowl. Discov.3
2017 An empirical comparison of techniques for the class imbalance problem in churn prediction
Bing Zhu 0005, Bart Baesens, Seppe K. L. M. vanden Broucke
Inf. Sci.3
2015 Profit maximizing logistic regression modeling for customer churn prediction
abstract
The selection of classifiers which are profitable is becoming more and more important in real-life situations such as customer churn management campaigns in the telecommunication sector. In previous works, the expected maximum profit (EMP) metric has been proposed, which explicitly takes the cost of offer and the customer lifetime value (CLV) of retained customers into account. It thus permits the selection of the most profitable classifier, which better aligns with business requirements of end-users and stake holders. However, modelers are currently limited to applying this metric in the evaluation step. Hence, we expand on the previous body of work and introduce a classifier that incorporates the EMP metric in the construction of a classification model. Our technique, called ProfLogit, explicitly takes profit maximization concerns into account during the training step, rather than the evaluation step. The technique is based on a logistic regression model which is trained using a genetic algorithm (GA). By means of an empirical benchmark study applied to real-life data sets, we show that ProfLogit generates substantial profit improvements compared to the classic logistic model for many data sets. In addition, profit-maximized coefficient estimates differ considerably in magnitude from the maximum likelihood estimates.
Eugen Stripling, Seppe K. L. M. vanden Broucke, Katrien Antonio, Bart Baesens, Monique Snoeck
DSAA2
2014 Determining Process Model Precision and Generalization with Weighted Artificial Negative Events
abstract
Process mining encompasses the research area which is concerned with knowledge discovery from event logs. One common process mining task focuses on conformance checking, comparing discovered or designed process models with actual real-life behavior as captured in event logs in order to assess the “goodness” of the process model. This paper introduces a novel conformance checking method to measure how well a process model performs in terms of precision and generalization with respect to the actual executions of a process as recorded in an event log. Our approach differs from related work in the sense that we apply the concept of so-called weighted artificial negative events toward conformance checking, leading to more robust results, especially when dealing with less complete event logs that only contain a subset of all possible process execution behavior. In addition, our technique offers a novel way to estimate a process model's ability to generalize. Existing literature has focused mainly on the fitness (recall) and precision (appropriateness) of process models, whereas generalization has been much more difficult to estimate. The described algorithms are implemented in a number of ProM plugins, and a Petri net conformance checking tool was developed to inspect process model conformance in a visual manner.
Seppe K. L. M. vanden Broucke, Jochen De Weerdt, Jan Vanthienen, Bart Baesens
IEEE Trans. Knowl. Data Eng.1
2013 A comprehensive benchmarking framework (CoBeFra) for conformance analysis between procedural process models and event logs in ProM
abstract
Process mining encompasses the research area which is concerned with knowledge discovery from information system event logs. Within the process mining research area, two prominent tasks can be discerned. First of all, process discovery deals with the automatic construction of a process model out of an event log. Secondly, conformance checking focuses on the assessment of the quality of a discovered or designed process model in respect to the actual behavior as captured in event logs. Hereto, multiple techniques and metrics have been developed and described in the literature. However, the process mining domain still lacks a comprehensive framework for assessing the goodness of a process model from a quantitative perspective. In this study, we describe the architecture of an extensible framework within ProM, allowing for the consistent, comparative and repeatable calculation of conformance metrics. For the development and assessment of both process discovery as well as conformance techniques, such a framework is considered greatly valuable.
Seppe K. L. M. vanden Broucke, Jochen De Weerdt, Jan Vanthienen, Bart Baesens
CIDM1
2013 Active Trace Clustering for Improved Process Discovery
abstract
Process discovery is the learning task that entails the construction of process models from event logs of information systems. Typically, these event logs are large data sets that contain the process executions by registering what activity has taken place at a certain moment in time. By far the most arduous challenge for process discovery algorithms consists of tackling the problem of accurate and comprehensible knowledge discovery from highly flexible environments. Event logs from such flexible systems often contain a large variety of process executions which makes the application of process mining most interesting. However, simply applying existing process discovery techniques will often yield highly incomprehensible process models because of their inaccuracy and complexity. With respect to resolving this problem, trace clustering is one very interesting approach since it allows to split up an existing event log so as to facilitate the knowledge discovery process. In this paper, we propose a novel trace clustering technique that significantly differs from previous approaches. Above all, it starts from the observation that currently available techniques suffer from a large divergence between the clustering bias and the evaluation bias. By employing an active learning inspired approach, this bias divergence is solved. In an assessment using four complex, real-life event logs, it is shown that our technique significantly outperforms currently available trace clustering techniques.
Jochen De Weerdt, Seppe K. L. M. vanden Broucke, Jan Vanthienen, Bart Baesens
IEEE Trans. Knowl. Data Eng.2
2012 Improved Artificial Negative Event Generation to Enhance Process Event Logs
Seppe K. L. M. vanden Broucke, Jochen De Weerdt, Bart Baesens, Jan Vanthienen
CAiSE1