Jochen De Weerdt

dblp:41/9119 · DBLP profile ↗
← Back
38ranked-venue papers in the field
3as first author
19since 2021 · last 2026
0000-0001-6151-0504ORCID · verified

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 14 (2 first)Data Mining & Knowledge Discovery · 12 (1 first)Business Process & Enterprise Data · 9Knowledge Engineering, Semantic Web & Information Systems · 2Information Retrieval & Web Search · 1
YearPublicationVenuePosition
2026 Time Series Foundation Models for Process Model Forecasting
Jari Peeperkorn, Johannes De Smedt, Jochen De Weerdt
CAiSE (1)4
2026 Model-driven stochastic trace clustering
Jari Peeperkorn, Johannes De Smedt, Jochen De Weerdt
Inf. Syst.3
2026 Object-centric process management: A research manifesto
abstract
Business process management employs process models and event logs to represent the behavior of the information systems under study. Traditional case-centric notions consider the order of activities and events in isolated process instances. The emerging field of object-centric processes challenges this assumption by putting objects in the center. Object-centric process mining and modeling approaches identify the structure of co-evolving data objects that influence the behavior of an information system to provide a comprehensive view of the system behavior. Object-centricity has been investigated independently in process modeling and in process mining, which resulted in the coexistence of seemingly contradictory assumptions and definitions. As a community effort, this research manifesto relates and aligns existing terminologies, definitions, and perspectives to provide a common ground for current and future research in object-centric business process management. Based on the current state of research, we propose a conceptualization that sets process models and event logs in relation to the information system’s behavior and the execution data it generates. The conceptualization aims at aligning different terminologies and, thus, providing a basis to model and analyze behavioral characteristics. Building on this common ground, we identify open research challenges along the most relevant research areas in object-centric process management. For each research area, its current status is investigated and an outline of the most relevant research challenges is presented.
Anjo Seidel, Mathias Weske, Marco Montali, Andrey Rivkin, Manfred Reichert, Jan Martijn E. M. van der Werf, Wil M. P. van der Aalst, Marius Breitmayer, Lukas Liß, Jan Niklas van Detten, Amin Jalali 0001, Shahrzad Khayatbashi, Maximilian König, Tom Lichtenstein, Stefanie Rinderle-Ma, Barbara Weber, Pnina Soffer, Lorenzo Rossi 0001, Daniel Calegari, Andrea Delgado 0001, Remco M. Dijkman, Sarah Winkler, Matthias Weidlich 0001, Sander J. J. Leemans, Dirk Fahland, Ava Swevels, Monique Snoeck, Giancarlo Guizzardi, Alessandro Gianola, Avigdor Gal, Ekkart Kindler, Irina A. Lomazova, Barbara Re 0001, Giovanni Meroni, Andrea Morichetta 0001, Alessandro Marcelletti, Sara Pettinari, Boudewijn F. van Dongen, Johannes De Smedt, Majid Rafiei, Julius Köpke, Thomas T. Hildebrandt, Francesca Zerbato, Luise Pufahl, Hajo A. Reijers, Artem Polyvyanyy, Chiara Di Francescomarino, Fabrizio Maria Maggi, Oscar Pastor 0001, Stephan Haarmann, Henderik A. Proper, Xixi Lu 0001, Hugo A. López 0001, Tijs Slaats, Jochen De Weerdt, Massimiliano de Leoni, Niels Martin, Karolin Winter, Nick R. T. P. van Beest, Orlenys López-Pintado, Sebastiaan J. van Zelst, Chiara Ghidini, Arik Senderovich
Inf. Syst.55
2025 On the Performance of LLMs for Real Estate Appraisal
Margot Geerts, Manon Reusens, Bart Baesens, Seppe K. L. M. vanden Broucke, Jochen De Weerdt
ECML/PKDD (9)5
2024 SuTraN: an Encoder-Decoder Transformer for Full-Context-Aware Suffix Prediction of Business Processes
abstract
Predictive Process Monitoring (PPM) in Process Mining (PM) uses predictive analytics to forecast business process progression. A key challenge is suffix prediction, forecasting future event sequences with activity labels, timestamps, and remaining runtime. Current techniques often focus on one-step-ahead predictions, rely on iterative feedback loops for suffix generation, and underutilize payload data. Additionally, many lag behind recent advancements in model architectures, sticking to LSTM-based models. Addressing these gaps, we propose SuTraN, a novel transformer architecture for PPM suffix prediction. SuTraN avoids iterative prediction loops and utilizes all available data, including event features, to forecast entire event suffixes in a single forward pass. Our approach integrates autoregressive suffix generation, data awareness, and seq2seq learning. Experimental results on real-life event logs demonstrate SuTraN’s superior performance in suffix prediction, highlighting its contributions often overlooked in current research.
Brecht Wuyts, Seppe K. L. M. vanden Broucke, Jochen De Weerdt
ICPM3
2024 GeoRF: a geospatial random forest
Margot Geerts, Seppe K. L. M. vanden Broucke, Jochen De Weerdt
Data Min. Knowl. Discov.3
2024 LS-ICE: A Load State Intercase Encoding framework for improved predictive monitoring of business processes
Björn Rafn Gunnarsson, Seppe K. L. M. vanden Broucke, Jochen De Weerdt
Inf. Syst.3
2024 Validation set sampling strategies for predictive process monitoring
Jari Peeperkorn, Seppe K. L. M. vanden Broucke, Jochen De Weerdt
Inf. Syst.3
2024 SHINE: A Scalable Heterogeneous Inductive Graph Neural Network for Large Imbalanced Datasets
abstract
Research interest in machine learning (ML) for graphs has skyrocketed in recent years. However, non-Euclidean graph structures inhibit the application of traditional ML algorithms. Consequently, scholars introduced graph learning algorithms tailored to network data, such as graph neural networks (GNNs). Most GNNs are designed for homogeneous and homophilous graphs and are evaluated on small, static, and balanced datasets, deviating from real-world conditions and industry applications. This paper introduces SHINE, a scalable heterogeneous inductive GNN for large imbalanced datasets. SHINE addresses four key challenges: scalability, network heterogeneity, inductive learning on dynamic graphs, and imbalanced node classification. SHINE comprises three core components: 1) a sampler based on nearest-neighbor (NN) search, 2) a heterogeneous GNN (HGNN) layer with a novel relationship aggregator, and 3) aggregator functions tailored to skewed class distributions. The components of SHINE are evaluated on benchmark datasets, while the integrated benefits of SHINE are demonstrated on two fraud detection datasets.
Rafaël Van Belle, Jochen De Weerdt
IEEE Trans. Knowl. Data Eng.2
2023 Manifold Learning for Adversarial Robustness in Predictive Process Monitoring
abstract
In recent years, many predictive models have been successfully applied to predictive process monitoring, enabling tasks such as predicting the next activity, remaining time, or the future state of a process instance (case). However, recent developments have shown the vulnerability of these models to adversarial attacks, causing algorithms to make incorrect predictions. This paper addresses this issue by leveraging adversarial examples to evaluate the predictive performance of predictive process monitoring models in the face of adversarial threats. Although augmenting training data with adversarial examples has proven effective in defending against specific adversarial attacks, it often remains insufficient in mitigating vulnerabilities to other types of attacks. Our proposed approach explores the use of manifold learning techniques to restrict these examples within the range of data on which the model is trained. By learning from these specifically engineered (hidden) attacks, we seek to develop models that maintain accuracy on new, unseen data while effectively improving adversarial robustness against potential threats.
Alexander Stevens, Jari Peeperkorn, Johannes De Smedt, Jochen De Weerdt
ICPM4
2023 A two-step anomaly detection based method for PU classification in imbalanced data sets
Carlos Ortega Vázquez, Seppe K. L. M. vanden Broucke, Jochen De Weerdt
Data Min. Knowl. Discov.3
2023 Process model forecasting and change exploration using time series analysis of event sequence data
abstract
Process analytics is a collection of data-driven techniques for, among others, making predictions for individual process instances or overall process models. At the instance level, various novel techniques have been recently devised, tackling analytical tasks such as next activity, remaining time, or outcome prediction. However, there is a notable void regarding predictions at the process model level. It is the ambition of this article to fill this gap. More specifically, we develop a technique to forecast the entire process model from historical event data. A forecasted model is a will-be process model representing a probable description of the overall process for a given period in the future. Such a forecast helps, for instance, to anticipate and prepare for the consequences of upcoming process drifts and emerging bottlenecks. Our technique builds on a representation of event data as multiple time series, each capturing the evolution of a behavioural aspect of the process model, such that corresponding time series forecasting techniques can be applied. Our implementation demonstrates the feasibility of process model forecasting using real-world event data. A user study using our Process Change Exploration tool confirms the usefulness and ease of use of the produced process model forecasts.
Johannes De Smedt, Anton Yeshchenko, Artem Polyvyanyy, Jochen De Weerdt, Jan Mendling
Data Knowl. Eng.4
2023 Can recurrent neural networks learn process model structure?
Jari Peeperkorn, Seppe K. L. M. vanden Broucke, Jochen De Weerdt
J. Intell. Inf. Syst.3
2022 Evaluation of Joint Modeling Techniques for Node Embedding and Community Detection on Graphs
abstract
Novel joint techniques capture both the microscopic context and the mesoscopic structure of networks by leveraging two previously separated fields of research: node representation learning (NRL) and community detection (CD). However, several limitations exist in the literature. First, a comprehensive comparison between these joint NRL-CD techniques is non-existent. Second, baseline techniques, datasets, evaluation metrics, and classification algorithms differ significantly between each method. Thirdly, the literature lacks a synchronized experimental approach, thus rendering comparison between these methods strenuous. To overcome these limitations, we present a uni-fied experimental setup mutually comparing six joint NRL-CD techniques and comparing them with corresponding NRL/CD baselines in three different settings: non-overlapping and over-lapping CD and node classification. Our results show that joint methods underperform on the node classification task but achieve relatively solid results for overlapping community detection. Our research contribution is two-fold: first, we show specific weaknesses of selected joint techniques in different tasks and data sets; and second, we suggest a more thorough experimental setup to benchmark joint techniques with simpler NRL and CD techniques.
Simon Hiel, Lore Nicolaers, Carlos Ortega Vázquez, Sandra Mitrovic, Bart Baesens, Jochen De Weerdt
ASONAM6
2022 Assessing the Robustness in Predictive Process Monitoring through Adversarial Attacks
abstract
As machine and deep learning models are increasingly leveraged in predictive process monitoring, the focus has shifted towards making these models explainable. The successful adoption of a model is dependent on whether decision-makers can trust the predictions and explanations made. However, recent studies have shown that deep learning models are vulnerable to adversarial attacks -small perturbations to the inputs-which trick deep learning algorithms into making incorrect predictions. An additional crucial property is that the explanations are robust against these adversarial attacks when the model decision was not affected. Therefore, this paper introduces a robustness assessment framework by investigating the impact of adversarial attacks on the robustness of predictive accuracy and explanations used in the field of predictive process monitoring. First, adversarial examples of cases in the independent test set are generated to examine the robustness of the predictive model against intentionally manipulated data. Next, the predictive models are compared with similar models trained on data imputed with adversarial attacks. We monitor the impact on predictive performance in terms of AUC at different stages of the case execution. Finally, the robustness of the explanations is calculated as the distance between the original explanations and the explanations extracted from the model trained on attacked data. We test multiple machine and deep learning techniques, namely the transparent logistic regression, random forests with Shapley values, and LSTM neural networks with attention. Results show that especially neural networks suffer from adversarial attacks, and the former two are mostly robust in terms of both predictive accuracy and explanations.
Alexander Stevens, Johannes De Smedt, Jari Peeperkorn, Jochen De Weerdt
ICPM4
2021 Process Model Forecasting Using Time Series Analysis of Event Sequence Data
Johannes De Smedt, Anton Yeshchenko, Artem Polyvyanyy, Jochen De Weerdt, Jan Mendling
ER4
2021 Conformance Checking in Process Mining
Mieke Jans, Jochen De Weerdt, Benoît Depaire, Marlon Dumas, Gert Janssenswillen
Inf. Syst.2
2021 tcc2vec: RFM-informed representation learning on call graphs for churn prediction
Sandra Mitrovic, Bart Baesens, Wilfried Lemahieu, Jochen De Weerdt
Inf. Sci.4
2021 Expert-driven trace clustering with instance-level constraints
Pieter De Koninck, Klaas Nelissen, Seppe K. L. M. vanden Broucke, Bart Baesens, Monique Snoeck, Jochen De Weerdt
Knowl. Inf. Syst.6
2020 Churn modeling with probabilistic meta paths-based representation learning
Sandra Mitrovic, Jochen De Weerdt
Inf. Process. Manag.2
2020 Mining Behavioral Sequence Constraints for Classification
abstract
Sequence classification deals with the task of finding discriminative and concise sequential patterns. To this purpose, many techniques have been proposed, which mainly resort to the use of partial orders to capture the underlying sequences in a database according to the labels. Partial orders, however, pose many limitations, especially on expressiveness, i.e., the aptitude towards capturing certain behavior, and on conciseness, i.e., doing so in a compact and informative way. These limitations can be addressed by using a better representation. In this paper, we present the interesting Behavioral Constraint Miner (iBCM), a sequence classification technique that discovers patterns using behavioral constraint templates. The templates comprise a variety of constraints and can express patterns ranging from simple occurrence, to looping and position-based behavior over a sequence. Furthermore, iBCM also captures negative constraints, i.e., absence of particular behavior. The constraints can be discovered by using simple string operations in an efficient way. Finally, deriving the constraints with a window-based approach allows to pinpoint where the constraints hold in a string, and to detect whether patterns are subject to concept drift. Through empirical evaluation, it is shown that iBCM is better capable of classifying sequences more accurately and concisely in a scalable manner.
Johannes De Smedt, Galina Deeva, Jochen De Weerdt
IEEE Trans. Knowl. Data Eng.3
2019 A comparison of methods for link sign prediction with signed network embeddings
abstract
In many real-world networks, it is important to explicitly differentiate between positive and negative links, thus considering the observed networks as signed. To derive useful features, just as in the case of unsigned networks, representation learning can be used to learn meaningful representations of a network that characterize its underlying topology. Several methods for learning representations on signed networks have already been proposed but have not been systematically benchmarked together before. Hence, in this paper, we bridge this literature gap providing a quantitative and qualitative benchmark of the four most prominent representation learning methods for signed networks. Results on three different datasets for link sign prediction showcase the superiority of the StEM method over its competitors both from a predictive performance and runtime perspective.
Sandra Mitrovic, Laurent Lecoutere, Jochen De Weerdt
ASONAM3
2019 Scalable Mixed-Paradigm Trace Clustering using Super-Instances
abstract
In process mining, one is often confronted with datasets that contain high degrees of control-flow variation. This causes the results of subsequent process mining steps to be less accurate or harder to understand. Trace clustering, or dividing the log into more homogeneous groups, can help limit this issue. Trace clustering techniques are commonly subdivided into two types: distance-based techniques and model-driven approaches. In this paper, a novel trace clustering technique is presented, that combines aspects of both paradigms. The core idea is to first learn so-called super-instances using a simple distance-driven technique and subsequently apply a model-driven technique to cluster the super-instances. Our technique not only shows qualitative improvements of the obtained clusterings, but also significantly improves the scalability, as shown in an experimental evaluation using real-life event logs.
Pieter De Koninck, Jochen De Weerdt
ICPM2
2018 Combining Temporal Aspects of Dynamic Networks with Node2Vec for a more Efficient Dynamic Link Prediction
abstract
In many real-life applications it is crucial to be able to, given a collection of link states of a network in a certain time period, accurately predict the link state of the network at a future time. This is known as dynamic link prediction, which compared to its static counterpart is more complex, as capturing the temporal characteristics is a non-trivial task. This explains while still majority of today's research in network representation learning focuses on static setting ignoring temporal information. In this work, we focus on one such case and aim at extending node2vec, representation learning method successfully applied for static link prediction, to a dynamic setup. This extended method is applied and validated on several real-life networks with different properties. Results show that taking into account dynamic aspect outperforms static approach. Additionally, based on the network properties, recommendations are given for the node2vec parameters.
Sam De Winter, Tim Decuypere, Sandra Mitrovic, Bart Baesens, Jochen De Weerdt
ASONAM5
2018 Discovering hidden dependencies in constraint-based declarative process models for improving understandability
Johannes De Smedt, Jochen De Weerdt, Estefanía Serral, Jan Vanthienen
Inf. Syst.2
2017 An Approach for Incorporating Expert Knowledge in Trace Clustering
Pieter De Koninck, Klaas Nelissen, Bart Baesens, Seppe K. L. M. vanden Broucke, Monique Snoeck, Jochen De Weerdt
CAiSE6
2017 Scalable RFM-enriched Representation Learning for Churn Prediction
abstract
Most of the recent studies on churn prediction in telco utilize social networks built on top of the call (and/or SMS) graphs to derive informative features. However, extracting features from large graphs, especially structural features, is an intricate process both from a methodological and computational perspective. Due to the former, feature extraction in the current literature has mainly been addressed in an ad-hoc and hand-crafted manner. Due to the latter, the full potential of the structural information is unexploited. In this work, we incorporate both interaction and structural information by devising two different ways of enriching original graphs with interaction information, delineated by the well-known RFM model. We circumvent the process of extensive manual feature engineering by enriching the networks and improving the scalability of the renowned node2vec approach to learn node representations. The obtained results demonstrate that our enriched network outperforms baseline RFM-based methods.
Sandra Mitrovic, Gaurav Singh 0001, Bart Baesens, Wilfried Lemahieu, Jochen De Weerdt
DSAA5
2017 Behavioral Constraint Template-Based Sequence Classification
Johannes De Smedt, Galina Deeva, Jochen De Weerdt
ECML/PKDD (2)3
2017 Explaining clusterings of process instances
Pieter De Koninck, Jochen De Weerdt, Seppe K. L. M. vanden Broucke
Data Min. Knowl. Discov.2
2017 Change visualisation: Analysing the resource and timing differences between two event logs
Wei Zhe Low, Wil M. P. van der Aalst, Arthur H. M. ter Hofstede, Moe Thandar Wynn, Jochen De Weerdt
Inf. Syst.5
2016 Improving Understandability of Declarative Process Models by Revealing Hidden Dependencies
Johannes De Smedt, Jochen De Weerdt, Estefanía Serral, Jan Vanthienen
CAiSE2
2014 Controlled automated discovery of collections of business process models
Luciano García-Bañuelos, Marlon Dumas, Marcello La Rosa, Jochen De Weerdt, Chathura C. Ekanayake
Inf. Syst.4
2014 Determining Process Model Precision and Generalization with Weighted Artificial Negative Events
abstract
Process mining encompasses the research area which is concerned with knowledge discovery from event logs. One common process mining task focuses on conformance checking, comparing discovered or designed process models with actual real-life behavior as captured in event logs in order to assess the “goodness” of the process model. This paper introduces a novel conformance checking method to measure how well a process model performs in terms of precision and generalization with respect to the actual executions of a process as recorded in an event log. Our approach differs from related work in the sense that we apply the concept of so-called weighted artificial negative events toward conformance checking, leading to more robust results, especially when dealing with less complete event logs that only contain a subset of all possible process execution behavior. In addition, our technique offers a novel way to estimate a process model's ability to generalize. Existing literature has focused mainly on the fitness (recall) and precision (appropriateness) of process models, whereas generalization has been much more difficult to estimate. The described algorithms are implemented in a number of ProM plugins, and a Petri net conformance checking tool was developed to inspect process model conformance in a visual manner.
Seppe K. L. M. vanden Broucke, Jochen De Weerdt, Jan Vanthienen, Bart Baesens
IEEE Trans. Knowl. Data Eng.2
2013 A comprehensive benchmarking framework (CoBeFra) for conformance analysis between procedural process models and event logs in ProM
abstract
Process mining encompasses the research area which is concerned with knowledge discovery from information system event logs. Within the process mining research area, two prominent tasks can be discerned. First of all, process discovery deals with the automatic construction of a process model out of an event log. Secondly, conformance checking focuses on the assessment of the quality of a discovered or designed process model in respect to the actual behavior as captured in event logs. Hereto, multiple techniques and metrics have been developed and described in the literature. However, the process mining domain still lacks a comprehensive framework for assessing the goodness of a process model from a quantitative perspective. In this study, we describe the architecture of an extensible framework within ProM, allowing for the consistent, comparative and repeatable calculation of conformance metrics. For the development and assessment of both process discovery as well as conformance techniques, such a framework is considered greatly valuable.
Seppe K. L. M. vanden Broucke, Jochen De Weerdt, Jan Vanthienen, Bart Baesens
CIDM2
2013 Active Trace Clustering for Improved Process Discovery
abstract
Process discovery is the learning task that entails the construction of process models from event logs of information systems. Typically, these event logs are large data sets that contain the process executions by registering what activity has taken place at a certain moment in time. By far the most arduous challenge for process discovery algorithms consists of tackling the problem of accurate and comprehensible knowledge discovery from highly flexible environments. Event logs from such flexible systems often contain a large variety of process executions which makes the application of process mining most interesting. However, simply applying existing process discovery techniques will often yield highly incomprehensible process models because of their inaccuracy and complexity. With respect to resolving this problem, trace clustering is one very interesting approach since it allows to split up an existing event log so as to facilitate the knowledge discovery process. In this paper, we propose a novel trace clustering technique that significantly differs from previous approaches. Above all, it starts from the observation that currently available techniques suffer from a large divergence between the clustering bias and the evaluation bias. By employing an active learning inspired approach, this bias divergence is solved. In an assessment using four complex, real-life event logs, it is shown that our technique significantly outperforms currently available trace clustering techniques.
Jochen De Weerdt, Seppe K. L. M. vanden Broucke, Jan Vanthienen, Bart Baesens
IEEE Trans. Knowl. Data Eng.1
2012 Improved Artificial Negative Event Generation to Enhance Process Event Logs
Seppe K. L. M. vanden Broucke, Jochen De Weerdt, Bart Baesens, Jan Vanthienen
CAiSE2
2012 A multi-dimensional quality assessment of state-of-the-art process discovery algorithms using real-life event logs
Jochen De Weerdt, Manu De Backer, Jan Vanthienen, Bart Baesens
Inf. Syst.1
2011 A robust F-measure for evaluating discovered process models
abstract
Within process mining research, one of the most important fields of study is process discovery, which can be defined as the extraction of control-flow models from audit trails or information system event logs. The evaluation of discovered process models is an essential but difficult task for any process discovery analysis. With this paper, we propose a novel approach for evaluating discovered process models based on artificially generated negative events. This approach allows for the definition of a behavioral F-measure for discovered process models, which is the main contribution of this paper.
Jochen De Weerdt, Manu De Backer, Jan Vanthienen, Bart Baesens
CIDM1