VLDB 2026 Research / reviewers in the wild / expert
Han van der Aa
dblp:132/6986
· DBLP profile ↗
41ranked-venue papers in the field
9as first author
27since 2021 · last 2026
0000-0002-4200-4937ORCID · verified
Domains — venue-derived; a paper can count in several
Business Process & Enterprise Data · 21 (4 first)Database Systems & Data Management · 19 (5 first)Knowledge Engineering, Semantic Web & Information Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enabling Small Language Models for Text-to-Process Extraction: Balancing Accuracy and Efficiency Through Distant Supervision
Julian Neuberger, Han van der Aa, Ivan Khrop, Hugo A. López 0001 |
CAiSE (2) | 2 |
| 2026 | Version Clustering: A Top-Down Approach for Process Concept Drift Detection
Bernold Rodrigo Abarca Zúñiga, Anton Yeshchenko, Han van der Aa |
CAiSE (1) | 3 |
| 2026 | Comprehensive characterization of concept drifts in process miningabstractBusiness processes are subject to changes due to the dynamic environments in which they are executed. These process changes can lead to concept drifts, which are situations when the characteristics of a business process have undergone significant changes, resulting in event logs that contain data on different versions of a process. The accuracy and usefulness of process mining results derived from such event logs may be compromised because they rely on historical data that no longer reflects the current process behavior, or because the results do not distinguish between different process versions. Therefore, concept drift detection in process mining aims to identify drifts recorded in an event log by detecting when they occurred, localizing process modifications, and characterizing how they manifest over time. This paper focuses on the latter task, i.e., drift characterization, which seeks to understand whether changes unfolded suddenly or gradually and if they form complex patterns like incremental or recurring drifts. However, current solutions for automatically detecting concept drifts from event logs lack comprehensive characterization capabilities. Instead, they mainly focus on drift detection and characterization of isolated process changes. This leads to an incomplete understanding of more complex concept drifts, like incremental and recurring drifts, when several process changes are inter-connected. This paper overcomes such limitations by introducing an improved taxonomy for characterizing concept drifts and a three-step framework that provides an automatic characterization of concept drifts from event logs. We evaluated our framework through elaborate evaluation experiments conducted using a large collection of synthetic event logs. The results highlight the effectiveness and accuracy of our proposed framework and show that it outperforms state-of-the-art techniques. Alexander Kraus 0001, Han van der Aa |
Inf. Syst. | 2 |
| 2026 | A framework for steady-state detection in process miningabstractSteady-state detection (SSD) is a crucial task in the analysis of complex and dynamic systems, as it enables the reliable assessment of system behavior by distinguishing between stable and unstable states. SSD techniques have been extensively studied and applied in various domains, including signal processing and industrial systems. However, their application within the information systems domain, particularly in process mining, has received little attention, even though business processes themselves can be regarded as complex socio-technical systems. In particular, event logs that capture the execution of business processes often contain data from both steady and non-steady states. Mixing up these states can significantly affect the accuracy and reliability of insights from common process mining tasks, such as process performance analysis and process discovery. To address this problem, we propose a dedicated SSD framework for process mining and demonstrate how differentiating between distinct process states can enhance the accuracy and reliability of process mining insights. The SSD framework takes an event log as input and identifies the existing steady and non-steady states along with their corresponding time periods. We evaluate the framework through two experiments: one assessing accuracy using simulated event logs and another demonstrating its impact on three key process mining tasks: process performance analysis, process discovery, and remaining time prediction. Alexander Kraus 0001, Keyvan Amiri Elyasi, Adrian Rebmann, Sherri Hadian, Han van der Aa |
Inf. Syst. | 5 |
| 2025 | On the Use of Steady-State Detection for Process Mining: Achieving More Accurate Insights
Alexander Kraus 0001, Keyvan Amiri Elyasi, Han van der Aa |
CAiSE (1) | 3 |
| 2025 | LLMs that Understand Processes: Instruction-tuning for Semantics-Aware Process MiningabstractProcess mining is increasingly using textual information associated with events to tackle tasks such as anomaly detection and process discovery. Such semantics-aware process mining focuses on what behavior should be possible in a process (i.e., expectations), thus providing an important complement to traditional, frequency-based techniques that focus on recorded behavior (i.e., reality). Large Language Models (LLMs) provide a powerful means for tackling semantics-aware tasks. However, the best performance is so far achieved through task-specific finetuning, which is computationally intensive and results in models that can only handle one specific task. To overcome this lack of generalization, we use this paper to investigate the potential of instruction-tuning for semantics-aware process mining. The idea of instruction-tuning here is to expose an LLM to promptanswer pairs for different tasks, e.g., anomaly detection and nextactivity prediction, making it more familiar with process mining, thus allowing it to also perform better at unseen tasks, such as process discovery. Our findings demonstrate a varied impact of instruction-tuning: while performance considerably improved on process discovery and prediction tasks, it varies across models on anomaly detection tasks, highlighting that the selection of tasks for instruction-tuning is critical to achieving desired outcomes. Vira Pyrih, Adrian Rebmann, Han van der Aa |
ICPM | 3 |
| 2024 | PGTNet: A Process Graph Transformer Network for Remaining Time Prediction of Business Process Instances
Keyvan Amiri Elyasi, Han van der Aa, Heiner Stuckenschmidt |
CAiSE | 2 |
| 2024 | A Universal Prompting Strategy for Extracting Process Model Information from Natural Language Text Using Large Language Models
Julian Neuberger, Lars Ackermann, Han van der Aa, Stefan Jablonski |
ER | 3 |
| 2024 | Privacy-Aware Analysis based on Data SeriesabstractData that is recorded about the operations of an organization constitutes a valuable source of information for monitoring and improvement. Specific use cases include the assessment of compliance to legal regulations, the analysis of performance bottlenecks, or the optimization of resource utilization. In recent years, a plethora of algorithms for operational analysis using data series, summarized as process mining, have been developed to support these use cases, e.g., by constructing models for simulation and prediction or by comparing the recorded data against a normative specification of a process. Data series often contain sensitive information, though, about the individuals that act as service consumers or service providers. Personal information is only partially hidden by obfuscation and pseudonymization and potential privacy breaches need to be prevented for ethical, legal, and economic reasons. This tutorial is devoted to methods for privacy-aware analysis using data series. It covers essential notions, reviews privacy-disclosure attacks, and outlines techniques to give formal privacy guarantees while largely maintaining the data's utility for operational analysis. The discussion is structured by the adopted perspective on the privacy of individuals, and the degree to which a data series contains contextual information. Stephan A. Fahrenkrog-Petersen, Han van der Aa, Matthias Weidlich 0001 |
ICDE | 2 |
| 2024 | A Context Framework for Sense-making of Process Mining ResultsabstractProcess mining research has made tremendous progress in analyzing, visualizing, and predicting the performance of business processes through computational techniques. However, little attention has been brought to understanding why and how business processes behave as they do. Process mining results alone are not sufficient to arrive at meaningful interpretations about the dynamics and changes of a given business process. Rather, we need to account for contextual factors that underlie and explain the behavior of processes. In this paper, we make two central contributions. First, we develop a framework that depicts relevant factors to make sense of process mining results. The framework is intended to help researchers and practitioners explain why and how processes change across a variety of contexts. Second, we demonstrate the application of our framework within a real-world case: a customer onboarding process in a European financial institution. Thomas Grisold, Han van der Aa, Sandro Franzoi, Sophie Hartl, Jan Mendling, Jan vom Brocke |
ICPM | 2 |
| 2024 | AgentSimulator: An Agent-based Approach for Data-driven Business Process SimulationabstractBusiness process simulation (BPS) is a versatile technique for estimating process performance across various scenarios. Traditionally, BPS approaches employ a control-flow-first perspective by enriching a process model with simulation parameters. Although such approaches can mimic the behavior of centrally orchestrated processes, such as those supported by workflow systems, current control-flow-first approaches cannot faithfully capture the dynamics of real-world processes that involve distinct resource behavior and decentralized decision-making. Recognizing this issue, this paper introduces AgentSimulator, a resource-first BPS approach that discovers a multi-agent system from an event log, modeling distinct resource behaviors and interaction patterns to simulate the underlying process. Our experiments show that AgentSimulator achieves state-of-the-art simulation accuracy with significantly lower computation times than existing approaches while providing high interpretability and adaptability to different types of process-execution scenarios. Lukas Kirchdorfer, Robert Blümel, Timotheus Kampik, Han van der Aa, Heiner Stuckenschmidt |
ICPM | 4 |
| 2024 | Evaluating the Ability of LLMs to Solve Semantics-Aware Process Mining TasksabstractThe process mining community has recently recognized the potential of large language models (LLMs) for tackling various process mining tasks. Initial studies report the capability of LLMs to support process analysis and even, to some extent, that they are able to reason about how processes work. This latter property suggests that LLMs could also be used to tackle process mining tasks that benefit from an understanding of process behavior. Examples of such tasks include (semantic) anomaly detection and next activity prediction, which both involve considerations of the meaning of activities and their interrelations. In this paper, we investigate the capabilities of LLMs to tackle such semantics-aware process mining tasks. Furthermore, whereas most works on the intersection of LLMs and process mining only focus on testing these models out of the box, we provide a more principled investigation of the utility of LLMs for process mining, including their ability to obtain process mining knowledge post-hoc by means of in-context learning and supervised fine-tuning. Concretely, we define three process mining tasks that benefit from an understanding of process semantics and provide extensive benchmarking datasets for each of them. Our evaluation experiments reveal that (1) LLMs fail to solve challenging process mining tasks out of the box and when provided only a handful of in-context examples, (2) but they yield strong performance when fine-tuned for these tasks, consistently surpassing smaller, encoder-based language models. Adrian Rebmann, Fabian David Schmidt, Goran Glavas, Han van der Aa |
ICPM | 4 |
| 2024 | Recognizing task-level events from user interaction dataabstractUser interaction data comprises events that capture individual actions that a user performs on their computer. Such events provide detailed records about how users carry out their tasks in a process, even when this involves different applications. Although the comprehensiveness of such data provides a promising basis for process mining, user interaction events cannot be used directly for this purpose, because they do not meet two essential requirements. In particular, they neither indicate their relation to a process-level activity nor their relation to a specific process execution. Therefore, user interaction data needs to be transformed so that it meets these requirements before process mining techniques can be applied. This transformation problem comprises identifying tasks and their types and determining the relation between tasks and process executions. While some existing approaches tackle parts of this problem, none address it comprehensively. Therefore, we propose an unsupervised approach for recognizing task-level events from user interaction data that addresses it in full. It segments user interaction data to identify tasks, categorizes these according to their type, and relates tasks to each other via object instances it extracts from the user interaction events. In this manner, our approach creates task-level events that meet the requirements of process mining settings. Our evaluation demonstrates the approach’s efficacy and shows that its combined consideration of control-flow, data, and semantic information allows it to outperform baseline approaches in both online and offline settings. Adrian Rebmann, Han van der Aa |
Inf. Syst. | 2 |
| 2023 | Unsupervised Task Recognition from User Interaction Streams
Adrian Rebmann, Han van der Aa |
CAiSE | 2 |
| 2023 | Activity Recommendation for Business Process Modeling with Pre-trained Language Models
Diana Sola, Han van der Aa, Christian Meilicke, Heiner Stuckenschmidt |
ESWC | 2 |
| 2023 | Optimal event log sanitization for privacy-preserving process mining
Stephan A. Fahrenkrog-Petersen, Han van der Aa, Matthias Weidlich 0001 |
Data Knowl. Eng. | 2 |
| 2023 | Semantics-aware mechanisms for control-flow anonymization in process mining
Stephan A. Fahrenkrog-Petersen, Martin Kabierski, Han van der Aa, Matthias Weidlich 0001 |
Inf. Syst. | 3 |
| 2022 | GECCO: Constraint-driven Abstraction of Low-level Event LogsabstractProcess mining enables the analysis of complex systems using event data recorded during the execution of processes. Specifically, models of these processes can be discovered from event logs, i.e., sequences of events. However, the recorded events are often too fine-granular and result in unstructured models that are not meaningful for analysis. Log abstraction therefore aims to group together events to obtain a higher-level representation of the event sequences. While such a transformation shall be driven by the analysis goal, existing techniques force users to define how the abstraction is done, rather than what the result shall be. In this paper, we propose GECCO, an approach for log abstraction that enables users to impose requirements on the resulting log in terms of constraints. GECCO then groups events so that the constraints are satisfied and the distance to the original log is minimized. Since exhaustive log abstraction suffers from an exponential runtime complexity, GECCO also offers a heuristic approach guided by behavioral dependencies found in the log. We show that the abstraction quality of GECCO is superior to baseline solutions and demonstrate the relevance of considering constraints during log abstraction in real-life settings. Adrian Rebmann, Matthias Weidlich 0001, Han van der Aa |
ICDE | 3 |
| 2022 | Sampling and approximation techniques for efficient process conformance checking
Martin Kabierski, Han van der Aa, Matthias Weidlich 0001 |
Inf. Syst. | 2 |
| 2022 | Enabling semantics-aware process mining through the automatic annotation of event logs
Adrian Rebmann, Han van der Aa |
Inf. Syst. | 2 |
| 2022 | Exploiting label semantics for rule-based activity recommendation in business process modeling
Diana Sola, Han van der Aa, Christian Meilicke, Heiner Stuckenschmidt |
Inf. Syst. | 2 |
| 2021 | Extracting Semantic Process Information from the Natural Language in Event Logs
Adrian Rebmann, Han van der Aa |
CAiSE | 2 |
| 2021 | Sketch2BPMN: Automatic Recognition of Hand-Drawn BPMN Models
Bernhard Schäfer, Han van der Aa, Henrik Leopold, Heiner Stuckenschmidt |
CAiSE | 2 |
| 2021 | A Rule-Based Recommendation Approach for Business Process Modeling
Diana Sola, Christian Meilicke, Han van der Aa, Heiner Stuckenschmidt |
CAiSE | 3 |
| 2021 | SaCoFa: Semantics-aware Control-flow Anonymization for Process MiningabstractPrivacy-preserving process mining enables the analysis of business processes using event logs, while giving guarantees on the protection of sensitive information on process stakeholders. To this end, existing approaches add noise to the results of queries that extract properties of an event log, such as the frequency distribution of trace variants, for analysis. Noise insertion neglects the semantics of the process, though, and may generate traces not present in the original log. This is problematic. It lowers the utility of the published data and makes noise easily identifiable, as some traces will violate well-known semantic constraints. In this paper, we therefore argue for privacy preservation that incorporates a process’ semantics. For common trace-variant queries, we show how, based on the exponential mechanism, semantic constraints are incorporated to ensure differential privacy of the query result. Experiments demonstrate that our semantics-aware anonymization yields event logs of significantly higher utility than existing approaches. Stephan A. Fahrenkrog-Petersen, Martin Kabierski, Fabian Rösel, Han van der Aa, Matthias Weidlich 0001 |
ICPM | 4 |
| 2021 | EIRES: Efficient Integration of Remote Data in Event Stream ProcessingabstractTo support reactive and predictive applications, complex event processing (CEP) systems detect patterns in event streams based on predefined queries. To determine the events that constitute a query match, their payload data may need to be assessed together with data from remote sources. Such dependencies are problematic, since waiting for remote data to be fetched interrupts the processing of the stream. Yet, without event selection based on remote data, the query state to maintain may grow exponentially. In either case, the performance of the CEP system degrades drastically. Bo Zhao 0019, Han van der Aa, Thanh Tam Nguyen, Nguyen Quoc Viet Hung, Matthias Weidlich 0001 |
SIGMOD Conference | 2 |
| 2021 | Natural language-based detection of semantic execution anomalies in event logs
Han van der Aa, Adrian Rebmann, Henrik Leopold |
Inf. Syst. | 1 |
| 2020 | Assessing the Compliance of Business Process Models with Regulatory Documents
Karolin Winter, Han van der Aa, Stefanie Rinderle-Ma, Matthias Weidlich 0001 |
ER | 2 |
| 2020 | Efficient Process Conformance Checking on the Basis of Uncertain Event-to-Activity MappingsabstractConformance checking enables organizations to automatically identify compliance violations based on the analysis of observed event data. A crucial requirement for conformance-checking techniques is that observed events can be mapped to normative process models used to specify allowed behavior. Without a mapping, it is not possible to determine if an observed event trace conforms to the specification or not. A considerable problem in this regard is that establishing a mapping between events and process model activities is an inherently uncertain task. Since the use of a particular mapping directly influences the conformance of an event trace to a specification, this uncertainty represents a major issue for conformance checking. To overcome this issue, we introduce a probabilistic conformance-checking technique that can deal with uncertain mappings. Our technique avoids the need to select a single mapping by taking the entire spectrum of possible mappings into account. A quantitative evaluation demonstrates that our technique can be applied on a considerable number of real-world processes where existing conformance-checking techniques fail. Han van der Aa, Henrik Leopold, Hajo A. Reijers |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2019 | Extracting Declarative Process Models from Natural Language
Han van der Aa, Claudio Di Ciccio, Henrik Leopold, Hajo A. Reijers |
CAiSE | 1 |
| 2019 | PRETSA: Event Log Sanitization for Privacy-aware Process DiscoveryabstractEvent logs that originate from information systems enable comprehensive analysis of business processes, e.g., by process model discovery. However, logs potentially contain sensitive information about individual employees involved in process execution that are only partially hidden by an obfuscation of the event data. In this paper, we therefore address the risk of privacy-disclosure attacks on event logs with pseudonymized employee information. To this end, we introduce PRETSA, a novel algorithm for event log sanitization that provides privacy guarantees in terms of k-anonymity and t-closeness. It thereby avoids disclosure of employee identities, their membership in the event log, and their characterization based on sensitive attributes, such as performance information. Through step-wise transformations of a prefix-tree representation of an event log, we maintain its high utility for discovery of a performance-annotated process model. Experiments with real-world data demonstrate that sanitization with PRETSA yields event logs of higher utility compared to methods that exploit frequency-based filtering, while providing the same privacy guarantees. Stephan A. Fahrenkrog-Petersen, Han van der Aa, Matthias Weidlich 0001 |
ICPM | 2 |
| 2019 | Using Hidden Markov Models for the accurate linguistic analysis of process model activity labels
Henrik Leopold, Han van der Aa, Jelmer Offenberg, Hajo A. Reijers |
Inf. Syst. | 2 |
| 2018 | A probabilistic evaluation procedure for process model matching techniques
Elena Kuss, Henrik Leopold, Han van der Aa, Heiner Stuckenschmidt, Hajo A. Reijers |
Data Knowl. Eng. | 3 |
| 2018 | Aligning textual and model-based process descriptions
Josep Sànchez-Ferreres, Han van der Aa, Josep Carmona 0001, Lluís Padró 0001 |
Data Knowl. Eng. | 2 |
| 2018 | Checking process compliance against natural language specifications using behavioral spaces
Han van der Aa, Henrik Leopold, Hajo A. Reijers |
Inf. Syst. | 1 |
| 2017 | Instance-Based Process Matching Using Event-Log Information
Han van der Aa, Avigdor Gal, Henrik Leopold, Hajo A. Reijers, Tomer Sagi, Roee Shraga |
CAiSE | 1 |
| 2017 | Checking Process Compliance on the Basis of Uncertain Event-to-Activity Mappings
Han van der Aa, Henrik Leopold, Hajo A. Reijers |
CAiSE | 1 |
| 2017 | Comparing textual descriptions to process models - The automatic detection of inconsistencies
Han van der Aa, Henrik Leopold, Hajo A. Reijers |
Inf. Syst. | 1 |
| 2017 | Transforming unstructured natural language descriptions into measurable process performance indicators using Hidden Markov Models
Han van der Aa, Henrik Leopold, Adela del-Río-Ortega, Manuel Resinas, Hajo A. Reijers |
Inf. Syst. | 1 |
| 2016 | Narrowing the Business-IT Gap in Process Performance Measurement
Han van der Aa, Adela del-Río-Ortega, Manuel Resinas, Henrik Leopold, Antonio Ruiz Cortés, Jan Mendling, Hajo A. Reijers |
CAiSE | 1 |
| 2016 | Probabilistic Evaluation of Process Model Matching Techniques
Elena Kuss, Henrik Leopold, Han van der Aa, Heiner Stuckenschmidt, Hajo A. Reijers |
ER | 3 |