VLDB 2026 Research / reviewers in the wild / expert
Joscha Cüppers
dblp:285/3656
· DBLP profile ↗
8ranked-venue papers
5as first author
7since 2021 · last 2026
0000-0001-6628-2192ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Causal Discovery from Interval-Based Event SequencesabstractIn this paper we address the problem of discovering causal relationships from observational event sequence data. Existing methods typically assume that events are instantaneous point events, however in many real-world settings, events have duration. For example, in healthcare, a patient's symptoms may persist over a time interval and influence clinical actions while ongoing. To address this, we introduce a causal model for interval-based event sequences that captures rich causal structures, including interactions between events and causal mechanisms that depend on whether other events are ongoing. We prove that our model is identifiable in the limit and present a practical causal discovery algorithm, Niagara, grounded in the algorithmic Markov condition. To select among candidate models, we employ a minimum description length (MDL) criterion, enabling robust inference even with limited data. We validate our approach on synthetic and real data and demonstrate its utility on a real-world medical case study, where it uncovers meaningful causal relationships from noisy, interval-based event data. Lénaïg Cornanguer, Joscha Cüppers, Jilles Vreeken |
AAAI | 2 |
| 2026 | SEQRET: Mining Rule Sets from Event SequencesabstractSummarizing event sequences is a key aspect of data mining. Most existing methods neglect conditional dependencies and focus on discovering sequential patterns only. In this paper, we study the problem of discovering both conditional and unconditional dependencies from event sequences. We do so by discovering rules of the form X --> Y where X and Y are sequential patterns. Rules like these are simple to understand and provide a clear description of the relation between the antecedent and the consequent. To discover succinct and non-redundant sets of rules we formalize the problem in terms of the Minimum Description Length principle. As the search space is enormous and does not exhibit helpful structure, we propose the SEQRET method to discover high-quality rule sets in practice. Through extensive empirical evaluation we show that unlike the state of the art, SEQRET ably recovers the ground truth on synthetic datasets and finds useful rules from real datasets. Aleena Siji, Joscha Cüppers, Osman Mian, Jilles Vreeken |
AAAI | 2 |
| 2025 | Succinct Interaction-Aware Explanations
Sascha Xu, Joscha Cüppers, Jilles Vreeken |
KDD (1) | 2 |
| 2024 | Discovering Sequential Patterns with Predictable Inter-event DelaysabstractSummarizing sequential data with serial episodes allows non-trivial insight into the data generating process. Existing methods penalize gaps in pattern occurrences equally, regardless of where in the pattern these occur. This results in a strong bias against patterns with long inter-event delays, and in addition that regularity in terms of delays is not rewarded or discovered---even though both aspects provide key insight. In this paper we tackle both these problems by explicitly modeling inter-event delay distributions. That is, we are not only interested in discovering the patterns, but also in describing how many times steps typically occur between their individual events. We formalize the problem in terms of the Minimum Description Length principle, by which we say the best set of patterns is the one that compresses the data best. The resulting optimization problem does not lend itself to exact optimization, and hence we propose Hopper to heuristically mine high quality patterns. Extensive experiments show that Hopper efficiently recovers the ground truth, discovers meaningful patterns from real-world data, and outperforms existing methods in discovering long-delay patterns. Joscha Cüppers, Paul Krieger, Jilles Vreeken |
AAAI | 1 |
| 2024 | Causal Discovery from Event Sequences by Local Cause-Effect AttributionabstractSequences of events, such as crashes in the stock market or outages in a network, contain strong temporal dependencies, whose understanding is crucial to react to and influence future events. In this paper, we study the problem of discovering the underlying causal structure from event sequences. To this end, we introduce a new causal model, where individual events of the cause trigger events of the effect with dynamic delays. We show that in contrast to existing methods based on Granger causality, our model is identifiable for both instant and delayed effects.
We base our approach on the Algorithmic Markov Condition, by which we identify the true causal network as the one that minimizes the Kolmogorov complexity. As the Kolmogorov complexity is not computable, we instantiate our model using Minimum Description Length and show that the resulting score identifies the causal direction. To discover causal graphs, we introduce the Cascade algorithm, which adds edges in topological order. Extensive evaluation shows that Cascade outperforms existing methods in settings with instantaneous effects, noise, and multiple colliders, and discovers insightful causal graphs on real-world data. Joscha Cüppers, Sascha Xu, Ahmed Musa, Jilles Vreeken |
NeurIPS | 1 |
| 2023 | Below the Surface: Summarizing Event Sequences with Generalized Sequential PatternsabstractWe study the problem of succinctly summarizing a database of event sequences in terms of generalized sequential patterns. That is, we are interested in patterns that are not exclusively defined over observed surface-level events, as is usual, but rather may additionally include generalized events that can match a set of events. To avoid spurious and redundant results we define the problem in terms of the Minimum Description Length principle, by which we are after that set of patterns and generalizations that together best compress the data without loss. The resulting optimization problem does not lend itself for exact search, which is why we propose the heuristic Flock algorithm to efficiently find high-quality models in practice. Extensive experiments on synthetic and real-world data show that Flock results in compact and easily interpretable models that accurately recover the ground truth, including rare instances of generalized patterns. Additionally Flock recovers how generalized events within patterns depend on each other, and overall provides clearer insight into the data-generating process than using state of the art algorithms that only consider surface-level patterns. Joscha Cüppers, Jilles Vreeken |
KDD | 1 |
| 2022 | Omen: discovering sequential patterns with reliable prediction delaysabstractAbstract Suppose we are given a discrete-valued time series $$X $$ X of observed events and an equally long binary sequence $$Y $$ Y that indicates whether something of interest happened at that particular point in time. We consider the problem of mining serial episodes, sequential patterns allowing for gaps, from $$X $$ X that reliably predict those interesting events. Withreliablewe mean patterns that not only predictthatan interesting event is likely to follow, but in particular that we can also accurately tell howhow longuntil that event will happen. In other words, we are specifically interested in patterns with a highly skewed distribution of delays between pattern occurrences and predicted events. As it is unlikely that a single pattern can explain a complex real-world progress, we are after the smallest, least redundant set of such patterns that together explain the interesting events well. We formally define this problem in terms of the Minimum Description Length principle, by which we identify the best patterns as those that describe the occurrences of interesting events $$Y $$ Y most succinctly given the data over $$X $$ X . As neither discovering the optimal explanation of $$Y $$ Y given a set of patterns, nor the discovery of optimal pattern set are problems that allow for straightforward optimization, we break the problem in two and propose effective heuristics for both. Through extensive empirical evaluation, we show that both our main method,Omen, and its fast approximationfOmen, work well in practice and both quantitatively and qualitatively beat the state of the art. Joscha Cüppers, Janis Kalofolias, Jilles Vreeken |
Knowl. Inf. Syst. | 1 |
| 2020 | Just Wait For It... Mining Sequential Patterns with Reliable Prediction DelaysabstractSuppose we are given an event sequence X of observed events and an equally long binary sequence Y that indicates whether something of interest happened at that particular point in time. We consider the problem of mining sequential patterns from X that reliably predict those interesting events. With reliable we mean those patterns that not only predict that an interesting event is likely to follow but especially those patterns for which we can with high precision tell how long until that event will happen. That is, we are after patterns that have highly skewed distributions of delays between pattern occurrences and predicted events. In particular, we are after the smallest, least redundant set of such patterns that together explain the interesting events well. We formally define this problem in terms of the Minimum Description Length principle, by which we identify the best patterns as those that describe the data most succinctly. As discovering the optimal explanation of Y given a set of patterns, as well as the discovery of optimal pattern set are both hard problems that do not allow for straightforward optimization, we propose the heuristic Omen algorithm. Through extensive empirical evaluation we show that Omen works well in practice and beats the state of the art both quantitatively and qualitatively. Joscha Cüppers, Jilles Vreeken |
ICDM | 1 |