Felix I. Stamm

dblp:277/0597 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2025
0000-0002-4469-7511ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Efficient Sampling of Temporal Networks with Preserved Causality Structure
abstract
In this paper, we extend the classical Color Refinement algorithm for static networks to temporal (undirected and directed) networks. This enables us to design an algorithm to sample synthetic networks that preserves the d-hop neighborhood structure of a given temporal network. The higher d is chosen, the better the temporal neighborhood structure of the original network is preserved. Specifically, we provide efficient algorithms that preserve time-respecting (”causal”) paths in the networks up to length d, and scale to real-world network sizes. We validate our approach theoretically (for Degree and Katz centrality) and experimentally (for edge persistence, causal triangles, and burstiness). An experimental comparison shows that our method retains these key temporal characteristics more effectively than existing randomization methods.
Felix I. Stamm, Mehdi Naima, Michael T. Schaub
SDM1
2023 Neighborhood Structure Configuration Models
abstract
We develop a new method to efficiently sample synthetic networks that preserve the d-hop neighborhood structure of a given network for any given d. The proposed algorithm trades off the diversity in network samples against the depth of the neighborhood structure that is preserved. Our key innovation is to employ a colored Configuration Model with colors derived from iterations of the so-called Color Refinement algorithm. We prove that with increasing iterations the preserved structural information increases: the generated synthetic networks and the original network become more and more similar, and are eventually indistinguishable in terms of centrality measures such as PageRank, HITS, Katz centrality and eigenvector centrality. Our work enables to efficiently generate samples with a precisely controlled similarity to the original network, especially for large networks.
Felix I. Stamm, Michael Scholkemper, Michael T. Schaub, Markus Strohmaier
WWW1
2022 Estimating the Pruned Search Space Size of Subgroup Discovery
abstract
Subgroup discovery (SD) is a well-established supervised pattern mining approach. A key practical challenge —in particular considering interactive mining strategies— is that it is difficult to estimate the runtime of an exhaustive search algorithm before actually running the algorithm even for experienced practitioners. This is due to the exponential explosion of the candidate search space, sophisticated pruning strategies, and implementation specifics that can all affect the runtime by orders of magnitude depending on the dataset and the exact mining task parameters. A subgroup discovery run could take mere minutes or literal years. We would not know until afterwards. In this paper, we study the estimation of the complexity and runtime of subgroup discovery algorithms by estimating the pruned search space size, i.e., the number of actually evaluated candidate subgroups. We propose a sampling-based algorithm called SDFASTEST. SDFASTEST can effectively estimate the pruned search space size of a search algorithm. In our extensive evaluation on 1026 different tasks with 2 search algorithms, SDFASTEST was able to reduce the average mean absolute log error of the search space size estimation by ca. 94% compared to the best baseline, a depth-based upper bound.
Lennart Purucker, Felix I. Stamm, Florian Lemmerich, Jöran Beel
ICDM2
2021 Redescription Model Mining
abstract
This paper introduces Redescription Model Mining, a novel approach to identify interpretable patterns across two datasets that share only a subset of attributes and have no common instances. In particular, Redescription Model Mining aims to find pairs of describable data subsets -- one for each dataset -- that induce similar exceptional models with respect to a prespecified model class. To achieve this, we combine two previously separate research areas: Exceptional Model Mining and Redescription Mining. For this new problem setting, we develop interestingness measures to select promising patterns, propose efficient algorithms, and demonstrate their potential on synthetic and real-world data. Uncovered patterns can hint at common underlying phenomena that manifest themselves across datasets, enabling the discovery of possible associations between (combinations of) attributes that do not appear in the same dataset.
Felix I. Stamm, Martin Becker 0003, Markus Strohmaier, Florian Lemmerich
KDD1