Stefan Wrobel

dblp:w/StefanWrobel · DBLP profile ↗
← Back
71ranked-venue papers
9as first author
12since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 55 · 6 first-author · 9 since 2021Databases, data management, data science and information retrieval · 30 · 5 first-author · 2 since 2021Theory of computation · 9 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Textual data bias detection and mitigation - an extensible pipeline with experimental evaluation
abstract
Textual data used to train large language models (LLMs) exhibits multifaceted bias manifestations, encompassing harmful language and skewed demographic distributions. Regulations such as the EU AI Act require identifying and mitigating biases against protected groups in data, with the ultimate goal of preventing unfair model outputs. However, practical guidance and operationalization are lacking. We propose a comprehensive data bias detection and mitigation pipeline comprising four components that address two data bias types, namely representation bias and (explicit) stereotypes for a configurable sensitive attribute. First, we leverage LLM-generated word lists, created according to defined quality criteria, to detect relevant group labels. Second, representation bias is quantified using the Demographic Representation Score. Third, we detect and mitigate stereotypes using sociolinguistically informed filtering. Finally, we mitigate representation bias through Grammar- and Context-Aware Counterfactual Data Augmentation. We conduct a twofold evaluation using gender, religion , and age as examples. First, we evaluate the effectiveness of each individual component on data debiasing through human validation and baseline comparison. The findings demonstrate that we successfully reduce representation bias and (explicit) stereotypes in a text dataset. Second, we evaluate the effect of data debiasing on model bias by benchmarking several models (0.6B-8B parameters) fine-tuned on the debiased text dataset. This evaluation reveals that LLMs fine-tuned on debiased data do not consistently show improved performance on bias benchmarks. These results expose critical gaps in current evaluation methodologies and highlight the need for targeted data interventions to address manifested model bias.
Rebekka Görge, Sujan Sai Gannamaneni, Tabea Naeven, Hammam Abdelwahab, Héctor Allende-Cid, Armin B. Cremers, Lennard Helmer, Michael Mock, Anna Schmitz 0001, Songkai Xue, Elif Yildirir, Maximilian Poretschkin, Stefan Wrobel
Expert Syst. Appl.13
2026 Learning Weakly Convex Sets in Metric Spaces
Eike Stadtländer, Tamás Horváth 0001, Stefan Wrobel
Mach. Learn.3
2025 The Local Convexification Method and Its Application to Learning Weakly Convex Boolean Functions
Eike Stadtländer, Tamás Horváth 0001, Stefan Wrobel
ECML/PKDD (4)3
2025 Improving graph neural networks through feature importance learning
abstract
Abstract Graph neural networks (GNNs) are among the most widely used methods for node classification in graphs. A common strategy to improve their predictive performance is to enrich nodes with additional features. A weakness of this method is that the set of appropriate features can vary from graph to graph. We address this shortcoming by proposing a novel method. In a preprocessing step, a first GNN is trained on a set of graphs with varying structural properties, using a candidate set of node features fixed in advance. The resulting GNN model is then used to predict the most relevant features from the candidate set for unseen target graphs, which are later processed for node classification. For each target graph, a second GNN is trained on the graph, which is enriched with the node feature vectors calculated for the features selected by the first GNN. A key advantage of the proposed method is that the features are selected without computing the candidate features for the target graph. Our experimental results on synthetic and real-world graphs show that even a few features selected in this way is sufficient to significantly improve the predictive performance of GNNs that use either none or all of the candidate features. Moreover, the time needed to learn the second GNN for the target graph can be reduced by up to two orders of magnitude.
Fouad Alkhoury, Tamás Horváth 0001, Christian Bauckhage, Stefan Wrobel
Mach. Learn.4
2024 Wasserstein dropout
abstract
Abstract Despite of its importance for safe machine learning, uncertainty quantification for neural networks is far from being solved. State-of-the-art approaches to estimate neural uncertainties are often hybrid, combining parametric models with explicit or implicit (dropout-based) ensembling. We take another pathway and propose a novel approach to uncertainty quantification for regression tasks, Wasserstein dropout, that is purely non-parametric. Technically, it captures aleatoric uncertainty by means of dropout-based sub-network distributions. This is accomplished by a new objective which minimizes the Wasserstein distance between the label distribution and the model distribution. An extensive empirical analysis shows that Wasserstein dropout outperforms state-of-the-art methods, on vanilla test data as well as under distributional shift in terms of producing more accurate and stable uncertainty estimates.
Joachim Sicking, Maram Akila, Maximilian Pintz, Tim Wirtz, Stefan Wrobel, Asja Fischer
Mach. Learn.5
2023 Maximal closed set and half-space separations in finite closure systems
Florian Seiffarth, Tamás Horváth 0001, Stefan Wrobel
Theor. Comput. Sci.3
2022 Graph Filtration Kernels
abstract
The majority of popular graph kernels is based on the concept of Haussler's R-convolution kernel and defines graph similarities in terms of mutual substructures. In this work, we enrich these similarity measures by considering graph filtrations: Using meaningful orders on the set of edges, which allow to construct a sequence of nested graphs, we can consider a graph at multiple granularities. A key concept of our approach is to track graph features over the course of such graph resolutions. Rather than to simply compare frequencies of features in graphs, this allows for their comparison in terms of when and for how long they exist in the sequences. In this work, we propose a family of graph kernels that incorporate these existence intervals of features. While our approach can be applied to arbitrary graph features, we particularly highlight Weisfeiler-Lehman vertex labels, leading to efficient kernels. We show that using Weisfeiler-Lehman labels over certain filtrations strictly increases the expressive power over the ordinary Weisfeiler-Lehman procedure in terms of deciding graph isomorphism. In fact, this result directly yields more powerful graph kernels based on such features and has implications to graph neural networks due to their close relationship to the Weisfeiler-Lehman method. We empirically validate the expressive power of our graph kernels and show significant improvements over state-of-the-art graph kernels in terms of predictive performance on various real-world benchmark datasets.
Till Hendrik Schulz, Pascal Welke, Stefan Wrobel
AAAI3
2022 A Fast Heuristic for Computing Geodesic Closures in Large Networks
Florian Seiffarth, Tamás Horváth 0001, Stefan Wrobel
DS3
2022 A generalized Weisfeiler-Lehman graph kernel
abstract
Abstract After more than one decade, Weisfeiler-Lehman graph kernels are still among the most prevalent graph kernels due to their remarkable predictive performance and time complexity. They are based on a fast iterative partitioning of vertices, originally designed for deciding graph isomorphism with one-sided error. The Weisfeiler-Lehman graph kernels retain this idea and compare such labels with respect to equality. This binary valued comparison is, however, arguably too rigid for defining suitable graph kernels for certain graph classes. To overcome this limitation, we propose a generalization of Weisfeiler-Lehman graph kernels which takes into account a more natural and finer grade of similarity between Weisfeiler-Lehman labels than equality. We show that the proposed similarity can be calculated efficiently by means of the Wasserstein distance between certain vectors representing Weisfeiler-Lehman labels. This and other facts give rise to the natural choice of partitioning the vertices with the Wasserstein k-means algorithm. We empirically demonstrate on the Weisfeiler-Lehman subtree kernel, which is one of the most prominent Weisfeiler-Lehman graph kernels, that our generalization significantly outperforms this and other state-of-the-art graph kernels in terms of predictive performance on datasets which contain structurally more complex graphs beyond the typically considered molecular graphs.
Till Hendrik Schulz, Tamás Horváth 0001, Pascal Welke, Stefan Wrobel
Mach. Learn.4
2021 Learning Weakly Convex Sets in Metric Spaces
abstract
Abstract One of the central problems studied in the theory of machine learning is the question of whether, for a given class of hypotheses, it is possible to efficiently find a consistent hypothesis, i.e., one with zero training error. While problems involving convex hypotheses have been extensively studied, the question of whether efficient learning is possible for non-convex hypotheses composed of possibly several disconnected regions is still not well understood. Although it has been shown quite a while ago that efficient learning of weakly convex hypotheses, a parameterized relaxation of convex hypotheses, is possible for the special case of Boolean functions, the question of whether this idea can be developed into a generic paradigm has not yet been studied. In this paper, we provide a positive answer and show that the consistent hypothesis finding problem can indeed be solved in polynomial time for a broad class of weakly convex hypotheses over metric spaces. To this end, we propose a general domain-independent algorithm for finding consistent weakly convex hypotheses and prove sufficient conditions for its efficiency that characterize the corresponding hypothesis classes. To illustrate our general algorithm and its properties, we discuss several non-trivial learning examples to demonstrate how it can be used to efficiently solve the corresponding consistent hypothesis finding problem. Without the weak convexity constraint, these problems are known to be computationally intractable. We then show that the general idea of our algorithm even extends to the extensional case, enabling applications, such as vertex classification in graphs. We prove that using our extended algorithm, the problem can be solved in polynomial time provided the distances in the domain can be computed efficiently.
Eike Stadtländer, Tamás Horváth 0001, Stefan Wrobel
ECML/PKDD (2)3
2021 Constructing Spaces and Times for Tactical Analysis in Football
abstract
A possible objective in analyzing trajectories of multiple simultaneously moving objects, such as football players during a game, is to extract and understand the general patterns of coordinated movement in different classes of situations as they develop. For achieving this objective, we propose an approach that includes a combination of query techniques for flexible selection of episodes of situation development, a method for dynamic aggregation of data from selected groups of episodes, and a data structure for representing the aggregates that enables their exploration and use in further analysis. The aggregation, which is meant to abstract general movement patterns, involves construction of new time-homomorphic reference systems owing to iterative application of aggregation operators to a sequence of data selections. As similar patterns may occur at different spatial locations, we also propose constructing new spatial reference systems for aligning and matching movements irrespective of their absolute locations. The approach was tested in application to tracking data from two Bundesliga games of the 2018/2019 season. It enabled detection of interesting and meaningful general patterns of team behaviors in three classes of situations defined by football experts. The experts found the approach and the underlying concepts worth implementing in tools for football analysts.
Gennady L. Andrienko, Natalia V. Andrienko, Gabriel Anzer, Pascal Bauer, Guido Budziak, Georg Fuchs, Dirk Hecker, Hendrik Weber, Stefan Wrobel
IEEE Trans. Vis. Comput. Graph.9
2021 A theoretical model for pattern discovery in visual analytics
abstract
The word ‘pattern’ frequently appears in the visualisation and visual analytics literature, but what do we mean when we talk about patterns? We propose a practicable definition of the concept of a pattern in a data distribution as a combination of multiple interrelated elements of two or more data components that can be represented and treated as a unified whole. Our theoretical model describes how patterns are made by relationships existing between data elements. Knowing the types of these relationships, it is possible to predict what kinds of patterns may exist. We demonstrate how our model underpins and refines the established fundamental principles of visualisation. The model also suggests a range of interactive analytical operations that can support visual analytics workflows where patterns, once discovered, are explicitly involved in further data analysis.
Natalia V. Andrienko, Gennady L. Andrienko, Silvia Miksch, Heidrun Schumann, Stefan Wrobel
Vis. Informatics5
2020 Decision Snippet Features
abstract
Decision trees excel at interpretability of their prediction results. To achieve required prediction accuracies, however, often large ensembles of decision trees - random forests - are considered, reducing interpretability due to large size. Additionally, their size slows down inference on modern hardware and restricts their applicability in low-memory embedded devices. We introduce Decision Snippet Features, which are obtained from small subtrees that appear frequently in trained random forests. We subsequently show that linear models on top of these features achieve comparable and sometimes even better predictive performance than the original random forest, while reducing the model size by up to two orders of magnitude.
Pascal Welke, Fouad Alkhoury, Christian Bauckhage, Stefan Wrobel
ICPR4
2020 HOPS: Probabilistic Subtree Mining for Small and Large Graphs
abstract
Frequent subgraph mining, i.e., the identification of relevant patterns in graph databases, is a well-known data mining problem with high practical relevance, since next to summarizing the data, the resulting patterns can also be used to define powerful domain-specific similarity functions for prediction. In recent years, significant progress has been made towards subgraph mining algorithms that scale to complex graphs by focusing on tree patterns and probabilistically allowing a small amount of incompleteness in the result. Nonetheless, the complexity of the pattern matching component used for deciding subtree isomorphism on arbitrary graphs has significantly limited the scalability of existing approaches. In this paper, we adapt sampling techniques from mathematical combinatorics to the problem of probabilistic subtree mining in arbitrary databases of many small to medium-size graphs or a single large graph. By restricting on tree patterns, we provide an algorithm that approximately counts or decides subtree isomorphism for arbitrary transaction graphs in sub-linear time with one-sided error. Our empirical evaluation on a range of benchmark graph datasets shows that the novel algorithm substantially outperforms state-of-the-art approaches both in the task of approximate counting of embeddings in single large graphs and in probabilistic frequent subtree mining in large databases of small to medium sized graphs.
Pascal Welke, Florian Seiffarth, Michael Kamp, Stefan Wrobel
KDD4
2020 Maximum Margin Separations in Finite Closure Systems
Florian Seiffarth, Tamás Horváth 0001, Stefan Wrobel
ECML/PKDD (1)3
2020 Adiabatic Quantum Computing for Max-Sum Diversification
abstract
The combinatorial problem of max-sum diversification asks for a maximally diverse subset of a given set of data. Here, we show that it can be expressed as an Ising energy minimization problem. Given this result, max-sum diversification can be solved on adiabatic quantum computers and we present proof of concept simulations which support this claim. This, in turn, suggests that quantum computing might play a role in data mining. We therefore discuss quantum computing in a tutorial like manner and elaborate on its current strengths and weaknesses for data analysis.
Christian Bauckhage, Rafet Sifa, Stefan Wrobel
SDM3
2020 Effective approximation of parametrized closure systems over transactional data streams
Daniel Trabold, Tamás Horváth 0001, Stefan Wrobel
Mach. Learn.3
2019 Leveraging Domain Knowledge for Reinforcement Learning Using MMC Architectures
Rajkumar Ramamurthy, Christian Bauckhage, Rafet Sifa, Jannis Schücker, Stefan Wrobel
ICANN (2)5
2019 Maximal Closed Set and Half-Space Separations in Finite Closure Systems
Florian Seiffarth, Tamás Horváth 0001, Stefan Wrobel
ECML/PKDD (1)3
2019 Probabilistic and exact frequent subtree mining in graphs beyond forests
Pascal Welke, Tamás Horváth 0001, Stefan Wrobel
Mach. Learn.3
2018 Policy Learning Using SPSA
Rajkumar Ramamurthy, Christian Bauckhage, Rafet Sifa, Stefan Wrobel
ICANN (3)4
2018 Efficient Decentralized Deep Learning by Dynamic Model Averaging
Michael Kamp, Linara Adilova, Joachim Sicking, Fabian Hüger, Peter Schlicht, Tim Wirtz, Stefan Wrobel
ECML/PKDD (1)7
2018 Mining Tree Patterns with Partially Injective Homomorphisms
Till Hendrik Schulz, Tamás Horváth 0001, Pascal Welke, Stefan Wrobel
ECML/PKDD (2)4
2018 Probabilistic frequent subtrees for efficient graph classification and retrieval
Pascal Welke, Tamás Horváth 0001, Stefan Wrobel
Mach. Learn.3
2017 Using Echo State Networks for Cryptography
Rajkumar Ramamurthy, Christian Bauckhage, Krisztián Búza, Stefan Wrobel
ICANN (2)4
2017 Co-Regularised Support Vector Regression
Katrin Ullrich, Michael Kamp, Thomas Gärtner 0001, Martin Vogt 0001, Stefan Wrobel
ECML/PKDD (2)5
2016 Min-Hashing for Probabilistic Frequent Subtree Feature Spaces
Pascal Welke, Tamás Horváth 0001, Stefan Wrobel
DS3
2015 Whole-body self-calibration via graph-optimization and automatic configuration selection
abstract
In this paper, we present a novel approach to accurately calibrate the kinematic model of a humanoid based on observations of its monocular camera. Our technique estimates the parameters of the complete model, consisting of the joint angle offsets of the whole body including the legs, as well as the camera extrinsic and intrinsic parameters. We cast the parameter estimation as a least-squares optimization problem. In the error function, we consider the residuals between camera observations of end-effector markers and their projections into the image based on the estimate of the calibration parameters. Furthermore, we developed an approach to automatically select a subset of configurations for the calibration process that yields a good trade-off between the number of observations and accuracy. As the experiments with a Nao humanoid show, we achieve an accurate calibration for this low-cost platform. Further, our approach to configuration selection yields substantially better optimization results compared to randomly chosen viable configurations. Hence, our system only requires a reduced number of configurations to achieve accurate results. Our optimization is general and the implementation, which is available online, can easily be applied to different humanoids.
Daniel Maier 0001, Stefan Wrobel, Maren Bennewitz
ICRA2
2014 On the Complexity of Frequent Subtree Mining in Very Simple Structures
Pascal Welke, Tamás Horváth 0001, Stefan Wrobel
ILP3
2013 Scalable Analysis of Movement Data for Extracting and Exploring Significant Places
abstract
Place-oriented analysis of movement data, i.e., recorded tracks of moving objects, includes finding places of interest in which certain types of movement events occur repeatedly and investigating the temporal distribution of event occurrences in these places and, possibly, other characteristics of the places and links between them. For this class of problems, we propose a visual analytics procedure consisting of four major steps: 1) event extraction from trajectories; 2) extraction of relevant places based on event clustering; 3) spatiotemporal aggregation of events or trajectories; 4) analysis of the aggregated data. All steps can be fulfilled in a scalable way with respect to the amount of the data under analysis; therefore, the procedure is not limited by the size of the computer's RAM and can be applied to very large data sets. We demonstrate the use of the procedure by example of two real-world problems requiring analysis at different spatial scales.
Gennady L. Andrienko, Natalia V. Andrienko, Christophe Hurter, Salvatore Rinzivillo, Stefan Wrobel
IEEE Trans. Vis. Comput. Graph.5
2012 Pedestrian Quantity Estimation with Trajectory Patterns
Thomas Liebig, Zhao Xu 0001, Michael May 0001, Stefan Wrobel
ECML/PKDD (2)4
2011 Introduction to the special issue on mining and learning with graphs
S. V. N. Vishwanathan, Samuel Kaski, Jennifer Neville, Stefan Wrobel
Mach. Learn.4
2010 Frequent subgraph mining in outerplanar graphs
Tamás Horváth 0001, Jan Ramon, Stefan Wrobel
Data Min. Knowl. Discov.3
2010 Listing closed sets of strongly accessible set systems with applications to data mining
Mario Boley, Tamás Horváth 0001, Axel Poigné, Stefan Wrobel
Theor. Comput. Sci.4
2009 A Logic-Based Approach to Relation Extraction from Texts
Tamás Horváth 0001, Gerhard Paass, Frank Reichartz, Stefan Wrobel
ILP4
2009 Efficient Discovery of Interesting Patterns Based on Strong Closedness
abstract
Finding patterns that are interesting to a user in a certain application context is one of the central goals of Data Mining research. Regarding all patterns above a certain frequency threshold as interesting is one way of defining interestingness. In this paper, however, we argue that in many applications, a different notion of interestingness is required in order to be able to capture “long”, and thus particularly informative, patterns that are correspondingly of low frequency. To identify such patterns, our proposed measure of interestingness is based on the degree or strength of closedness of the patterns. We show that (a) indeed this definition selects long interesting patterns that are difficult to identify with frequency-based approaches, and (b) that it selects patterns that are robust against noise and/or dynamic changes. We prove that the family of interesting patterns proposed here forms a closure system and use the corresponding closure operator to design a mining algorithm listing these patterns in amortized quadratic time. In particular, for non-sparse datasets its time complexity is O(nm) per pattern, where n denotes the number of items and m the size of the database. This is equal to the best known time bound for listing ordinary closed frequent sets, which is a special case of our problem. We also report empirical results with real-world datasets.
Mario Boley, Tamás Horváth 0001, Stefan Wrobel
SDM3
2008 Tight Optimistic Estimates for Fast Subgroup Discovery
Henrik Grosskreutz, Stefan Rüping 0001, Stefan Wrobel
ECML/PKDD (1)3
2007 Efficient Closed Pattern Mining in Strongly Accessible Set Systems (Extended Abstract)
Mario Boley, Tamás Horváth 0001, Axel Poigné, Stefan Wrobel
PKDD4
2007 Geovisual analytics for spatial decision support: Setting the research agenda
abstract
This article summarizes the results of the workshop on Visualization, Analytics & Spatial Decision Support, which took place at the GIScience conference in September 2006. The discussions at the workshop and analysis of the state of the art have revealed a need in concerted cross‐disciplinary efforts to achieve substantial progress in supporting space‐related decision making. The size and complexity of real‐life problems together with their ill‐defined nature call for a true synergy between the power of computational techniques and the human capabilities to analyze, envision, reason, and deliberate. Existing methods and tools are yet far from enabling this synergy. Appropriate methods can only appear as a result of a focused research based on the achievements in the fields of geovisualization and information visualization, human‐computer interaction, geographic information science, operations research, data mining and machine learning, decision science, cognitive science, and other disciplines. The name ‘Geovisual Analytics for Spatial Decision Support’ suggested for this new research direction emphasizes the importance of visualization and interactive visual interfaces and the link with the emerging research discipline of Visual Analytics. This article, as well as the whole special issue, is meant to attract the attention of scientists with relevant expertise and interests to the major challenges requiring multidisciplinary efforts and to promote the establishment of a dedicated research community where an appropriate range of competences is combined with an appropriate breadth of thinking.
Gennady L. Andrienko, Natalia V. Andrienko, Piotr Jankowski 0001, Daniel A. Keim, Menno-Jan Kraak, Alan M. MacEachren, Stefan Wrobel
Int. J. Geogr. Inf. Sci.7
2006 Multi-class Ensemble-Based Active Learning
Christine Kopp, Stefan Wrobel
ECML2
2006 Efficient co-regularised least squares regression
abstract
In many applications, unlabelled examples are inexpensive and easy to obtain. Semi-supervised approaches try to utilise such examples to reduce the predictive error. In this paper, we investigate a semi-supervised least squares regression algorithm based on the co-learning approach. Similar to other semi-supervised algorithms, our base algorithm has cubic runtime complexity in the number of unlabelled examples. To be able to handle larger sets of unlabelled examples, we devise a semi-parametric variant that scales linearly in the number of unlabelled examples. Experiments show a significant error reduction by co-regularisation and a large runtime improvement for the semi-parametric approximation. Last but not least, we propose a distributed procedure that can be applied without collecting all data at a single site.
Ulf Brefeld, Thomas Gärtner 0001, Tobias Scheffer, Stefan Wrobel
ICML4
2006 Frequent subgraph mining in outerplanar graphs
abstract
In recent years there has been an increased interest in algorithms that can perform frequent pattern discovery in large databases of graph structured objects. While the frequent connected subgraph mining problem for tree datasets can be solved in incremental polynomial time, it becomes intractable for arbitrary graph databases. Existing approaches have therefore resorted to various heuristic strategies and restrictions of the search space, but have not identified a practically relevant tractable graph class beyond trees. In this paper, we define the class of so called tenuous outerplanar graphs, a strict generalization of trees, develop a frequent subgraph mining algorithm for tenuous outerplanar graphs that works in incremental polynomial time, and evaluate the algorithm empirically on the NCI molecular graph dataset.
Tamás Horváth 0001, Jan Ramon, Stefan Wrobel
KDD3
2006 Bias-Free Hypothesis Evaluation in Multirelational Domains
Christine Kopp, Stefan Wrobel
PAKDD2
2004 A comparative study on methods for reducing myopia of hill-climbing search in multirelational learning
abstract
Hill-climbing search is the most commonly used search algorithm in ILP systems because it permits the generation of theories in short running times. However, a well known drawback of this greedy search strategy is its myopia. Macro-operators (or macros for short), a recently proposed technique to reduce the search space explored by exhaustive search, can also be argued to reduce the myopia of hill-climbing search by automatically performing a variable-depth look-ahead in the search space. Surprisingly, macros have not been employed in a greedy learner. In this paper, we integrate macros into a hill-climbing learner. In a detailed comparative study in several domains, we show that indeed a hill-climbing learner using macros performs significantly better than current state-of-the-art systems involving other techniques for reducing myopia, such as fixed-depth look-ahead, template-based look-ahead, beam-search, or determinate literals. In addition, macros, in contrast to some of the other approaches, can be computed fully automatically and do not require user involvement nor special domain properties such as determinacy.
Lourdes Peña Castillo, Stefan Wrobel
ICML2
2004 Cyclic pattern kernels for predictive graph mining
abstract
S.158-167
Tamás Horváth 0001, Thomas Gärtner 0001, Stefan Wrobel
KDD3
2003 Learning Minesweeper with Multirelational Learning
Lourdes Peña Castillo, Stefan Wrobel
IJCAI2
2003 A Comparative Evaluation of Feature Set Evolution Strategies for Multirelational Boosting
Susanne Hoche, Stefan Wrobel
ILP2
2003 Comparative Evaluation of Approaches to Propositionalization
Mark-A. Krogel, Simon Alan Rawles, Filip Zelezný, Peter A. Flach, Nada Lavrac, Stefan Wrobel
ILP6
2002 Feature Selection for Propositionalization
Mark-A. Krogel, Stefan Wrobel
Discovery Science2
2002 Macro-Operators in Multirelational Learning: A Search-Space Reduction Technique
Lourdes Peña Castillo, Stefan Wrobel
ECML2
2002 Scaling Boosting by Margin-Based Inclusionof Features and Relations
Susanne Hoche, Stefan Wrobel
ECML2
2002 A Scalable Constant-Memory Sampling Algorithm for Pattern Discovery in Large Databases
Tobias Scheffer, Stefan Wrobel
PKDD2
2002 Finding the Most Interesting Patterns in a Database Quickly by Using Sequential Sampling
Tobias Scheffer, Stefan Wrobel
J. Mach. Learn. Res.2
2001 Towards Discovery of Deep and Wide First-Order Structures: A Case Study in the Domain of Mutagenicity
Tamás Horváth 0001, Stefan Wrobel
Discovery Science2
2001 Scalability, Search, and Sampling: From Smart Algorithms to Active Discovery
Stefan Wrobel
ECML1
2001 Mining the Web with Active Hidden Markov Models
abstract
Given the enormous amounts of information available only in unstructured or semi-structured textual documents, tools for information extraction (IE) have become enormously important. IE tools identify the relevant information in such documents and convert it into a structured format such as a database or an XML document. While first IE algorithms were hand-crafted sets of rules, researchers soon turned to learning extraction rules from hand-labeled documents. Unfortunately, rule-based approaches sometimes fail to provide the necessary robustness against the inherent variability of document, structure, which has led to the recent interest in using hidden Markov models (HMMs). By using additional unlabeled documents as they are usually readily available in most applications, we can perform active learning of HMMs. The idea of active learning algorithms is to identify unlabeled observations that would be most useful when labeled by the user. Such algorithms are known for classification, clustering, and regression; we present the first algorithm for active learning of hidden Markov models.
Tobias Scheffer, Christian Decomain, Stefan Wrobel
ICDM3
2001 Incremental Maximization of Non-Instance-Averaging Utility Functions with Applications to Knowledge Discovery Problems
Tobias Scheffer, Stefan Wrobel
ICML2
2001 Active Hidden Markov Models for Information Extraction
Tobias Scheffer, Christian Decomain, Stefan Wrobel
IDA3
2001 Relational Learning Using Constrained Confidence-Rated Boosting
Susanne Hoche, Stefan Wrobel
ILP2
2001 Transformation-Based Learning Using Multirelational Aggregation
Mark-A. Krogel, Stefan Wrobel
ILP2
2001 Scalability, Search, and Sampling: From Smart Algorithms to Active Discovery
Stefan Wrobel
PKDD1
2001 Relational Instance-Based Learning with Lists and Terms
Tamás Horváth 0001, Stefan Wrobel, Uta Bohnebeck
Mach. Learn.2
2000 Extending K-Means Clustering to First-Order Representations
Mathias Kirsten, Stefan Wrobel
ILP2
2000 A sequential sampling algorithm for a general class of utility criteria
abstract
Many discovery problems, e.g., subgroup or association rule discovery, can naturally be cast as n-best hypothesis problems where the goal is to nd the n hypotheses from a given hypothesis space that score best according to a given utility function. We present a sampling algorithm that solves this problem by issuing a small number of database queries while guaranteeing precise bounds on condence and quality of solutions. Known sampling algorithms assume that the utility be the average (over the examples) of some function, which is not the case for many frequently used utility functions. We show that our algorithm works for all utilities that can be estimated with bounded error. We provide such error bounds and resulting worst-case sample bounds for some of the most frequently used utilities, and prove that there is no sampling algorithm for another popular class of utility functions. The algorithm is sequential in the sense that it starts to return (or discard) hypotheses that already...
Tobias Scheffer, Stefan Wrobel
KDD2
1998 Scalability Issues in Inductive Logic Programming
Stefan Wrobel
ALT1
1997 An Algorithm for Multi-relational Discovery of Subgroups
Stefan Wrobel
PKDD1
1996 Extensibility in Data Mining Systems
Stefan Wrobel, Dietrich Wettschereck, Edgar Sommer, Werner Emde
KDD1
1994 Concept Formation During Interactive Theory Revision
Stefan Wrobel
Mach. Learn.1
1993 On the Proper Definition of Minimality in Specialization and Theory Revision
Stefan Wrobel
ECML1
1991 Towards a Model of Grounded Concept Formation
Stefan Wrobel
IJCAI1
1988 Design Goals for Sloppy Modeling Systems
Stefan Wrobel
Int. J. Man Mach. Stud.1