EDBT 2026 Demo / reviewers in the wild / expert
Bart Goethals
dblp:g/BartGoethals
· DBLP profile ↗
72ranked-venue papers in the field
8as first author
15since 2021 · last 2025
0000-0001-9327-9554ORCID · verified
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 45 (8 first)Information Retrieval & Web Search · 19Database Systems & Data Management · 5Big Data, Cloud & Distributed Data Systems · 2Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | What Data is Really Necessary? A Feasibility Study of Inference Data Minimization for Recommender SystemsabstractData minimization is a legal principle requiring personal data processing to be limited to what is necessary for a specified purpose. Operationalizing this principle for recommender systems, which rely on extensive personal data, remains a significant challenge. This paper conducts a feasibility study on minimizing implicit feedback inference data for such systems. We propose a novel problem formulation, analyze various minimization techniques, and investigate key factors influencing their effectiveness. We demonstrate that substantial inference data reduction is technically feasible without significant performance loss. However, its practicality is critically determined by two factors: the technical setting (e.g., performance targets, choice of model) and user characteristics (e.g., history size, preference complexity). Thus, while we establish its technical feasibility, we conclude that data minimization remains practically challenging and its dependence on the technical and user context makes a universal standard for data 'necessity' difficult to implement. Jens Leysen, Marco Favier, Bart Goethals |
CIKM | 3 |
| 2025 | APS Explorer: Navigating Algorithm Performance Spaces for Informed Dataset SelectionabstractDataset selection is crucial for offline recommender system experiments, as mismatched data (e.g., sparse interaction scenarios require datasets with low user-item density) can lead to unreliable results. Yet, 86\% of ACM RecSys 2024 papers provide no justification for their dataset choices, with most relying on just four datasets: Amazon (38\%), MovieLens (34\%), Yelp (15\%), and Gowalla (12\%). While Algorithm Performance Spaces (APS) were proposed to guide dataset selection, their adoption has been limited due to the absence of an intuitive, interactive tool for APS exploration. Therefore, we introduce the APS Explorer, a web-based visualization tool for interactive APS exploration, enabling data-driven dataset selection. The APS Explorer provides three interactive features: (1) an interactive PCA plot showing dataset similarity via performance patterns, (2) a dynamic meta-feature table for dataset comparisons, and (3) a specialized visualization for pairwise algorithm performance. Tobias Vente, Michael Heep, Abdullah Abbas, Theodor Sperle, Jöran Beel, Bart Goethals |
RecSys | 6 |
| 2024 | Session-based News Recommendation Using Cohesive PatternsabstractIn the rapidly evolving field of news recommendation, where user preferences are highly dynamic and content quickly becomes obsolete, providing timely and relevant recommendations presents a significant challenge. Traditional recommender systems typically rely on complex collaborative filtering models that depend on extensive user histories. In the news domain, however, such histories are often scarce due to the high prevalence of anonymous users. To address these challenges, we introduce a novel session-based recommendation method that leverages cohesive sequential pattern mining. Rather than relying on traditional frequency-based pattern utility metrics, our approach prioritizes pattern cohesiveness, which captures the temporal proximity of item interactions within a pattern, resulting in recommendations that align more closely with the user’s ongoing session.We conduct a comprehensive empirical evaluation of our approach using four large-scale real-world news datasets. The results demonstrate that our method, SeQcsp, significantly outperforms state-of-the-art session-based recommendation algorithms in terms of accuracy, ranking quality, as well as diversity. Furthermore, SeQcsp provides recommendations faster than most existing methods and is effective for both short and long user sessions, highlighting its robustness, adaptability, and efficiency. Mozhgan Karimi, Len Feremans, Boris Cule, Bart Goethals |
IEEE Big Data | 4 |
| 2024 | The Role of Unknown Interactions in Implicit Matrix Factorization - A Probabilistic ViewabstractMatrix factorization is a well-known and effective methodology for top-k list recommendation. It became widely known during the Netflix challenge in 2006, and since then, many adapted and improved versions have been published. A particularly interesting matrix factorization algorithm called iALS (for implicit Alternating Least Squares) adapts the method for implicit feedback, i.e. a setting where only a very small amount of positive labels are available along with a majority of unknown labels. Compared to the classical task of rating prediction, learning from implicit feedback is applicable to many more domains, as the data is more abundant and requires less effort to elicit from users. However, the sparsity, imbalance, and implicit nature of the signal also pose unique challenges to retrieving the most relevant items to recommend. Joey De Pauw, Bart Goethals |
RecSys | 2 |
| 2024 | A Framework and Toolkit for Testing the Correctness of Recommendation AlgorithmsabstractEvaluating recommender systems adequately and thoroughly is an important task. Significant efforts are dedicated to proposing metrics, methods, and protocols for doing so. However, there has been little discussion in the recommender systems’ literature on the topic of testing. In this work, we adopt and adapt concepts from the software testing domain, e.g., code coverage, metamorphic testing, or property-based testing, to help researchers to detect and correct faults in recommendation algorithms. We propose a test suite that can be used to validate the correctness of a recommendation algorithm, and thus identify and correct issues that can affect the performance and behavior of these algorithms. Our test suite contains both black box and white box tests at every level of abstraction, i.e., system, integration, and unit. To facilitate adoption, we release RecPack Tests , an open-source Python package containing template test implementations. We use it to test four popular Python packages for recommender systems: RecPack , PyLensKit , Surprise , and Cornac . Despite the high test coverage of each of these packages, we find that we are still able to uncover undocumented functional requirements and even some bugs. This validates our thesis that testing the correctness of recommendation algorithms can complement traditional methods for evaluating recommendation algorithms. Lien Michiels, Robin Verachtert, Andres Ferraro, Kim Falk, Bart Goethals |
Trans. Recomm. Syst. | 5 |
| 2023 | The Impact of a Popularity Punishing Hyperparameter on ItemKNN Recommendation Performance
Robin Verachtert, Jeroen Craps, Lien Michiels, Bart Goethals |
ECIR (2) | 4 |
| 2023 | How Should We Measure Filter Bubbles? A Regression Model and Evidence for Online NewsabstractNews media play an important role in democratic societies. Central to fulfilling this role is the premise that users should be exposed to diverse news. However, news recommender systems are gaining popularity on news websites, which has sparked concerns over filter bubbles. More specifically, editors, policy-makers and scholars are worried that these news recommender systems may expose users to less diverse content over time. To the best of our knowledge, this hypothesis has not been tested in a longitudinal observational study of real users that interact with a real news website. Such observational studies require the use of research methods that are robust and can account for the many covariates that may influence the diversity of recommendations at any given time. In this work, we propose an analysis model to study whether the variety of articles recommended to a user decreases over time in such an observational study design. Further, we present results from two case studies using aggregated and anonymized data that were collected by two western European news websites employing a collaborative filtering-based news recommender system to serve (personalized) recommendations to their users. Through these case studies we validate empirically that our modeling assumptions are sound and supported by the data, and that our model obtains more reliable and interpretable results than analysis methods used in prior empirical work on filter bubbles. Our case studies provide evidence of a small decrease in the topic variety of a user’s recommendations in the first weeks after they sign up, but no evidence of a decrease in political variety. Lien Michiels, Jorre T. A. Vannieuwenhuyze, Jens Leysen, Robin Verachtert, Annelien Smets, Bart Goethals |
RecSys | 6 |
| 2023 | Leveraging Sequential Episode Mining for Session-Based News Recommendation
Mozhgan Karimi, Boris Cule, Bart Goethals |
WISE | 3 |
| 2023 | Pessimistic Decision-Making for Recommender SystemsabstractModern recommender systems are often modelled under the sequential decision-making paradigm, where the systemdecideswhich recommendations to show in order to maximise some notion of either imminent or long-term reward. Such methods often require an explicit model of the reward a certain context-action pair will yield – for example, the probability of a click on a recommendation. This common machine learning task is highly non-trivial, as the data-generating process for contexts and actions can be skewed by the recommender system itself. Indeed, when the deployed recommendation policy at data collection time does not pick its actions uniformly-at-random, this leads to a selection bias that can impede effective reward modelling. This in turn makes off-policy learning – the typical setup in industry – particularly challenging. Existing approaches for value-based learning break down in such environments. In this work, we propose and validate a generalpessimisticreward modelling approach for off-policy learning in recommendation. Bayesian uncertainty estimates allow us to express scepticism about our own reward model, which can in turn be used to generate a conservative decision rule. We show how it alleviates a well-known decision making phenomenon known as the Optimiser’s Curse, and draw parallels with existing work on pessimistic policy learning. Leveraging the available closed-form expressions for both the posterior mean and variance when a ridge regressor models the reward, we show how to apply pessimism effectively and efficiently to an off-policy recommendation use-case. Empirical observations in a wide range of simulated environments show that pessimistic decision-making leads to a significant and robust increase in recommendation performance. The merits of our approach are most outspoken in realistic settings with limited logging randomisation, limited training samples, and larger action spaces. We discuss the impact of our contributions in the context of related applications like computational advertising, and present a scope for future research based on hybrid off-/on-policy bandit learning methods for recommendation. Olivier Jeunen, Bart Goethals |
Trans. Recomm. Syst. | 2 |
| 2022 | RecPack: An(other) Experimentation Toolkit for Top-N Recommendation using Implicit Feedback DataabstractRecPack is an easy-to-use, flexible and extensible toolkit for top-N recommendation with implicit feedback data. Its goal is to support researchers with the development of their recommendation algorithms, from similarity-based to deep learning algorithms, and allow for correct, reproducible and reusable experimentation. In this demo, we give an overview of the package and show how researchers can use it to their advantage when developing recommendation algorithms. Lien Michiels, Robin Verachtert, Bart Goethals |
RecSys | 3 |
| 2022 | Who do you think I am? Interactive User Modelling with Item MetadataabstractRecommender systems are used in many different applications and contexts, however their main goal can always be summarised as “connecting relevant content to interested users”. Explanations have been found to help recommender systems achieve this goal by giving users a look under the hood that helps them understand why they are recommended certain items. Furthermore, explanations can be considered to be the first step towards interacting with the system. Indeed, for a user to give feedback and guide the system towards better understanding her preferences, it helps if the user has a better idea of what the system has already learned. Joey De Pauw, Koen Ruymbeek, Bart Goethals |
RecSys | 3 |
| 2022 | PETSC: pattern-based embedding for time series classification
Len Feremans, Boris Cule, Bart Goethals |
Data Min. Knowl. Discov. | 3 |
| 2021 | Pessimistic Reward Models for Off-Policy Learning in RecommendationabstractMethods for bandit learning from user interactions often require a model of the reward a certain context-action pair will yield – for example, the probability of a click on a recommendation. This common machine learning task is highly non-trivial, as the data-generating process for contexts and actions is often skewed by the recommender system itself. Indeed, when the deployed recommendation policy at data collection time does not pick its actions uniformly-at-random, this leads to a selection bias that can impede effective reward modelling. This in turn makes off-policy learning – the typical setup in industry – particularly challenging. Olivier Jeunen, Bart Goethals |
RecSys | 2 |
| 2021 | Top-K Contextual Bandits with Equity of ExposureabstractThe contextual bandit paradigm provides a general framework for decision-making under uncertainty. It is theoretically well-defined and well-studied, and many personalisation use-cases can be cast as a bandit learning problem. Because this allows for the direct optimisation of utility metrics that rely on online interventions (such as click-through-rate (CTR)), this framework has become an attractive choice to practitioners. Historically, the literature on this topic has focused on a one-sided, user-focused notion of utility, overall disregarding the perspective of content providers in online marketplaces (for example, musical artists on streaming services). If not properly taken into account – recommendation systems in such environments are known to lead to unfair distributions of attention and exposure, which can directly affect the income of the providers. Recent work has shed a light on this, and there is now a growing consensus that some notion of “equity of exposure” might be preferable to implement in many recommendation use-cases. Olivier Jeunen, Bart Goethals |
RecSys | 2 |
| 2021 | High-dimensional Sparse Embeddings for Collaborative FilteringabstractA widely adopted paradigm in the design of recommender systems is to represent users and items as vectors, often referred to as latent factors or embeddings. Embeddings can be obtained using a variety of recommendation models and served in production using a variety of data engineering solutions. Embeddings also facilitate transfer learning, where trained embeddings from one model are reused in another. In contrast, some of the best-performing collaborative filtering models today are high-dimensional linear models that do not rely on factorization, and so they do not produce embeddings [27, 28]. They also require pruning, amounting to a trade-off between the model size and the density of the predicted affinities. This paper argues for the use of high-dimensional, sparse latent factor models, instead. We propose a new recommendation model based on a full-rank factorization of the inverse Gram matrix. The resulting high-dimensional embeddings can be made sparse while still factorizing a dense affinity matrix. We show how the embeddings combine the advantages of latent representations with the performance of high-dimensional linear models. Jan Van Balen, Bart Goethals |
WWW | 2 |
| 2020 | Closed-Form Models for Collaborative Filtering with Side-InformationabstractRecent work has shown that, despite their simplicity, item-based models optimised through ridge regression can attain highly competitive results on collaborative filtering tasks. As these models are analytically computable and thus forgo the need for often expensive iterative optimisation procedures, they are an attractive choice for practitioners. We study the applicability of such closed-form models to implicit-feedback collaborative filtering when additional side-information or metadata about items is available. Two complementary extensions to the easer paradigm are proposed, based on collective and additive models. Through an extensive empirical analysis on several large-scale datasets, we show that our methods can effectively exploit side-information whilst retaining a closed-form solution, and improve upon the state-of-the-art without increasing the computational complexity of the original easer approach. Additionally, empirical results demonstrate that the use of side-information leads to more “long tail” items being recommended, benefiting the recommendations’ coverage of the item catalogue. Olivier Jeunen, Jan Van Balen, Bart Goethals |
RecSys | 3 |
| 2019 | Pattern-Based Anomaly Detection in Mixed-Type Time Series
Len Feremans, Vincent Vercruyssen, Boris Cule, Wannes Meert, Bart Goethals |
ECML/PKDD (1) | 5 |
| 2019 | Efficient similarity computation for collaborative filtering in dynamic environmentsabstractThe problem of computing all pairwise similarities in a large collection of vectors is a well-known and common data mining task. As the number and dimensionality of these vectors keeps increasing, however, currently existing approaches are often unable to meet the strict efficiency requirements imposed by the environments they need to perform in. Real-time neighbourhood-based collaborative filtering (CF) is one example of such an environment in which performance is critical. Olivier Jeunen, Koen Verstrepen, Bart Goethals |
RecSys | 3 |
| 2019 | Interactive evaluation of recommender systems with SNIPER: an episode mining approachabstractRecommender systems are typically evaluated using either offline methods, online methods, or through user studies. In this paper we take an episode mining approach to analysing recommender system data and we demonstrate how we can use SNIPER, a tool for interactive pattern mining, to analyse and understand the behaviour of recommender systems. We describe the required data format, and present a useful scenario of how a user can interact with the system to answer questions about the quality of recommendations. Sandy Moens, Olivier Jeunen, Bart Goethals |
RecSys | 3 |
| 2019 | Efficiently mining cohesion-based patterns and rules in event sequences
Boris Cule, Len Feremans, Bart Goethals |
Data Min. Knowl. Discov. | 3 |
| 2019 | Proximity Forest: an effective and scalable distance-based classifier for time series
Benjamin Lucas, Ahmed Shifaz, Charlotte Pelletier, Lachlan O'Neill, Nayyar Abbas Zaidi, Bart Goethals, François Petitjean, Geoffrey I. Webb |
Data Min. Knowl. Discov. | 6 |
| 2018 | Mining Top-k Quantile-based Cohesive Sequential PatternsabstractFinding patterns in long event sequences is an important data mining task. Two decades ago research focused on finding all frequent patterns, where the anti-monotonic property of support was used to design efficient algorithms. Recent research focuses on producing a smaller output containing only the most interesting patterns. To achieve this goal, we introduce a new interestingness measure by computing the proportion of the occurrences of a pattern that are cohesive. This measure is robust to outliers, and is applicable to sequential patterns. We implement an efficient algorithm based on constrained prefix-projected pattern growth and pruning based on an upper bound to uncover the set of top-k quantile-based cohesive sequential patterns. We run experiments to compare our method with existing state-of-the-art methods for sequential pattern mining and show that our algorithm is efficient and produces qualitatively interesting patterns on large event sequences. Len Feremans, Boris Cule, Bart Goethals |
SDM | 3 |
| 2018 | Analyzing concept drift and shift from sample data
Geoffrey I. Webb, Loong Kuan Lee, Bart Goethals, François Petitjean |
Data Min. Knowl. Discov. | 3 |
| 2017 | Combining Instance and Feature Neighbors for Efficient Multi-label ClassificationabstractMulti-label classification problems occur naturally in different domains. For example, within text categorization the goal is to predict a set of topics for a document, and within image scene classification the goal is to assign labels to different objects in an image. In this work we propose a combination of two variations of k nearest neighborhoods (kNN) where the first neighborhood is computed instance (or row) based and the second neighborhood is feature (or column) based. Instance based kNN is inspired by user-based collaborative filtering, while feature kNN is inspired by item-based collaborative filtering. Finally we apply a linear combination of instance and feature neighbors scores and apply a single threshold to predict the set of labels. Experiments on various multi-label datasets show that our algorithm outperforms other state-of-the-art methods such as ML-kNN, IBLR and Binary Relevance with SVM, on different evaluation metrics. Finally our algorithm uses an inverted index during neighborhood search and scales to extreme datasets that have millions of instances, features and labels. Len Feremans, Boris Cule, Celine Vens, Bart Goethals |
DSAA | 4 |
| 2017 | Cleaning Data with Forbidden ItemsetsabstractMethods for cleaning dirty data typically rely on additional information about the data, such as user-specified constraints that specify when a database is dirty. These constraints often involve domain restrictions and illegal value combinations. Traditionally, a database is considered clean if all constraints are satisfied. However, many real-world scenario's only have a dirty database available. In such a context, we adopt a dynamic notion of data quality, in which the data is clean if an error discovery algorithm does not find any errors. We introduce forbidden itemsets which capture unlikely value co-occurrences in dirty data, and we derive properties of the lift measure to provide an efficient algorithm for mining low lift forbidden itemsets. We further introduce a repair method which guarantees that the repaired database does not contain any low lift forbidden itemsets. The algorithm uses nearest neighbor imputation to suggest possible repairs. Optional user interaction can easily be integrated into the proposed cleaning method. Evaluation on real-world data shows that errors are typically discovered with high precision, while the suggested repairs are of good quality and do not introduce new forbidden itemsets, as desired. Joeri Rammelaere, Floris Geerts, Bart Goethals |
ICDE | 3 |
| 2016 | Efficient Discovery of Sets of Co-occurring Items in Event Sequences
Boris Cule, Len Feremans, Bart Goethals |
ECML/PKDD (1) | 3 |
| 2016 | Pattern Based Sequence ClassificationabstractSequence classification is an important task in data mining. We address the problem of sequence classification using rules composed of interesting patterns found in a dataset of labelled sequences and accompanying class labels. We measure the interestingness of a pattern in a given class of sequences by combining the cohesion and the support of the pattern. We use the discovered patterns to generate confident classification rules, and present two different ways of building a classifier. The first classifier is based on an improved version of the existing method of classification based on association rules, while the second ranks the rules by first measuring their value specific to the new data object. Experimental results show that our rule based classifiers outperform existing comparable classifiers in terms of accuracy and stability. Additionally, we test a number of pattern feature based models that use different kinds of patterns as features to represent each sequence as a feature vector. We then apply a variety of machine learning algorithms for sequence classification, experimentally demonstrating that the patterns we discover represent the sequences well, and prove effective for the classification task. Boris Cule, Bart Goethals |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2015 | Efficient Cluster Detection by Ordered Neighborhoods
Emin Aksehirli, Bart Goethals, Emmanuel Müller |
DaWaK | 2 |
| 2015 | Cohesion based co-location pattern miningabstractBecause of a wide range of applications, e.g., GPS applications and location based services, spatial pattern discovery is an important task in data mining. A co-location pattern is defined as a subset of spatial items whose instances are often located together in spatial proximity. Current co-location mining algorithms are unable to quantify the spatial proximity of a co-location pattern. We propose a co-location pattern miner aiming to discover co-location patterns in a multidimensional spatial structure by measuring the cohesion of a pattern. We present two ways to build the co-location pattern miner, FromOne and FromAll, in an attempt to find a balance between accuracy and runtime. Additionally, we propose a method named Fre-ball to transform a structure into a transaction database, after which any existing itemset mining algorithm can be used to find the co-location patterns. An experimental evaluation shows that FromOne and Fre-ball are more efficient than existing methods. The usefulness of our methods is demonstrated by applying them on the publicly available geographical data of the city of Antwerp in Belgium. Boris Cule, Bart Goethals |
DSAA | 3 |
| 2015 | Determining the Presence of Political Parties in Social Circles
Christophe Van Gysel, Bart Goethals, Maarten de Rijke |
ICWSM | 2 |
| 2015 | Mining Association Rules in Graphs Based on Frequent Cohesive Itemsets
Tayena Hendrickx, Boris Cule, Pieter Meysman, Stefan Naulaerts, Kris Laukens, Bart Goethals |
PAKDD (2) | 6 |
| 2015 | Top-N Recommendation for Shared AccountsabstractStandard collaborative filtering recommender systems assume that every account in the training data represents a single user. However, multiple users often share a single account. A typical example is a single shopping account for the whole family. Traditional recommender systems fail in this situation. If contextual information is available, context aware recommender systems are the state-of-the-art solution. Yet, often no contextual information is available. Therefore, we introduce the challenge of recommending to shared accounts in the absence of contextual information. We propose a solution to this challenge for all cases in which the reference recommender system is an item-based top-N collaborative filtering recommender system, generating recommendations based on binary, positive-only feedback. We experimentally show the advantages of our proposed solution for tackling the problems that arise from the existence of shared accounts on multiple datasets. Koen Verstrepen, Bart Goethals |
RecSys | 2 |
| 2014 | Welcome from DSAA 2014 chairsabstractData driven scientific discovery approach has already been agreed to be an important emerging paradigm for computing in areas including social, service, Internet of Things (or sensor networks), and cloud. Under this paradigm, Big Data is the core that drives new researches in many areas, from environmental to social. There are many new scientific challenges when facing this big data phenomenon, ranging from capture, creation, storage, search, sharing, analysis, and visualization. The complication here is not just the storage, I/O, query, and performance, but also the integration across heterogeneous, interdependent complex data resources for real-time decision-making, collaboration, and ultimately value co-creation. Data sciences encompass the larger areas of data analytics, machine learning and managing big data. Advanced data analytics has become essential to glean a deep understanding of large data sets and to convert data into actionable intelligence. With the rapid growth in the volumes of data available to enterprises, Government and on the web, automated techniques for analyzing the data have become essential. Philip S. Yu, Masaru Kitsuregawa, Hiroshi Motoda, Bart Goethals, Minyi Guo, Longbing Cao, George Karypis, Irwin King, Wei Wang 0379 |
DSAA | 4 |
| 2014 | Unifying nearest neighbors collaborative filteringabstractWe study collaborative filtering for applications in which there exists for every user a set of items about which the user has given binary, positive-only feedback (one-class collaborative filtering). Take for example an on-line store that knows all past purchases of every customer. An important class of algorithms for one-class collaborative filtering are the nearest neighbors algorithms, typically divided into user-based and item-based algorithms. We introduce a reformulation that unifies user- and item-based nearest neighbors algorithms and use this reformulation to propose a novel algorithm that incorporates the best of both worlds and outperforms state-of-the-art algorithms. Additionally, we propose a method for naturally explaining the recommendations made by our algorithm and show that this method is also applicable to existing user-based nearest neighbors methods. Koen Verstrepen, Bart Goethals |
RecSys | 2 |
| 2014 | Mining frequent itemsets in a stream
Toon Calders, Nele Dexters, Joris J. M. Gillis, Bart Goethals |
Inf. Syst. | 4 |
| 2013 | Frequent Itemset Mining for Big DataabstractFrequent Itemset Mining (FIM) is one of the most well known techniques to extract knowledge from data. The combinatorial explosion of FIM methods become even more problematic when they are applied to Big Data. Fortunately, recent improvements in the field of parallel programming already provide good tools to tackle this problem. However, these tools come with their own technical challenges, e.g. balanced data distribution and inter-communication costs. In this paper, we investigate the applicability of FIM techniques on the MapReduce platform. We introduce two new methods for mining large datasets: Dist-Eclat focuses on speed while BigFIM is optimized to run on really large datasets. In our experiments we show the scalability of our methods. Sandy Moens, Emin Aksehirli, Bart Goethals |
IEEE BigData | 3 |
| 2013 | Cartification: A Neighborhood Preserving Transformation for Mining High Dimensional DataabstractThe analysis of high dimensional data comes with many intrinsic challenges. In particular, cluster structures become increasingly hard to detect when the data includes dimensions irrelevant to the individual clusters. With increasing dimensionality, distances between pairs of objects become very similar, and hence, meaningless for knowledge discovery. In this paper we propose Cartification, a new transformation to circumvent this problem. We transform each object into an item set, which represents the neighborhood of the object. We do this for multiple views on the data, resulting in multiple neighborhoods per object. This transformation enables us to preserve the essential pair wise-similarities of objects over multiple views, and hence, to improve knowledge discovery in high dimensional data. Our experiments show that frequent item set mining on the certified data outperforms competing clustering approaches on the original data space, including traditional clustering, random projections, principle component analysis, subspace clustering, and clustering ensemble. Emin Aksehirli, Bart Goethals, Emmanuel Müller, Jilles Vreeken |
ICDM | 2 |
| 2013 | Mining Interesting Itemsets in Graph Datasets
Boris Cule, Bart Goethals, Tayena Hendrickx |
PAKDD (1) | 2 |
| 2013 | ClaSP: An Efficient Algorithm for Mining Frequent Closed Sequences
Antonio Gomariz, Manuel Campos, Roque Marín, Bart Goethals |
PAKDD (1) | 4 |
| 2013 | Itemset Based Sequence Classification
Boris Cule, Bart Goethals |
ECML/PKDD (1) | 3 |
| 2012 | MARBLES: Mining Association Rules Buried in Long Event SequencesabstractSequential pattern discovery is a well-studied field in data mining. Episodes are sequential patterns that describe events that often occur in the vicinity of each other. Episodes can impose restrictions on the order of the events, which makes them a versatile technique for describing complex patterns in the sequence. Most of the research on episodes deals with special cases such as serial and parallel episodes, while discovering general episodes is surprisingly understudied. This is particularly true when it comes to discovering association rules between them. In this paper we propose an algorithm that mines association rules between two general episodes. On top of the traditional definitions of frequency and confidence, we introduce two novel confidence measures for the rules. The major challenge in mining these association rules is pattern explosion. To limit the output, we aim to eliminate all redundant rules. We define the class of closed association rules, and show that this class contains all non-redundant output. To make the algorithm efficient, we use further pruning steps along the way. First of all, we generate only free and closed frequent episodes from which we create candidate rules, we speed up the evaluation of the rules, and finally prune the remaining non-closed rules from the output. Boris Cule, Nikolaj Tatti, Bart Goethals |
SDM | 3 |
| 2012 | An inductive database system based on virtual mining views
Hendrik Blockeel, Toon Calders, Élisa Fromont, Bart Goethals, Adriana Prado, Céline Robardet |
Data Min. Knowl. Discov. | 4 |
| 2012 | Mining frequent conjunctive queries in relational databases through dependency discovery
Bart Goethals, Dominique Laurent 0001, Wim Le Page, Cheikh Tidiane Dieng |
Knowl. Inf. Syst. | 1 |
| 2011 | Mining Train Delays
Boris Cule, Bart Goethals, Sven Tassenoy, Sabine Verboven |
IDA | 2 |
| 2011 | GaMuSo: Graph Base Music Recommendation in a Social Bookmarking Service
Jeroen De Knijf, Anthony M. L. Liekens, Bart Goethals |
IDA | 3 |
| 2011 | MIME: a framework for interactive visual pattern miningabstractWe present a framework for interactive visual pattern mining. Our system enables the user to browse through the data and patterns easily and intuitively, using a toolbox consisting of interestingness measures, mining algorithms and post-processing algorithms to assist in identifying interesting patterns. By mining interactively, we enable the user to combine their subjective interestingness measure and background knowledge with a wide variety of objective measures to easily and quickly mine the most important and interesting patterns. Basically, we enable the user to become an essential part of the mining algorithm. Our demo currently applies to mining interesting itemsets and association rules, and its extension to episodes and decision trees is ongoing. Bart Goethals, Sandy Moens, Jilles Vreeken |
KDD | 1 |
| 2011 | MIME: A Framework for Interactive Visual Pattern Mining
Bart Goethals, Sandy Moens, Jilles Vreeken |
ECML/PKDD (3) | 1 |
| 2010 | Discovery and Application of Functional Dependencies in Conjunctive Query Mining
Bart Goethals, Dominique Laurent 0001, Wim Le Page |
DaWak | 1 |
| 2010 | Approximation of Frequentness Probability of Itemsets in Uncertain DataabstractMining frequent item sets from transactional datasets is a well known problem with good algorithmic solutions. Most of these algorithms assume that the input data is free from errors. Real data, however, is often affected by noise. Such noise can be represented by uncertain datasets in which each item has an existence probability. Recently, Bernecker et al. (2009) proposed the frequentness probability, i.e., the probability that a given item set is frequent, to select item sets in an uncertain database. A dynamic programming approach to evaluate this measure was given as well. We argue, however, that for the setting of Bernecker et al. (2009), that assumes independence between the items, already well-known statistical tools exist. We show how the frequentness probability can be approximated extremely accurately using a form of the central limit theorem. We experimentally evaluated our approximation and compared it to the dynamic programming approach. The evaluation shows that our approximation method is extremely accurate even for very small databases while at the same time it has much lower memory overhead and computation time. Toon Calders, Calin Garboni, Bart Goethals |
ICDM | 3 |
| 2010 | Predicting the severity of a reported bugabstractThe severity of a reported bug is a critical factor in deciding how soon it needs to be fixed. Unfortunately, while clear guidelines exist on how to assign the severity of a bug, it remains an inherent manual process left to the person reporting the bug. In this paper we investigate whether we can accurately predict the severity of a reported bug by analyzing its textual description using text mining algorithms. Based on three cases drawn from the open-source community (Mozilla, Eclipse and GNOME), we conclude that given a training set of sufficient size (approximately 500 reports per severity), it is possible to predict the severity with a reasonable accuracy (both precision and recall vary between 0.65-0.75 with Mozilla and Eclipse; 0.70-0.85 in the case of GNOME). Ahmed Lamkanfi, Serge Demeyer, Emanuel Giger, Bart Goethals |
MSR | 4 |
| 2010 | Efficient Pattern Mining of Uncertain Data with Sampling
Toon Calders, Calin Garboni, Bart Goethals |
PAKDD (1) | 3 |
| 2010 | Mining Association Rules in Long Sequences
Boris Cule, Bart Goethals |
PAKDD (1) | 2 |
| 2009 | A New Constraint for Mining Sets in SequencesabstractDiscovering interesting patterns in event sequences is a popular task in the field of data mining. Most existing methods try to do this based on some measure of cohesion to determine an occurrence of a pattern, and a frequency threshold to determine if the pattern occurs often enough. We introduce a new constraint based on a new interestingness measure combining the cohesion and the frequency of a pattern. For a dataset consisting of a single sequence, the cohesion is measured as the average length of the smallest intervals containing the pattern for each occurrence of its events, and the frequency is measured as the probability of observing an event of that pattern. We present a similar constraint for datasets consisting of multiple sequences. We present algorithms to efficiently identify the thus defined interesting patterns, given a dataset and a user-defined threshold. After applying our method to both synthetic and real-life data, we conclude that it indeed gives intuitive results in a number of applications. Boris Cule, Bart Goethals, Céline Robardet |
SDM | 2 |
| 2008 | Mining Views: Database Views for Data MiningabstractWe present a system towards the integration of data mining into relational databases. To this end, a relational database model is proposed, based on the so called virtual mining views. We show that several types of patterns and models over the data, such as itemsets, association rules and decision trees, can be represented and queried using a unifying framework. Hendrik Blockeel, Toon Calders, Élisa Fromont, Bart Goethals, Adriana Prado |
ICDE | 4 |
| 2008 | Sequence Mining Automata: A New Technique for Mining Frequent Sequences under Regular ExpressionsabstractIn this paper we study the problem of mining frequent sequences satisfying a given regular expression. Previous approaches to solve this problem were focusing on its search space, pushing (in some way) the given regular expression to prune unpromising candidate patterns. On the contrary, we focus completely on the given input data and regular expression. We introduce sequence mining automata (SMA), a specialized kind of Petri Net that while reading input sequences, it produces for each sequence all and only the patterns contained in the sequence and that satisfy the given regular expression. Based on this automaton, we develop a family of algorithms. Our thorough experimentation on different datasets and application domains confirms that in many cases our methods outperform the current state of the art of frequent sequence mining algorithms using regular expressions (in some cases of orders of magnitude). Roberto Trasarti, Francesco Bonchi, Bart Goethals |
ICDM | 3 |
| 2008 | An inductive database prototype based on virtual mining viewsabstractWe present a prototype of an inductive database. Our system enables the user to query not only the data stored in the database but also generalizations (e.g. rules or trees) over these data through the use of virtual mining views. The mining views are relational tables that virtually contain the complete output of data mining algorithms executed over a given dataset. The prototype implemented into PostgreSQL currently integrates frequent itemset, association rule and decision tree mining. We illustrate the interactive and iterative capabilities of our system with a description of a complete data mining scenario. Hendrik Blockeel, Toon Calders, Élisa Fromont, Bart Goethals, Adriana Prado, Céline Robardet |
KDD | 4 |
| 2008 | Mining Association Rules of Simple Conjunctive QueriesabstractWe present an algorithm for mining association rules in arbitrary relational databases.We define association rules over a simple, but appealing subclass of conjunctive queries, and show that many interesting patterns can be found.We propose an efficient algorithm and a database-oriented implementation in SQL, together with several promising and convincing experimental results. Bart Goethals, Wim Le Page, Heikki Mannila |
SDM | 1 |
| 2008 | Guest Editors' Introduction: Special issue of Selected Papers from ECML PKDD 2008
Walter Daelemans, Bart Goethals, Katharina Morik |
Data Min. Knowl. Discov. | 2 |
| 2007 | Mining Frequent Itemsets in a StreamabstractWe study the problem of finding frequent itemsets in a continuous stream of transactions. The current frequency of an itemset in a stream is defined as its maximal frequency over all possible windows in the stream from any point in the past until the current state that satisfy a minimal length constraint. Properties of this new measure are studied and an incremental algorithm that allows, at any time, to immediately produce the current frequencies of all frequent itemsets is proposed. Experimental and theoretical analysis show that the space requirements for the algorithm are extremely small for many realistic data distributions. Toon Calders, Nele Dexters, Bart Goethals |
ICDM | 3 |
| 2007 | Non-derivable itemset miningabstractAll frequent itemset mining algorithms rely heavily on the monotonicity principle for pruning. This principle allows for excluding candidate itemsets from the expensive counting phase. In this paper, we present sound and complete deduction rules to derive bounds on the support of an itemset. Based on these deduction rules, we construct a condensed representation of all frequent itemsets, by removing those itemsets for which the support can be derived, resulting in the so called Non-Derivable Itemsets (NDI) representation. We also present connections between our proposal and recent other proposals for condensed representations of frequent itemsets. Experiments on real-life datasets show the effectiveness of the NDI representation, making the search for frequent non-derivable itemsets a useful and tractable alternative to mining all frequent itemsets. Toon Calders, Bart Goethals |
Data Min. Knowl. Discov. | 2 |
| 2006 | Mining rank-correlated sets of numerical attributesabstractWe study the mining of interesting patterns in the presence of numerical attributes. Instead of the usual discretization methods, we propose the use of rank based measures to score the similarity of sets of numerical attributes. New support measures for numerical data are introduced, based on extensions of Kendall's tau, and Spearman's Footrule and rho. We show how these support measures are related. Furthermore, we introduce a novel type of pattern combining numerical and categorical attributes. We give efficient algorithms to find all frequent patterns for the proposed support measures, and evaluate their performance on real-life datasets. Toon Calders, Bart Goethals, Szymon Jaroszewicz |
KDD | 2 |
| 2006 | Integrating Pattern Mining in Relational Databases
Toon Calders, Bart Goethals, Adriana Prado |
PKDD | 2 |
| 2005 | Mining tree queries in a graphabstractWe present an algorithm for mining tree-shaped patterns in a large graph. Novel about our class of patterns is that they can contain constants, and can contain existential nodes which are not counted when determining the number of occurrences of the pattern in the graph. Our algorithm has a number of provable optimality properties, which are based on the theory of conjunctive database queries. We propose a database-oriented implementation in SQL, and report upon some initial experimental results obtained with our implementation on graph data about food webs, about protein interactions, and about citation analysis. Bart Goethals, Eveline Hoekx, Jan Van den Bussche |
KDD | 1 |
| 2005 | Depth-First Non-Derivable Itemset MiningabstractMining frequent itemsets is one of the main problems in data mining. Much effort went into developing efficient and scalable algorithms for this problem. When the support threshold is set too low, however, or the data is highly correlated, the number of frequent itemsets can become too large, independently of the algorithm used. Therefore, it is often more interesting to mine a reduced collection of interesting itemsets, i.e., a condensed representation. Recently, in this context, the non-derivable itemsets were proposed as an important class of itemsets. An itemset is called derivable when its support is completely determined by the support of its subsets. As such, derivable itemsets represent redundant information and can be pruned from the collection of frequent itemsets. It was shown both theoretically and experimentally that the collection of non-derivable frequent itemsets is in general much smaller than the complete set of frequent itemsets. A breadth-first, Apriori-based algorithm, called NDI, to find all non-derivable itemsets was proposed. In this paper we present a depth-first algorithm, dfNDI, that is based on Eclat for mining the non-derivable itemsets. dfNDI is evaluated on real-life datasets, and experiments show that dfNDI outperforms NDI with an order of magnitude. Toon Calders, Bart Goethals |
SDM | 2 |
| 2005 | Mining Non-Derivable Association RulesabstractAssociation rule mining typically results in large amounts of redundant rules. We introduce efficient methods for deriving tight bounds for confidences of association rules, given their subrules. If the lower and upper bounds of a rule coincide, the confidence is uniquely determined by the subrules and the rule can be pruned as redundant, or derivable, without any loss of information. Experiments on real, dense benchmark data sets show that, depending on the case, up to 99–99.99 % of rules are derivable. A lossy pruning strategy, where those rules are removed for which the width of the bounded confidence interval is 1 percentage point, reduced the number of rules by a furher order of magnitude. The novelty of our work is twofold. First, it gives absolute bounds for the confidence instead of relying on point estimates or heuristics. Second, no specific inference system is assumed for computing the bounds; instead, the bounds follow from the definition of association rules. Our experimental results demonstrate that the bounds are usually narrow and the approach has great practical significance, also in comparison to recent related approaches. 1 Bart Goethals, Juho Muhonen, Hannu Toivonen |
SDM | 1 |
| 2005 | Tight upper bounds on the number of candidate patternsabstractIn the context of mining for frequent patterns using the standard levelwise algorithm, the following question arises: given the current level and the current set of frequent patterns, what is the maximal number of candidate patterns that can be generated on the next level? We answer this question by providing tight upper bounds, derived from a combinatorial result from the sixties by Kruskal and Katona. Our result is useful to secure existing algorithms from a combinatorial explosion of the number of candidate patterns. Floris Geerts, Bart Goethals, Jan Van den Bussche |
ACM Trans. Database Syst. | 2 |
| 2004 | FP-Bonsai: The Art of Growing and Pruning Small FP-Trees
Francesco Bonchi, Bart Goethals |
PAKDD | 2 |
| 2003 | Minimal k-Free Representations of Frequent Sets
Toon Calders, Bart Goethals |
PKDD | 2 |
| 2002 | Mining All Non-derivable Frequent Itemsets
Toon Calders, Bart Goethals |
PKDD | 2 |
| 2001 | A Tight Upper Bound on the Number of Candidate PatternsabstractIn the context of mining for frequent patterns using the standard level-wise algorithm, the following question arises: given the current level and the current set of frequent patterns, what is the maximal number of candidate patterns that can be generated on the next level? We answer this question by providing a tight upper bound, derived from a combinatorial result by J. Kruskal (1963) and G. Katona (1968). Our result is useful for reducing the number of database scans. Floris Geerts, Bart Goethals, Jan Van den Bussche |
ICDM | 2 |
| 2000 | On Supporting Interactive Association Rule Mining
Bart Goethals, Jan Van den Bussche |
DaWaK | 1 |
| 2000 | A data mining framework for optimal product selection in retail supermarket data: the generalized PROFSET modelabstractArticle Free Access Share on A data mining framework for optimal product selection in retail supermarket data: the generalized PROFSET model Authors: Tom Brijs Limburg University Centre, Universitaire Campus, B-3590 Diepenbeek Limburg University Centre, Universitaire Campus, B-3590 DiepenbeekView Profile , Bart Goethals Limburg University Centre, Universitaire Campus, B-3590 Diepenbeek Limburg University Centre, Universitaire Campus, B-3590 DiepenbeekView Profile , Gilbert Swinnen Limburg University Centre, Universitaire Campus, B-3590 Diepenbeek Limburg University Centre, Universitaire Campus, B-3590 DiepenbeekView Profile , Koen Vanhoof Limburg University Centre, Universitaire Campus, B-3590 Diepenbeek Limburg University Centre, Universitaire Campus, B-3590 DiepenbeekView Profile , Geert Wets Limburg University Centre, Universitaire Campus, B-3590 Diepenbeek Limburg University Centre, Universitaire Campus, B-3590 DiepenbeekView Profile Authors Info & Claims KDD '00: Proceedings of the sixth ACM SIGKDD international conference on Knowledge discovery and data miningAugust 2000 Pages 300–304https://doi.org/10.1145/347090.347156Online:01 August 2000Publication History 56citation1,552DownloadsMetricsTotal Citations56Total Downloads1,552Last 12 Months49Last 6 weeks5 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Tom Brijs, Bart Goethals, Gilbert Swinnen, Koen Vanhoof, Geert Wets |
KDD | 2 |