VLDB 2026 Research / reviewers in the wild / expert
Céline Robardet
dblp:r/CRobardet
· DBLP profile ↗
44ranked-venue papers in the field
4as first author
15since 2021 · last 2026
0000-0002-8583-9408ORCID · verified
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 38 (4 first)Database Systems & Data Management · 4Information Retrieval & Web Search · 1Knowledge Engineering, Semantic Web & Information Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Differentiable parameter-less co-clustering using graph neural networksabstractAbstract Co-clustering refers to the simultaneous clustering of rows and columns in a data matrix, uncovering joint patterns between two distinct sets, such as documents and terms or users and products. Traditional co-clustering algorithms typically rely on discrete optimization techniques based on enumeration, which can limit both scalability and flexibility. In this paper, we introduce a differentiable programming approach to co-clustering that enables the continuous optimization of co-partitions using graph neural networks. Our method is grounded in an associative co-clustering quality measure that is independent of the number of clusters and dynamically adjusts this parameter by jointly considering both partitions. By leveraging automatic differentiation and graph neural networks, our approach scales to very large datasets while maintaining high-quality co-cluster structures. We evaluate our method using different types of graph neural networks and initialization strategies. Furthermore, when compared with recent state-of-the-art methods for co-clustering and graph clustering, our approach achieves competitive or superior results in terms of accuracy. Most importantly, it is the only algorithm that successfully completes on the largest benchmark dataset. Alessio Ragno, Pierre-Angelo Peyrie, Marc Plantevit, Ruggero G. Pensa, Céline Robardet |
Data Min. Knowl. Discov. | 5 |
| 2025 | Explainability of Molecular Graph Neural NetworkabstractGraph Neural Networks (GNNs) have demonstrated strong performance in molecular interaction prediction, but their interpretability remains limited, especially in domain-specific applications like ligand-receptor modeling. This paper presents a model-agnostic explainer for GNN-CLS, a specialized GNN model designed to predict interactions between molecules and olfactory receptor proteins. The proposed method uses cooperative game theory to identify influential molecular substructures and receptor sequence regions, offering faithful and theoretically grounded explanations of model predictions. This approach enhances transparency by revealing which features drive predictive outcomes, helping bridge the gap between model performance and chemical insight. The contributions include a formal framework for relevance attribution and interaction analysis, positioning this work at the intersection of explainable AI and computational chemistry. Ataollah Kamal, Matej Hladis, Jérémie Topin, Marc Plantevit, Sébastien Fiorucci, Céline Robardet |
DSAA | 6 |
| 2025 | Diffusion for Explainable Unsupervised Anomaly DetectionabstractStatistical anomaly detection is critical across various domains, including healthcare, finance, industry, and cybersecurity. While supervised methods often achieve high performance, the limited availability of labeled data requires effective unsupervised techniques. In this paper, we introduce Dataset Sampling Iterative Learning (DSIL), a novel iterative learning framework for unsu-pervised anomaly detection leveraging generative modeling with diffusion. Our approach progressively refines an unlabeled dataset by identifying and removing anomalies, effectively approximating a semi-supervised setup. We demonstrate the efficiency of our framework with Diffusion Time Estimation (DTE). Furthermore, it enables better explainability through a novel approach of noised-feature discovery. Extensive experiments against unsupervised methods on both synthetic and real-world datasets demonstrate improved state-of-the-art performance. Finally, we suggest a novel usage of existing metrics to evaluate the explainability of anomaly detection models. Elouan Vincent, Alexandre Dréan, Julien Perez, Marc Plantevit, Céline Robardet |
DSAA | 5 |
| 2025 | Faithful Explanations for Graph Classification Using Logic
Alessio Ragno, Marc Plantevit, Céline Robardet |
ECML/PKDD (4) | 3 |
| 2025 | Leveraging internal representations of GNNs with Shapley values
Ataollah Kamal, Alessio Ragno, Marc Plantevit, Céline Robardet |
Data Min. Knowl. Discov. | 4 |
| 2024 | DiffVersify: a Scalable Approach to Differentiable Pattern Mining with Coverage Regularization
Thibaut Chataing, Julien Perez, Marc Plantevit, Céline Robardet |
ECML/PKDD (6) | 4 |
| 2024 | On GNN explainability with activation rules
Luca Veyrin-Forrer, Ataollah Kamal, Stefan Duffner, Marc Plantevit, Céline Robardet |
Data Min. Knowl. Discov. | 5 |
| 2023 | Electricity Price Forecasting based on Order Books: a differentiable optimization approachabstractWe consider day-ahead electricity price forecasting on the European market. In this market, participants can offer electricity for sale or purchase for a specific price by submitting overnight orders. Market operators determine the market clearing price – the price at which the amount of electricity supplied equals the amount of electricity demanded – using the Euphemia balancing algorithm. EUPHEMIA is a quadratic optimization problem that maximizes the social welfare defined as the sum of the supplier surplus and consumer surplus while ensuring a null energy balance. This mechanism deeply influences the price calculation, but has so far been little considered in electricity price forecasting algorithms. Existing models are generally based on identifying relationships between exogenous characteristics (consumption and production forecasts) and the market clearing price to be predicted. A few studies have examined the EUPHEMIA mechanism during prediction, by doing costly manual transformations on order books. In this article, we overcome this limitation by considering the pricing mechanism during model training. For this, we use a predict-and-optimize strategy with differentiable optimization. We design a fully differentiable and scalable solving method for the EUPHEMIA optimization problem and apply it on real-life data from the European Power Exchange (EPEX). We design different model architectures using our differentiable solver and empirically study the impact of taking into account the optimal calculation of prices within the training of the neural network. Léonard Tschora, Tias Guns, Erwan Pierre, Marc Plantevit, Céline Robardet |
DSAA | 5 |
| 2023 | Forecasting Electricity Prices: An Optimize Then Predict-Based Approach
Léonard Tschora, Erwan Pierre, Marc Plantevit, Céline Robardet |
IDA | 4 |
| 2023 | Methods for explaining Top-N recommendations through subgroup discovery
Mouloud Iferroudjene, Corentin Lonjarret, Céline Robardet, Marc Plantevit, Martin Atzmüller |
Data Min. Knowl. Discov. | 3 |
| 2022 | Towards a better identification of Bitcoin actors by supervised learning
Rafael Ramos Tubino, Céline Robardet, Rémy Cazabet |
Data Knowl. Eng. | 2 |
| 2022 | In pursuit of the hidden features of GNN's internal representations
Luca Veyrin-Forrer, Ataollah Kamal, Stefan Duffner, Marc Plantevit, Céline Robardet |
Data Knowl. Eng. | 5 |
| 2021 | Interpretable Summaries of Black Box Incident Triaging with Subgroup DiscoveryabstractThe need of predictive maintenance comes with an increasing number of incidents reported by monitoring systems and equipment/software users. In the front line, on-call engineers (OCEs) have to quickly assess the degree of severity of an incident and decide which service to contact for corrective actions. To automate these decisions, several predictive models have been proposed, but the most efficient models are opaque (say, black box), strongly limiting their adoption. In this paper, we propose an efficient black box model based on 170K incidents reported to our company over the last 7 years and emphasize on the need of automating triage when incidents are massively reported on thousands of servers running our product, an ERP. Recent developments in eXplainable Artificial Intelligence (XAI) help in providing global explanations to the model, but also, and most importantly, with local explanations for each model prediction/outcome. Sadly, providing a human with an explanation for each outcome is not conceivable when dealing with an important number of daily predictions. To address this problem, we propose an original data-mining method rooted in Subgroup Discovery, a pattern mining technique with the natural ability to group objects that share similar explanations of their black box predictions and provide a description for each group. We evaluate this approach and present our preliminary results which give us good hope towards an effective OCE's adoption. We believe that this approach provides a new way to address the problem of model agnostic outcome explanation. Youcef Remil, Ahmed Anes Bendimerad, Marc Plantevit, Céline Robardet, Mehdi Kaytoue-Uberall |
DSAA | 4 |
| 2021 | Sequential recommendation with metric models based on frequent sequences
Corentin Lonjarret, Roch Auburtin, Céline Robardet, Marc Plantevit |
Data Min. Knowl. Discov. | 3 |
| 2021 | User-Driven Geolocated Event Detection in Social MediaabstractEvent detection is one of the most important research topics in social media analysis. Despite this interest, few researchers have addressed the problem of identifying geolocated events in an unsupervised way, and none includes user interests during the process. In this paper, we tackle the problem of local event detection from social media data. We present a method to automatically identify events by evaluating the burstiness of hashtags in a geographical area and a time interval, and at the same time integrating user feedback. We devise two algorithms to discover user-driven events. The first one relies on an exact enumeration process, while the other directly samples the space of events. In our empirical study, we provide evidence that geolocated events cannot be detected by non location-aware methods. We also show that our methods (i) outperform by a factor of two to several orders of magnitude state-of-the-art methods designed to discover geolocated events, (ii) are more robust to noise, and (iii) produce high quality events with respect to user interests. Ahmed Anes Bendimerad, Marc Plantevit, Céline Robardet, Sihem Amer-Yahia |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2020 | Why Should I Trust This Item? Explaining the Recommendations of any ModelabstractExplainable AI has received a lot of attention over the past decade, with the proposal of many methods explaining black box classifiers such as neural networks. Despite the ubiquity of recommender systems in the digital world, only few researchers have attempted to explain their functioning, whereas it raises e.g., ethical issues. Indeed, recommender systems direct user choices to a large extent and their impact is important as they give access to only a small part of the range of items (e.g., products and/or services), as the submerged part of the iceberg. Consequently, they limit access to other resources. The potentially negative effects of these systems have been pointed out as phenomena like echo chambers and winner-take-all effects, because the internal logic of these systems is to likely enclose the consumer in a "dej́ a vu" loop. Therefore, it is crucial to provide explanations' of such recommender systems and to identify the user data that led the system to make a specific recommendation. This makes it possible to evaluate recommender systems not only regarding their efficiency (i.e., their capability to recommend an item that was actually chosen by the user), but also w.r.t. the diversity, relevance and timeliness of the active data used to make the recommendation. In this paper, we propose a deep analysis of 7 state-of-the-art models learnt on 6 datasets based on the identification of the items or the sequences of items actively used by the models. The proposed method, which is based on subgroup discovery with different pattern languages (i.e., itemsets and sequences), provides interpretable explanations of the recommendations - useful to compare different models and explain the reasons behind the recommendation to the user. Corentin Lonjarret, Céline Robardet, Marc Plantevit, Roch Auburtin, Martin Atzmüller |
DSAA | 2 |
| 2020 | Gibbs Sampling Subjectively Interesting TilesabstractThe local pattern mining literature has long struggled with the so-called pattern explosion problem: the size of the set of patterns found exceeds the size of the original data. This causes computational problems (enumerating a large set of patterns will inevitably take a substantial amount of time) as well as problems for interpretation and usability (trawling through a large set of patterns is often impractical). Two complementary research lines aim to address this problem. The first aims to develop better measures of interestingness, in order to reduce the number of uninteresting patterns that are returned [ 6 , 10 ]. The second aims to avoid an exhaustive enumeration of all ‘interesting’ patterns (where interestingness is quantified in a more traditional way, e.g. frequency), by directly sampling from this set in a way that more ‘interesting’ patterns are sampled with higher probability [ 2 ]. Unfortunately, the first research line does not reduce computational cost, while the second may miss out on the most interesting patterns. In this paper, we combine the best of both worlds for mining interesting tiles [ 8 ] from binary databases. Specifically, we propose a new pattern sampling approach based on Gibbs sampling, where the probability of sampling a pattern is proportional to their subjective interestingness [ 6 ]—an interestingness measure reported to better represent true interestingness. The experimental evaluation confirms the theory, but also reveals an important weakness of the proposed approach which we speculate is shared with any other pattern sampling approach. We thus conclude with a broader discussion of this issue, and a forward look. Ahmed Anes Bendimerad, Jefrey Lijffijt, Marc Plantevit, Céline Robardet, Tijl De Bie |
IDA | 4 |
| 2020 | SIAS-miner: mining subjectively interesting attributed subgraphsabstractAbstract Data clustering, local pattern mining, and community detection in graphs are three mature areas of data mining and machine learning. In recent years, attributed subgraph mining has emerged as a new powerful data mining task in the intersection of these areas. Given a graph and a set of attributes for each vertex, attributed subgraph mining aims to find cohesive subgraphs for which (some of) the attribute values have exceptional values. The principled integration of graph and attribute data poses two challenges: (1) the definition of a pattern syntax (the abstract form of patterns) that is intuitive and lends itself to efficient search, and (2) the formalization of the interestingness of such patterns. We propose an integrated solution to both of these challenges. The proposed pattern syntax improves upon prior work in being both highly flexible and intuitive. Plus, we define an effective and principled algorithm to enumerate patterns of this syntax. The proposed approach for quantifying interestingness of these patterns is rooted in information theory, and is able to account for background knowledge on the data. While prior work quantified the interestingness for the cohesion of the subgraph and for the exceptionality of its attributes separately, then combining these in a parameterized trade-off, we instead handle this trade-off implicitly in a principled, parameter-free manner. Empirical results confirm we can efficiently find highly interesting subgraphs. Ahmed Anes Bendimerad, Ahmad Mel, Jefrey Lijffijt, Marc Plantevit, Céline Robardet, Tijl De Bie |
Data Min. Knowl. Discov. | 5 |
| 2019 | FSSD - A Fast and Efficient Algorithm for Subgroup Set DiscoveryabstractSubgroup discovery (SD) is the task of discovering interpretable patterns in the data that stand out w.r.t. some property of interest. Discovering patterns that accurately discriminate a class from the others is one of the most common SD tasks. Standard approaches of the literature are based on local pattern discovery, which is known to provide an overwhelmingly large number of redundant patterns. To solve this issue, pattern set mining has been proposed: instead of evaluating the quality of patterns separately, one should consider the quality of a pattern set as a whole. The goal is to provide a small pattern set that is diverse and well-discriminant to~the target class. In this work, we introduce a novel formulation of the task of diverse subgroup set discovery where both discriminative power and diversity of the subgroup set are incorporated in the same quality measure. We propose an efficient and parameter-free algorithm dubbed FSSD and based on a greedy scheme. FSSD uses several optimization strategies that enable to efficiently provide a high quality pattern set in a short amount of time. Adnene Belfodil, Aimene Belfodil, Ahmed Anes Bendimerad, Philippe Lamarre, Céline Robardet, Mehdi Kaytoue-Uberall, Marc Plantevit |
DSAA | 5 |
| 2019 | Contrastive Antichains in HierarchiesabstractConcepts are often described in terms of positive integer-valued attributes that are organized in a hierarchy. For example, cities can be described in terms of how many places there are of various types (e.g. nightlife spots, residences, food venues), and these places are organized in a hierarchy (e.g. a Portuguese restaurant is a type of food venue). This hierarchy imposes particular constraints on the values of related attributes---e.g. there cannot be more Portuguese restaurants than food venues. Moreover, knowing that a city has many food venues makes it less surprising that it also has many Portuguese restaurants, and vice versa. In the present paper, we attempt to characterize such concepts in terms of so-called contrastive antichains: particular kinds of subsets of their attributes and their values. We address the question of when a contrastive antichain is interesting, in the sense that it concisely describes the unique aspects of the concept, and this while duly taking into account the known attribute dependencies implied by the hierarchy. Our approach is capable of accounting for previously identified contrastive antichains, making iterative mining possible. Besides the interestingness measure, we also present an algorithm that scales well in practice, and demonstrate the usefulness of the method in an extensive empirical results section. Ahmed Anes Bendimerad, Jefrey Lijffijt, Marc Plantevit, Céline Robardet, Tijl De Bie |
KDD | 4 |
| 2019 | Rank correlated subgroup discovery
Mohamed-Ali Hammal, Hélène Mathian, Luc Merchez, Marc Plantevit, Céline Robardet |
J. Intell. Inf. Syst. | 5 |
| 2018 | Mining exceptional closed patterns in attributed graphs
Ahmed Anes Bendimerad, Marc Plantevit, Céline Robardet |
Knowl. Inf. Syst. | 3 |
| 2016 | Unsupervised Exceptional Attributed Sub-Graph Mining in Urban DataabstractGeo-located social media provide a wealth of information that describes urban areas based on user descriptions and comments. Such data makes possible to identify meaningful city neighborhoods on the basis of the footprints left by a large and diverse population that uses this type of media. In this paper, we present some methods to exhibit the predominant activities and their associated urban areas to automatically describe a whole city. Based on a suitable attributed graph model, our approach identifies neighborhoods with homogeneous and exceptional characteristics. We introduce the novel problem of exceptional sub-graph mining in attributed graphs and propose a complete algorithm that takes benefits from new upper bounds and pruning properties. We also propose an approach to sample the space of exceptional sub-graphs within a given time-budget. Experiments performed on 10 real datasets are reported and demonstrate the relevancy and the limits of both approaches. Ahmed Anes Bendimerad, Marc Plantevit, Céline Robardet |
ICDM | 3 |
| 2016 | Guest editors' introduction to the EcmlPkdd 2016 journal track special issue of Machine Learning
Thomas Gärtner 0001, Mirco Nanni, Andrea Passerini, Céline Robardet |
Data Min. Knowl. Discov. | 4 |
| 2015 | Gazouille: Detecting and Illustrating Local Events from Geolocalized Social Media Streams
Pierre Houdyer, Albrecht Zimmermann, Mehdi Kaytoue-Uberall, Marc Plantevit, Céline Robardet |
ECML/PKDD (3) | 6 |
| 2014 | Triggering patterns of topology changes in dynamic graphsabstractTo describe the dynamics taking place in networks that structurally change over time, we propose an approach to search for attributes whose value changes impact the topology of the graph. In several applications, it appears that the variations of a group of attributes are often followed by some structural changes in the graph that one may assume they generate. We formalize the triggering pattern discovery problem as a method jointly rooted in sequence mining and graph analysis. We apply our approach on three real-world dynamic graphs of different natures - a co-authoring network, an airline network, and a social bookmarking system - assessing the relevancy of the triggering pattern mining approach. Mehdi Kaytoue-Uberall, Yoann Pitarch, Marc Plantevit, Céline Robardet |
ASONAM | 4 |
| 2014 | Granularity of Co-evolution Patterns in Dynamic Attributed Graphs
Elise Desmier, Marc Plantevit, Céline Robardet, Jean-François Boulicaut |
IDA | 3 |
| 2013 | When TEDDY meets GrizzLY: temporal dependency discovery for triggering road deicing operationsabstractTemporal dependencies between multiple sensor data sources link two types of events if the occurrence of one is repeatedly followed by the appearance of the other in a certain time interval. TEDDY algorithm aims at discovering such dependencies, identifying the statically significant time intervals with a chi2 test. We present how these dependencies can be used within the GrizzLY project to tackle an environmental and technical issue: the deicing of the roads. This project aims to wisely organize the deicing operations of an urban area, based on several sensor network measures of local atmospheric phenomena. A spatial and temporal dependency-based model is built from these data to predict freezing alerts. Céline Robardet, Vasile-Marian Scuturici, Marc Plantevit, Antoine Fraboulet |
KDD | 1 |
| 2013 | Trend Mining in Dynamic Attributed Graphs
Elise Desmier, Marc Plantevit, Céline Robardet, Jean-François Boulicaut |
ECML/PKDD (1) | 3 |
| 2013 | Parameter-less co-clustering for star-structured heterogeneous data
Dino Ienco, Céline Robardet, Ruggero G. Pensa, Rosa Meo |
Data Min. Knowl. Discov. | 2 |
| 2013 | Mining Graph Topological Patterns: Finding Covariations among Vertex DescriptorsabstractWe propose to mine the graph topology of a large attributed graph by finding regularities among vertex descriptors. Such descriptors are of two types: 1) the vertex attributes that convey the information of the vertices themselves and 2) some topological properties used to describe the connectivity of the vertices. These descriptors are mostly of numerical or ordinal types and their similarity can be captured by quantifying their covariation. Mining topological patterns relies on frequent pattern mining and graph topology analysis to reveal the links that exist between the relation encoded by the graph and the vertex attributes. We propose three interestingness measures of topological patterns that differ by the pairs of vertices considered while evaluating up and down co-variations between vertex descriptors. An efficient algorithm that combines search and pruning strategies to look for the most relevant topological patterns is presented. Besides a classical empirical study, we report case studies on four real-life networks showing that our approach provides valuable knowledge. Adriana Prado, Marc Plantevit, Céline Robardet, Jean-François Boulicaut |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2012 | An inductive database system based on virtual mining views
Hendrik Blockeel, Toon Calders, Élisa Fromont, Bart Goethals, Adriana Prado, Céline Robardet |
Data Min. Knowl. Discov. | 6 |
| 2012 | Complex systems science: Dreams of universality, interdisciplinarity realityabstractUsing a large database (∼215,000 records) of relevant articles, we empirically study the complex systems field and its claims to find universal principles applying to systems in general. The study of references shared by the articles allows us to obtain a global point of view on the structure of this highly interdisciplinary field. We show that its overall coherence does not arise from a universal theory, but instead from computational techniques and fruitful adaptations of the idea of self‐organization to specific systems. We also find that communication between different disciplines goes through specific “trading zones,” i.e., subcommunities that create an interface around specific tools (a DNA microchip) or concepts (a network). Sebastian Grauwin, Guillaume Beslon, Eric Fleury, Sara Franceschelli, Céline Robardet, Jean-Baptiste Rouquier, Pablo Jensen |
J. Assoc. Inf. Sci. Technol. | 5 |
| 2009 | Constraint-Based Pattern Mining in Dynamic GraphsabstractDynamic graphs are used to represent relationships between entities that evolve over time. Meaningful patterns in such structured data must capture strong interactions and their evolution over time. In social networks, such patterns can be seen as dynamic community structures, i.e., sets of individuals who strongly and repeatedly interact. In this paper, we propose a constraint-based mining approach to uncover evolving patterns. We propose to mine dense and isolated subgraphs defined by two user-parameterized constraints. The temporal evolution of such patterns is captured by associating a temporal event type to each identified subgraph. We consider five basic temporal events: The formation, dissolution, growth, diminution and stability of subgraphs from one time stamp to the next. We propose an algorithm that finds such subgraphs in a time series of graphs processed incrementally. The extraction is feasible due to efficient patterns and data pruning strategies. We demonstrate the applicability of our method on several real-world dynamic graphs and extract meaningful evolving communities. Céline Robardet |
ICDM | 1 |
| 2009 | A New Constraint for Mining Sets in SequencesabstractDiscovering interesting patterns in event sequences is a popular task in the field of data mining. Most existing methods try to do this based on some measure of cohesion to determine an occurrence of a pattern, and a frequency threshold to determine if the pattern occurs often enough. We introduce a new constraint based on a new interestingness measure combining the cohesion and the frequency of a pattern. For a dataset consisting of a single sequence, the cohesion is measured as the average length of the smallest intervals containing the pattern for each occurrence of its events, and the frequency is measured as the probability of observing an event of that pattern. We present a similar constraint for datasets consisting of multiple sequences. We present algorithms to efficiently identify the thus defined interesting patterns, given a dataset and a user-defined threshold. After applying our method to both synthetic and real-life data, we conclude that it indeed gives intuitive results in a number of applications. Boris Cule, Bart Goethals, Céline Robardet |
SDM | 3 |
| 2009 | Constraint-Based Subspace ClusteringabstractIn high dimensional data, the general performance of traditional clustering algorithms decreases.This is partly because the similarity criterion used by these algorithms becomes inadequate in high dimensional space.Another reason is that some dimensions are likely to be irrelevant or contain noisy data, thus hiding a possible clustering.To overcome these problems, subspace clustering techniques, which can automatically find clusters in relevant subsets of dimensions, have been developed.However, due to the huge number of subspaces to consider, these techniques often lack efficiency.In this paper we propose to extend the framework of bottomup subspace clustering algorithms by integrating background knowledge and, in particular, instance-level constraints to speed up the enumeration of subspaces.We show how this new framework can be applied to both density and distancebased bottom-up subspace clustering techniques.Our experiments on real datasets show that instance-level constraints cannot only increase the efficiency of the clustering process but also the accuracy of the resultant clustering. Élisa Fromont, Adriana Prado, Céline Robardet |
SDM | 3 |
| 2009 | Closed patterns meet n-ary relationsabstractSet pattern discovery from binary relations has been extensively studied during the last decade. In particular, many complete and efficient algorithms for frequent closed set mining are now available. Generalizing such a task to n -ary relations ( n ≥ 2) appears as a timely challenge. It may be important for many applications, for example, when adding the time dimension to the popular objects × features binary case. The generality of the task (no assumption being made on the relation arity or on the size of its attribute domains) makes it computationally challenging. We introduce an algorithm called Data-Peeler. From an n -ary relation, it extracts all closed n -sets satisfying given piecewise (anti) monotonic constraints. This new class of constraints generalizes both monotonic and antimonotonic constraints. Considering the special case of ternary relations, Data-Peeler outperforms the state-of-the-art algorithms CubeMiner and Trias by orders of magnitude. These good performances must be granted to a new clever enumeration strategy allowing to efficiently enforce the closeness property. The relevance of the extracted closed n -sets is assessed on real-life 3-and 4-ary relations. Beyond natural 3-or 4-ary relations, expanding a relation with an additional attribute can help in enforcing rather abstract constraints such as the robustness with respect to binarization. Furthermore, a collection of closed n -sets is shown to be an excellent starting point to compute a tiling of the dataset. Loïc Cerf, Jérémy Besson, Céline Robardet, Jean-François Boulicaut |
ACM Trans. Knowl. Discov. Data | 3 |
| 2008 | An inductive database prototype based on virtual mining viewsabstractWe present a prototype of an inductive database. Our system enables the user to query not only the data stored in the database but also generalizations (e.g. rules or trees) over these data through the use of virtual mining views. The mining views are relational tables that virtually contain the complete output of data mining algorithms executed over a given dataset. The prototype implemented into PostgreSQL currently integrates frequent itemset, association rule and decision tree mining. We illustrate the interactive and iterative capabilities of our system with a description of a complete data mining scenario. Hendrik Blockeel, Toon Calders, Élisa Fromont, Bart Goethals, Adriana Prado, Céline Robardet |
KDD | 6 |
| 2008 | Data Peeler: Contraint-Based Closed Pattern Mining in n-ary RelationsabstractSet pattern discovery from binary relations has been extensively studied during the last decade. In particular, many complete and efficient algorithms which extract frequent closed sets are now available. Generalizing such a task to n-ary relations (n ≥ 2) appears as a timely challenge. It may be important for many applications, e.g., when adding the time dimension to the popular objects × features binary case. The generality of the task — no assumption being made on the relation arity or on the size of its attribute domains — makes it computationally challenging. We introduce an algorithm called Data-Peeler. From a n-ary relation, it extracts all closed n-sets satisfying given piecewise (anti)-monotonic constraints. This new class of constraints generalizes both monotonic and anti-monotonic constraints. Considering the special case of ternary relations, Data-Peeler outperforms the state-of-the-art algorithms CubeMiner and Trias by orders of magnitude. These good performances must be granted to a new clever enumeration strategy allowing an efficient closeness checking. An original application on a real-life 4-ary relation is used to assess the relevancy of closed n-sets constraint-based mining. Loïc Cerf, Jérémy Besson, Céline Robardet, Jean-François Boulicaut |
SDM | 3 |
| 2007 | A New Way to Aggregate Preferences: Application to Eurovision Song Contests
Jérémy Besson, Céline Robardet |
IDA | 2 |
| 2005 | A Bi-clustering Framework for Categorical Data
Ruggero G. Pensa, Céline Robardet, Jean-François Boulicaut |
PKDD | 2 |
| 2004 | Constraint-Based Mining of Formal Concepts in Transactional Data
Jérémy Besson, Céline Robardet, Jean-François Boulicaut |
PAKDD | 2 |
| 2001 | Comparison of Three Objective Functions for Conceptual Clustering
Céline Robardet, Fabien Feschet |
PKDD | 1 |
| 2000 | An Experimental Study of Partition Quality Indices in Clustering
Céline Robardet, Fabien Feschet, Nicolas Nicoloyannis |
PKDD | 1 |