Arnaud Soulet

dblp:84/2151 · DBLP profile ↗
← Back
35ranked-venue papers in the field
10as first author
10since 2021 · last 2026
0000-0001-8335-6069ORCID · verified

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 20 (8 first)Database Systems & Data Management · 9Knowledge Engineering, Semantic Web & Information Systems · 4 (2 first)Big Data, Cloud & Distributed Data Systems · 2
YearPublicationVenuePosition
2026 Rankingdom: A cooperative architecture for the on-demand analysis of Wikidata
abstract
Large knowledge graphs such as DBpedia, Wikidata, and YAGO are rich sources of structured information, widely used in domains like information retrieval and recommendation. However, their potential for supporting knowledge workers remains underexploited. As in other complex networks where specialized metrics have emerged (e.g., bibliometrics), knowledge graphs offer promising opportunities for domain-specific analysis – particularly in areas lacking established quantitative tools. In this paper, we present Rankingdom , a web application for knowledge workers to analyze any entity from Wikidata despite its heterogeneity and sheer volume. For this purpose, we introduce complementary indicators to position or evaluate any Wikidata entity within its domain, addressing analysis tasks that are challenging for traditional methods. We propose an on-demand analysis architecture that distributes computation to clients while centralizing results for frugality, by benefiting from a client-side SPARQL parallelization engine ( SParaQL ). We demonstrate the effectiveness of SParaQL through performance tests on DBpedia, Wikidata, and YAGO with two analytical queries, as well as via a real-world deployment including caching of over 10,000 entities and a user study.
Hassan Abdallah 0002, Béatrice Bouchou-Markhoff, Manon Ovide, Louise Parkin, Arnaud Soulet
Data Knowl. Eng.5
2024 RPS: A Generic Reservoir Patterns Sampler
abstract
Efficient learning from streaming data is important for modern data analysis due to the continuous and rapid evolution of data streams. Despite significant advancements in stream pattern mining, challenges persist, particularly in managing complex data streams like sequential and weighted itemsets. While reservoir sampling serves as a fundamental method for randomly selecting fixed-size samples from data streams, its application to such complex patterns remains largely unexplored. In this study, we introduce an approach that harnesses a weighted reservoir to facilitate direct pattern sampling from streaming batch data, thus ensuring scalability and efficiency. We present a generic algorithm capable of addressing temporal biases and handling various pattern types, including sequential, weighted, and unweighted itemsets. Through comprehensive experiments conducted on real-world datasets, we evaluate the effectiveness of our method, showcasing its ability to construct accurate incremental online classifiers for sequential data.
Lamine Diop, Marc Plantevit, Arnaud Soulet
IEEE Big Data3
2024 Ranking Indicator Discovery from Very Large Knowledge Graphs
abstract
Ranking indicators are essential tools for comparing the importance of various entities such as cities or scientists. While extensively used in fields like econometrics and scientometrics, many other domains lack systematic approaches for developing these indicators. In this paper, we introduce a novel method for automatically discovering ranking indicators from very large knowledge graphs. To this end, we formalize the notion of counting graph pattern (CG) as a special SPARQL query, and the concept of ideal ranking indicator as a CG whose result induces a strict total order on a set of entities. To assess the interestingness of ranking indicators, we employ the proportion of covered entities along with an inequality measure, namely the Gini coefficient. We further present Algorithm Ranking Indicator Pattern Miner (RIPM) , to efficiently identify interesting ranking indicators for a given field, thanks to pruning techniques for handling the very large search space. Our experimental study shows the effectiveness of our optimizations. It also validates that RIPM extracts transparent, diverse, and understandable indicators through a user survey and a comparison with two baselines. This work has significant implications for fields lacking dedicated communities working on ranking tasks, providing a robust tool to automatically produce ranking indicators, and the associated rankings.
Hassan Abdallah 0002, Béatrice Bouchou-Markhoff, Arnaud Soulet
Proc. VLDB Endow.3
2023 Identifying Survival-Changing Sequential Patterns for Employee Attrition Analysis
abstract
Employee attrition is a pervasive problem for many organizations, and reducing it has become a key goal in the business world. Although there is a substantial body of literature on predicting customer attrition, the literature on employee attrition is comparatively limited. Moreover, even studies that do address employee attrition often fail to consider the impact of time and duration on attrition rates. In this context, the present paper aims to fill this gap in the literature by combining frequent pattern mining in sequences of events and survival analysis with Kaplan-Meier to examine how event sequences affect employee attrition. We introduce the notion of survival-changing sequential patterns that highlight events that significantly impact the survival estimator. Our findings suggest that certain patterns are associated with a higher rate of employee retention, while the addition of specific events can have a positive or negative impact on employee survival. This research highlights the importance of analyzing event sequences and duration when attempting to reduce employee attrition rates. The practical implications of this research are significant, as it provides a framework for organizations seeking to retain their employees and enhance their overall performance.
Youssef Oubelmouh, Frédéric Fargon, Cyril de Runz, Arnaud Soulet, Cyril Veillon
DSAA4
2023 Should We Consider On-Demand Analysis in Scale-Free Networks?
Arnaud Soulet
IDA1
2022 Trie-based Output Space Itemset Sampling
abstract
Pattern sampling algorithms produce interesting patterns with a probability proportional to a given utility measure. Utility changes need quick repreprocessing when sampling patterns from large databases. In this context, existing sampling techniques require storing all data in memory, which is costly. To tackle these issues, this work enriches D. Knuth’s trie structure, avoiding 1) the need to access the database to sample since patterns are drawn directly from the enriched trie and 2) the necessity to reprocess the whole dataset when utility changes. We define the trie of occurrences that our first algorithm TPSpace (Trie-based Pattern Space) uses to materialize all of the database patterns. Factorizing transaction prefixes compresses the transactional database. TPSampling (Trie-based Pattern Sampling), our second algorithm, draws patterns from a trie of occurrences under a length-based utility measure. Experiments show that TPSampling produces thousands of patterns in seconds.
Lamine Diop, Cheikh Talibouya Diop, Arnaud Giacometti, Arnaud Soulet
IEEE Big Data4
2022 Preface
Sihem Amer-Yahia, Arnaud Soulet
Data Knowl. Eng.2
2022 Pattern on demand in transactional distributed databases
Lamine Diop, Cheikh Talibouya Diop, Arnaud Giacometti, Arnaud Soulet
Inf. Syst.4
2021 Comparison Table Generation from Knowledge Bases
Arnaud Giacometti, Béatrice Bouchou-Markhoff, Arnaud Soulet
ESWC3
2021 Reservoir Pattern Sampling in Data Streams
Arnaud Giacometti, Arnaud Soulet
ECML/PKDD (1)2
2020 Pattern Sampling in Distributed Databases
Lamine Diop, Cheikh Talibouya Diop, Arnaud Giacometti, Arnaud Soulet
ADBIS4
2020 Sequential pattern sampling with norm-based utility
Lamine Diop, Cheikh Talibouya Diop, Arnaud Giacometti, Dominique Li, Arnaud Soulet
Knowl. Inf. Syst.5
2019 Mining Significant Maximum Cardinalities in Knowledge Bases
Arnaud Giacometti, Béatrice Bouchou-Markhoff, Arnaud Soulet
ISWC (1)3
2019 Anytime Large-Scale Analytics of Linked Open Data
Arnaud Soulet, Fabian M. Suchanek
ISWC (1)1
2018 Sequential Pattern Sampling with Norm Constraints
abstract
In recent years, the field of pattern mining has shifted to user-centered methods. In such a context, it is necessary to have a tight coupling between the system and the user where mining techniques provide results at any time or within a short response time of only few seconds. Pattern sampling is a non-exhaustive method for instantly discovering relevant patterns that ensures a good interactivity while providing strong statistical guarantees due to its random nature. Curiously, such an approach investigated for itemsets and subgraphs has not yet been applied to sequential patterns, which are useful for a wide range of mining tasks and application fields. In this paper, we propose the first method for sequential pattern sampling. In addition to address sequential data, the originality of our approach is to introduce a constraint on the norm to control the length of the drawn patterns and to avoid the pitfall of the "long tail" where the rarest patterns flood the user. We propose a new constrained two-step random procedure, named CSSampling, that randomly draws sequential patterns according to frequency with an interval constraint on the norm. We demonstrate that this method performs an exact sampling. Moreover, despite the use of rejection sampling, the experimental study shows that CSSampling remains efficient and the constraint helps to draw general patterns of the "head". We also illustrate how to benefit from these sampled patterns to instantly build an associative classifier dedicated to sequences. This classification approach rivals state of the art proposals showing the interest of constrained sequential pattern sampling.
Lamine Diop, Cheikh Talibouya Diop, Arnaud Giacometti, Dominique Li, Arnaud Soulet
ICDM5
2018 How Your Supporters and Opponents Define Your Interestingness
Bruno Crémilleux, Arnaud Giacometti, Arnaud Soulet
ECML/PKDD (1)3
2018 Dense Neighborhood Pattern Sampling in Numerical Data
abstract
Pattern mining in numerical data remains a challenging task due to the pattern search space that becomes potentially infinite with real-valued dimensions. Most approaches reluctantly reduced the expressiveness of mined patterns to make possible extraction. Despite this expressiveness loss, they do not provide results within a short response time of a few seconds. This paper addresses the instant discovery of patterns in numerical data based on sampling techniques. Instead of splitting each dimension into intervals, we use a metric to introduce the density as new interestingness measure, and to define neighborhood patterns. The language of neighborhood patterns is semantically rich but in return, its size is infinite. We then present a new exact and non-enumerative random procedure to sample this infinite language according to density. An experimental study demonstrates the good compromise between precision and diversity of neighborhood patterns. Finally, in the context of associative classification, we show that a sample of neighborhood patterns is as accurate as traditional methods that traverses the entire search space.
Arnaud Giacometti, Arnaud Soulet
SDM2
2018 Representativeness of Knowledge Bases with the Generalized Benford's Law
Arnaud Soulet, Arnaud Giacometti, Béatrice Bouchou-Markhoff, Fabian M. Suchanek
ISWC (1)1
2017 MapFIM: Memory Aware Parallelized Frequent Itemset Mining in Very Large Datasets
Khanh-Chuong Duong, Mostafa Bamha, Arnaud Giacometti, Dominique Li, Arnaud Soulet, Christel Vrain
DEXA (1)5
2017 Interactive Pattern Sampling for Characterizing Unlabeled Data
Arnaud Giacometti, Arnaud Soulet
IDA2
2016 Frequent Pattern Outlier Detection Without Exhaustive Mining
Arnaud Giacometti, Arnaud Soulet
PAKDD (2)2
2015 Contextual preference mining for user profile construction
Sandra de Amo, Mouhamadou Saliou Diallo, Cheikh Talibouya Diop, Arnaud Giacometti, Dominique Li, Arnaud Soulet
Inf. Syst.6
2014 Balancing the Analysis of Frequent Patterns
Arnaud Giacometti, Dominique Li, Arnaud Soulet
PAKDD (1)3
2014 Efficiently Depth-First Minimal Pattern Mining
Arnaud Soulet, François Rioult
PAKDD (1)1
2012 Mining Contextual Preference Rules for Building User Profiles
Sandra de Amo, Mouhamadou Saliou Diallo, Cheikh Talibouya Diop, Arnaud Giacometti, Dominique Li, Arnaud Soulet
DaWaK6
2011 A Relational View of Pattern Discovery
Arnaud Giacometti, Patrick Marcel, Arnaud Soulet
DASFAA (1)3
2011 Mining Dominant Patterns in the Sky
abstract
Pattern discovery is at the core of numerous data mining tasks. Although many methods focus on efficiency in pattern mining, they still suffer from the problem of choosing a threshold that influences the final extraction result. The goal of our study is to make the results of pattern mining useful from a user-preference point of view. To this end, we integrate into the pattern discovery process the idea of skyline queries in order to mine skyline patterns in a threshold-free manner. Because the skyline patterns satisfy a formal property of dominations, they not only have a global interest but also have semantics that are easily understood by the user. In this work, we first establish theoretical relationships between pattern condensed representations and skyline pattern mining. We also show that it is possible to compute automatically a subset of measures involved in the user query which allows the patterns to be condensed and thus facilitates the computation of the skyline patterns. This forms the basis for a novel approach to mining skyline patterns. We illustrate the efficiency of our approach over several data sets including a use case from chemo informatics and show that small sets of dominant patterns are produced under various measures.
Arnaud Soulet, Chedy Raïssi, Marc Plantevit, Bruno Crémilleux
ICDM1
2010 Cube Based Summaries of Large Association Rule Sets
Marie N'diaye, Cheikh Talibouya Diop, Arnaud Giacometti, Patrick Marcel, Arnaud Soulet
ADMA (1)5
2009 Query recommendations for OLAP discovery driven analysis
abstract
Recommending database queries is an emerging and promising field of investigation. This is of particular interest in the domain of OLAP systems where the user is left with the tedious process of navigating large datacubes. In this paper we present a framework for a recommender system for OLAP users, that leverages former users' investigations to enhance discovery driven analysis. The main idea is to recommend to the user the discoveries detected in those former sessions that investigated the same unexpected data as the current session.
Arnaud Giacometti, Patrick Marcel, Elsa Nègre, Arnaud Soulet
DOLAP4
2008 Adequate Condensed Representations of Patterns
Arnaud Soulet, Bruno Crémilleux
ECML/PKDD (1)1
2008 Adequate condensed representations of patterns
Arnaud Soulet, Bruno Crémilleux
Data Min. Knowl. Discov.1
2005 Average Number of Frequent (Closed) Patterns in Bernouilli and Markovian Databases
abstract
In data mining, enumerate the frequent or the closed patterns is often the first difficult task leading to the association rules discovery. The number of these patterns represents a great interest. The lower bound is known to be constant whereas the upper bound is exponential, but both situations correspond to pathological cases. For the first time, we give an average analysis of the number of frequent or closed patterns. Average analysis is often closer to real situations and gives more information about the role of the parameters. In this paper, two probabilistic models are studied: a Bernoulli and a Markovian. In both models and for large databases, we prove that the number of frequent patterns, for a fixed frequency threshold, is exponential in the number of items and polynomial in the number of transactions. On the other hand, for a proportional frequency threshold, the number of frequent patterns is polynomial in the number of items and does not involve the number of transactions. Finally, we prove in the Bernoulli model that the number of closed patterns, for a proportional frequency threshold, is polynomial in the number of items.
Loïck Lhote, François Rioult, Arnaud Soulet
ICDM3
2005 Optimizing Constraint-Based Mining by Automatically Relaxing Constraints
abstract
In constraint-based mining, the monotone and anti-monotone properties are exploited to reduce the search space. Even if a constraint has not such suitable properties, existing algorithms can be re-used thanks to an approximation, called relaxation. In this paper, we automatically compute monotone relaxations of primitive-based constraints. First, we show that the latter are a superclass of combinations of both kinds of monotone constraints. Second, we add two operators to detect the properties of monotonicity of such constraints. Finally, we define relaxing operators to obtain monotone relaxations of them.
Arnaud Soulet, Bruno Crémilleux
ICDM1
2005 An Efficient Framework for Mining Flexible Constraints
Arnaud Soulet, Bruno Crémilleux
PAKDD1
2004 Condensed Representation of Emerging Patterns
Arnaud Soulet, Bruno Crémilleux, François Rioult
PAKDD1