Arnaud Giacometti

dblp:g/AGiacometti · DBLP profile ↗
← Back
29ranked-venue papers
12as first author
4since 2021 · last 2022
0000-0003-0270-5146ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 25 · 11 first-author · 4 since 2021Artificial intelligence and machine learning · 13 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2022 Trie-based Output Space Itemset Sampling
abstract
Pattern sampling algorithms produce interesting patterns with a probability proportional to a given utility measure. Utility changes need quick repreprocessing when sampling patterns from large databases. In this context, existing sampling techniques require storing all data in memory, which is costly. To tackle these issues, this work enriches D. Knuth’s trie structure, avoiding 1) the need to access the database to sample since patterns are drawn directly from the enriched trie and 2) the necessity to reprocess the whole dataset when utility changes. We define the trie of occurrences that our first algorithm TPSpace (Trie-based Pattern Space) uses to materialize all of the database patterns. Factorizing transaction prefixes compresses the transactional database. TPSampling (Trie-based Pattern Sampling), our second algorithm, draws patterns from a trie of occurrences under a length-based utility measure. Experiments show that TPSampling produces thousands of patterns in seconds.
Lamine Diop, Cheikh Talibouya Diop, Arnaud Giacometti, Arnaud Soulet
IEEE Big Data3
2022 Pattern on demand in transactional distributed databases
Lamine Diop, Cheikh Talibouya Diop, Arnaud Giacometti, Arnaud Soulet
Inf. Syst.3
2021 Comparison Table Generation from Knowledge Bases
Arnaud Giacometti, Béatrice Bouchou-Markhoff, Arnaud Soulet
ESWC1
2021 Reservoir Pattern Sampling in Data Streams
Arnaud Giacometti, Arnaud Soulet
ECML/PKDD (1)1
2020 Pattern Sampling in Distributed Databases
Lamine Diop, Cheikh Talibouya Diop, Arnaud Giacometti, Arnaud Soulet
ADBIS3
2020 MAPK-means: A clustering algorithm with quantitative preferences on attributes
abstract
This paper describes a new semi-supervised clustering algorithm as part of a more general framework of interactive exploratory clustering, that favors the exploration of possible clustering solutions so that an expert tailors the best clustering according to her domain knowledge and preferences. Co ntrary to most existing approaches, the novel algorithm considers the feature space as a first class citizen for the exploration of alternative solutions. Our proposal represents and integrates quantitative preferences on attributes that will guide the exploration of possible solutions by learning an appropriate space metric. It also achieves a compromise clustering based on expert confidence, between a data-driven and a user-driven solution and converges with a good complexity. We show experimentally that our method is also able to deal with irrelevant user preferences and correct those choices in order to achieve a better solution. Experiments show that the best results may be achieved only with the addition of preferences to traditional metric learning algorithms and that our approach performs better than state-of-the-art algorithms.
Adnan El Moussawi, Arnaud Giacometti, Nicolas Labroche, Arnaud Soulet
Intell. Data Anal.2
2020 Sequential pattern sampling with norm-based utility
Lamine Diop, Cheikh Talibouya Diop, Arnaud Giacometti, Dominique Li, Arnaud Soulet
Knowl. Inf. Syst.3
2019 Mining Significant Maximum Cardinalities in Knowledge Bases
Arnaud Giacometti, Béatrice Bouchou-Markhoff, Arnaud Soulet
ISWC (1)1
2018 Sequential Pattern Sampling with Norm Constraints
abstract
In recent years, the field of pattern mining has shifted to user-centered methods. In such a context, it is necessary to have a tight coupling between the system and the user where mining techniques provide results at any time or within a short response time of only few seconds. Pattern sampling is a non-exhaustive method for instantly discovering relevant patterns that ensures a good interactivity while providing strong statistical guarantees due to its random nature. Curiously, such an approach investigated for itemsets and subgraphs has not yet been applied to sequential patterns, which are useful for a wide range of mining tasks and application fields. In this paper, we propose the first method for sequential pattern sampling. In addition to address sequential data, the originality of our approach is to introduce a constraint on the norm to control the length of the drawn patterns and to avoid the pitfall of the "long tail" where the rarest patterns flood the user. We propose a new constrained two-step random procedure, named CSSampling, that randomly draws sequential patterns according to frequency with an interval constraint on the norm. We demonstrate that this method performs an exact sampling. Moreover, despite the use of rejection sampling, the experimental study shows that CSSampling remains efficient and the constraint helps to draw general patterns of the "head". We also illustrate how to benefit from these sampled patterns to instantly build an associative classifier dedicated to sequences. This classification approach rivals state of the art proposals showing the interest of constrained sequential pattern sampling.
Lamine Diop, Cheikh Talibouya Diop, Arnaud Giacometti, Dominique Li, Arnaud Soulet
ICDM3
2018 How Your Supporters and Opponents Define Your Interestingness
Bruno Crémilleux, Arnaud Giacometti, Arnaud Soulet
ECML/PKDD (1)2
2018 Dense Neighborhood Pattern Sampling in Numerical Data
abstract
Pattern mining in numerical data remains a challenging task due to the pattern search space that becomes potentially infinite with real-valued dimensions. Most approaches reluctantly reduced the expressiveness of mined patterns to make possible extraction. Despite this expressiveness loss, they do not provide results within a short response time of a few seconds. This paper addresses the instant discovery of patterns in numerical data based on sampling techniques. Instead of splitting each dimension into intervals, we use a metric to introduce the density as new interestingness measure, and to define neighborhood patterns. The language of neighborhood patterns is semantically rich but in return, its size is infinite. We then present a new exact and non-enumerative random procedure to sample this infinite language according to density. An experimental study demonstrates the good compromise between precision and diversity of neighborhood patterns. Finally, in the context of associative classification, we show that a sample of neighborhood patterns is as accurate as traditional methods that traverses the entire search space.
Arnaud Giacometti, Arnaud Soulet
SDM1
2018 Representativeness of Knowledge Bases with the Generalized Benford's Law
Arnaud Soulet, Arnaud Giacometti, Béatrice Bouchou-Markhoff, Fabian M. Suchanek
ISWC (1)2
2017 MapFIM: Memory Aware Parallelized Frequent Itemset Mining in Very Large Datasets
Khanh-Chuong Duong, Mostafa Bamha, Arnaud Giacometti, Dominique Li, Arnaud Soulet, Christel Vrain
DEXA (1)3
2017 Interactive Pattern Sampling for Characterizing Unlabeled Data
Arnaud Giacometti, Arnaud Soulet
IDA1
2016 Clustering with Quantitative User Preferences on Attributes
abstract
This paper proposes a new semi-supervised clustering framework to represent and integrate quantitative preferences on attributes. A new metric learning algorithm is derived that achieves a compromise clustering between a data-driven and a user-driven solution and converges with a good complexity. We observe experimentally that the addition of preferences may be essential to achieve a better clustering. We also show that our approach performs better than the state-of-the art algorithms.
Adnan El Moussawi, Ahmed Cheriat, Arnaud Giacometti, Nicolas Labroche, Arnaud Soulet
ICTAI3
2016 Frequent Pattern Outlier Detection Without Exhaustive Mining
Arnaud Giacometti, Arnaud Soulet
PAKDD (2)1
2015 Contextual preference mining for user profile construction
Sandra de Amo, Mouhamadou Saliou Diallo, Cheikh Talibouya Diop, Arnaud Giacometti, Dominique Li, Arnaud Soulet
Inf. Syst.4
2014 Balancing the Analysis of Frequent Patterns
Arnaud Giacometti, Dominique Li, Arnaud Soulet
PAKDD (1)1
2012 Mining Contextual Preference Rules for Building User Profiles
Sandra de Amo, Mouhamadou Saliou Diallo, Cheikh Talibouya Diop, Arnaud Giacometti, Dominique Li, Arnaud Soulet
DaWaK4
2011 A Relational View of Pattern Discovery
Arnaud Giacometti, Patrick Marcel, Arnaud Soulet
DASFAA (1)1
2010 Cube Based Summaries of Large Association Rule Sets
Marie N'diaye, Cheikh Talibouya Diop, Arnaud Giacometti, Patrick Marcel, Arnaud Soulet
ADMA (1)3
2009 Recommending Multidimensional Queries
Arnaud Giacometti, Patrick Marcel, Elsa Nègre
DaWaK1
2009 Query recommendations for OLAP discovery driven analysis
abstract
Recommending database queries is an emerging and promising field of investigation. This is of particular interest in the domain of OLAP systems where the user is left with the tedious process of navigating large datacubes. In this paper we present a framework for a recommender system for OLAP users, that leverages former users' investigations to enhance discovery driven analysis. The main idea is to recommend to the user the discoveries detected in those former sessions that investigated the same unexpected data as the current session.
Arnaud Giacometti, Patrick Marcel, Elsa Nègre, Arnaud Soulet
DOLAP1
2009 A Framework for Pattern-Based Global Models
Arnaud Giacometti, Eynollah Khanjari, Patrick Marcel, Arnaud Soulet
IDEAL1
2008 A framework for recommending OLAP queries
abstract
An OLAP analysis session can be defined as an interactive session during which a user launches queries to navigate within a cube. Very often choosing which part of the cube to navigate further, and thus designing the forthcoming query, is a difficult task. In this paper, we propose to use what the OLAP users did during their former exploration of the cube as a basis for recommending OLAP queries to the user. We present a generic framework that allows to recommend OLAP queries based on the OLAP server query log. This framework is generic in the sense that changing its parameters changes the way the recommendations are computed. We show how to use this framework for recommending simple MDX queries and we provide some experimental results to validate our approach.
Arnaud Giacometti, Patrick Marcel, Elsa Nègre
DOLAP1
2007 Mining First-Order Temporal Interval Patterns with Regular Expression Constraints
Sandra de Amo, Arnaud Giacometti, Waldecir Pereira Junior
DaWaK2
2007 Temporal Conditional Preferences over Sequences of Objects
abstract
Most research on preference elicitation, preference reasoning and preference query languages design focus mainly on preferences over single objects represented by relational tuples. An increasing interest on preferences over more complex structures like sets of objects has arised in recent papers. However, most recent applications deal with more sophisticated complex structured objects, like sequences, trees and graphs. In this paper, we introduce TPref, a formalism to reason with qualitative conditional preferences over sequences of objects indexed in time. TPref generalizes the CP-Nets formalism by allowing temporal conditional preferences besides the static rules used in CP-Nets. An algorithm for computing optimal sequences of objects satisfying a set of temporal constraints is also presented.
Sandra de Amo, Arnaud Giacometti
ICTAI (2)2
2005 A personalization framework for OLAP queries
abstract
OLAP users heavily rely on visualization of query answers for their interactive analysis of massive amounts of data. Very often, these answers cannot be visualized entirely and the user has to navigate through them to find relevant facts.In this paper, we propose a framework for personalizing OLAP queries. In this framework, the user is asked to give his (her) preferences and a visualization constraint, that can be for instance the limitations imposed by the device used to display the answer to a query. Given this, for each query, our method computes the part of the answer that respects both the user preferences and the visualization constraint. In addition, a personalized structure for the visualization is proposed.
Ladjel Bellatreche, Arnaud Giacometti, Patrick Marcel, Hassina Mouloudi, Dominique Laurent 0001
DOLAP2
2002 Composition of Mining Contexts for Efficient Extraction of Association Rules
Cheikh Talibouya Diop, Arnaud Giacometti, Dominique Laurent 0001, Nicolas Spyratos
EDBT2