Alexandre Termier

dblp:57/1585 · DBLP profile ↗
← Back
28ranked-venue papers in the field
4as first author
5since 2021 · last 2024
0000-0003-1784-0017ORCID · corroborated

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 22 (3 first)Database Systems & Data Management · 4 (1 first)Information Retrieval & Web Search · 1Knowledge Engineering, Semantic Web & Information Systems · 1
YearPublicationVenuePosition
2024 Sky-signatures: detecting and characterizing recurrent behavior in sequential data
Clément Gautrais, Peggy Cellier, Thomas Guyet, Rene Quiniou, Alexandre Termier
Data Min. Knowl. Discov.5
2023 Generating Robust Counterfactual Explanations
Victor Guyomard, Françoise Fessant, Thomas Guyet, Tassadit Bouadi, Alexandre Termier
ECML/PKDD (3)5
2022 VCNet: A Self-explaining Model for Realistic Counterfactual Generation
Victor Guyomard, Françoise Fessant, Thomas Guyet, Tassadit Bouadi, Alexandre Termier
ECML/PKDD (1)5
2022 XEM: An explainable-by-design ensemble method for multivariate time series classification
Kevin Fauvel, Élisa Fromont, Véronique Masson, Philippe Faverdin, Alexandre Termier
Data Min. Knowl. Discov.5
2021 HiPaR: Hierarchical Pattern-Aided Regression
Luis Galárraga, Olivier Pelgrin, Alexandre Termier
PAKDD (1)3
2020 Widening for MDL-Based Retail Signature Discovery
abstract
Signature patterns have been introduced to model repetitive behavior, e.g., of customers repeatedly buying the same set of products in consecutive time periods. A disadvantage of existing approaches to signature discovery, however, is that the required number of occurrences of a signature needs to be manually chosen. To address this limitation, we formalize the problem of selecting the best signature using the minimum description length (MDL) principle. To this end, we propose an encoding for signature models and for any data stream given such a signature model. As finding the MDL-optimal solution is unfeasible, we propose a novel algorithm that is an instance of widening , i.e., a diversified beam search that heuristically explores promising parts of the search space. Finally, we demonstrate the effectiveness of the problem formalization and the algorithm on a real-world retail dataset, and show that our approach yields relevant signatures.
Clément Gautrais, Peggy Cellier, Matthijs van Leeuwen, Alexandre Termier
IDA4
2019 Statistically Significant Discriminative Patterns Searching
Hoang-Son Pham, Gwendal Virlet, Dominique Lavenier, Alexandre Termier
DaWaK4
2019 Towards Sustainable Dairy Management - A Machine Learning Enhanced Method for Estrus Detection
abstract
Our research tackles the challenge of milk production resource use efficiency in dairy farms with machine learning methods. Reproduction is a key factor for dairy farm performance since cows milk production begin with the birth of a calf. Therefore, detecting estrus, the only period when the cow is susceptible to pregnancy, is crucial for farm efficiency. Our goal is to enhance estrus detection (performance, interpretability), especially on the currently undetected silent estrus (35% of total estrus), and allow farmers to rely on automatic estrus detection solutions based on affordable data (activity, temperature). In this paper, we first propose a novel approach with real-world data analysis to address both behavioral and silent estrus detection through machine learning methods. Second, we present LCE, a local cascade based algorithm that significantly outperforms a typical commercial solution for estrus detection, driven by its ability to detect silent estrus. Then, our study reveals the pivotal role of activity sensors deployment in estrus detection. Finally, we propose an approach relying on global and local (behavioral versus silent) algorithm interpretability (SHAP) to reduce the mistrust in estrus detection solutions.
Kevin Fauvel, Véronique Masson, Élisa Fromont, Philippe Faverdin, Alexandre Termier
KDD5
2018 Are your data gathered?
abstract
Understanding data distributions is one of the most fundamental research topic in data analysis. The literature provides a great deal of powerful statistical learning algorithms to gain knowledge on the underlying distribution given multivariate observations. We are likely to find out a dependence between features, the appearance of clusters or the presence of outliers. Before such deep investigations, we propose the folding test of unimodality. As a simple statistical description, it allows to detect whether data are gathered or not (unimodal or multimodal). To the best of our knowledge, this is the first multivariate and purely statistical unimodality test. It makes no distribution assumption and relies only on a straightforward p-value. Through real world data experiments, we show its relevance and how it could be useful for clustering.
Alban Siffer, Pierre-Alain Fouque, Alexandre Termier, Christine Largouët
KDD3
2018 Mining Periodic Patterns with a MDL Criterion
Esther Galbrun, Peggy Cellier, Nikolaj Tatti, Alexandre Termier, Bruno Crémilleux
ECML/PKDD (2)4
2017 Anomaly Detection in Streams with Extreme Value Theory
abstract
Anomaly detection in time series has attracted considerable attention due to its importance in many real-world applications including intrusion detection, energy management and finance. Most approaches for detecting outliers rely on either manually set thresholds or assumptions on the distribution of data according to Chandola, Banerjee and Kumar.
Alban Siffer, Pierre-Alain Fouque, Alexandre Termier, Christine Largouët
KDD3
2017 Purchase Signatures of Retail Customers
Clément Gautrais, Rene Quiniou, Peggy Cellier, Thomas Guyet, Alexandre Termier
PAKDD (1)5
2017 TopPI: An efficient algorithm for item-centric mining
Vincent Leroy 0001, Martin Kirchgessner, Alexandre Termier, Sihem Amer-Yahia
Inf. Syst.3
2016 TopPI: An Efficient Algorithm for Item-Centric Mining
Martin Kirchgessner, Vincent Leroy 0001, Alexandre Termier, Sihem Amer-Yahia, Marie-Christine Rousset
DaWaK3
2016 Understanding Customer Attrition at an Individual Level: a New Model in Grocery Retail Context
abstract
This paper presents a new model to detect and explain customer defection in a grocery retail context. This new model analyzes the evolution of each customer basket content. It therefore provides actionable knowledge for the retailer at an individual scale. In addition, this model is able to identify customers that are likely to defect in the future months.
Clément Gautrais, Peggy Cellier, Thomas Guyet, Rene Quiniou, Alexandre Termier
EDBT5
2015 Interactive User Group Analysis
abstract
User data is becoming increasingly available in multiple domains ranging from phone usage traces to data on the social Web. The analysis of user data is appealing to scientists who work on population studies, recommendations, and large-scale data analytics. We argue for the need for an interactive analysis to understand the multiple facets of user data and address different analytics scenarios. Since user data is often sparse and noisy, we propose to produce labeled groups that describe users with common properties and develop IUGA, an interactive framework based on group discovery primitives to explore the user space. At each step of IUGA, an analyst visualizes group members and may take an action on the group (add/remove members) and choose an operation (exploit/explore) to discover more groups and hence more users. Each discovery operation results in k most relevant and diverse groups. We formulate group exploitation and exploration as optimization problems and devise greedy algorithms to enable efficient group discovery. Finally, we design a principled validation methodology and run extensive experiments that validate the effectiveness of IUGA on large datasets for different user space analysis scenarios.
Behrooz Omidvar-Tehrani, Sihem Amer-Yahia, Alexandre Termier
CIKM3
2015 Selecting representative instances from datasets
abstract
We propose in this paper a new, alternative approach for the problem of finding a set of representative objects in large datasets. To do so, we first formulate the general Instance Selection Problem (ISP) and then study three variants of that in order to select instances from different regions of the data. These variants aim at finding the objects located in three very different locations of the data: the inner frontier, the central area and the outer frontier. Solutions to these problems have been discussed and their complexities have been studied. To illustrate the effectiveness of the proposed techniques, we first use a small, synthetic dataset for visualization purpose. We then study them on the Reuters dataset and show that the integration of instances selected by the ISP techniques is able to provide a good representation of the data and can be considered as a complementary approach for the state-of-the-art methods. Finally, we examine the quality of the selected objects by applying a topic-based analysis in order to show how well the selected documents cover the topics in the Reuters dataset.
Seyed Hamid Mirisaee, Ahlame Douzal Chouakria, Alexandre Termier
DSAA3
2015 PGLCM: efficient parallel mining of closed frequent gradual itemsets
Trong Dinh Thac Do, Alexandre Termier, Anne Laurent, Benjamin Négrevergne, Behrooz Omidvar-Tehrani, Sihem Amer-Yahia
Knowl. Inf. Syst.2
2014 Itemset approximation using Constrained Binary Matrix Factorization
abstract
We address in this paper the problem of efficiently finding a few number of representative frequent itemsets in transaction matrices. To do so, we propose to rely on matrix decomposition techniques, and more precisely on Constrained Binary Matrix Factorization (CBMF) which decomposes a given binary matrix into the product of two lower dimensional binary matrices, called factors. We first show, under binary constraints, that one can interpret the first factor as a transaction matrix operating on packets of items, whereas the second factor indicates which item belongs to which packet. We then formally prove that one can directly mine the CBMF factors in order to find (approximate) itemsets of a given size and support in the original transaction matrix. Then through a detailed experimental study, we show that the frequent itemsets produced by our method represent a significant portion of the set of all frequent itemsets according to existing metrics, while being up to several orders of magnitude less numerous.
Seyed Hamid Mirisaee, Éric Gaussier, Alexandre Termier
DSAA3
2014 Para Miner: a generic pattern mining algorithm for multi-core architectures
Benjamin Négrevergne, Alexandre Termier, Marie-Christine Rousset, Jean-François Méhaut
Data Min. Knowl. Discov.2
2013 Efficiently rewriting large multimedia application execution traces with few event sequences
abstract
The analysis of multimedia application traces can reveal important information to enhance program execution comprehension. However typical size of traces can be in gigabytes, which hinders their effective exploitation by application developers. In this paper, we study the problem of finding a set of sequences of events that allows a reduced-size rewriting of the original trace. These sequences of events, that we call blocks, can simplify the exploration of large execution traces by allowing application developers to see an abstraction instead of low-level events.
Christiane Kamdem Kengne, Léon Constantin Fopa, Alexandre Termier, Noha Ibrahim, Marie-Christine Rousset, Takashi Washio, Miguel Santana
KDD3
2010 PGP-mc: Towards a Multicore Parallel Approach for Mining Gradual Patterns
Anne Laurent, Benjamin Négrevergne, Nicolas Sicard, Alexandre Termier
DASFAA (1)4
2010 PGLCM: Efficient Parallel Mining of Closed Frequent Gradual Itemsets
abstract
Numerical data (e.g., DNA micro-array data, sensor data) pose a challenging problem to existing frequent pattern mining methods which hardly handle them. In this framework, gradual patterns have been recently proposed to extract covariations of attributes, such as: "When X increases, Y decreases". There exist some algorithms for mining frequent gradual patterns, but they cannot scale to real-world databases. We present in this paper GLCM, the first algorithm for mining closed frequent gradual patterns, which proposes strong complexity guarantees: the mining time is linear with the number of closed frequent gradual item sets. Our experimental study shows that GLCM is two orders of magnitude faster than the state of the art, with a constant low memory usage. We also present PGLCM, a parallelization of GLCM capable of exploiting multicore processors, with good scale-up properties on complex datasets. These algorithms are the first algorithms capable of mining large real world datasets to discover gradual patterns.
Trong Dinh Thac Do, Anne Laurent, Alexandre Termier
ICDM3
2010 Combining Logic and Probabilities for Discovering Mappings between Taxonomies
Rémi Tournaire, Jean-Marc Petit, Marie-Christine Rousset, Alexandre Termier
KSEM4
2008 DryadeParent, An Efficient and Robust Closed Attribute Tree Mining Algorithm
abstract
In this paper, we present a new tree mining algorithm, DryadeParent, based on the hooking principle first introduced in DRYADE. In the experiments, we demonstrate that the branching factor and depth of the frequent patterns to find are key factors of complexity for tree mining algorithms, even if often overlooked in previous work. We show that DryadeParent outperforms the current fastest algorithm, CMTreeMiner, by orders of magnitude on data sets where the frequent tree patterns have a high branching factor.
Alexandre Termier, Marie-Christine Rousset, Michèle Sebag, Kouzou Ohara, Takashi Washio, Hiroshi Motoda
IEEE Trans. Knowl. Data Eng.1
2005 Efficient Mining of High Branching Factor Attribute Trees
abstract
In this paper, we present a new tree mining algorithm, DryadeParent, based on the hooking principle first introduced in Dryade (Termier et al, 2004). In the experiments, we demonstrate that the branching factor and depth of the frequent patterns to find are key factor of complexity for tree mining algorithms. We show that DryadeParent outperforms the current fastest algorithm, CMTreeMiner, by orders of magnitude on datasets where the frequent patterns have a high branching factor.
Alexandre Termier, Marie-Christine Rousset, Michèle Sebag, Kouzou Ohara, Takashi Washio, Hiroshi Motoda
ICDM1
2004 DRYADE: A New Approach for Discovering Closed Frequent Trees in Heterogeneous Tree Databases
abstract
In this paper we present a novel algorithm for discovering tree patterns in a tree database. This algorithm uses a relaxed tree inclusion definition, making the problem more complex (checking tree inclusion is NP-complete), but allowing to mine highly heterogeneous databases. To obtain good performances, our DRYADE algorithm, discovers only closed frequent tree patterns.
Alexandre Termier, Marie-Christine Rousset, Michèle Sebag
ICDM1
2002 TreeFinder: a First Step towards XML Data Mining
abstract
In this paper we consider the problem of searching frequent trees from a collection of tree-structured data modeling XML data. The TreeFinder algorithm aims at finding trees, such that their exact or perturbed copies are frequent in a collection of labelled trees. To cope with complexity issues, TreeFinder is correct but not complete: it finds a subset of actually frequent trees. The default of completeness is experimentally investigated on artificial medium size datasets; it is shown that TreeFinder reaches completeness or falls short for a range of experimental settings.
Alexandre Termier, Marie-Christine Rousset, Michèle Sebag
ICDM1