Bruno Crémilleux

dblp:c/BrunoCremilleux · DBLP profile ↗
← Back
38ranked-venue papers in the field
1as first author
10since 2021 · last 2026
0000-0001-8294-9049ORCID · verified

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 30 (1 first)Database Systems & Data Management · 3Information Retrieval & Web Search · 3Knowledge Engineering, Semantic Web & Information Systems · 2
YearPublicationVenuePosition
2026 Exploiting Treatment Similarities for Enhanced Multi-treatment Uplift Prediction
Nathan Le Boudec, Nicolas Voisine, Bruno Crémilleux
IDA3
2026 Heterogeneous Pattern Sampling According to Frequency
Rayane Lachache, Djawad Bekkoucha, Abdelkader Ouali, Bruno Crémilleux, Thi-Bich-Hanh Dao, Christel Vrain
IDA4
2026 Efficiently sampling interval patterns from numerical databases
abstract
Pattern sampling has emerged as a promising approach for information discovery in large databases, allowing analysts to focus on a manageable subset of patterns. In this approach, patterns are randomly drawn based on an interestingness measure, such as frequency or hyper-volume. This paper presents the first sampling approach designed to handle interval patterns in numerical databases. This approach, named Fips , samples interval patterns proportionally to their frequency. It uses a multi-step sampling procedure and addresses a key challenge in numerical data: accurately determining the number of interval patterns that cover each object. We extend this work with HFips , which samples interval patterns proportionally to both their frequency and hyper-volume. These methods efficiently tackle the well-known long-tail phenomenon in pattern sampling. We formally prove that Fips and HFips sample interval patterns in proportion to their frequency and the product of hyper-volume and frequency, respectively. Through experiments on several databases, we demonstrate the quality of the obtained patterns and their robustness against the long-tail phenomenon.
Djawad Bekkoucha, Lamine Diop, Abdelkader Ouali, Bruno Crémilleux, Patrice Boizumault
Data Knowl. Eng.4
2026 Multi-treatment uplift evaluation on non-random assignment biased data
Nathan Le Boudec, Nicolas Voisine, Bruno Crémilleux
Data Knowl. Eng.3
2024 WaveLSea: helping experts interactively explore pattern mining search spaces
Etienne Lehembre, Bruno Crémilleux, Albrecht Zimmermann, Bertrand Cuissart, Abdelkader Ouali
Data Min. Knowl. Discov.2
2023 Towards a (Semi-)Automatic Urban Planning Rule Identification in the French Language
abstract
One of the objectives of the Hérelles project is to find new mechanisms to facilitate the labeling (or semantization) of clusters from time series of satellite images. To achieve this, a proposed solution is to associate textual elements of interest with satellite data. The first step in this process consists of an automatic extraction of the information in the form of rules from urban planning documents composed in the French language. To address this challenge, we propose a method which is based on the multi-label classification of textual segments. It includes a special format for representing segments, in which each segment has a title and a subtitle. In addition, we propose a cascade approach aiming to deal with hierarchy of class labels. Finally, we develop several text augmentation techniques for the texts in French, which are able to improve the prediction results. We demonstrate experimentally that the resulting framework correctly classifies each type of segment with more than 90% of accuracy.
Maksim Koptelov, Margaux Holveck, Bruno Crémilleux, Justine Reynaud, Mathieu Roche, Maguelonne Teisseire
DSAA3
2023 Parameter-Free Bayesian Decision Trees for Uplift Modeling
Mina Rafla, Nicolas Voisine, Bruno Crémilleux
PAKDD (2)3
2022 Selecting Outstanding Patterns Based on Their Neighbourhood
Etienne Lehembre, Ronan Bureau, Bruno Crémilleux, Bertrand Cuissart, Jean Luc Lamotte, Alban Lepailleur, Abdelkader Ouali, Albrecht Zimmermann
IDA3
2022 Evaluation of Uplift Models with Non-Random Assignment Bias
Mina Rafla, Nicolas Voisine, Bruno Crémilleux
IDA3
2022 A Non-parametric Bayesian Approach for Uplift Discretization and Feature Selection
Mina Rafla, Nicolas Voisine, Bruno Crémilleux, Marc Boullé
ECML/PKDD (5)3
2020 Using Data Science to Improve the Identification of Plant Nutritional Status
abstract
Developing products for improving plant nutritional status, e.g. fertilizers or plant growth regulators, is an important topic to move towards sustainability in agriculture and to ensure to feed the world population. A key challenge is to identify when, what, and how much nutrients to add to plants' growth environment. In this paper, we study a use case on how to characterize rapeseed plant nutrient deficiencies during their growth. A promising approach consists of deriving data from spectroscopy of leaves, and using this representation to predict what kind of deficiency (if any) plants are undergoing. We are considering three research questions: 1) from which day after onset of a nutrient deficiency we can identify it, 2) whether leaves that have sprouted under nutrient-rich conditions can still help in identifying problems. Third, and most importantly, performing the spectroscopy on the full range of wavelengths is expensive, which under production conditions allows for only relatively few samples. We therefore explore how to perform dimensionality reduction and preprocessing to achieve good predictive accuracy. We show that 1) deficiencies can be identified early on, 2) leaf generations help to predict nutrient deficiencies, and 3) that preprocessing increases the accuracy and dimensionality reduction can be performed without loss of accuracy. Along the way, we find that our some of our industry partners' assumptions about the data do not seem to be borne out by our empirical results, and that the subset of data they initially selected turns out to be too easy to model. The full data leads to more informative insights.
David Condaminet, Albrecht Zimmermann, Bastien Billiot, Bruno Crémilleux, Sylvain Pluchon
DSAA4
2018 Link Prediction in Multi-layer Networks and Its Application to Drug Design
Maksim Koptelov, Albrecht Zimmermann, Bruno Crémilleux
IDA3
2018 PrePeP: A Tool for the Identification and Characterization of Pan Assay Interference Compounds
abstract
Pan Assays Interference Compounds (PAINS) are a significant problem in modern drug discovery: compounds showing non-target specific activity in high-throughput screening can mislead medicinal chemists during hit identification, wasting time and resources. Recent work has shown that existing structural alerts are not up to the task of identifying PAINS. To address this short-coming, we are in the process of developing a tool, PrePeP, that predicts PAINS, and allows experts to visually explore the reasons for the prediction. In the paper, we discuss the different aspects that are involved in developing a functional tool: systematically deriving structural descriptors, addressing the extreme imbalance of the data, offering visual information that pharmacological chemists are familiar with. We evaluate the quality of the approach using benchmark data sets from the literature and show that we correct several short-comings of existing PAINS alerts that have recently been pointed out.
Maksim Koptelov, Albrecht Zimmermann, Pascal Bonnet, Ronan Bureau, Bruno Crémilleux
KDD5
2018 How Your Supporters and Opponents Define Your Interestingness
Bruno Crémilleux, Arnaud Giacometti, Arnaud Soulet
ECML/PKDD (1)1
2018 Mining Periodic Patterns with a MDL Criterion
Esther Galbrun, Peggy Cellier, Nikolaj Tatti, Alexandre Termier, Bruno Crémilleux
ECML/PKDD (2)5
2018 Constrained distance based clustering for time-series: a comparative and experimental study
Thomas Andrew Lampert, Thi-Bich-Hanh Dao, Baptiste Lafabregue, Nicolas Serrette, Germain Forestier, Bruno Crémilleux, Christel Vrain, Pierre Gançarski
Data Min. Knowl. Discov.6
2017 Integer Linear Programming for Pattern Set Mining; with an Application to Tiling
Abdelkader Ouali, Albrecht Zimmermann, Samir Loudni, Yahia Lebbah, Bruno Crémilleux, Patrice Boizumault, Lakhdar Loukil
PAKDD (2)5
2015 Minimal Jumping Emerging Patterns: Computation and Practical Assessment
Bamba Kane, Bertrand Cuissart, Bruno Crémilleux
PAKDD (1)3
2015 Soft constraints for pattern mining
Willy Ugarte, Patrice Boizumault, Samir Loudni, Bruno Crémilleux, Alban Lepailleur
J. Intell. Inf. Syst.4
2014 Sequence Classification Based on Delta-Free Sequential Patterns
abstract
Sequential pattern mining is one of the most studied and challenging tasks in data mining. However, the extension of well-known methods from many other classical patterns to sequences is not a trivial task. In this paper we study the notion of δ-freeness for sequences. While this notion has extensively been discussed for itemsets, this work is the first to extend it to sequences. We define an efficient algorithm devoted to the extraction of δ-free sequential patterns. Furthermore, we show the advantage of the δ-free sequences and highlight their importance when building sequence classifiers, and we show how they can be used to address the feature selection problem in statistical classifiers, as well as to build symbolic classifiers which optimizes both accuracy and earliness of predictions.
Pierre Holat, Marc Plantevit, Chedy Raïssi, Nadi Tomeh, Thierry Charnois, Bruno Crémilleux
ICDM6
2014 Image re-ranking based on statistics of frequent patterns
abstract
Text-based image retrieval is a popular and simple framework consisting in using text annotations (e.g. image names, tags) to perform image retrieval, allowing to handle efficiently very large image collections. Even if the set of images retrieved using text annotations is noisy, it constitutes a reasonable initial set of images that can be considered as a bootstrap and improved further by analyzing image content. In this context, this paper introduces an approach for improving this initial set by re-ranking the so-obtained images, assuming that non-relevant images are scattered (i.e. they do not form clusters), unlike the relevant ones. More specifically, the approach consists in computing efficiently and on the fly frequent closed patterns, and in re-ranking images based on the number of patterns they contain. To do this, the paper introduces a simple but powerful new scoring function. The approach is validated on three different datasets for which state-of-the-art results are obtained.
Winn Voravuthikunchai, Bruno Crémilleux, Frédéric Jurie
ICMR2
2013 Parameter-free classification in multi-class imbalanced data sets
Loïc Cerf, Dominique Gay, Nazha Selmaoui-Folcher, Bruno Crémilleux, Jean-François Boulicaut
Data Knowl. Eng.4
2012 Constrained Clustering Using SAT
Jean-Philippe Métivier, Patrice Boizumault, Bruno Crémilleux, Mehdi Khiari, Samir Loudni
IDA3
2011 Mining Dominant Patterns in the Sky
abstract
Pattern discovery is at the core of numerous data mining tasks. Although many methods focus on efficiency in pattern mining, they still suffer from the problem of choosing a threshold that influences the final extraction result. The goal of our study is to make the results of pattern mining useful from a user-preference point of view. To this end, we integrate into the pattern discovery process the idea of skyline queries in order to mine skyline patterns in a threshold-free manner. Because the skyline patterns satisfy a formal property of dominations, they not only have a global interest but also have semantics that are easily understood by the user. In this work, we first establish theoretical relationships between pattern condensed representations and skyline pattern mining. We also show that it is possible to compute automatically a subset of measures involved in the user query which allows the patterns to be condensed and thus facilitates the computation of the skyline patterns. This forms the basis for a novel approach to mining skyline patterns. We illustrate the efficiency of our approach over several data sets including a use case from chemo informatics and show that small sets of dominant patterns are produced under various measures.
Arnaud Soulet, Chedy Raïssi, Marc Plantevit, Bruno Crémilleux
ICDM4
2011 Extracting and summarizing the frequent emerging graph patterns from a dataset of graphs
Guillaume Poezevara, Bertrand Cuissart, Bruno Crémilleux
J. Intell. Inf. Syst.3
2010 Recursive Sequence Mining to Discover Named Entity Relations
Peggy Cellier, Thierry Charnois, Marc Plantevit, Bruno Crémilleux
IDA4
2009 Missing Values: Proposition of a Typology and Characterization with an Association Rule-Based Model
Leila Ben Othman, François Rioult, Sadok Ben Yahia, Bruno Crémilleux
DaWaK4
2009 Condensed Representation of Sequential Patterns According to Frequency-Based Measures
Marc Plantevit, Bruno Crémilleux
IDA2
2008 Adequate Condensed Representations of Patterns
Arnaud Soulet, Bruno Crémilleux
ECML/PKDD (1)2
2008 Adequate condensed representations of patterns
Arnaud Soulet, Bruno Crémilleux
Data Min. Knowl. Discov.2
2007 Exclusion-inclusion based text categorization of biomedical articles
abstract
In this paper, we propose a new approach based on two original principles to categorize biomedical articles. On the one hand, we combine linguistic, structural and metric descriptors to build patterns stemming from data mining techniques. On the other hand, we take into account the importance of the absence of patterns to the categorization task by using an exclusion-inclusion method. To avoid a crisp effect between the absence and the presence of a pattern, the exclusion-inclusion method uses two regret measures to quantify the interest of a weak pattern according to the other classes and among patterns from a same class. The global decision is based on the generalization of the local patterns, firstly by using patterns excluding classes, then according to the regret ratios. Experiments show the effectiveness of the approach.
Nadia Zerida, Nadine Lucas, Bruno Crémilleux
ACM Symposium on Document Engineering3
2006 Optimized Rule Mining Through a Unified Framework for Interestingness Measures
Céline Hébert, Bruno Crémilleux
DaWaK2
2006 Combining linguistic and structural descriptors for mining biomedical literature
abstract
This work proposes an original combination of linguistic and structural descriptors to represent the content of biomedical papers. The objective is to show the effectiveness of descriptors taking into account the structure of documents to characterise three kinds of biomedical texts (reviews, research and clinical papers). The description of text is made at various levels, from the global level to the local one. The contexts makes it possible to characterise the three classes. The characterisation of the textual resources is carried out quantitatively by using the discriminating capacity of techniques of data mining based on emerging patterns.
Nadia Zerida, Nadine Lucas, Bruno Crémilleux
ACM Symposium on Document Engineering3
2005 Optimizing Constraint-Based Mining by Automatically Relaxing Constraints
abstract
In constraint-based mining, the monotone and anti-monotone properties are exploited to reduce the search space. Even if a constraint has not such suitable properties, existing algorithms can be re-used thanks to an approximation, called relaxation. In this paper, we automatically compute monotone relaxations of primitive-based constraints. First, we show that the latter are a superclass of combinations of both kinds of monotone constraints. Second, we add two operators to detect the properties of monotonicity of such constraints. Finally, we define relaxing operators to obtain monotone relaxations of them.
Arnaud Soulet, Bruno Crémilleux
ICDM2
2005 An Efficient Framework for Mining Flexible Constraints
Arnaud Soulet, Bruno Crémilleux
PAKDD2
2004 Condensed Representation of Emerging Patterns
Arnaud Soulet, Bruno Crémilleux, François Rioult
PAKDD2
2003 Condensed Representations in Presence of Missing Values
François Rioult, Bruno Crémilleux
IDA2
1998 Treatment of Missing Values for Association Rules
Arnaud Ragel, Bruno Crémilleux
PAKDD2