Marc Boullé

dblp:94/5206 · DBLP profile ↗
← Back
41ranked-venue papers
17as first author
4since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 38 · 17 first-author · 3 since 2021Databases, data management, data science and information retrieval · 15 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 first-authorSoftware engineering, systems software and programming languages · 3 · 3 first-authorTheory of computation · 2
YearPublicationVenuePosition
2026 Selection of secondary features from multi-table data for classification
Nicolas Voisine, Lou-Anne Quellet, Marc Boullé, Fabrice Clérot, Anais Collin
Data Knowl. Eng.3
2024 Floating-point histograms for exploratory analysis of large scale real-world data sets
abstract
Histograms are among the most popular methods used in exploratory analysis to summarize univariate distributions. In particular, irregular histograms are good non-parametric density estimators that require very few parameters: the number of bins with their lengths and frequencies. Although many approaches have been proposed in the literature to infer these parameters, most existing histogram methods are difficult to exploit for exploratory analysis in the case of real-world data sets, with scalability issues, truncated data, outliers or heavy-tailed distributions. In this paper, we focus on the G-Enum histogram method, which exploits the Minimum Description Length (MDL) principle to build histograms without any user parameter. We then propose to extend this method by exploiting a new modeling space based on floating-point representation, with the objective of building histograms resistant to outliers or heavy-tailed distributions. We also suggest several heuristics and a methodology suitable for the exploratory analysis of large scale real-world data sets, whose underlying patterns are difficult to recover for digitization reasons. Extensive experiments show the benefits of the approach, evaluated with a dual objective: the accuracy of density estimation in the case of outliers or heavy-tailed distributions, and the effectiveness of the approach for exploratory data analysis.
Marc Boullé
Intell. Data Anal.1
2022 A Non-parametric Bayesian Approach for Uplift Discretization and Feature Selection
Mina Rafla, Nicolas Voisine, Bruno Crémilleux, Marc Boullé
ECML/PKDD (5)4
2021 Interpretable Feature Construction for Time Series Extrinsic Regression
Dominique Gay, Alexis Bondu, Vincent Lemaire 0001, Marc Boullé
PAKDD (1)4
2020 Multivariate Time Series Classification: A Relational Way
Dominique Gay, Alexis Bondu, Vincent Lemaire 0001, Marc Boullé, Fabrice Clérot
DaWaK4
2019 FEARS: a Feature and Representation Selection approach for Time Series Classification
abstract
This paper presents a method which extracts informative features while selecting simultaneously adequate representations for Time Series Classification. This method simultaneously (i) selects alternative representations, such as derivatives, cumulative integrals, power spectrum … (ii) and extracts informative features (via automatic variable construction) from the selected set of representations. The suggested approach is decomposed in three steps: (i) the original time series are transformed into several representations which are stored as relational data; (ii) then, a {regularized} propositionalisation method is applied in order to generate informative aggregate features; (iii) finally, a selective Naive Bayes classifier is learned from the outcoming feature-value data table. The previous steps are repeated by a forward backward selection algorithm in order to select the most informative subset of representations. The suggested approach proves to be highly competitive when compared with state-of-the-art methods while extracting interpretable features. Furthermore, the suggested approach is almost parameter free and only requires few hardware resources.
Alexis Bondu, Dominique Gay, Vincent Lemaire 0001, Marc Boullé, Eole Cervenka
ACML4
2019 A scalable robust and automatic propositionalization approach for Bayesian classification of large mixed numerical and categorical data
Marc Boullé, Clément Charnay, Nicolas Lachiche
Mach. Learn.1
2018 Hierarchical two-part MDL code for multinomial distributions
Marc Boullé
Int. J. Approx. Reason.1
2017 MiSeRe-Hadoop: A Large-Scale Robust Sequential Classification Rules Mining Framework
Elias Egho, Dominique Gay, Romain Trinquart, Marc Boullé, Nicolas Voisine, Fabrice Clérot
DaWaK4
2017 A user parameter-free approach for mining robust sequential classification rules
Elias Egho, Dominique Gay, Marc Boullé, Nicolas Voisine, Fabrice Clérot
Knowl. Inf. Syst.3
2016 Predicting Dangerous Seismic Events in Coal Mines under Distribution Drift
abstract
DOAJ is a unique and extensive index of diverse open access journals from around the world, driven by a growing community, committed to ensuring quality content is freely available online for everyone.
Marc Boullé
FedCSIS1
2015 TESS: Temporal event sequence summarization
abstract
We suggest a novel method of clustering and exploratory analysis of temporal event sequences data (also known as categorical time series) based on three-dimensional data grid models. A data set of temporal event sequences can be represented as a data set of three-dimensional points, each point is defined by three variables: a sequence identifier, a time value and an event value. Instantiating data grid models to the 3D-points turns the problem into 3D-coclustering. The sequences are partitioned into clusters, the time variable is discretized into intervals and the events are partitioned into clusters. The cross-product of the univariate partitions forms a multivariate partition of the representation space, i.e., a grid of cells and it also represents a nonparametric estimator of the joint distribution of the sequences, time and events dimensions. Thus, the sequences are grouped together because they have similar joint distribution of time and events, i.e., similar distribution of events along the time dimension. The best data grid is computed using a parameter-free Bayesian model selection approach. We also suggest several criteria for exploiting the resulting grid through agglomerative hierarchies, for interpreting the clusters of sequences and characterizing their components through insightful visualizations. Extensive experiments on both synthetic and real-world data sets demonstrate that our approach is efficient, effective and discover meaningful underlying patterns in sets of temporal event sequences.
Dominique Gay, Romain Guigourès, Marc Boullé, Fabrice Clérot
DSAA3
2015 Tagging fireworkers activities from body sensors under distribution drift
abstract
We describe our submission to the AAIA'15 Data Mining Competition, where the objective is to tag the activity of firefighters based on vital functions and movement sensor readings.Our solution exploits a selective naive Bayes classifier, with optimal preprocessing, variable selection and model averaging, together with an automatic variable construction method that builds many variables from time series records.The most challenging part of the challenge is that the input variables are not independent and identically distributed (i.i.d.) between the train and test datasets.We suggest a methodology to alleviate this problem, that enabled to get a final score of 0.76 (team marcb).
Marc Boullé
FedCSIS1
2015 A Parameter-Free Approach for Mining Robust Sequential Classification Rules
abstract
Sequential data is generated in many domains of science and technology. Although many studies have been carried out for sequence classification in the past decade, the problem is still a challenge, particularly for pattern-based methods. We identify two important issues related to pattern-based sequence classification which motivate the present work: the curse of parameter tuning and the instability of common interestingness measures. To alleviate these issues, we suggest a new approach and framework for mining sequential rule patterns for classification purpose. We introduce a space of rule pattern models and a prior distribution defined on this model space. From this model space, we define a Bayesian criterion for evaluating the interest of sequential patterns. We also develop a parameter-free algorithm to efficiently mine sequential patterns from the model space. Extensive experiments show that (i) the new criterion identifies interesting and robust patterns, (ii) the direct use of the mined rules as new features in a classification process demonstrates higher inductive performance than the state-of-the-art sequential pattern based classifiers.
Elias Egho, Dominique Gay, Marc Boullé, Nicolas Voisine, Fabrice Clérot
ICDM3
2015 Concept drift detection using supervised bivariate grids
abstract
We present an on-line method for concept change detection on labeled data streams. Our detection method uses a bivariate supervised criterion to determine if the data in two windows come from the same distribution. Our method has no assumption neither on data distribution nor on change type. It has the ability to detect changes of different kinds (mean, variance...). Experiments show that our method performs better than well-known methods from the literature. Additionally, except from the window sizes, no user parameter is required in our method.
Christophe Salperwyck, Marc Boullé, Vincent Lemaire 0001
IJCNN2
2015 Country-Scale Exploratory Analysis of Call Detail Records Through the Lens of Data Grid Models
Romain Guigourès, Dominique Gay, Marc Boullé, Fabrice Clérot, Fabrice Rossi
ECML/PKDD (3)3
2014 Parsimonious Naive Bayes
abstract
We describe our submission to the AAIA'14 Data Mining Competition, where the objective was to reach good predictive performance on text mining classification problems while using a small number of variables.Our submission was ranked 6 th , less than 1% behind the winner.We also present an empirical study on the trade-off between parsimony of the representation and accuracy, and show how good performance can be obtained quickly and efficiently.
Marc Boullé
FedCSIS1
2014 Towards Automatic Feature Construction for Supervised Classification
Marc Boullé
ECML/PKDD (1)1
2014 Khiops CoViz: A Tool for Visual Exploratory Analysis of k-Coclustering Results
Bruno Guerraz, Marc Boullé, Dominique Gay, Fabrice Clérot
ECML/PKDD (3)2
2013 SAXO: An optimized data-driven symbolic representation of time series
abstract
In France, the currently emerging “smart grid” and more particularly the 35 millions of “smart meters” will produce a large amount of daily updated metering data. The main french provider of electricity (EDF) is interested by compact and generic representations of time series which allow to accelerate the processing of data. This article proposes a new data-driven symbolic representation of time series named SAXO, where each symbol represents a typical distribution of data points. Furthermore, the time dimension is optimally discretized into intervals by using a parameter free Bayesian coclustering approach (MODL). SAXO is favorably compared with the SAX representation by evaluating a classifier trained from recoded datasets. Our experiments highlight a significant gap in performance between both approaches.
Alexis Bondu, Marc Boullé, Benoît Grossin
IJCNN2
2012 Itemset-Based Variable Construction in Multi-relational Supervised Learning
Dhafer Lahbib, Marc Boullé, Dominique Laurent 0001
ILP2
2012 A Bayesian Approach for Classification Rule Mining in Quantitative Databases
Dominique Gay, Marc Boullé
ECML/PKDD (2)2
2012 Functional data clustering via piecewise constant nonparametric density estimation
Marc Boullé
Pattern Recognit.1
2011 A supervised approach for change detection in data streams
abstract
In recent years, the amount of data to process has increased in many application areas such as network monitoring, web click and sensor data analysis. Data stream mining answers to the challenge of massive data processing, this paradigm allows for treating pieces of data on the fly and overcomes exhaustive data storage. The detection of changes in a data stream distribution is an important issue which application area is wide. In this article, change detection problem is turned into a supervised learning task. We chose to exploit the supervised discretization method “MODL” given its interesting properties. Our approach is favorably compared with an alternative method on artificial data streams, and is applied on real data streams.
Alexis Bondu, Marc Boullé
IJCNN2
2010 Modelling Complex Data by Learning Which Variable to Construct
Françoise Fessant, Aurélie Le Cam, Marc Boullé, Raphaël Féraud
DaWak3
2010 Exploration vs. exploitation in active learning : A Bayesian approach
abstract
The labeling of training examples could be a costly task in numerous cases of supervised learning. Active learning strategies address this problem and select unlabeled examples which are considered as the most useful for the training of a predictive model. The choice of examples to be labeled can be considered as a dilemma between the exploration and the exploitation of the input data space. In this article, a new active learning strategy that manages this compromise is proposed. This strategy is based on a Bayesian formalism that minimizes assumptions on data. An experimental validation is conducted on a unidimensional dataset, the objective is to assess the position of a step function from noisy examples. Our approach is favorably compared to an ad hoc strategy : the probabilistic dichotomy.
Alexis Bondu, Vincent Lemaire 0001, Marc Boullé
IJCNN3
2010 A method to build a representation using a classifier and its use in a K Nearest Neighbors-based deployment
abstract
The K Nearest Neighbors (KNN) is strongly dependent on the quality of the distance metric used. For supervised classification problems, the aim of metric learning is to learn a distance metric for the input data space from a given collection of pair of similar/dissimilar points. A crucial point is the distance metric used to measure the closeness of instances. In the industrial context of this paper the key point is that a very interesting source of knowledge is available : a classifier to be deployed. The knowledge incorporated in this classifier is used to guide the choice (or the construction) of a distance adapted to the situation Then a KNN-based deployment is elaborated to speed up the deployment of the classifier compared to a direct deployment.
Vincent Lemaire 0001, Marc Boullé, Fabrice Clérot, Pascal Gouzien
IJCNN2
2010 A non-parametric semi-supervised discretization method
Alexis Bondu, Marc Boullé, Vincent Lemaire 0001
Knowl. Inf. Syst.2
2010 Bayesian instance selection for the nearest neighbor rule
Sylvain Ferrandiz, Marc Boullé
Mach. Learn.2
2009 A Parameter-Free Classification Method for Large Scale Learning
Marc Boullé
J. Mach. Learn. Res.1
2008 A Non-parametric Semi-supervised Discretization Method
abstract
Semi-supervised classification methods aim to exploit labelled and unlabelled examples to train a predictive model. Most of these approaches make assumptions on the distribution of classes. This article first proposes a new semi-supervised discretization method which adopts very low informative prior on data. This method discretizes the numerical domain of a continuous input variable, while keeping the information relative to the prediction of classes. Then, an in-depth comparison of this semi-supervised method with the original supervised MODL approach is presented. We demonstrate that the semi-supervised approach is asymptotically equivalent to the supervised approach, improved with a post-optimization of the intervals bounds location.
Alexis Bondu, Marc Boullé, Vincent Lemaire 0001, Stéphane Loiseau, Béatrice Duval
ICDM2
2007 Report on Preliminary Experiments with Data Grid Models in the Agnostic Learning vs. Prior Knowledge Challenge
abstract
This paper introduces a new method1to automatically, rapidly and reliably evaluate the class conditional information of any subset of variables in supervised learning. It is based on a partitioning of each input variable, in intervals in the numerical case and in groups of values in the categorical case. The cross-product of the univariate partitions forms a multivariate partition of the input representation space into a set of cells. This multivariate partition, called data grid, allows to evaluate the correlation between the input variables and the output variable. The best data grid is searched owing to a Bayesian model selection approach and to combinatorial algorithms. Three classification techniques exploiting data grids differently are presented and evaluated in the Agnostic Learning vs. Prior Knowledge Challenge. These preliminary experiments demonstrate the interest of using data grid in machine learning tasks.
Marc Boullé
IJCNN1
2007 Compression-Based Averaging of Selective Naive Bayes Classifiers
Marc Boullé
J. Mach. Learn. Res.1
2007 A New Probabilistic Approach in Rank Regression with Optimal Bayesian Partitioning
Carine Hue, Marc Boullé
J. Mach. Learn. Res.2
2006 Optimal Bayesian 2D-Discretization for Variable Ranking in Regression
Marc Boullé, Carine Hue
Discovery Science1
2006 Regularization and Averaging of the Selective Naive Bayes classifier
abstract
The Nai've Bayes classifier has proved to be very effective on many real data applications. Its performances usually benefit from an accurate estimation of univariate conditional probabilities and from variable selection. However, although variable selection is a desirable feature, it is prone to overfitting. In this paper, we introduce a new regularization technique to select the most probable subset of variables and propose a new model averaging method. The weighting scheme on the models reduces to a weighting scheme on the variables, and finally results in a Naive Bayes with "soft variable selection". Extensive experimental results show that the averaged regularized classifier outperforms the initial selective Naive Bayes classifier.
Marc Boullé
IJCNN1
2006 Supervised evaluation of Voronoi partitions
Sylvain Ferrandiz, Marc Boullé
Intell. Data Anal.2
2006 MODL: A Bayes optimal discretization method for continuous attributes
Marc Boullé
Mach. Learn.1
2005 Optimal bin number for equal frequency discretizations in supervized learning
Marc Boullé
Intell. Data Anal.1
2005 A Bayes Optimal Approach for Partitioning the Values of Categorical Attributes
abstract
In supervised machine learning, the partitioning of the values (also called grouping) of a categorical attribute aims at constructing a new synthetic attribute which keeps the information of the initial attribute and reduces the number of its values. In this paper, we propose a new grouping method MODL founded on a Bayesian approach. The method relies on a model space of grouping models and on a prior distribution defined on this model space. This results in an evaluation criterion of grouping, which is minimal for the most probable grouping given the data, i.e. the Bayes optimal grouping. We propose new super-linear optimization heuristics that yields near-optimal groupings. Extensive comparative experiments demonstrate that the MODL grouping method builds high quality groupings in terms of predictive quality, robustness and small number of groups.
Marc Boullé
J. Mach. Learn. Res.1
2004 Khiops: A Statistical Discretization Method of Continuous Attributes
Marc Boullé
Mach. Learn.1