Alexis Bondu

dblp:49/477 · DBLP profile ↗
← Back
22ranked-venue papers
8as first author
10since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 7 first-author · 9 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-author · 3 since 2021Theory of computation · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Automatic Feature Engineering for Time Series Extrinsic Regression: A Comparative Study of Signal Processing Libraries
abstract
Extrinsic regression of time series data consists in predicting the value of a numerical target variable using an input vector which is a time series. The target variable is considered as “extrinsic” as it is not of the same nature as the series values and may not necessarily follow the temporal continuity of the series. This formalization addresses a wide range of problems in different application areas, such as environmental, health or sentiment analysis. In line with the literature on supervised classification of time series, some classification methods have been adapted to the task of regression. Existing regression methods are diverse and use different paradigms, e.g. distance-based methods, interval-based or neural network-based approaches. In parallel to these developments, several libraries for unsupervised feature extraction from time series data have been developed, primarily for descriptive analysis and visualization purposes. In this paper, we combine existing regression methods with signal processing libraries that extract features from time series. To that purpose, the potential of 10 libraries, for the extrinsic regression task, across a set of 61 datasets and six usual regressors is evaluated. The comparative analysis of results from over 3,000 learning ex-periments suggests that unsupervised feature extraction achieves competitive performance for extrinsic regression.
Aurélien Renault, Dominique Gay, Noureddine Yassine Nair Benrekia, Vincent Lemaire 0001, Alexis Bondu
DSAA5
2023 Automatic Feature Engineering for Time Series Classification: Evaluation and Discussion
abstract
Time Series Classification (TSC) has received much attention in the past two decades and is still a crucial and challenging problem in data science and knowledge engineering. Indeed, along with the increasing availability of time series data, many TSC algorithms have been suggested by the research community in the literature. Besides state-of-the-art methods based on similarity measures, intervals, shapelets, dictionaries, deep learning methods or hybrid ensemble methods, several tools for extracting unsupervised informative summary statistics, aka features, from time series have been designed in the recent years. Originally designed for descriptive analysis and visualization of time series with informative and interpretable features, very few of these feature engineering tools have been benchmarked for TSC problems and compared with state-of-the-art TSC algorithms in terms of predictive performance. In this article, we aim at filling this gap and propose a simple TSC process to evaluate the potential predictive performance of the feature sets obtained with existing feature engineering tools. Thus, we present an empirical study of 11 feature engineering tools branched with 9 supervised classifiers over 112 time series data sets. The analysis of the results of more than 10000 learning experiments indicate that feature-based methods perform as accurately as current state-of-the-art TSC algorithms, and thus should rightfully be considered further in the TSC literature.
Aurélien Renault, Alexis Bondu, Vincent Lemaire 0001, Dominique Gay
IJCNN2
2023 Biquality learning: a framework to design algorithms dealing with closed-set distribution shifts
Pierre Nodet, Vincent Lemaire 0001, Alexis Bondu, Antoine Cornuéjols
Mach. Learn.3
2022 When to Classify Events in Open Times Series?
Youssef Achenchabe, Alexis Bondu, Antoine Cornuéjols, Vincent Lemaire 0001
ACML2
2022 Early and Revocable Time Series Classification
abstract
International audience
Youssef Achenchabe, Alexis Bondu, Antoine Cornuéjols, Vincent Lemaire 0001
IJCNN2
2021 Early Classification of Time Series: Cost-based multiclass Algorithms
abstract
Early classification of time series assigns each time series to one of a set of pre-defined classes using as few measurements as possible while preserving a high accuracy. This implies solving online the trade-off between the earliness and the prediction accuracy. This has been formalized in previous work where a cost-based framework taking into account both the cost of misclassification and the cost of delaying the decision has been proposed. The best resulting method, called Economy-$\gamma$, is unfortunately so far limited to binary classification problems. This paper presents a set of six new methods that extend the Economy-$\gamma$method in order to solve multiclass classification problems. Extensive experiments on 33 datasets allowed us to compare the performance of the six proposed approaches to the state-of-the-art one. The results show that: (i) all proposed methods perform significantly better than the state of the art one; (ii) the best way to extend Economy-$\gamma$to multiclass problems is to use a confidence score, either the Gini index or the maximum probability.
Paul-Emile Zafar, Youssef Achenchabe, Alexis Bondu, Antoine Cornuéjols, Vincent Lemaire 0001
DSAA3
2021 Importance Reweighting for Biquality Learning
abstract
The field of Weakly Supervised Learning (WSL) has recently seen a surge of popularity, with numerous papers addressing different types of “supervision deficiencies”, namely: poor quality, non adaptability, and insufficient quantity of labels. Regarding quality, label noise can be of different types, including completely-at-random, at-random or even not-at-random. All these kinds of label noise are addressed separately in the literature, leading to highly specialized approaches. This paper proposes an original, encompassing, view of Weakly Supervised Learning, which results in the design of generic approaches capable of dealing with any kind of label noise. For this purpose, an alternative setting called “Biquality data” is used. It assumes that a small trusted dataset of correctly labeled examples is available, in addition to an untrusted dataset of noisy examples. In this paper, we propose a new reweigthing scheme capable of identifying noncorrupted examples in the untrusted dataset. This allows one to learn classifiers using both datasets. Extensive experiments that simulate several types of label noise and that vary the quality and quantity of untrusted examples, demonstrate that the proposed approach outperforms baselines and state-of-the-art approaches.
Pierre Nodet, Vincent Lemaire 0001, Alexis Bondu, Antoine Cornuéjols, Adam Ouorou
IJCNN3
2021 From Weakly Supervised Learning to Biquality Learning: an Introduction
abstract
The field of Weakly Supervised Learning (WSL) has recently seen a surge of popularity, with numerous papers addressing different types of “supervision deficiencies”. In WSL use cases, a variety of situations exists where the collected “information” is imperfect. The paradigm of WSL attempts to list and cover these problems with associated solutions. In this paper, we review the research progress on WSL with the aim to make it as a brief introduction to this field. We present the three axis of WSL cube and an overview of most of all the elements of their facets. We propose three measurable quantities that acts as coordinates in the previously defined cube namely: Quality, Adaptability and Quantity of information. Thus we suggest that Biquality Learning framework can be defined as a plan of the WSL cube and propose to re-discover previously unrelated patches in WSL literature as a unified Biquality Learning literature.
Pierre Nodet, Vincent Lemaire 0001, Alexis Bondu, Antoine Cornuéjols, Adam Ouorou
IJCNN3
2021 Interpretable Feature Construction for Time Series Extrinsic Regression
Dominique Gay, Alexis Bondu, Vincent Lemaire 0001, Marc Boullé
PAKDD (1)2
2021 Early classification of time series
Youssef Achenchabe, Alexis Bondu, Antoine Cornuéjols, Asma Dachraoui
Mach. Learn.2
2020 Multivariate Time Series Classification: A Relational Way
Dominique Gay, Alexis Bondu, Vincent Lemaire 0001, Marc Boullé, Fabrice Clérot
DaWaK2
2019 FEARS: a Feature and Representation Selection approach for Time Series Classification
abstract
This paper presents a method which extracts informative features while selecting simultaneously adequate representations for Time Series Classification. This method simultaneously (i) selects alternative representations, such as derivatives, cumulative integrals, power spectrum … (ii) and extracts informative features (via automatic variable construction) from the selected set of representations. The suggested approach is decomposed in three steps: (i) the original time series are transformed into several representations which are stored as relational data; (ii) then, a {regularized} propositionalisation method is applied in order to generate informative aggregate features; (iii) finally, a selective Naive Bayes classifier is learned from the outcoming feature-value data table. The previous steps are repeated by a forward backward selection algorithm in order to select the most informative subset of representations. The suggested approach proves to be highly competitive when compared with state-of-the-art methods while extracting interpretable features. Furthermore, the suggested approach is almost parameter free and only requires few hardware resources.
Alexis Bondu, Dominique Gay, Vincent Lemaire 0001, Marc Boullé, Eole Cervenka
ACML1
2019 Toward a Framework for Seasonal Time Series Forecasting Using Clustering
Colin Leverger, Simon Malinowski, Thomas Guyet, Vincent Lemaire 0001, Alexis Bondu, Alexandre Termier
IDEAL (1)5
2015 Realistic and very fast simulation of individual electricity consumptions
abstract
The incoming smart grid represents a significant break for the European utilities in terms of data volume to be processed. In France, one year of individual consumptions represents more than 600 billion data points. Since real data is not yet available, our objective consists in simulating realistic individual consumptions. A new generative model of time series is proposed which combines the MODL coclustering approach with Markov chains. This approach is evaluated on a real dataset provided by the Irish CER. Our experiments demonstrate the ability of the generative model to efficiently reproduce the dynamic of the original time series.
Alexis Bondu, Asma Dachraoui
IJCNN1
2015 Early Classification of Time Series as a Non Myopic Sequential Decision Making Problem
Asma Dachraoui, Alexis Bondu, Antoine Cornuéjols
ECML/PKDD (1)2
2014 Evaluation Protocol of Early Classifiers over Multiple Data Sets
Asma Dachraoui, Alexis Bondu, Antoine Cornuéjols
ICONIP (2)2
2013 SAXO: An optimized data-driven symbolic representation of time series
abstract
In France, the currently emerging “smart grid” and more particularly the 35 millions of “smart meters” will produce a large amount of daily updated metering data. The main french provider of electricity (EDF) is interested by compact and generic representations of time series which allow to accelerate the processing of data. This article proposes a new data-driven symbolic representation of time series named SAXO, where each symbol represents a typical distribution of data points. Furthermore, the time dimension is optimally discretized into intervals by using a parameter free Bayesian coclustering approach (MODL). SAXO is favorably compared with the SAX representation by evaluating a classifier trained from recoded datasets. Our experiments highlight a significant gap in performance between both approaches.
Alexis Bondu, Marc Boullé, Benoît Grossin
IJCNN1
2011 A supervised approach for change detection in data streams
abstract
In recent years, the amount of data to process has increased in many application areas such as network monitoring, web click and sensor data analysis. Data stream mining answers to the challenge of massive data processing, this paradigm allows for treating pieces of data on the fly and overcomes exhaustive data storage. The detection of changes in a data stream distribution is an important issue which application area is wide. In this article, change detection problem is turned into a supervised learning task. We chose to exploit the supervised discretization method “MODL” given its interesting properties. Our approach is favorably compared with an alternative method on artificial data streams, and is applied on real data streams.
Alexis Bondu, Marc Boullé
IJCNN1
2010 Exploration vs. exploitation in active learning : A Bayesian approach
abstract
The labeling of training examples could be a costly task in numerous cases of supervised learning. Active learning strategies address this problem and select unlabeled examples which are considered as the most useful for the training of a predictive model. The choice of examples to be labeled can be considered as a dilemma between the exploration and the exploitation of the input data space. In this article, a new active learning strategy that manages this compromise is proposed. This strategy is based on a Bayesian formalism that minimizes assumptions on data. An experimental validation is conducted on a unidimensional dataset, the objective is to assess the position of a step function from noisy examples. Our approach is favorably compared to an ad hoc strategy : the probabilistic dichotomy.
Alexis Bondu, Vincent Lemaire 0001, Marc Boullé
IJCNN1
2010 A non-parametric semi-supervised discretization method
Alexis Bondu, Marc Boullé, Vincent Lemaire 0001
Knowl. Inf. Syst.1
2008 A Non-parametric Semi-supervised Discretization Method
abstract
Semi-supervised classification methods aim to exploit labelled and unlabelled examples to train a predictive model. Most of these approaches make assumptions on the distribution of classes. This article first proposes a new semi-supervised discretization method which adopts very low informative prior on data. This method discretizes the numerical domain of a continuous input variable, while keeping the information relative to the prediction of classes. Then, an in-depth comparison of this semi-supervised method with the original supervised MODL approach is presented. We demonstrate that the semi-supervised approach is asymptotically equivalent to the supervised approach, improved with a post-optimization of the intervals bounds location.
Alexis Bondu, Marc Boullé, Vincent Lemaire 0001, Stéphane Loiseau, Béatrice Duval
ICDM1
2008 Adaptive curiosity for emotions detection in speech
abstract
Exploratory activities seem to be crucial for our cognitive development. According to psychologists, exploration is an intrinsically rewarding behaviour. The developmental robotics aims to design computational systems that are endowed with such an intrinsic motivation mechanism. There are possible links between developmental robotics and machine learning. Affective computing takes into account emotions in human machine interactions for intelligent system design. The main difficulty to implement automatic detection of emotions in speech is the prohibitive labelling cost of data. Active learning tries to select the most informative examples to build a training set for a predictive model. In this article, the adaptive curiosity framework is used in terms of active learning terminology, and directly compared with existing algorithms on an emotion detection problem.
Alexis Bondu, Vincent Lemaire 0001
IJCNN1