Vincent Lemaire 0001

dblp:96/4840-1 · DBLP profile ↗
← Back
11ranked-venue papers in the field
0as first author
8since 2021 · last 2026
0000-0002-6030-2356ORCID · conflict

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 9Big Data, Cloud & Distributed Data Systems · 2
YearPublicationVenuePosition
2026 Calibration Improves Detection of Mislabeled Examples
Ilies Chibane, Thomas George, Pierre Nodet, Vincent Lemaire 0001
DaWaK4
2025 Operational Evaluation of Algorithms for Online Streaming Continual Multi-Label Classification
abstract
Recent progress has been made in the field of multilabel streaming classification, where an instance can be associated with several labels simultaneously. Most recent research has focused on adapting models to the dynamic distribution of nonstationary data streams. However, continual learning is not only adaptation to concept drift: phenomena such as catastrophic forgetting, as well as forward and backward transfers, appear when new classification tasks are introduced in the data stream. This paper aims to develop a standardized and operational evaluation protocol specifically adapted to the study of these phenomena, in order to identify the most promising strategies for this new multi-label, multi-task learning problem on tabular data stream. This protocol, which is reproducible, fair, and flexible enough to allow the simulation of a large number of scenarios, includes the creation of multi-label and multi-task streams, and an evaluation protocol for measuring (a) online performance, (b) phenomena linked to continual learning and (c) resources consumed. It is tested to compare 18 continual multi-label classification strategies on 6 open literature datasets and 3 simulated datasets. This exploratory analysis has enabled us to identify the promising nature of neural networks coupled with data replay.
Hugo Peuzet, Pascale Kuntz, Frank Meyer, Vincent Lemaire 0001, Killian Le Mau
DSAA4
2025 Automatic Feature Engineering for Time Series Extrinsic Regression: A Comparative Study of Signal Processing Libraries
abstract
Extrinsic regression of time series data consists in predicting the value of a numerical target variable using an input vector which is a time series. The target variable is considered as “extrinsic” as it is not of the same nature as the series values and may not necessarily follow the temporal continuity of the series. This formalization addresses a wide range of problems in different application areas, such as environmental, health or sentiment analysis. In line with the literature on supervised classification of time series, some classification methods have been adapted to the task of regression. Existing regression methods are diverse and use different paradigms, e.g. distance-based methods, interval-based or neural network-based approaches. In parallel to these developments, several libraries for unsupervised feature extraction from time series data have been developed, primarily for descriptive analysis and visualization purposes. In this paper, we combine existing regression methods with signal processing libraries that extract features from time series. To that purpose, the potential of 10 libraries, for the extrinsic regression task, across a set of 61 datasets and six usual regressors is evaluated. The comparative analysis of results from over 3,000 learning ex-periments suggests that unsupervised feature extraction achieves competitive performance for extrinsic regression.
Aurélien Renault, Dominique Gay, Noureddine Yassine Nair Benrekia, Vincent Lemaire 0001, Alexis Bondu
DSAA4
2024 COCALITE: A Hybrid Model COmbining CAtch22 and LITE for Time Series Classification
abstract
Time series classification has achieved significant advancements through deep learning models; however, these models often suffer from high complexity and computational costs. To address these challenges while maintaining effectiveness, we introduce COCALITE, an innovative hybrid model that combines the efficient LITE model with an augmented version incorporating Catch22 features during training. COCALITE operates with only 4.7% of the parameters of the state-of-the-art Inception model, significantly reducing computational overhead. By integrating these complementary approaches, COCALITE leverages both effective feature engineering and deep learning techniques to enhance classification accuracy. Our extensive evaluation across 128 datasets from the UCR archive demonstrates that COCALITE achieves competitive performance, offering a compelling solution for resource-constrained environments.
Oumaima Badi, Maxime Devanne, Ali Ismail-Fawaz, Javidan Abdullayev, Vincent Lemaire 0001, Stefano Berretti, Jonathan Weber, Germain Forestier
IEEE Big Data5
2024 A practical approach to novel class discovery in tabular data
Colin Troisemaine, Alexandre Reiffers, Stéphane Gosselin, Vincent Lemaire 0001, Sandrine Vaton
Data Min. Knowl. Discov.4
2022 Progressive prediction of hospitalisation and patient disposition in the emergency department
abstract
Hospitals face high occupation rates resulting in a longer boarding time and more complex bed management. This task could be facilitated by anticipating the unscheduled admissions. We study the capability of information from French electronic health records of an emergency department (ED) to predict patient disposition decisions. We compare the performances of five learning models in predicting the admission of a patient visiting an emergency department and in predicting the patient’s place of admission at two progressive time points throughout the ED care process: triage and initial assessment. Medical and administrative data were retrospectively collected on 53,608 visits to the Groupe Hospitalier Bretagne Sud, France, from July 2020 to June 2021. Our best model achieve a ROC-AUC equal to 88% and F1-score equal to 75% for admission prediction. Regarding medical unit admission prediction, the global ROC-AUC equals to 87% and F1-score ranges from 38% to 77% for the four admission classes, i.e., intensive care unit (6% of the dataset), medicine units (45%), surgery units (13.2%), and observation unit (35.8%). A validation with a posterior dataset indicates constant results.
Laura Uhl, Vincent Augusto, Vincent Lemaire 0001, Youenn Alexandre, Fanny Jardinaud, Paolo Bercelli, Saber Aloui
IEEE Big Data3
2021 Early Classification of Time Series: Cost-based multiclass Algorithms
abstract
Early classification of time series assigns each time series to one of a set of pre-defined classes using as few measurements as possible while preserving a high accuracy. This implies solving online the trade-off between the earliness and the prediction accuracy. This has been formalized in previous work where a cost-based framework taking into account both the cost of misclassification and the cost of delaying the decision has been proposed. The best resulting method, called Economy-$\gamma$, is unfortunately so far limited to binary classification problems. This paper presents a set of six new methods that extend the Economy-$\gamma$method in order to solve multiclass classification problems. Extensive experiments on 33 datasets allowed us to compare the performance of the six proposed approaches to the state-of-the-art one. The results show that: (i) all proposed methods perform significantly better than the state of the art one; (ii) the best way to extend Economy-$\gamma$to multiclass problems is to use a confidence score, either the Gini index or the maximum probability.
Paul-Emile Zafar, Youssef Achenchabe, Alexis Bondu, Antoine Cornuéjols, Vincent Lemaire 0001
DSAA5
2021 Interpretable Feature Construction for Time Series Extrinsic Regression
Dominique Gay, Alexis Bondu, Vincent Lemaire 0001, Marc Boullé
PAKDD (1)3
2020 Multivariate Time Series Classification: A Relational Way
Dominique Gay, Alexis Bondu, Vincent Lemaire 0001, Marc Boullé, Fabrice Clérot
DaWaK3
2010 A non-parametric semi-supervised discretization method
Alexis Bondu, Marc Boullé, Vincent Lemaire 0001
Knowl. Inf. Syst.3
2008 A Non-parametric Semi-supervised Discretization Method
abstract
Semi-supervised classification methods aim to exploit labelled and unlabelled examples to train a predictive model. Most of these approaches make assumptions on the distribution of classes. This article first proposes a new semi-supervised discretization method which adopts very low informative prior on data. This method discretizes the numerical domain of a continuous input variable, while keeping the information relative to the prediction of classes. Then, an in-depth comparison of this semi-supervised method with the original supervised MODL approach is presented. We demonstrate that the semi-supervised approach is asymptotically equivalent to the supervised approach, improved with a post-optimization of the intervals bounds location.
Alexis Bondu, Marc Boullé, Vincent Lemaire 0001, Stéphane Loiseau, Béatrice Duval
ICDM3