VLDB 2026 Research / reviewers in the wild / expert
Dominique Gay
dblp:30/2311
· DBLP profile ↗
19ranked-venue papers
7as first author
3since 2021 · last 2025
0000-0002-0671-4616ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 16 · 7 first-author · 2 since 2021Artificial intelligence and machine learning · 13 · 3 first-author · 2 since 2021Theory of computation · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Automatic Feature Engineering for Time Series Extrinsic Regression: A Comparative Study of Signal Processing LibrariesabstractExtrinsic regression of time series data consists in predicting the value of a numerical target variable using an input vector which is a time series. The target variable is considered as “extrinsic” as it is not of the same nature as the series values and may not necessarily follow the temporal continuity of the series. This formalization addresses a wide range of problems in different application areas, such as environmental, health or sentiment analysis. In line with the literature on supervised classification of time series, some classification methods have been adapted to the task of regression. Existing regression methods are diverse and use different paradigms, e.g. distance-based methods, interval-based or neural network-based approaches. In parallel to these developments, several libraries for unsupervised feature extraction from time series data have been developed, primarily for descriptive analysis and visualization purposes. In this paper, we combine existing regression methods with signal processing libraries that extract features from time series. To that purpose, the potential of 10 libraries, for the extrinsic regression task, across a set of 61 datasets and six usual regressors is evaluated. The comparative analysis of results from over 3,000 learning ex-periments suggests that unsupervised feature extraction achieves competitive performance for extrinsic regression. Aurélien Renault, Dominique Gay, Noureddine Yassine Nair Benrekia, Vincent Lemaire 0001, Alexis Bondu |
DSAA | 2 |
| 2023 | Automatic Feature Engineering for Time Series Classification: Evaluation and DiscussionabstractTime Series Classification (TSC) has received much attention in the past two decades and is still a crucial and challenging problem in data science and knowledge engineering. Indeed, along with the increasing availability of time series data, many TSC algorithms have been suggested by the research community in the literature. Besides state-of-the-art methods based on similarity measures, intervals, shapelets, dictionaries, deep learning methods or hybrid ensemble methods, several tools for extracting unsupervised informative summary statistics, aka features, from time series have been designed in the recent years. Originally designed for descriptive analysis and visualization of time series with informative and interpretable features, very few of these feature engineering tools have been benchmarked for TSC problems and compared with state-of-the-art TSC algorithms in terms of predictive performance. In this article, we aim at filling this gap and propose a simple TSC process to evaluate the potential predictive performance of the feature sets obtained with existing feature engineering tools. Thus, we present an empirical study of 11 feature engineering tools branched with 9 supervised classifiers over 112 time series data sets. The analysis of the results of more than 10000 learning experiments indicate that feature-based methods perform as accurately as current state-of-the-art TSC algorithms, and thus should rightfully be considered further in the TSC literature. Aurélien Renault, Alexis Bondu, Vincent Lemaire 0001, Dominique Gay |
IJCNN | 4 |
| 2021 | Interpretable Feature Construction for Time Series Extrinsic Regression
Dominique Gay, Alexis Bondu, Vincent Lemaire 0001, Marc Boullé |
PAKDD (1) | 1 |
| 2020 | Multivariate Time Series Classification: A Relational Way
Dominique Gay, Alexis Bondu, Vincent Lemaire 0001, Marc Boullé, Fabrice Clérot |
DaWaK | 1 |
| 2019 | FEARS: a Feature and Representation Selection approach for Time Series ClassificationabstractThis paper presents a method which extracts informative features while selecting simultaneously adequate representations for Time Series Classification. This method simultaneously (i) selects alternative representations, such as derivatives, cumulative integrals, power spectrum … (ii) and extracts informative features (via automatic variable construction) from the selected set of representations. The suggested approach is decomposed in three steps: (i) the original time series are transformed into several representations which are stored as relational data; (ii) then, a {regularized} propositionalisation method is applied in order to generate informative aggregate features; (iii) finally, a selective Naive Bayes classifier is learned from the outcoming feature-value data table. The previous steps are repeated by a forward backward selection algorithm in order to select the most informative subset of representations. The suggested approach proves to be highly competitive when compared with state-of-the-art methods while extracting interpretable features. Furthermore, the suggested approach is almost parameter free and only requires few hardware resources. Alexis Bondu, Dominique Gay, Vincent Lemaire 0001, Marc Boullé, Eole Cervenka |
ACML | 2 |
| 2017 | MiSeRe-Hadoop: A Large-Scale Robust Sequential Classification Rules Mining Framework
Elias Egho, Dominique Gay, Romain Trinquart, Marc Boullé, Nicolas Voisine, Fabrice Clérot |
DaWaK | 2 |
| 2017 | A user parameter-free approach for mining robust sequential classification rules
Elias Egho, Dominique Gay, Marc Boullé, Nicolas Voisine, Fabrice Clérot |
Knowl. Inf. Syst. | 2 |
| 2015 | TESS: Temporal event sequence summarizationabstractWe suggest a novel method of clustering and exploratory analysis of temporal event sequences data (also known as categorical time series) based on three-dimensional data grid models. A data set of temporal event sequences can be represented as a data set of three-dimensional points, each point is defined by three variables: a sequence identifier, a time value and an event value. Instantiating data grid models to the 3D-points turns the problem into 3D-coclustering. The sequences are partitioned into clusters, the time variable is discretized into intervals and the events are partitioned into clusters. The cross-product of the univariate partitions forms a multivariate partition of the representation space, i.e., a grid of cells and it also represents a nonparametric estimator of the joint distribution of the sequences, time and events dimensions. Thus, the sequences are grouped together because they have similar joint distribution of time and events, i.e., similar distribution of events along the time dimension. The best data grid is computed using a parameter-free Bayesian model selection approach. We also suggest several criteria for exploiting the resulting grid through agglomerative hierarchies, for interpreting the clusters of sequences and characterizing their components through insightful visualizations. Extensive experiments on both synthetic and real-world data sets demonstrate that our approach is efficient, effective and discover meaningful underlying patterns in sets of temporal event sequences. Dominique Gay, Romain Guigourès, Marc Boullé, Fabrice Clérot |
DSAA | 1 |
| 2015 | A Parameter-Free Approach for Mining Robust Sequential Classification RulesabstractSequential data is generated in many domains of science and technology. Although many studies have been carried out for sequence classification in the past decade, the problem is still a challenge, particularly for pattern-based methods. We identify two important issues related to pattern-based sequence classification which motivate the present work: the curse of parameter tuning and the instability of common interestingness measures. To alleviate these issues, we suggest a new approach and framework for mining sequential rule patterns for classification purpose. We introduce a space of rule pattern models and a prior distribution defined on this model space. From this model space, we define a Bayesian criterion for evaluating the interest of sequential patterns. We also develop a parameter-free algorithm to efficiently mine sequential patterns from the model space. Extensive experiments show that (i) the new criterion identifies interesting and robust patterns, (ii) the direct use of the mined rules as new features in a classification process demonstrates higher inductive performance than the state-of-the-art sequential pattern based classifiers. Elias Egho, Dominique Gay, Marc Boullé, Nicolas Voisine, Fabrice Clérot |
ICDM | 2 |
| 2015 | Country-Scale Exploratory Analysis of Call Detail Records Through the Lens of Data Grid Models
Romain Guigourès, Dominique Gay, Marc Boullé, Fabrice Clérot, Fabrice Rossi |
ECML/PKDD (3) | 2 |
| 2014 | Khiops CoViz: A Tool for Visual Exploratory Analysis of k-Coclustering Results
Bruno Guerraz, Marc Boullé, Dominique Gay, Fabrice Clérot |
ECML/PKDD (3) | 3 |
| 2013 | Parameter-free classification in multi-class imbalanced data sets
Loïc Cerf, Dominique Gay, Nazha Selmaoui-Folcher, Bruno Crémilleux, Jean-François Boulicaut |
Data Knowl. Eng. | 2 |
| 2012 | A Bayesian Approach for Classification Rule Mining in Quantitative Databases
Dominique Gay, Marc Boullé |
ECML/PKDD (2) | 1 |
| 2012 | Application-independent feature construction based on almost-closedness properties
Dominique Gay, Nazha Selmaoui-Folcher, Jean-François Boulicaut |
Knowl. Inf. Syst. | 1 |
| 2011 | A clustering-based visualization of colocation patternsabstractExtraction of interesting colocations in geo-referenced data is one of the major tasks in spatial pattern mining. The goal is to find sets of spatial object-types with instances located in the same neighborhood. In this context, the main drawback is the visualization and interpretation of extracted patterns by domain experts. Indeed, common textual representation of colocations loses important spatial information such as the position, the orientation or the spatial distribution of the patterns. To overcome this problem, we propose a new clustering-based visualization technique deeply integrated in the colocation mining algorithm. This new simple, concise and intuitive cartographic visualization considers both spatial information and expert practices. This proposition has been integrated in a Geographic Information System and experimented on a real-world geological data set. Domain experts confirm the added-value of this visualization approach. Elise Desmier, Frédéric Flouvat, Dominique Gay, Nazha Selmaoui-Folcher |
IDEAS | 3 |
| 2009 | Application-Independent Feature Construction from Noisy Samples
Dominique Gay, Nazha Selmaoui-Folcher, Jean-François Boulicaut |
PAKDD | 1 |
| 2008 | A Parameter-Free Associative Classification Method
Loïc Cerf, Dominique Gay, Nazha Selmaoui-Folcher, Jean-François Boulicaut |
DaWaK | 2 |
| 2008 | Feature Construction Based on Closedness Properties Is Not That Simple
Dominique Gay, Nazha Selmaoui-Folcher, Jean-François Boulicaut |
PAKDD | 1 |
| 2006 | Feature Construction and delta-Free Sets in 0/1 Samples
Nazha Selmaoui-Folcher, Claire Leschi, Dominique Gay, Jean-François Boulicaut |
Discovery Science | 3 |