Max Chevalier

dblp:c/MaxChevalier · DBLP profile ↗
← Back
15ranked-venue papers in the field
3as first author
8since 2021 · last 2026
0000-0001-5402-6255ORCID · verified

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 6 (1 first)Data Mining & Knowledge Discovery · 5 (1 first)Information Retrieval & Web Search · 3 (1 first)Knowledge Engineering, Semantic Web & Information Systems · 1
YearPublicationVenuePosition
2026 Incremental data alignment for evolving datasets
abstract
Information systems face significant challenges in today’s constantly evolving digital environments, including dynamic data, heterogeneous sources, and analytical complexities, which directly impact decision-making processes and organizational competitiveness. Data alignment, the process of aligning different sources using their schema and instances, has become a vital solution for ensuring data consistency and enabling effective data exploration. However, existing methods often rely on static approaches which lack adaptability to dynamic data environments and require full recomputation with every change. This study provides an extended evaluation of our previously proposed incremental alignment approach, IDAGEmb, which leverages dynamic graph embedding techniques to refine alignments progressively. Unlike traditional static methods, IDAGEmb adapts to changes in real time, efficiently handling schema modifications and evolving data instances. Our evaluation highlights significant improvements in managing heterogeneous data, optimizing resource usage, and maintaining alignment accuracy in dynamic environments. By integrating incremental graph embeddings, this approach offers a solution for dynamic data environments, providing organizations with consistent and actionable insights. This work builds upon our earlier results, offering a new perspective on data alignment for evolving datasets and emphasizing the effectiveness of dynamic embedding techniques.
Oumaima El Haddadi, Max Chevalier, Bernard Dousset, Ahmad El Allaoui, Anass El Haddadi, Olivier Teste
Data Knowl. Eng.2
2024 Towards Regional Explanations with Validity Domains for Local Explanations
Robin Cugny, Julien Aligon, Max Chevalier, Geoffrey Roman-Jimenez, Olivier Teste
DaWaK3
2024 IDAGEmb: An Incremental Data Alignment Based on Graph Embedding
Oumaima El Haddadi, Max Chevalier, Bernard Dousset, Ahmad El Allaoui, Anass El Haddadi, Olivier Teste
DaWaK2
2024 Empowering CamemBERT Legal Entity Extraction With LLM Boostrapping
Julien Breton, Mokhtar Boumedyen Billami, Max Chevalier, Cássia Trojahn dos Santos
EKAW3
2024 Similarity Measures Recommendation for Mixed Data Clustering
abstract
Clustering is an important data mining task which is widely spread in various domains such as biology, finance, marketing, healthcare, and social sciences. It allows the end user to discover, through built clusters, relationships within data. Many non-expert users perceive clustering as an "easy" task because it always produces a result. However, choosing a clustering algorithm at random, without proper parameter tuning, often leads to poor results. In particular, an important choice when applying a clustering algorithm to a specific dataset is the similarity measure. Since clustering algorithms rely on similarities between data points to build clusters, the chosen similarity measure should fit the data as accurately as possible in order to form the best clusters. Mixed Data are data that are characterized by numerical as well as categorical attributes. When clustering mixed data, the same similarity measure cannot be used for the two attribute types. Commonly a pair of similarity measures is used, one dedicated to numerical attributes and one dedicated to categorical attributes. The choice of these two most appropriate similarity measures is very important in mixed data, as it significantly affects the clustering performance.
Abdoulaye Diop, Nabil El Malki, Max Chevalier, André Péninou, Geoffrey Roman-Jimenez, Olivier Teste
SSDBM3
2022 AutoXAI: A Framework to Automatically Select the Most Adapted XAI Solution
abstract
A large number of XAI (eXplainable Artificial Intelligence) solutions have been proposed in recent years. Recently, thanks to new XAI evaluation metrics, it has become possible to compare these XAI solutions. However, selecting the most relevant XAI solution among all this diversity is still a tedious task, especially if a user has specific needs and constraints. In this paper, we propose AutoXAI, a framework that recommends the best XAI solution and its hyperparameters according to specified XAI evaluation metrics while considering the user's context (dataset, machine learning model, XAI needs and constraints). It adapts approaches from context-aware recommender systems on one side and strategies of optimization and evaluation from AutoML (Automated Machine Learning) on the other. Through two use cases, we show that AutoXAI recommends XAI solutions adapted to the user's needs with the best hyperparameters matching the user's constraints.
Robin Cugny, Julien Aligon, Max Chevalier, Geoffrey Roman-Jimenez, Olivier Teste
CIKM3
2022 Impact of similarity measures on clustering mixed data
abstract
In many domains, we face heterogeneous data with both numeric and categorical attributes. Clustering such data is challenging because the notion of similarity is not well defined due to the multiple data types. Existing clustering algorithms for these data are mainly based on two strategies: the homogenization one where all attributes are converted to a single type and the mixed one where similarity measures for the different data types are combined to define a similarity measure for heterogeneous data. We propose a framework in which we evaluate and compare several clustering algorithms using these two strategies on many real-world data sets. Then, motivated by the importance of similarity in clustering and the diversity of similarity measures for each data type, we proposed as a second study, to evaluate how their choice affects the performance of clustering algorithms using the mixed strategy. Our results suggest that the mixed strategy is preferable to the homogenization one since it uses adapted similarity measures for the different data types. Furthermore, the choice of similarity measures is very important for most of used mixed methods and an optimal choice may lead to great improvements compared to classically used similarity measures.
Abdoulaye Diop, Nabil El Malki, Max Chevalier, André Péninou, Olivier Teste
SSDBM3
2021 Designing a Business View of Enterprise Data: An approach based on a Decentralised Enterprise Knowledge Graph
abstract
Nowadays, companies manage a large volume of data usually organised in ”silos”. Each ”data silo” contains data related to a specific Business Unit, or a project. This scattering of data does not facilitate decision-making requiring the use and cross-checking of data coming from different silos. So, a challenge remains: the construction of a Business View of all data in a company. In this paper, we introduce the concepts of Enterprise Knowledge Graph (EKG) and Decentralised EKG (DEKG). Our DEKG aims at generating a Business View corresponding to a synthetic view of data sources. We first define and model a DEKG with an original process to generate a Business View before presenting the possible implementation of a DEKG.
Bastien Vidé, Joan Marty, Franck Ravat, Max Chevalier
IDEAS4
2018 Querying Heterogeneous Data in Graph-Oriented NoSQL Systems
Mohammed El Malki, Hamdi Ben Hamadou, Max Chevalier, André Péninou, Olivier Teste
DaWaK3
2015 Implementation of Multidimensional Databases in Column-Oriented NoSQL Systems
Max Chevalier, Mohammed El Malki, Arlind Kopliku, Olivier Teste, Ronan Tournier
ADBIS1
2015 Implementation of Multidimensional Databases with Document-Oriented NoSQL
Max Chevalier, Mohammed El Malki, Arlind Kopliku, Olivier Teste, Ronan Tournier
DaWaK1
2010 Social validation of collective annotations: Definition and experiment
abstract
Abstract People taking part in argumentative debates through collective annotations face a highly cognitive task when trying to estimate the group's global opinion. In order to reduce this effort, we propose in this paper to model such debates prior to evaluating their “social validation.” Computing the degree of global confirmation (or refutation) enables the identification of consensual (or controversial) debates. Readers as well as prominent information systems may thus benefit from this information. The accuracy of the social validation measure was tested through an online study conducted with 121 participants. We compared their human perception of consensus in argumentative debates with the results of the three proposed social validation algorithms. Their efficiency in synthesizing opinions was demonstrated by the fact that they achieved an accuracy of up to 84%.
Guillaume Cabanac, Max Chevalier, Claude Chrisment, Christine Julien 0002
J. Assoc. Inf. Sci. Technol.2
2008 Zdravko Markov and Daniel T. Larose, Data Mining the Web: Uncovering Patterns in Web Content, Structure, and Usage
Max Chevalier
Inf. Retr.1
2007 An Annotation Management System for Multidimensional Databases
Guillaume Cabanac, Max Chevalier, Franck Ravat, Olivier Teste
DaWaK2
2007 An Original Usage-Based Metrics for Building a Unified View of Corporate Documents
Guillaume Cabanac, Max Chevalier, Claude Chrisment, Christine Julien 0002
DEXA2