Nabil El Malki

dblp:218/4766 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
3since 2021 · last 2024
0000-0002-5823-8412ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2024 Similarity Measures Recommendation for Mixed Data Clustering
abstract
Clustering is an important data mining task which is widely spread in various domains such as biology, finance, marketing, healthcare, and social sciences. It allows the end user to discover, through built clusters, relationships within data. Many non-expert users perceive clustering as an "easy" task because it always produces a result. However, choosing a clustering algorithm at random, without proper parameter tuning, often leads to poor results. In particular, an important choice when applying a clustering algorithm to a specific dataset is the similarity measure. Since clustering algorithms rely on similarities between data points to build clusters, the chosen similarity measure should fit the data as accurately as possible in order to form the best clusters. Mixed Data are data that are characterized by numerical as well as categorical attributes. When clustering mixed data, the same similarity measure cannot be used for the two attribute types. Commonly a pair of similarity measures is used, one dedicated to numerical attributes and one dedicated to categorical attributes. The choice of these two most appropriate similarity measures is very important in mixed data, as it significantly affects the clustering performance.
Abdoulaye Diop, Nabil El Malki, Max Chevalier, André Péninou, Geoffrey Roman-Jimenez, Olivier Teste
SSDBM2
2022 Impact of similarity measures on clustering mixed data
abstract
In many domains, we face heterogeneous data with both numeric and categorical attributes. Clustering such data is challenging because the notion of similarity is not well defined due to the multiple data types. Existing clustering algorithms for these data are mainly based on two strategies: the homogenization one where all attributes are converted to a single type and the mixed one where similarity measures for the different data types are combined to define a similarity measure for heterogeneous data. We propose a framework in which we evaluate and compare several clustering algorithms using these two strategies on many real-world data sets. Then, motivated by the importance of similarity in clustering and the diversity of similarity measures for each data type, we proposed as a second study, to evaluate how their choice affects the performance of clustering algorithms using the mixed strategy. Our results suggest that the mixed strategy is preferable to the homogenization one since it uses adapted similarity measures for the different data types. Furthermore, the choice of similarity measures is very important for most of used mixed methods and an optimal choice may lead to great improvements compared to classically used similarity measures.
Abdoulaye Diop, Nabil El Malki, Max Chevalier, André Péninou, Olivier Teste
SSDBM2
2021 A New Accurate Clustering Approach for Detecting Different Densities in High Dimensional Data
Nabil El Malki, Robin Cugny, Olivier Teste, Franck Ravat
DaWaK1
2020 DECWA: Density-Based Clustering using Wasserstein Distance
abstract
Clustering is a data analysis method for extracting knowledge by discovering groups of data called clusters. Among these methods, state-of-the-art density-based clustering methods have proven to be effective for arbitrary-shaped clusters. Despite their encouraging results, they suffer to find low-density clusters, near clusters with similar densities, and high-dimensional data. Our proposals are a new characterization of clusters and a new clustering algorithm based on spatial density and probabilistic approach. First of all, sub-clusters are built using spatial density represented as probability density function (p.d.f) of pairwise distances between points. A method is then proposed to agglomerate similar sub-clusters by using both their density (p.d.f) and their spatial distance. The key idea we propose is to use the Wasserstein metric, a powerful tool to measure the distance between p.d.f of sub-clusters. We show that our approach outperforms other state-of-the-art density-based clustering methods on a wide variety of datasets.
Nabil El Malki, Robin Cugny, Olivier Teste, Franck Ravat
CIKM1
2020 KD-means: Clustering Method for Massive Data based on KD-tree
Nabil El Malki, Franck Ravat, Olivier Teste
DOLAP1