VLDB 2026 Research / reviewers in the wild / expert
Mustapha Lebbah
dblp:72/2625
· DBLP profile ↗
83ranked-venue papers
10as first author
31since 2021 · last 2026
0000-0001-7245-6371ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 72 · 10 first-author · 25 since 2021Databases, data management, data science and information retrieval · 13 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6Systems, architecture and hardware · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Energy-Based Dropout with Patch-Level RegularizationabstractDropout is a widely used stochastic regularization technique, yet it overlooks structural dependencies within feature maps.We introduce PB-EDropout, an energy-based approach that preserves low-energy spatial patches within each channel while suppressing the rest.During training, candidate masks are sampled from Gibbs distributions and refined using genetic operators, and a running exponential moving average yields deterministic masks for inference.Experiments on shallow CNNs demonstrate that PB-EDropout consistently improves test accuracy over standard dropout, remains effective even with frozen masks, generates interpretable visualizations of discriminative features and are available here https://github.com/Tom-Dvk/PB-EDropout/tree/main. Tom Devynck, Bilal Faye, Djamel Bouchaffra, Nadjib Lazaar, Hanene Azzag, Mustapha Lebbah |
ESANN | 6 |
| 2026 | Game Theory Meets Statistical Physics: A Novel Deep Neural Networks DesignabstractWe introduce a novel deep graphical representation that integrates game theory (GT) principles with the laws of statistical physics (SP), enabling feature extraction and pattern classification within a unified learning framework. In our approach, neurons in a network are analogous to players in a GT model. Each neuron, viewed as a classical particle governed by the laws of SP, corresponds to a set of actions that represent specific activation values. The feed-forward process in deep learning (DL) is interpreted as a sequential game with each game involving multiple players. During training, neurons are evaluated iteratively and filtered based on their contributions to a payoff function, which is quantified using the Shapley value driven by a Gaussian-Boltzmann energy model. To mitigate the computational burden of exact Shapley value computations, we employ Monte-Carlo (MC) sampling, reducing the algorithmic complexity from exponential to polynomial. This approximation significantly improves scalability, making our framework suitable for larger networks. Neurons that significantly contribute to the payoff form strong coalitions, and only these neurons are allowed to propagate information to the next layers. Using the Shapley value, we devised a new model regularization technique, thereby improving overall performance. We applied this framework to facial age estimation and gender classification tasks. Experimental results show that our approach outperforms several traditional and recent machine learning models in terms of accuracy, precision, recall, and $F1$ -score. Djamel Bouchaffra, Fayçal Ykhlef, Bilal Faye, Mustapha Lebbah, Hanene Azzag |
IEEE Trans. Cybern. | 4 |
| 2025 | Leveraging Text-to-Text Transformers as Classifier Chain for Few-Shot Multi-Label ClassificationabstractMulti-label text classification (MLTC) is an essential task in NLP applications.Traditional methods require extensive labeled data and are limited to fixed label sets.Extracting labels with large language models (LLMs) is more effective and universal, but incurs high computational costs.In this work, we introduce a distillation-based T5 generalist model for zero-shot MLTC and few-shot fine-tuning.Our model accommodates variable label sets with general domain-agnostic pretraining, while modeling dependency between labels.Experiments show that our approach outperforms baselines of similar size on three few-shot tasks.Our code is available at repository. Quang Anh Nguyen, Nadi Tomeh, Mustapha Lebbah, Thierry Charnois, Hanene Azzag |
EMNLP | 3 |
| 2025 | Performance monitoring and wear comprehension through Neural NetworkabstractIn this paper, we present a novel approach to modeling the wear of complex dynamic systems, exemplified by aircraft engines, through the construction of a structured latent space.Unlike traditional methods, our model does not rely on explicit wear data but instead leverages supervised training to minimize the error on observable system parameters.Beyond wear forecasting, this work offers a foundation for unsupervised diagnosis, risk prevention, and the quantification of repair impacts. Thomas Binet, Hanene Azzag, Mustapha Lebbah, Jérôme Lacaille |
ESANN | 3 |
| 2025 | Context normalization: A new approach for the stability and improvement of neural network performanceabstractDeep neural networks face challenges with distribution shifts across layers, affecting model convergence and performance. While Batch Normalization (BN) addresses these issues, its reliance on a single Gaussian distribution assumption limits adaptability. To overcome this, alternatives like Layer Normalization, Group Normalization, and Mixture Normalization emerged, yet struggle with dynamic activation distributions. We propose ”Context Normalization” (CN), introducing contexts constructed from domain knowledge. CN normalizes data within the same context, enabling local representation. During backpropagation, CN learns normalized parameters and model weights for each context, ensuring efficient convergence and superior performance compared to BN and MN. This approach emphasizes context utilization, offering a fresh perspective on activation normalization in neural networks. We release our code at https://github.com/b-faye/Context-Normalization . Bilal Faye, Hanene Azzag, Mustapha Lebbah, Fangchen Feng |
Data Knowl. Eng. | 3 |
| 2025 | A distributed inference algorithm for Dirichlet process mixture models with exponential family componentsabstractThis article extends our earlier work published in [24] . Dirichlet Process Mixture Models (DPMMs) are widely used for clustering due to their ability to infer the number of clusters via a Bayesian non-parametric framework. However, MCMC inference with DPMMs becomes computationally intensive on large datasets. To address this, we propose a distributed inference algorithm, DisCGS, which approximates the collapsed Gibbs sampler using sufficient statistics and is specifically designed for horizontally distributed data across independent and potentially heterogeneous machines, making it well-suited for federated learning scenarios. Our contributions are threefold: first, we develop and evaluate DisCGS for the multivariate Gaussian DPMM, demonstrating significant computational gains—for example, reducing runtime from approximately 12 h to 3 min for 100 iterations on 100 K data points, a 200 × speedup—while maintaining comparable clustering quality as measured by Adjusted Rand Index (ARI), Normalized Mutual Information (NMI), and Accuracy. Second, we introduce a multinomial DPMM formulation within the DisCGS framework tailored for discrete data such as text documents, showcasing its applicability to categorical and count-based data common in natural language processing and genomics. Third, we prove that DisCGS generalizes to all exponential family distributions, thereby extending its flexibility and applicability beyond Gaussian and multinomial models. The source code for the continuous (Gaussian) model is available at https://github.com/redakhoufache/DisCGS , and the discrete (multinomial) version for text clustering is provided at https://github.com/redakhoufache/DisCGS-for-Discrete-data . Reda Khoufache, Mustapha Lebbah, Hanene Azzag, Étienne Goffinet, Djamel Bouchaffra |
Neurocomputing | 2 |
| 2025 | OneEncoder: a lightweight framework for efficient multimodal training
Bilal Faye, Hanene Azzag, Mustapha Lebbah, Djamel Bouchaffra |
Neural Comput. Appl. | 3 |
| 2025 | Redesigning deep neural networks: Bridging game theory and statistical physics
Djamel Bouchaffra, Fayçal Ykhlef, Bilal Faye, Mustapha Lebbah, Hanene Azzag |
Neural Networks | 4 |
| 2024 | Enhancing Few-Shot Topic Classification with Verbalizers. a Study on Automatic Verbalizer and Ensemble MethodsabstractAs pretrained language model emerge and consistently develop, prompt-based training has become a well-studied paradigm to improve the exploitation of models for many natural language processing tasks. Furthermore, prompting demonstrates great performance compared to conventional fine-tuning in scenarios with limited annotated data, such as zero-shot or few-shot situations. Verbalizers are crucial in this context, as they help interpret masked word distributions generated by language models into output predictions. This study introduces a benchmarking approach to assess three common baselines of verbalizers for topic classification in few-shot learning scenarios. Additionally, we find that increasing the number of label words for automatic label word searching enhances model performance. Moreover, we investigate the effectiveness of template assembling with various aggregation strategies to develop stronger classifiers that outperform models trained with individual templates. Our approach achieves comparable results to prior research while using significantly fewer resources. Our code is available at https://github.com/quang-anh-nguyen/verbalizer_benchmark.git. Quang Anh Nguyen, Nadi Tomeh, Mustapha Lebbah, Thierry Charnois, Hanene Azzag, Santiago Cordoba Muñoz |
LREC/COLING | 3 |
| 2024 | HDBSCAN for 3-rd order tensorabstractSeveral methods for tensor clustering require hyperparameters such as the cluster size or the number of clusters per mode.These methods present a challenge because, for real datasets, such inputs cannot be determined without incurring significant costs.Recently, Multi-Slice Clustering (MSC) has addressed this issue by utilizing a threshold parameter to perform data clustering.MSC identifies signal slices that reside in a lower-dimensional subspace within a 3rd-order rank-1 tensor dataset.However, determining the tensor rank remains a complex task.The current work introduces a new approach to tensor clustering that can extract clusters of similar slices and is also capable of finding co-clustering and triclustering in 3rd-order tensors of any rank.Our algorithm is based on the density of the data. Dina Faneva Andriantsiory, Joseph Ben Geloun, Mustapha Lebbah |
ESANN | 3 |
| 2024 | Lightweight Cross-Modal Representation LearningabstractLow-cost cross-modal representation learning is crucial for deriving semantic representations across diverse modalities such as text, audio, images, and video.Traditional approaches typically depend on large specialized models trained from scratch, requiring extensive datasets and resulting in high resource and time costs.To overcome these challenges, we introduce a novel approach named Lightweight Cross-Modal Representation Learning (LightCRL).This method uses a single neural network titled Deep Fusion Encoder (DFE), which projects data from multiple modalities into a shared latent representation space.This reduces the overall parameter count while still delivering robust performance comparable to more complex systems.The code is available via https://github.com/b Bilal Faye, Hanene Azzag, Mustapha Lebbah, Djamel Bouchaffra |
ESANN | 3 |
| 2024 | From Data to Simulation: Capturing Aircraft Engine Degradation DynamicsabstractThe analysis and simulation of aircraft engine behavior have garnered significant attention in the aeronautical industry, primarily due to its implications for performance, maintenance, safety, and sustainability.Our work successfully showcases the efficacy of utilizing time series data collected from our aircraft engines to construct a digital twin capable of dynamically emulating their real-time behavior.We then introduce a new methodology to model the physical engine's degradation and meticulously monitor its evolution over time.By continuously analyzing the simulated data against real-world performance measurements, our approach offers valuable insights into the engine's long-term behavior and health trajectory. Abdellah Madane, Florent Forest, Hanene Azzag, Mustapha Lebbah, Jérôme Lacaille |
ESANN | 4 |
| 2024 | Adaptative Context Normalization: A Boost for Deep Learning in Image ProcessingabstractDeep Neural network learning for image processing faces major challenges related to changes in distribution across layers, which disrupt model convergence and performance. Activation normalization methods, such as Batch Normalization (BN), have revolutionized this field, but they rely on the simplified assumption that data distribution can be modelled by a single Gaussian distribution. To overcome these limitations, Mixture Normalization (MN) introduced an approach based on a Gaussian Mixture Model (GMM), assuming multiple components to model the data. However, this method entails substantial computational requirements associated with the use of Expectation-Maximization algorithm to estimate parameters of each Gaussian components. To address this issue, we introduce Adaptative Context Normalization (ACN), a novel supervised approach that introduces the concept of “context”, which groups together a set of data with similar characteristics. Data belonging to the same context are normalized using the same parameters, enabling local representation based on contexts. For each context, the normalized parameters, as the model weights are learned during the backpropagation phase. ACN not only ensures speed, convergence, and superior performance compared to BN and MN but also presents a fresh perspective that underscores its particular efficacy in the field of image processing. We release our code at https://github.com/b-faye/Adaptative-Context-Normalization. Bilal Faye, Hanene Azzag, Mustapha Lebbah, Djamel Bouchaffra |
ICIP | 3 |
| 2024 | AESim: A Data-Driven Aircraft Engine Simulator
Abdellah Madane, Florent Forest, Hanene Azzag, Mustapha Lebbah, Jérôme Lacaille |
IJCAI | 4 |
| 2024 | UAN: Unsupervised Adaptive NormalizationabstractDeep neural networks have become a staple in solving intricate problems, proving their mettle in a wide array of applications. However, their training process is often hampered by shifting activation distributions during backpropagation, resulting in unstable gradients. Batch Normalization (BN) addresses this issue by normalizing activations, which allows for the use of higher learning rates. Despite its benefits, BN is not without drawbacks, including its dependence on mini-batch size and the presumption of a uniform distribution of samples. To overcome this, several alternatives have been proposed, such as Layer Normalization, Group Normalization, and Mixture Normalization. These methods may still struggle to adapt to the dynamic distributions of neuron activations during the learning process. To bridge this gap, we introduce Unsupervised Adaptive Normalization (UAN), an innovative algorithm that seamlessly integrates clustering for normalization with deep neural network learning in a singular process. UAN executes clustering using the Gaussian mixture model, determining parameters for each identified cluster, by normalizing neuron activations. These parameters are concurrently updated as weights in the deep neural network, aligning with the specific requirements of the target task during backpropagation. This unified approach of clustering and normalization, underpinned by neuron activation normalization, fosters an adaptive data representation that is specifically tailored to the target task. This adaptive feature of UAN enhances gradient stability, resulting in faster learning and augmented neural network performance. UAN outperforms the classical methods by adapting to the target task and is effective in classification, and domain adaptation. We release our code at github repository Bilal Faye, Hanene Azzag, Mustapha Lebbah, Fangchen Fang |
IJCNN | 3 |
| 2024 | One-Pass Generation of Multivariate Time Series through Conditional Multivariate ModelingabstractIn recent years, exploring deep generative models for generating time series has garnered significant interest within the research community. These models have found wide-ranging applications in areas such as data augmentation, scenario simulation, and the imputation of missing data. The authenticity of the generated time series has seen remarkable advancements with the integration of recurrent neural networks (RNNs) and generative adversarial networks (GANs). RNNs used to represent the state-of-the-art (SOA) in processing sequence dependencies until the advent of Transformers, which redefined the SOA, especially in Natural Language Processing and Computer Vision. The introduction of a transformer-based GAN represented an innovative step forward, aiming to address the limitations inherent in RNNs. However, this model’s efficacy is constrained when faced with unimodal data distribution assumptions, leading to arbitrary outputs in complex distribution scenarios. This paper introduces a novel Multivariate Time Series Conditional GAN (MTS-CGAN), that leverages transformer-based architectures in generator and discriminator networks. MTS-CGAN conditions the generation process on a specific encoded context (categorical and MTS inputs), enabling one-pass generation of multivariate time series, and accommodating mixed distribution frameworks, outperforming existing models. We evaluate MTS-CGAN using quantitative metrics across multiple multivariate time series datasets. Furthermore, we propose also an innovative adaptation of the Frechet Inception Distance (FID), tailored for time series, to assess the quality of the generated data. This research demonstrates the potential of MTS-CGAN in generating high-fidelity multivariate time series. Abdellah Madane, Florent Forest, Hanene Azzag, Mustapha Lebbah, Jérôme Lacaille |
IJCNN | 4 |
| 2024 | Distributed MCMC Inference for Bayesian Non-parametric Latent Block Model
Reda Khoufache, Anisse Belhadj, Hanene Azzag, Mustapha Lebbah |
PAKDD (1) | 4 |
| 2024 | Distributed Collapsed Gibbs Sampler for Dirichlet Process Mixture Models in Federated LearningabstractDirichlet Process Mixture Models (DPMMs) are widely used to address clustering problems. Their main advantage lies in their ability to automatically estimate the number of clusters during the inference process through the Bayesian non-parametric framework. However, the inference becomes considerably slow as the dataset size increases. This paper proposes a new distributed Markov Chain Monte Carlo (MCMC) inference method for DPMMs (DisCGS) using sufficient statistics. Our approach uses the collapsed Gibbs sampler and is specifically designed to work on distributed data across independent and heterogeneous machines, which habilitates its use in horizontal federated learning. Our method achieves highly promising results and notable scalability. For instance, with a dataset of 100K data points, the centralized algorithm requires approximately 12 hours to complete 100 iterations while our approach achieves the same number of iterations in just 3 minutes, reducing the execution time by a factor of 200 without compromising clustering performance. The code source is publicly available at https://github.com/redakhoufache/DisCGS. Reda Khoufache, Mustapha Lebbah, Hanene Azzag, Étienne Goffinet, Djamel Bouchaffra |
SDM | 2 |
| 2024 | Exploring accuracy and interpretability trade-off in tabular learning with novel attention-based models
Kodjo Mawuena Amekoe, Hanene Azzag, Zaineb Chelly Dagdia, Mustapha Lebbah, Gregoire Jaffre |
Neural Comput. Appl. | 4 |
| 2023 | TabSRA: An Attention based Self-Explainable Model for Tabular LearningabstractWe propose TabSRA, a novel self-explainable, and accurate model for tabular learning.TabSRA is based on SRA (Self-Reinforcement Attention), new attention mechanism that helps to learn an intelligible representation of the raw input data through element-wise vector multiplication.The learned representation is aggregated by a highly transparent function (e.g linear), which produces the final output.Experimental results on synthetic and real-world classification problems show that the proposed TabSRA solution outperforms existing widely used self-explainable models and performs comparably to full complexity state-of-the-art models in term of accuracy while providing a faithful feature attribution.Source code is available at https://github.com/anselmeamekoe/TabSRA. Kodjo Mawuena Amekoe, Mohamed Djallel Dilmi, Hanene Azzag, Zaineb Chelly Dagdia, Mustapha Lebbah, Gregoire Jaffre |
ESANN | 5 |
| 2023 | Selecting the Number of Clusters K with a Stability Trade-off: An Internal Validation Criterion
Alex Mourer, Florent Forest, Mustapha Lebbah, Hanene Azzag, Jérôme Lacaille |
PAKDD (1) | 3 |
| 2023 | Regions of interest selection in histopathological images using subspace and multi-objective stream clustering
Mohammed Oualid Attaoui, Nassima Dif, Hanene Azzag, Mustapha Lebbah |
Vis. Comput. | 4 |
| 2022 | Transfer learning from synthetic labels for histopathological images classification
Nassima Dif, Mohammed Oualid Attaoui, Zakaria Elberrichi, Mustapha Lebbah, Hanene Azzag |
Appl. Intell. | 4 |
| 2021 | A New Subspace Multi-Objective Approach for the Clustering and Selection of Regions of Interests in Histopathological ImagesabstractHistopathology images represent a source of assistance for pathologists when diagnosing Cancer. However, in histopathology or in cancer image analysis, pathologists mostly diagnose the pathology as positive if a small part of it is considered cancer tissue. These small parts are called regions of interest (ROI) or patches. Finding the relevant patches is crucial as it can save computation time and memory. Subspace clustering discovers clusters embedded in multiple, overlapping subspaces of high dimensional data. It is an extension of feature selection, which tries to identify relevant subsets of features that are relevant to the clustering process. However, subspace clustering algorithms provide a partition of the data based on one cluster validity measure, assuming a homogeneous similarity measure over the entire data set makes the algorithms not robust to variations in the data characteristics. Therefore, it is beneficial to optimize multiple validity indices simultaneously to capture different aspects of the datasets. The goal of Multi-Objective clustering methods (MOC) is to derive significant clusters by applying two or more objective functions. This paper proposes a new clustering algorithm for patch selection based on subspace and multi-objective clustering to find the data's best partitioning and the images' most relevant patches. Mohammed Oualid Attaoui, Hanene Azzag, Nabil Keskes, Mustapha Lebbah |
CEC | 4 |
| 2021 | Multivariate Time Series Multi-Coclustering. Application to Advanced Driving Assistance System ValidationabstractDriver assistance systems development remains a technical challenge for car manufacturers.Validating these systems requires to assess the assistance systems performances in a considerable number of driving contexts.Groupe Renault uses massive simulation for this task, which allows reproducing the complexity of physical driving conditions precisely and produces large volumes of multivariate time series.We present the operational constraints and scientific challenges related to these datasets and our proposal of an adapted model-based multiple coclustering approach, which creates several independent partitions by grouping redundant variables.This method natively performs model selection, missing values inference, noisy samples handling, confidence interval production, while keeping a sparse parameter numbers.The proposed model is evaluated on a synthetic dataset, and applied to a driver assistance system validation use-case. Étienne Goffinet, Mustapha Lebbah, Hanene Azzag, Loïc Giraldi, Anthony Coutant |
ESANN | 2 |
| 2021 | A New Nearest Neighbor Median Shift Clustering for Binary Data
Gaël Beck, Mustapha Lebbah, Hanene Azzag, Tarn Duong |
ICANN (5) | 2 |
| 2021 | Multi-Slice Clustering for 3-order TensorabstractSeveral methods of triclustering of three dimensional data require the specification of the cluster size in each dimension. This introduces a certain degree of arbitrariness. To address this issue, we propose a new method, namely the multi-slice clustering (MSC) for a 3-order tensor data set. We analyse, in each dimension or tensor mode, the spectral decomposition of each tensor slice, i.e. a matrix. Thus, we define a similarity measure between matrix slices up to a threshold (precision) parameter, and from that, identify a cluster. The intersection of all partial clusters provides the desired triclustering. The effectiveness of our algorithm is shown on both synthetic and real-world data sets. Dina Faneva Andriantsiory, Joseph Ben Geloun, Mustapha Lebbah |
ICMLA | 3 |
| 2021 | Experience feedback using Representation Learning for Few-Shot Object Detection on Aerial ImagesabstractThis paper proposes a few-shot method based on Faster R-CNN and representation learning for object detection in aerial images. The two classification branches of Faster R-CNN are replaced by prototypical networks for online adaptation to new classes. These networks produce embeddings vectors for each generated box, which are then compared with class prototypes. The distance between an embedding and a prototype determines the corresponding classification score. The networks are trained in an episodic manner. A new detection task is randomly sampled at each epoch, consisting in detecting only a subset of the classes annotated in the dataset. This strategy encourages the network to adapt to new classes as it would at test time. In addition, several ideas are explored to improve the proposed method such as a hard negative examples mining strategy and self-supervised clustering for background objects. The performance of our method is assessed on DOTA, a large-scale remote sensing images dataset. The experiments conducted provide a broader understanding of the capabilities of representation learning. It highlights in particular some intrinsic weaknesses for the few-shot object detection task. Finally, some suggestions and perspectives are formulated according to these insights. Pierre Le Jeune, Mustapha Lebbah, Anissa Zergaïnoh-Mokraoui, Hanene Azzag |
ICMLA | 2 |
| 2021 | An Evolutionary Computing-Based Efficient Hybrid Task Scheduling Approach for Heterogeneous Computing Environment
Muhammad Sulaiman 0003, Zahid Halim, Mustapha Lebbah, Muhammad Waqas 0001, Shanshan Tu |
J. Grid Comput. | 3 |
| 2021 | Subspace data stream clustering with global and local weighting models
Mohammed Oualid Attaoui, Hanene Azzag, Mustapha Lebbah, Nabil Keskes |
Neural Comput. Appl. | 3 |
| 2021 | Deep embedded self-organizing maps for joint representation learning and topology-preserving clustering
Florent Forest, Mustapha Lebbah, Hanene Azzag, Jérôme Lacaille |
Neural Comput. Appl. | 2 |
| 2020 | An Invariance-guided Stability Criterion for Time Series Clustering ValidationabstractTime series clustering is a challenging task due to the specificities of this type of data. Temporal correlation and invariance to transformations such as shifting, warping or noise prevent the use of standard data mining methods. Time series clustering has been mostly studied under the angle of finding efficient algorithms and distance metrics adapted to the specific nature of time series data. Much less attention has been devoted to the general problem of model selection. Clustering stability has emerged as a universal and model-agnostic principle for clustering model selection. This principle can be stated as follows: an algorithm should find a structure in the data that is resilient to perturbation by sampling or noise. We propose to apply stability analysis to time series by leveraging prior knowledge on the nature and invariances of the data. These invariances determine the perturbation process used to assess stability. Based on a recently introduced criterion combining between-cluster and within-cluster stability, we propose an invariance-guided method for model selection, applicable to a wide range of clustering algorithms. Experiments conducted on artificial and benchmark data sets demonstrate the ability of our criterion to discover structure and select the correct number of clusters, whenever data invariances are known beforehand. Florent Forest, Alex Mourer, Mustapha Lebbah, Hanene Azzag, Jérôme Lacaille |
ICPR | 3 |
| 2020 | Autonomous Driving Validation with Model-Based Dictionary Clustering
Étienne Goffinet, Mustapha Lebbah, Hanene Azzag, Loïc Giraldi |
ECML/PKDD (4) | 2 |
| 2020 | A scalable and effective rough set theory-based approach for big data pre-processingabstractAbstract A big challenge in the knowledge discovery process is to perform data pre-processing, specifically feature selection, on a large amount of data and high dimensional attribute set. A variety of techniques have been proposed in the literature to deal with this challenge with different degrees of success as most of these techniques need further information about the given input data for thresholding, need to specify noise levels or use some feature ranking procedures. To overcome these limitations, rough set theory (RST) can be used to discover the dependency within the data and reduce the number of attributes enclosed in an input data set while using the data alone and requiring no supplementary information. However, when it comes to massive data sets, RST reaches its limits as it is highly computationally expensive. In this paper, we propose a scalable and effective rough set theory-based approach for large-scale data pre-processing, specifically for feature selection, under the Spark framework. In our detailed experiments, data sets with up to 10,000 attributes have been considered, revealing that our proposed solution achieves a good speedup and performs its feature selection task well without sacrificing performance. Thus, making it relevant to big data. Zaineb Chelly Dagdia, Christine Zarges, Gaël Beck, Mustapha Lebbah |
Knowl. Inf. Syst. | 4 |
| 2019 | Deep Embedded SOM: joint representation learning and self-organization
Florent Forest, Mustapha Lebbah, Hanene Azzag, Jérôme Lacaille |
ESANN | 2 |
| 2019 | Soft Subspace Growing Neural Gas for Data Stream Clustering
Mohammed Oualid Attaoui, Mustapha Lebbah, Nabil Keskes, Hanene Azzag, Mohammed Ghesmoune |
ICANN (4) | 2 |
| 2019 | A distributed approximate nearest neighbors algorithm for efficient large scale mean shift clustering
Gaël Beck, Tarn Duong, Mustapha Lebbah, Hanene Azzag, Christophe Cérin |
J. Parallel Distributed Comput. | 3 |
| 2018 | A Distributed Rough Set Theory Algorithm based on Locality Sensitive Hashing for an Efficient Big Data Pre-processingabstractA big challenge in the knowledge discovery process is to perform big data pre-processing; specifically feature selection. To handle this challenge, Rough Set Theory (RST) has been considered as one of the most powerful techniques as it has much to offer for feature selection. To extend its applicability to big data, a distributed version of RST was developed. However, one of its key challenges is the partitioning of the feature search space in the distributed environment while guaranteeing data dependency. In this paper, we propose a new distributed version of RST based on Locality Sensitive Hashing (LSH), named LSH-dRST, for big data pre-processing. LSH-dRST uses LSH to match similar features into the same bucket and maps the generated buckets into partitions to enable the splitting of the universe in a more appropriate way. We compare LSH-dRST to the standard distributed RST technique which is based on a random partitioning of the universe and demonstrate that our LSH-dRST is not only scalable but also more reliable for feature selection; making it more relevant to big data pre-processing. We also demonstrate that our LSH-dRST ensures the partitioning of the high dimensional feature search space in a more reliable way. Hence, guarantees data dependency in the distributed environment, and ensures a lower computational cost. Zaineb Chelly Dagdia, Christine Zarges, Gaël Beck, Hanene Azzag, Mustapha Lebbah |
IEEE BigData | 5 |
| 2018 | A Generic and Scalable Pipeline for Large-Scale Analytics of Continuous Aircraft Engine DataabstractA major application of data analytics for aircraft engine manufacturers is engine health monitoring, which consists in improving availability and operation of engines by leveraging operational data and past events. Traditional tools can no longer handle the increasing volume and velocity of data collected on modern aircraft. We propose a generic and scalable pipeline for large-scale analytics of operational data from a recent type of aircraft engine, oriented towards health monitoring applications. Based on Hadoop and Spark, our approach enables domain experts to scale their algorithms and extract features from tens of thousands of flights stored on a cluster. All computations are performed using the Spark framework, however custom functions and algorithms can be integrated without knowledge of distributed programming. Unsupervised learning algorithms are integrated for clustering and dimensionality reduction of the flight features, in order to allow efficient visualization and interpretation through a dedicated web application. The use case guiding our work is a methodology for engine fleet monitoring with a self-organizing map. Finally, this pipeline is meant to be end-to-end, fully customizable and ready for use in an industrial setting. Florent Forest, Jérôme Lacaille, Mustapha Lebbah, Hanene Azzag |
IEEE BigData | 3 |
| 2018 | A Complete Data Science Work-flow For Insurance FieldabstractIn recent years, "Big Data" has become a new ubiquitous term. Big Data is transforming science, engineering, medicine, health-care, finance, business, and ultimately our society itself. Learning from Big Data has become a significant challenge and requires development of new types of algorithms. Most machine learning algorithms can not easily scale up to Big Data. MapReduce is a simplified programming model for processing large datasets in a distributed and parallel manner. In this paper, we present our work carried in a big data project1which is dedicated to the insurance sector. This allows us to validate our method on real-world data for insurance. We present the complete pipeline or work-flow going from data collection to visualization, passing by data fusion, data analysis, clustering, and prediction tasks. The insurance dataset is enriched with data collected from heterogeneous sources. A predictive and analysis system is proposed by combining the clustering result with decision trees. We use the topological approach, especially the SOM method, for its interest in being able to cluster and visualize the data at the same time. We make the source code of our SOM-MapReduce algorithm, written with Spark using the MapReduce paradigm, publicly available2. Mohammed Ghesmoune, Mustapha Lebbah, Hanene Azzag, Salima Benbernou, Mourad Ouziri, Tarn Duong |
IEEE BigData | 2 |
| 2018 | Hierarchical Laplacian Score for unsupervised feature selectionabstractIn this paper, we address the problem of unsupervised feature selection. This is an important challenge due to the absence of class labels that would guide the search for relevant information. Motivated by this challenge, we define the new method named Hierarchical Laplacian Score (HLS) that constrains the Laplacian Score using a tree topology structure. The purpose of using this structure is to automatically discover local data structure and local nearest neighbors for each data object. Experimental results on various datasets have demonstrated the effectiveness of the proposed algorithm in clustering and classification applications. Nhat-Quang Doan, Hanene Azzag, Mustapha Lebbah |
IJCNN | 3 |
| 2017 | Return of experience on the mean-shift clustering for heterogeneous architecture use caseabstractThe exponential increment in data size poses new challenges for computer scientists, giving rise to a new set of methodologies under the term Big Data. Many efficient algorithms for machine learning have been proposed, facing up time and memory requirements. Nevertheless, with hardware acceleration, multiple software instructions can be integrated and executed into a single hardware die. Current researches aim at eliminating the burden for the user in using multiple processor types. In this paper we propose our return of experience on a new way of implementing machine learning algorithms on heterogeneous hardware. To explore our vision, we use a parallel Mean-shift algorithm, developed at LIPN as our case study to investigate issues in building efficient Machine Learning libraries for heterogeneous systems. The ultimate goal is to provide a core set of building blocks for Machine Learning programming that could serve either to build new applications on heterogeneous architectures or to control the evolution of the underlying platform. We thus examine the difficulties encountered during the implementation of the algorithm with the aim to discover methodologies for building systems based on heterogeneous hardware. We also discover issues and building blocks for solving concrete machine learning (ML) problems on the Chisel software stack we use for this purpose. Christophe Cérin, Jean-Luc Gaudiot, Mustapha Lebbah, Foutse Yuehgoh |
IEEE BigData | 3 |
| 2017 | A distributed rough set theory based algorithm for an efficient big data pre-processing under the spark frameworkabstractBig Data reduction is a main point of interest across a wide variety of fields. This domain was further investigated when the difficulty in quickly acquiring the most useful information from the huge amount of data at hand was encountered. To achieve the task of data reduction, specifically feature selection, several state-of-the-art methods were proposed. However, most of them require additional information about the given data for thresholding, noise levels to be specified or they even need a feature ranking procedure. Thus, it seems necessary to think about a more adequate feature selection technique which can extract features using information contained within the dataset alone. Rough Set Theory (RST) can be used as such a technique to discover data dependencies and to reduce the number of features contained in a dataset using the data alone, requiring no additional information. However, despite being a powerful feature selection technique, RST is computationally expensive and only practical for small datasets. Therefore, in this paper, we present a novel efficient distributed Rough Set Theory based algorithm for large-scale data pre-processing under the Spark framework. Our experimental results show the efficient applicability of our RST solution to Big Data without any significant information loss. Zaineb Chelly Dagdia, Christine Zarges, Gaël Beck, Mustapha Lebbah |
IEEE BigData | 4 |
| 2017 | An Instance Based Model for Scalable Theta -SubsumptionabstractThe θ-subsumption test is known to be a bottleneck in Inductive Logic Programming. The state-of-the-art learning systems in this field are hardly scalable. Last year, we have created a distributed θ-subsumption process based on an Actor Model, with the aim of being able to decide subsumption on very large clauses. This model was correct and complete, but was also very slow. This is why we introduce ANTS (Actor Network based Theta-Subsumption), a new model also based on an actor network, which is significantly faster than the previous one. Hippolyte Léger, Dominique Bouthinon, Mustapha Lebbah, Hanene Azzag |
ICTAI | 3 |
| 2017 | Big Data: from collection to visualization
Mohammed Ghesmoune, Hanene Azzag, Salima Benbernou, Mustapha Lebbah, Tarn Duong, Mourad Ouziri |
Mach. Learn. | 4 |
| 2016 | Distributed mean shift clustering with approximate nearest neighboursabstractWe introduce an efficient distributed implementation of nearest neighbour mean shift clustering (NNMS). The computationally intensive nature of NNMS has so far restricted its application to complex data sets where a flexible clustering with non-ellipsoidal clusters would be beneficial. A parallel implementation of the standard serial NNMS algorithm on its own brings insufficient performance gains so we introduce two further algorithmic improvements: a normal scale (NS) choice of the optimal number of nearest neighbours, and locality sensitive hashing (LSH) to approximate nearest neighbour searches. Combining these improvements into a single distributed algorithm DNNMS offers the potential for an efficient method for Big Data Clustering. Gaël Beck, Tarn Duong, Hanene Azzag, Mustapha Lebbah |
IJCNN | 4 |
| 2016 | GTM Mixture through time for sequential dataabstractGenerative Topographic Mapping (GTM) is a popular probabilistic framework for modeling non-linear relationships in high-dimensional data as well as for unsupervised learning and visualization of such data. It is also known as to provide a principled probabilistic alternative to the well-known Self-Organizing Map (SOM) in the neural networks community, thanks to its flexible mixture model formulation and the desirable properties of the expectation-maximization (EM) algorithm. However, much attention has been focused on the use of GTM for multivariate data, in general assumed to be independent and identically distributed (i.i.d) and the problem of modeling sequences using GTM is less investigated. In this paper, we focus on GTM for unsupervised modeling and visualization of sequential data. We consider modeling sequences of continuous multidimensional observations and we propose a GTM through time (GTM-TT) approach based on hidden Markov models (HMM) where the observations are a sent of independent sequences, rather than a signle sequence. We further extend the model to the clustering of multiple sequences by proposing a GTM-TT mixture model. The model parameters are estimated by maximum likelihood via the EM algorithm. The proposed approach is evaluated using simulated data and real-world data. Rakia Jaziri, Faicel Chamroukhi, Mustapha Lebbah, Younès Bennani |
IJCNN | 3 |
| 2016 | CL-AntInc Algorithm for Clustering Binary Data Streams Using the Ants BehaviorabstractIn this paper, we present a new approach using a non-hierarchical method in graph environment and the concept of artificial ants for both clustering and visualization using Tulip framework. This model can be presented to take into account data in blocks in an incremental way. It seems especially interesting to process binary data streaming. In this algorithm, we also suggest to apply swarm intelligence techniques for the incremental processing of this new challenging data type. The main novelty of this research work resides on the adaptation of CL-AntInc to perform clustering binary data streams and building growing graphs increasingly for this type of data. The proposed algorithm performance is evaluated using real world data sets extracted from Machine Learning Repository. Our algorithm is competitive when compared with other stream clustering methods. Nesrine Masmoudi, Hanene Azzag, Mustapha Lebbah, Cyrille Bertelle, Maher Ben Jemaa |
KES | 3 |
| 2016 | A new Growing Neural Gas for clustering data streams
Mohammed Ghesmoune, Mustapha Lebbah, Hanene Azzag |
Neural Networks | 2 |
| 2016 | Nearest neighbour estimators of density derivatives, with application to mean shift clustering
Tarn Duong, Gaël Beck, Hanene Azzag, Mustapha Lebbah |
Pattern Recognit. Lett. | 4 |
| 2015 | How to use ants for data stream clusteringabstractWe present in this paper a new bio-inspired algorithm that dynamically creates groups of data. This algorithm is based on the concept of artificial ants that move together in a complex manner with simple localization rules. Each ant represents one datum in the algorithm. The moves of ants aim at creating homogeneous groups of data that evolve together in a graph environment. We also suggest an extension to this algorithm to treat data streaming. The extended algorithm has been tested on real-world data. Our algorithms yielded competitive results as compared to K-means and Ascending Hierarchical Clustering (AHC), two well known methods. Nesrine Masmoudi, Hanene Azzag, Mustapha Lebbah, Cyrille Bertelle, Maher Ben Jemaa |
CEC | 3 |
| 2015 | Clustering of Binary Data Sets Using Artificial Ants Algorithm
Nesrine Masmoudi, Hanene Azzag, Mustapha Lebbah, Cyrille Bertelle, Maher Ben Jemaa |
ICONIP (1) | 3 |
| 2015 | Growing Hierarchical Trees for Data Stream clustering and visualizationabstractData stream clustering aims at studying large volumes of data that arrive continuously and the objective is to build a good clustering of the stream, using a small amount of memory and time. Visualization is still a big challenge for large data streams. In this paper we present a new approach using a hierarchical and topological structure (or network) for both clustering and visualization. The topological network is represented by a graph in which each neuron represents a set of similar data points and neighbor neurons are connected by edges. The hierarchical component consists of multiple tree-like hierarchic of clusters which allow to describe the evolution of data stream, and then analyze explicitly their similarity. This adaptive structure can be exploited by descending top-down from the topological level to any hierarchical level. The performance of the proposed algorithm is evaluated on both synthetic and real-world datasets. Nhat-Quang Doan, Mohammed Ghesmoune, Hanene Azzag, Mustapha Lebbah |
IJCNN | 4 |
| 2015 | Clustering Over Data Streams Based on Growing Neural Gas
Mohammed Ghesmoune, Mustapha Lebbah, Hanene Azzag |
PAKDD (2) | 2 |
| 2015 | Probabilistic Self-Organizing Map for Clustering and Visualizing non-i.i.d DataabstractWe present a generative approach to train a new probabilistic self-organizing map (PrSOMS) for dependent and nonidentically distributed data sets. Our model defines a low-dimensional manifold allowing friendly visualizations. To yield the topology preserving maps, our model has the SOM like learning behavior with the advantages of probabilistic models. This new paradigm uses hidden Markov models (HMM) formalism and introduces relationships between the states. This allows us to take advantage of all the known classical views associated to topographic map. The objective function optimization has a clear interpretation, which allows us to propose expectation-maximization (EM) algorithm, based on the forward–backward algorithm, to train the model. We demonstrate our approach on two data sets: The real-world data issued from the "French National Audiovisual Institute" and handwriting data captured using a WACOM tablet. Mustapha Lebbah, Rakia Jaziri, Younès Bennani, Jean-Hugues Chenot |
Int. J. Comput. Intell. Appl. | 1 |
| 2014 | Biclustering using Spark-MapReduceabstractBiclustering approaches are more complex compared to the traditional clustering particularly those requiring large dataset and Mapreduce platforms. We propose a new approach of biclustering based on popular self-organizing maps, which is one of the famous unsupervised learning algorithms. We have designed scalable implementations of the new topological biclustering algorithm using MapReduce with the Spark platform. Tugdual Sarazin, Mustapha Lebbah, Hanene Azzag |
IEEE BigData | 2 |
| 2014 | G-Stream: Growing Neural Gas over Data Stream
Mohammed Ghesmoune, Hanene Azzag, Mustapha Lebbah |
ICONIP (1) | 3 |
| 2014 | Feature Group Weighting and Topological Biclustering
Tugdual Sarazin, Mustapha Lebbah, Hanene Azzag, Amine Chaibi |
ICONIP (2) | 2 |
| 2013 | A new bi-clustering approach using topological mapsabstractIn this paper, we propose a new bi-clustering algorithm based on self-organizing maps titled BiTM (Bi-clustering using Topological Map). BiTM provides a simultaneous clustering of rows and columns of the data matrix in order to increase the homogeneity of bi-clusters by respecting neighborhood relationship and using a single map. BiTM maps provide a new topological visualization of the bi-clusters. Experimental results and comparison studies show that BiTM improves the results in term of bi-clustering and visualization. Amine Chaibi, Mustapha Lebbah, Hanene Azzag |
IJCNN | 2 |
| 2013 | Self-organizing trees for visualizing protein datasetabstractClustering and visualizing multidimensional or structured data are important tasks for data analysis, especially in bioinformatics. Self-organizing models are often used to address both of these issues. In this paper we introduce a hierarchical and topological visualization technique called Self-organizing Trees (SoT) which is able to represent data in hierarchical and topological structure. The experiment is conducted on a real-world protein data set. Nhat-Quang Doan, Hanene Azzag, Mustapha Lebbah, Guillaume Santini |
IJCNN | 3 |
| 2013 | A New Visualization of Group-Outliers in Unsupervised LearningabstractThis paper presents a new method for computing a quantitative score which can help in detecting cluster outliers using visualisation task. Self-organising map is incorporated in the proposed approach. The proposed method is evaluated on a number of datasets from UCI. Visualizations and experimental results show that GOF sensibly improves the results in term of cluster-outlier detection. The development of the SOM based visualization tool intends to provide additional exploratory data analysis techniques by offering a tool that allows effective extraction and exploration of patterns. Amine Chaibi, Mustapha Lebbah, Hanene Azzag |
IV | 2 |
| 2013 | Group Outlier factor: a New Score using Self-Organising Map for Group-Outlier and Novelty DetectionabstractThis paper describe a new concept of "cluster outlier-ness". In order to quantify it, we propose a relative isolation score named group outlier factor (GOF). GOF is a score, which is computed during a clustering process using self-organizing maps. The main difference between GOF and existing methods is that, being an outlier is not associated to a single pattern but to a cluster. Thus, an outlier factor (OF) with respect to each cluster is computed for each new sample and compared to the GOF score associated for each cluster. OF is used as a novelty detection classifier. This approach allows to identify meaningful outlier-clusters and detects novel patterns that previous approaches could not find. Experimental results and comparison studies show that the use of GOF sensibly improves the results in term of cluster-outlier and novelty detection. Amine Chaibi, Mustapha Lebbah, Hanene Azzag |
Int. J. Comput. Intell. Appl. | 2 |
| 2012 | Automatic Group-Outlier Detection
Amine Chaibi, Hanene Azzag, Mustapha Lebbah |
ESANN | 3 |
| 2012 | Self-Organizing Map and Tree Topology for Graph Summarization
Nhat-Quang Doan, Hanene Azzag, Mustapha Lebbah |
ICANN (2) | 3 |
| 2012 | Novelty Detection Using a New Group Outlier Factor
Amine Chaibi, Mustapha Lebbah, Hanene Azzag |
ICONIP (3) | 2 |
| 2012 | Growing Self-organizing Trees for knowledge discovery from dataabstractIn this paper, we propose a new unsupervised learning method based on growing neural gas and using self-assembly rules to build hierarchical structures. Our method named GSoT (Growing Self-organizing Trees) depicts data in topological and hierarchical organization. This makes GSoT a good tool for data clustering and knowledge discovery. Experiments conducted on real data sets demonstrate the good performance of GSoT. Nhat-Quang Doan, Hanene Azzag, Mustapha Lebbah |
IJCNN | 3 |
| 2012 | Graph Decomposition Using Self-organizing TreesabstractIn this paper, we present a new approach for graph decomposition using topological and hierarchical partitioning of data. Our method called GD-SOM-Tree (Graph Decomposition using Self-Organizing Trees) is based on self-organizing models. The benefit of this novel approach is to represent and visualize hierarchical relations which replace the original graph with a summary and gives a good understanding of the underlying problem. Nhat-Quang Doan, Hanene Azzag, Mustapha Lebbah |
IV | 3 |
| 2011 | SOS-HMM: Self-Organizing Structure of Hidden Markov Model
Rakia Jaziri, Mustapha Lebbah, Younès Bennani, Jean-Hugues Chenot |
ICANN (2) | 2 |
| 2011 | Probabilistic Self-Organizing Maps for multivariate sequencesabstractThis paper describes a new algorithm to learn a new probabilistic Self-Organizing Map for not independent and not identically distributed data set. This new paradigm probabilistic self-organizing map uses HMM (Hidden Markov Models) formalism and introduces relationships between the states of the map. The map structure is integrated in the parameter estimation of Markov model using a neighborhood function to learn a topographic clustering. We have applied this novel model to cluster and to reconstruct the data captured using a WACOM tablet. Rakia Jaziri, Mustapha Lebbah, Nicoleta Rogovschi, Younès Bennani |
IJCNN | 2 |
| 2010 | Map-TreeMaps: A New Approach for Hierarchical and Topological ClusteringabstractWe present in this paper a new clustering method which provides self-organization of hierarchical clustering. This method represents large datasets on a forest of original trees which are projected on a simple 2D geometric relationship using tree map representation. The obtained partition is represented by a map of tree maps, which define a tree of data. In this paper, we provide the rules that build a tree of node/data by using distance between data in order to decide where connect nodes. Visual and empirical results based on both synthetic and real datasets from the UCI repository, are given and discussed. Hanene Azzag, Mustapha Lebbah, Aymen Arfaoui |
ICMLA | 2 |
| 2010 | Topological Hierarchical Tree Using Artificial Ants
Mustapha Lebbah, Hanene Azzag |
ICONIP (1) | 1 |
| 2010 | Topographic under-sampling for unbalanced distributionsabstractSeveral aspects could affect the existing machine learning algorithms. One of these aspects is related to unbalanced classes in which the number of observations belonging to a class, greatly exceeds the observations in other classes. We propose in this paper an under-sampling method which uses self-organizing map to cluster the majority class guided with minority class. The proposed approach has been validated on multiple data sets using decision trees as a classifier with cross validation. The experimental results showed that elimination from majority class by integrating Neighborhood Cleaning Rule in SOM algorithm, produce high and very promising performance. Fatma Hamdi, Mustapha Lebbah, Younès Bennani |
IJCNN | 2 |
| 2010 | Visualization and clustering of categorical data with probabilistic self-organizing map
Mustapha Lebbah, Khalid Benabdeslem |
Neural Comput. Appl. | 1 |
| 2009 | From variable weighting to cluster characterization in topographic unsupervised learningabstractWe introduce a new learning approach, which provides simultaneously self-organizing map (SOM) and local weight vector for each cluster. The proposed approach is computationally simple, and learns a different features vector weights for each cell (relevance vector). Based on the self-organizing map approach, we present two new simultaneously clustering and weighting algorithms: local weighting observation lwo-SOM and local weighting distance lwd-SOM. Both algorithms achieve the same goal by minimizing different cost functions. After learning phase, a selection method with weight vectors is used to prune the irrelevant variables and thus we can characterize the clusters. We illustrate the performance of the proposed approach using different data sets. A number of synthetic and real data are experimented on to show the benefits of the proposed local weighting using self-organizing models. Nistor Grozavu, Younès Bennani, Mustapha Lebbah |
IJCNN | 3 |
| 2008 | Clustering of Self-Organizing Map
Hanene Azzag, Mustapha Lebbah |
ESANN | 2 |
| 2008 | Relational Analysis for Consensus Clustering from Multiple PartitionsabstractThis paper deals with the problem of combining multiple clustering algorithms using the same data set to get a single consensus clustering. Our contribution is to formally define the cluster consensus problem as an optimization problem. to reach this goal, we propose an original existing algorithm but still relatively unknown method named relational analysis (RA). This method has several advantages among which we can quote: its low computational complexity, it does not require a number of clusters and does not neglect the weak clustering result. The unsupervised clustering consensus method implemented in this work is quite general. We evaluate the effectiveness of cluster consensus in three qualitatively different data sets. Promising results are provided in all three situations for synthetic as well as real data sets. Mustapha Lebbah, Younès Bennani, Hamid Benhadda |
ICMLA | 1 |
| 2008 | Probabilistic Mixed Topological Map for Categorical and Continuous DataabstractThis paper introduces a new probabilistic topological map as generative model that includes mixture of Gaussian and Bernoulli distribution. This model is dedicated to cluster mixed data with continuous and categorical variables. This model is fitted by maximum likelihood using the EM algorithm. Examples using real data set allow to validate our model. The proposed approach has the advantage comparing to existing topological map of providing a set of prototype with the same coding as the learning data. More information is produced with this model that could be used in practical applications. Nicoleta Rogovschi, Mustapha Lebbah, Younès Bennani |
ICMLA | 2 |
| 2008 | A Probabilistic Self-Organizing Map for Binary Data Topographic ClusteringabstractThis paper introduces a probabilistic self-organizing map for topographic clustering, analysis and visualization of multivariate binary data or categorical data using binary coding. We propose a probabilistic formalism dedicated to binary data in which cells are represented by a Bernoulli distribution. Each cell is characterized by a prototype with the same binary coding as used in the data space and the probability of being different from this prototype. The learning algorithm, Bernoulli on self-organizing map, that we propose is an application of the EM standard algorithm. We illustrate the power of this method with six data sets taken from a public data set repository. The results show a good quality of the topological ordering and homogenous clustering. Mustapha Lebbah, Younès Bennani, Nicoleta Rogovschi |
Int. J. Comput. Intell. Appl. | 1 |
| 2007 | BeSOM : Bernoulli on Self-Organizing MapabstractThis paper introduces a probabilistic self-organizing map for clustering, analysis and visualization of multivariate binary data. We propose a probabilistic formalism dedicated to binary data in which cells are represented by a Bernoulli distribution. Each cell is characterized by a prototype with the same binary coding as used in the data space and the probability of being different from this prototype. The learning algorithm, BeSOM, that we propose is an application of the EM standard algorithm. We illustrate the power of this method with two data sets taken from a public data set repository: a handwritten digit data set and a zoo data set. The results show a good quality of the topological ordering and homogenous clustering. Mustapha Lebbah, Nicoleta Rogovschi, Younès Bennani |
IJCNN | 1 |
| 2005 | Mixed Topological Map
Mustapha Lebbah, Aymeric Chazottes, Fouad Badran, Sylvie Thiria |
ESANN | 1 |
| 2004 | Visualization and classification with categorical topological map
Mustapha Lebbah, Fouad Badran, Sylvie Thiria |
ESANN | 1 |
| 2002 | Categorical Topological Map
Mustapha Lebbah, Christian Chabanon, Fouad Badran, Sylvie Thiria |
ICANN | 1 |
| 2000 | Topological map for binary data
Mustapha Lebbah, Fouad Badran, Sylvie Thiria |
ESANN | 1 |