EDBT 2026 Demo / reviewers in the wild / expert
Hanene Azzag
dblp:a/HaneneAzzag · also Hanane Azzag, Hanene Azzag-Khelif
· DBLP profile ↗
13ranked-venue papers in the field
2as first author
5since 2021 · last 2026
0000-0001-6876-0688ORCID · verified
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 5Big Data, Cloud & Distributed Data Systems · 4Information Retrieval & Web Search · 2 (2 first)Database Systems & Data Management · 1Knowledge Engineering, Semantic Web & Information Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Special issue on intelligent systems, ISMIS'24 selected papers
Annalisa Appice, Hanene Azzag, Mohand-Said Hacid, Allel HadjAli |
J. Intell. Inf. Syst. | 2 |
| 2025 | Context normalization: A new approach for the stability and improvement of neural network performanceabstractDeep neural networks face challenges with distribution shifts across layers, affecting model convergence and performance. While Batch Normalization (BN) addresses these issues, its reliance on a single Gaussian distribution assumption limits adaptability. To overcome this, alternatives like Layer Normalization, Group Normalization, and Mixture Normalization emerged, yet struggle with dynamic activation distributions. We propose ”Context Normalization” (CN), introducing contexts constructed from domain knowledge. CN normalizes data within the same context, enabling local representation. During backpropagation, CN learns normalized parameters and model weights for each context, ensuring efficient convergence and superior performance compared to BN and MN. This approach emphasizes context utilization, offering a fresh perspective on activation normalization in neural networks. We release our code at https://github.com/b-faye/Context-Normalization . Bilal Faye, Hanene Azzag, Mustapha Lebbah, Fangchen Feng |
Data Knowl. Eng. | 2 |
| 2024 | Distributed MCMC Inference for Bayesian Non-parametric Latent Block Model
Reda Khoufache, Anisse Belhadj, Hanene Azzag, Mustapha Lebbah |
PAKDD (1) | 3 |
| 2024 | Distributed Collapsed Gibbs Sampler for Dirichlet Process Mixture Models in Federated LearningabstractDirichlet Process Mixture Models (DPMMs) are widely used to address clustering problems. Their main advantage lies in their ability to automatically estimate the number of clusters during the inference process through the Bayesian non-parametric framework. However, the inference becomes considerably slow as the dataset size increases. This paper proposes a new distributed Markov Chain Monte Carlo (MCMC) inference method for DPMMs (DisCGS) using sufficient statistics. Our approach uses the collapsed Gibbs sampler and is specifically designed to work on distributed data across independent and heterogeneous machines, which habilitates its use in horizontal federated learning. Our method achieves highly promising results and notable scalability. For instance, with a dataset of 100K data points, the centralized algorithm requires approximately 12 hours to complete 100 iterations while our approach achieves the same number of iterations in just 3 minutes, reducing the execution time by a factor of 200 without compromising clustering performance. The code source is publicly available at https://github.com/redakhoufache/DisCGS. Reda Khoufache, Mustapha Lebbah, Hanene Azzag, Étienne Goffinet, Djamel Bouchaffra |
SDM | 3 |
| 2023 | Selecting the Number of Clusters K with a Stability Trade-off: An Internal Validation Criterion
Alex Mourer, Florent Forest, Mustapha Lebbah, Hanene Azzag, Jérôme Lacaille |
PAKDD (1) | 4 |
| 2020 | Autonomous Driving Validation with Model-Based Dictionary Clustering
Étienne Goffinet, Mustapha Lebbah, Hanene Azzag, Loïc Giraldi |
ECML/PKDD (4) | 3 |
| 2018 | A Distributed Rough Set Theory Algorithm based on Locality Sensitive Hashing for an Efficient Big Data Pre-processingabstractA big challenge in the knowledge discovery process is to perform big data pre-processing; specifically feature selection. To handle this challenge, Rough Set Theory (RST) has been considered as one of the most powerful techniques as it has much to offer for feature selection. To extend its applicability to big data, a distributed version of RST was developed. However, one of its key challenges is the partitioning of the feature search space in the distributed environment while guaranteeing data dependency. In this paper, we propose a new distributed version of RST based on Locality Sensitive Hashing (LSH), named LSH-dRST, for big data pre-processing. LSH-dRST uses LSH to match similar features into the same bucket and maps the generated buckets into partitions to enable the splitting of the universe in a more appropriate way. We compare LSH-dRST to the standard distributed RST technique which is based on a random partitioning of the universe and demonstrate that our LSH-dRST is not only scalable but also more reliable for feature selection; making it more relevant to big data pre-processing. We also demonstrate that our LSH-dRST ensures the partitioning of the high dimensional feature search space in a more reliable way. Hence, guarantees data dependency in the distributed environment, and ensures a lower computational cost. Zaineb Chelly Dagdia, Christine Zarges, Gaël Beck, Hanene Azzag, Mustapha Lebbah |
IEEE BigData | 4 |
| 2018 | A Generic and Scalable Pipeline for Large-Scale Analytics of Continuous Aircraft Engine DataabstractA major application of data analytics for aircraft engine manufacturers is engine health monitoring, which consists in improving availability and operation of engines by leveraging operational data and past events. Traditional tools can no longer handle the increasing volume and velocity of data collected on modern aircraft. We propose a generic and scalable pipeline for large-scale analytics of operational data from a recent type of aircraft engine, oriented towards health monitoring applications. Based on Hadoop and Spark, our approach enables domain experts to scale their algorithms and extract features from tens of thousands of flights stored on a cluster. All computations are performed using the Spark framework, however custom functions and algorithms can be integrated without knowledge of distributed programming. Unsupervised learning algorithms are integrated for clustering and dimensionality reduction of the flight features, in order to allow efficient visualization and interpretation through a dedicated web application. The use case guiding our work is a methodology for engine fleet monitoring with a self-organizing map. Finally, this pipeline is meant to be end-to-end, fully customizable and ready for use in an industrial setting. Florent Forest, Jérôme Lacaille, Mustapha Lebbah, Hanene Azzag |
IEEE BigData | 4 |
| 2018 | A Complete Data Science Work-flow For Insurance FieldabstractIn recent years, "Big Data" has become a new ubiquitous term. Big Data is transforming science, engineering, medicine, health-care, finance, business, and ultimately our society itself. Learning from Big Data has become a significant challenge and requires development of new types of algorithms. Most machine learning algorithms can not easily scale up to Big Data. MapReduce is a simplified programming model for processing large datasets in a distributed and parallel manner. In this paper, we present our work carried in a big data project1which is dedicated to the insurance sector. This allows us to validate our method on real-world data for insurance. We present the complete pipeline or work-flow going from data collection to visualization, passing by data fusion, data analysis, clustering, and prediction tasks. The insurance dataset is enriched with data collected from heterogeneous sources. A predictive and analysis system is proposed by combining the clustering result with decision trees. We use the topological approach, especially the SOM method, for its interest in being able to cluster and visualize the data at the same time. We make the source code of our SOM-MapReduce algorithm, written with Spark using the MapReduce paradigm, publicly available2. Mohammed Ghesmoune, Mustapha Lebbah, Hanene Azzag, Salima Benbernou, Mourad Ouziri, Tarn Duong |
IEEE BigData | 3 |
| 2015 | Clustering Over Data Streams Based on Growing Neural Gas
Mohammed Ghesmoune, Mustapha Lebbah, Hanene Azzag |
PAKDD (2) | 3 |
| 2014 | Biclustering using Spark-MapReduceabstractBiclustering approaches are more complex compared to the traditional clustering particularly those requiring large dataset and Mapreduce platforms. We propose a new approach of biclustering based on popular self-organizing maps, which is one of the famous unsupervised learning algorithms. We have designed scalable implementations of the new topological biclustering algorithm using MapReduce with the Spark platform. Tugdual Sarazin, Mustapha Lebbah, Hanene Azzag |
IEEE BigData | 3 |
| 2007 | On building graphs of documents with artificial antsabstractWe present an incremental algorithm for building a neighborhood graph from a set of documents. This algorithm is based on a population of artificial agents that imitate the way real ants build structures with self-assembly behaviors. We show that our method outperforms standard algorithms for building such neighborhood graphs (up to 2230 times faster on the tested databases with equal quality) and how the user may interactively explore the graph. Hanene Azzag, Julien Lavergne, Christiane Guinot, Gilles Venturini |
WWW | 1 |
| 2006 | Generating maps of web pages using cellular automataabstractThe aim of web pages visualization is to present in a very informative and interactive way a set of web documents to the user in order to let him or her navigate through these documents. In the web context, this may correspond to several user's tasks: displaying the results of a search engine, or visualizing a graph of pages such as a hypertext or a surf map. In addition to web pages visualization, web pages clustering also greatly improves the amount of information presented to the user by highlighting the similarities between the documents [6]. In this paper we explore the use of a cellular automata (CA) to generate such maps of web pages. Hanene Azzag, David Ratsimba, David Da Costa, Gilles Venturini, Christiane Guinot |
WWW | 1 |