VLDB 2026 Research / reviewers in the wild / expert
Zaineb Chelly Dagdia
dblp:37/8336 · also Zeineb Chelly
· DBLP profile ↗
28ranked-venue papers
12as first author
13since 2021 · last 2026
0000-0002-2551-6586ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 8 first-author · 7 since 2021Databases, data management, data science and information retrieval · 9 · 5 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 3 since 2021Security and privacy · 3 · 2 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Privacy Preserving Personalized Next Location Prediction Via Encrypted Shuffled Federated Learning and Fuzzy ClusteringabstractInternational audience Saloua Bouabba, Karine Zeitouni, Bassem Haidar, Nazim Agoulmine, Zaineb Chelly Dagdia |
MDM | 5 |
| 2026 | Locally Differentially Private Synthesis of Decentralised Heterogeneous Social Graphs via Spectral Embeddings and Bayesian Optimisation
Manel Jerbi, Zaineb Chelly Dagdia, Sjouke Mauw |
SECRYPT (1) | 2 |
| 2026 | Rethinking fairness in unsupervised healthcare AI: A methodological scoping reviewabstractOBJECTIVE: Fairness in machine learning has been extensively studied in supervised settings, where labeled outcomes allow direct assessment of bias. In contrast, fairness in unsupervised learning-particularly in healthcare remains insufficiently examined. In the absence of labels, it is unclear how fairness should be defined or evaluated for discovered structures such as patient subgroups, disease subtypes, or trajectories, despite their growing influence on clinical understanding and decision-making. This review aims to systematically examine how fairness is conceptualized, operationalized, and evaluated in unsupervised healthcare AI. METHODS: We conducted a PRISMA-guided methodological scoping review of the literature on fairness in unsupervised learning applied to healthcare data. The review focused on identifying algorithmic mechanisms, evaluation strategies, and methodological assumptions rather than comparing predictive performance or quantitatively synthesizing results across the reviewed literature. Records were analyzed with respect to data modalities, unsupervised learning techniques, and the underlying definitions of fairness they employed. RESULTS: The review reveals rapid growth in interest in fairness-aware unsupervised healthcare AI, accompanied by substantial heterogeneity and conceptual inconsistency. Fairness is addressed across diverse data types and methodological approaches, often without explicit alignment to clinical or ethical objectives. To structure this fragmented landscape, we propose a taxonomy of fairness approaches organized into five families: Individual Fairness, Performance Dependence, Welfare-Anchored approaches, Statistical Inference-based approaches, and Representation Parity. Each family embodies a distinct conception of equity and entails specific ethical and methodological trade-offs. We further identify recurring challenges, including defining fairness without labeled outcomes, limited incorporation of clinical expertise, and weak alignment between fairness objectives and medical validity. CONCLUSION: Fairness in unsupervised healthcare AI is an emerging but conceptually unsettled field. Current approaches reflect diverse and sometimes incompatible notions of equity, underscoring the need for clearer theoretical grounding. Progress will require explicit articulation of fairness goals, stronger integration of domain expertise and participatory evaluation, and closer alignment between algorithmic fairness criteria and clinically meaningful structures. This review provides a conceptual and methodological foundation to support more rigorous and transparent development of fair unsupervised healthcare AI systems. Malek Adouani, Djillali Annane, Zaineb Chelly Dagdia |
J. Biomed. Informatics | 3 |
| 2025 | Fair and Privacy-Preserving Synthetic Data Generation via Clustering-Based Variational Autoencoder and Adversarially Debiased Wasserstein Generative Adversarial Networks with Gradient Penalty
Malek Adouani, Zaineb Chelly Dagdia |
ECML/PKDD (1) | 2 |
| 2024 | Federated TimeGAN for Privacy Preserving Synthetic Trajectory GenerationabstractMobility datasets are crucial for various applications. However, sharing this data raises privacy concerns due to the sensitive nature of geolocation information. Synthetic data generation has recently emerged as a promising solution to protect geo-privacy of trajectory data. Current approaches rely on having a large set of authentic trajectories collected from individual users to train generative networks. However, this assumption proves impractical in many real-world scenarios due to the sensitive personal information typically embedded within trajectories. Our approach leverages federated learning to generate privacy-preserving synthetic trajectories without the need for centralized data collection. Experimental results demonstrate that our distributed framework effectively produces synthetic trajectories with distributions comparable to baseline, offering a privacy-conscious alternative for geo-privacy protection in mobility datasets. Saloua Bouabba, Karine Zeitouni, Bassem Haidar, Nazim Agoulmine, Zaineb Chelly Dagdia |
MDM | 5 |
| 2024 | Exploring accuracy and interpretability trade-off in tabular learning with novel attention-based models
Kodjo Mawuena Amekoe, Hanene Azzag, Zaineb Chelly Dagdia, Mustapha Lebbah, Gregoire Jaffre |
Neural Comput. Appl. | 3 |
| 2023 | Immune-Based System to Enhance Malware DetectionabstractMalicious apps use various methods to spread viruses, take control of computers and/or IoT devices, and steal sensitive data such as credit card numbers or other personal information. Despite the numerous existing means of intrusion detection, malware code is not easily detectable. The primary issue with current malware detection approaches is their inability to identify novel attacks and obfuscated malware, as they rely on static bases of malware examples, making them susceptible to new unseen malware behaviors. To address this, we propose a new method for malware recognition, which consists of two processes: the first process creates new instances of malware using a memetic algorithm, and the second process detects these new instances of attacks through solid detectors produced by an artificial immune system-based algorithm. Our new malware recognition method has proven its merits through thorough experiments on widely used datasets and evaluation metrics, and has been compared to prominent state-of-the-art methods. Manel Jerbi, Zaineb Chelly Dagdia, Slim Bechikh, Lamjed Ben Said |
CEC | 2 |
| 2023 | TabSRA: An Attention based Self-Explainable Model for Tabular LearningabstractWe propose TabSRA, a novel self-explainable, and accurate model for tabular learning.TabSRA is based on SRA (Self-Reinforcement Attention), new attention mechanism that helps to learn an intelligible representation of the raw input data through element-wise vector multiplication.The learned representation is aggregated by a highly transparent function (e.g linear), which produces the final output.Experimental results on synthetic and real-world classification problems show that the proposed TabSRA solution outperforms existing widely used self-explainable models and performs comparably to full complexity state-of-the-art models in term of accuracy while providing a faithful feature attribution.Source code is available at https://github.com/anselmeamekoe/TabSRA. Kodjo Mawuena Amekoe, Mohamed Djallel Dilmi, Hanene Azzag, Zaineb Chelly Dagdia, Mustapha Lebbah, Gregoire Jaffre |
ESANN | 4 |
| 2023 | Clustering Corticosteroids Responsiveness in Sepsis Patients using Game-Theoretic Rough SetsabstractPerforming data mining tasks in the medical domain poses a significant challenge, mainly due to the uncertainty present in patients' data, such as incompleteness or missingness.In this paper, we focus on the data mining task of clustering corticosteroid (CS) responsiveness in sepsis patients.We address the issue and challenge of missing data by applying Game-Theoretic Rough Sets (GTRS) as a three-way decision approach.Our study considers the APROCCHS cohort, comprising 1240 sepsis patients, provided by the Assistance Publique-Hôpitaux de Paris (AP-HP), France.Our experimental results on the APROCCHS cohort indicate that GTRS maintains the trade-off between accuracy and generality, demonstrating its effectiveness even when increasing the number of missing values. Rahma Hellali, Zaineb Chelly Dagdia, Karine Zeitouni |
FedCSIS | 2 |
| 2022 | Malware Evolution and Detection Based on the Variable Precision Rough Set ModelabstractTo offer innovative malware evolution techniques, it is appealing to integrate approaches that handle imperfect data and knowledge.In fact, malware writers tend to target some precise features within the app's code to camouflage the malicious content.Those features may sometimes present conflictual information about the true nature of the content of the app (malicious/benign).In this paper, we show how the Variable Precision Rough Set (VPRS) model can be combined with optimization techniques, in particular Bilevel-Optimization-Problems (BLOPs), in order to establish a detection model capable of following the crazy race of malware evolution initiated among malware-developers.We propose a new malware detection technique, based on such hybridization, named Variable Precision Rough set Malware Detection (ProRSDet), that offers robust detection rules capable of revealing the new nature of a given app.ProRSDet attains encouraging results when tested against various state-of-the-art malware detection systems using common evaluation metrics. Manel Jerbi, Zaineb Chelly Dagdia, Slim Bechikh, Lamjed Ben Said |
FedCSIS | 2 |
| 2022 | Android malware detection as a Bi-level problem
Manel Jerbi, Zaineb Chelly Dagdia, Slim Bechikh, Lamjed Ben Said |
Comput. Secur. | 2 |
| 2021 | Malware Detection Using Rough Set Based Evolutionary Optimization
Manel Jerbi, Zaineb Chelly Dagdia, Slim Bechikh, Lamjed Ben Said |
ICONIP (5) | 2 |
| 2021 | A Detailed Study of the Distributed Rough Set Based Locality Sensitive Hashing Feature Selection TechniqueabstractInternational audience Zaineb Chelly Dagdia, Christine Zarges |
Fundam. Informaticae | 1 |
| 2020 | Automatic Rule Extraction from Access Rules Using Genetic Programming
Paloma de las Cuevas, Pablo García-Sánchez, Zaineb Chelly Dagdia, Maribel García Arenas, Juan Julián Merelo Guervós |
EvoApplications | 3 |
| 2020 | On the use of artificial malicious patterns for android malware detection
Manel Jerbi, Zaineb Chelly Dagdia, Slim Bechikh, Lamjed Ben Said |
Comput. Secur. | 2 |
| 2020 | A scalable and effective rough set theory-based approach for big data pre-processingabstractAbstract A big challenge in the knowledge discovery process is to perform data pre-processing, specifically feature selection, on a large amount of data and high dimensional attribute set. A variety of techniques have been proposed in the literature to deal with this challenge with different degrees of success as most of these techniques need further information about the given input data for thresholding, need to specify noise levels or use some feature ranking procedures. To overcome these limitations, rough set theory (RST) can be used to discover the dependency within the data and reduce the number of attributes enclosed in an input data set while using the data alone and requiring no supplementary information. However, when it comes to massive data sets, RST reaches its limits as it is highly computationally expensive. In this paper, we propose a scalable and effective rough set theory-based approach for large-scale data pre-processing, specifically for feature selection, under the Spark framework. In our detailed experiments, data sets with up to 10,000 attributes have been considered, revealing that our proposed solution achieves a good speedup and performs its feature selection task well without sacrificing performance. Thus, making it relevant to big data. Zaineb Chelly Dagdia, Christine Zarges, Gaël Beck, Mustapha Lebbah |
Knowl. Inf. Syst. | 1 |
| 2018 | A Distributed Rough Set Theory Algorithm based on Locality Sensitive Hashing for an Efficient Big Data Pre-processingabstractA big challenge in the knowledge discovery process is to perform big data pre-processing; specifically feature selection. To handle this challenge, Rough Set Theory (RST) has been considered as one of the most powerful techniques as it has much to offer for feature selection. To extend its applicability to big data, a distributed version of RST was developed. However, one of its key challenges is the partitioning of the feature search space in the distributed environment while guaranteeing data dependency. In this paper, we propose a new distributed version of RST based on Locality Sensitive Hashing (LSH), named LSH-dRST, for big data pre-processing. LSH-dRST uses LSH to match similar features into the same bucket and maps the generated buckets into partitions to enable the splitting of the universe in a more appropriate way. We compare LSH-dRST to the standard distributed RST technique which is based on a random partitioning of the universe and demonstrate that our LSH-dRST is not only scalable but also more reliable for feature selection; making it more relevant to big data pre-processing. We also demonstrate that our LSH-dRST ensures the partitioning of the high dimensional feature search space in a more reliable way. Hence, guarantees data dependency in the distributed environment, and ensures a lower computational cost. Zaineb Chelly Dagdia, Christine Zarges, Gaël Beck, Hanene Azzag, Mustapha Lebbah |
IEEE BigData | 1 |
| 2018 | Distributed Rough Set Based Feature Selection Approach to Analyse Deep and Hand-crafted Features for Mammography Mass ClassificationabstractBreast cancer has a high incidence among women worldwide. This, together with the recent developments in deep learning based convolutional networks, have motivated research towards the enhancement of Computer Aided Diagnosis (CAD) systems. In this paper, the performance of a densely connected convolutional network (DenseNet) for breast cancer was investigated for the malignant/benign classification of mammographic masses. Different mammography data sets were collected to investigate the capacity of this network for learning a combination of these databases. To achieve this, internal low-level, mid-level and high-level features/abstracts were extracted from the model together with hand-crafted features, generating a vast amount of data. Using the distributed rough set based feature selection approach (Sp-RST), significant features were selected from both deep learning based features and hand-crafted ones, and fed into a learning model with separate and combined data approaches for the classification of mammographic masses. Results show that by using Sp-RST as a powerful technique capable of performing big data preprocessing, DenseNet had the representational capacity to learn mammographic abnormalities. Azam Hamidinekoo, Zaineb Chelly Dagdia, Zobia Suhail, Reyer Zwiggelaar |
IEEE BigData | 2 |
| 2018 | Rough Set Theory as a Data Mining Technique: A Case Study in Epidemiology and Cancer Incidence Prediction
Zaineb Chelly Dagdia, Christine Zarges, Benjamin Schannes, Martin Micalef, Lino Galiana, Benoît Rolland, Olivier de Fresnoye, Mehdi Benchoufi |
ECML/PKDD (3) | 1 |
| 2017 | A distributed rough set theory based algorithm for an efficient big data pre-processing under the spark frameworkabstractBig Data reduction is a main point of interest across a wide variety of fields. This domain was further investigated when the difficulty in quickly acquiring the most useful information from the huge amount of data at hand was encountered. To achieve the task of data reduction, specifically feature selection, several state-of-the-art methods were proposed. However, most of them require additional information about the given data for thresholding, noise levels to be specified or they even need a feature ranking procedure. Thus, it seems necessary to think about a more adequate feature selection technique which can extract features using information contained within the dataset alone. Rough Set Theory (RST) can be used as such a technique to discover data dependencies and to reduce the number of features contained in a dataset using the data alone, requiring no additional information. However, despite being a powerful feature selection technique, RST is computationally expensive and only practical for small datasets. Therefore, in this paper, we present a novel efficient distributed Rough Set Theory based algorithm for large-scale data pre-processing under the Spark framework. Our experimental results show the efficient applicability of our RST solution to Big Data without any significant information loss. Zaineb Chelly Dagdia, Christine Zarges, Gaël Beck, Mustapha Lebbah |
IEEE BigData | 1 |
| 2016 | From the General to the Specific: Inducing a Novel Dendritic Cell Algorithm from a Detailed State-of-the-Art ReviewabstractConsidered as one of the emerging evolutionary algorithms, the Dendritic Cell Algorithm (DCA) is based on the behavior of specific immune agents; known as Dendritic Cells (DCs). Studies related with DCA are increasingly becoming popular and this is due to the worthy characteristics expressed by the algorithm as it exhibits several potentially beneficial features for binary classification problems. Yet, according to our best knowledge, there is no study summarizing the basic features of the DCA developed versions all in one paper. Hence, in this paper we aim at summarizing the powerful characteristics of the DCA while making a general review of the algorithm. We aim at studying the various versions of DCA while highlighting their characteristics, advantages and limitations. Based on this study and from the conducted conclusions, we intend to generate a well-studied novel DC classifier based on the positive aspects reflected by the previously proposed DCA versions. Results show that our proposed algorithm succeeds in obtaining significantly improved classification accuracy. Zaineb Chelly Dagdia, Zied Elouedi |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2016 | A survey of the dendritic cell algorithm
Zaineb Chelly Dagdia, Zied Elouedi |
Knowl. Inf. Syst. | 1 |
| 2015 | A New Version of the Dendritic Cell Immune Algorithm Based on the K-Nearest Neighbors
Kaouther Ben Ali, Zaineb Chelly Dagdia, Zied Elouedi |
ICONIP (1) | 2 |
| 2014 | A study of the data pre-processing module of the dendritic cell evolutionary algorithmabstractData reduction as a critical step in the process of data pre-processing presents a central point of interest across a wide variety of fields. Data pre-processing has a significant impact on the performance of any machine learning algorithm. In this context, we focus our research paper on investigating the data pre-processing phase of a recent evolutionary algorithm named the Dendritic Cell Algorithm (DCA). We aim at reviewing the data pre-processing phase of the DCA while making a comparative study of the used data reduction techniques within the DCA. This is needed to clarify the differences, the advantages and the characteristics of the previously proposed techniques with the DCA. The output of the comparison will facilitate the task of the developer to select the most useful technique to be adopted and integrated in the DCA data pre-processing module. Zaineb Chelly Dagdia, Zied Elouedi |
CoDIT | 1 |
| 2014 | A two-leveled hybrid dendritic cell algorithm under imprecise reasoningabstractThe Dendritic Cell Algorithm(DCA) is a bio-inspired algorithm based on the behavior of Dendritic Cells(DCs). The DCA performance relies on its data pre-processing phase where feature extraction and signal categorization are performed and which are based on the use of the Principal Component Analysis(PCA) technique. However, using PCA presents a limitation as it destroys the underlying semantics of the features after reduction. To overcome this limitation, Rough Set Theory(RST) was applied as a pre-processor; but, still the developed rough approach presents an information loss as data should be discretized beforehand. Indeed, DCA was known to be sensitive to the input class data order. This is due to the crisp separation between the two DCs contexts; semi-mature and mature. Thus, the aim of this paper is to develop a novel DCA version based on a two-leveled hybrid model handling the mentioned DCA shortcomings. In the top-level, our proposed algorithm applies a more adequate feature extraction technique based on Fuzzy Rough Set Theory(FRST) to build a solid data pre-processing phase. At the bottom level, our algorithm applies Fuzzy Set Theory to smooth the crisp separation between the DCs contexts. Results show that our proposed algorithm succeeds in obtaining significantly improved classification accuracy. Zaineb Chelly Dagdia, Zied Elouedi |
GECCO | 1 |
| 2013 | A Fuzzy-Rough Data Pre-processing Approach for the Dendritic Cell Classifier
Zaineb Chelly Dagdia, Zied Elouedi |
ECSQARU | 1 |
| 2013 | Supporting Fuzzy-Rough Sets in the Dendritic Cell Algorithm Data Pre-processing Phase
Zaineb Chelly Dagdia, Zied Elouedi |
ICONIP (2) | 1 |
| 2012 | RST-DCA: A Dendritic Cell Algorithm Based on Rough Set Theory
Zaineb Chelly Dagdia, Zied Elouedi |
ICONIP (3) | 1 |