Nadia Essoussi

dblp:62/797 · DBLP profile ↗
← Back
25ranked-venue papers
0as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 2 since 2021Databases, data management, data science and information retrieval · 10 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 3 since 2021Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1
YearPublicationVenuePosition
2025 Balancing Explainability and Accuracy in Credit Risk Classification Using Neuro-Fuzzy Model
abstract
Credit risk classification using financial and transactional data improves decision-making and enhances the accuracy of risk assessment. While deep learning models have shown strong predictive capabilities in credit risk classification, their black-box nature often limits interpretability and trust, especially in the banking sector. This paper investigates how to develop credit risk classification models that balance between accuracy and explainability, promoting transparency and confidence for financial decision makers. To achieve this, we propose a twophase explainable credit classification approach. The first phase uses neuro-fuzzy modeling to learn predictive models from data in the form of IF-THEN rules. The second phase introduces a pruning technique to reduce the number of generated rules by removing redundant or less important ones. Experiments conducted on two real credit risk datasets demonstrate that the proposed method maintains high predictive accuracy while enhancing explainability.
Sirine Ben Ghozzi, Mohamed Aymen Ben HajKacem, Nadia Essoussi
AICCSA3
2024 Explainable Ensemble Machine Learning Method for Credit Risk Classification
abstract
Credit risk classification (CRC) is a crucial task for banks to determine the financial position of the client for credit. Several machine learning models were proposed to deal with credit risk classification. However, conventional methods operate as black-box models that provide only the classification of clients without providing further explanations. To address this issue, we propose an explainable ensemble machine learning method for credit risk classification named EEML. The proposed method is based on combining five different machine learning models into an aggregate-learner, creating a standalone ensemble model. The interpretability and explainability of EEML outputs are enhanced by leveraging the capabilities of Shapley Additive exPlanations (SHAP). Experiments conducted on two real credit risk datasets have shown the performance of EEML compared to existing explainable credit risk classification methods. The EEML method gives high accuracy in classifying clients while providing explainability.
Sirine Ben Ghozzi, Mohamed Aymen Ben HajKacem, Nadia Essoussi
INISTA3
2024 Multi-view subspace text clustering
Maha Fraj, Mohamed Aymen Ben HajKacem, Nadia Essoussi
J. Intell. Inf. Syst.3
2022 Detection of Hot Topics Using Multi-view Text Clustering
Maha Fraj, Mohamed Aymen Ben HajKacem, Nadia Essoussi
iiWAS3
2021 Spark Based Text Clustering Method Using Hashing
Mohamed Aymen Ben HajKacem, Chiheb-Eddine Ben N'cir, Nadia Essoussi
DaWaK3
2020 Self-Organizing Map for Multi-view Text Clustering
Maha Fraj, Mohamed Aymen Ben HajKacem, Nadia Essoussi
DaWaK3
2020 Parallel K-Prototypes Clustering with High Efficiency and Accuracy
Hiba Jridi, Mohamed Aymen Ben HajKacem, Nadia Essoussi
DaWaK3
2019 Ensemble Method for Multi-view Text Clustering
Maha Fraj, Mohamed Aymen Ben HajKacem, Nadia Essoussi
ICCCI (1)3
2019 STiMR k-Means: An Efficient Clustering Method for Big Data
abstract
Big Data clustering has become an important challenge in data analysis since several applications require scalable clustering methods to organize such data into groups of similar objects. Given the computational cost of most of the existing clustering methods, we propose in this paper a new clustering method, referred to as STiMR [Formula: see text]-means, able to provide good tradeoff between scalability and clustering quality. The proposed method is based on the combination of three acceleration techniques: sampling, triangle inequality and MapReduce. Sampling is used to reduce the number of data points when building cluster prototypes, triangle inequality is used to reduce the number of comparisons when looking for nearest clusters and MapReduce is used to configure a parallel framework for running the proposed method. Experiments performed on simulated and real datasets have shown the effectiveness of the proposed method, with the existing ones, in terms of running time, scalability and internal validity measures.
Mohamed Aymen Ben HajKacem, Chiheb-Eddine Ben N'cir, Nadia Essoussi
Int. J. Pattern Recognit. Artif. Intell.3
2019 One-pass MapReduce-based clustering method for mixed large scale data
abstract
Big data is often characterized by a huge volume and a mixed types of attributes namely, numeric and categorical. K-prototypes has been fitted into MapReduce framework and hence it has become a solution for clustering mixed large scale data. However, k-prototypes requires computing all distances between each of the cluster centers and the data points. Many of these distance computations are redundant, because data points usually stay in the same cluster after first few iterations. Also, k-prototypes is not suitable for running within MapReduce framework: the iterative nature of k-prototypes cannot be modeled through MapReduce since at each iteration of k-prototypes, the whole data set must be read and written to disks and this results a high input/output (I/O) operations. To deal with these issues, we propose a new one-pass accelerated MapReduce-based k-prototypes clustering method for mixed large scale data. The proposed method reads and writes data only once which reduces largely the I/O operations compared to existing MapReduce implementation of k-prototypes. Furthermore, the proposed method is based on a pruning strategy to accelerate the clustering process by reducing the redundant distance computations between cluster centers and data points. Experiments performed on simulated and real data sets show that the proposed method is scalable and improves the efficiency of the existing k-prototypes methods.
Mohamed Aymen Ben HajKacem, Chiheb-Eddine Ben N'cir, Nadia Essoussi
J. Intell. Inf. Syst.3
2018 A Novel Tweets Clustering Method using Word Embeddings
abstract
Twitter users share a variety of information discussing multiple topics. Clustering-based methods have become an effective solution to group together tweets related to the same topics. However, due to the lack of context, short texts are challenging to model. Most of the existing clustering methods use the Vector Space Model (VSM) to transform tweets into a structured form. However, this representation do not consider the semantic relationships between words and suffers from high dimensionality and sparsity. To deal with these issues, we propose a new approach that aims to group tweets into topically coherent clusters by preserving the semantic links between words using word embeddings and low dimensional vector representations. The experimental results show the superior performance of the proposed method compared to existing ones.
Maha Fraj, Mohamed Aymen Ben HajKacem, Nadia Essoussi
AICCSA3
2018 Scalable Random Sampling K-Prototypes Using Spark
Mohamed Aymen Ben HajKacem, Chiheb-Eddine Ben N'cir, Nadia Essoussi
DaWaK3
2018 A New Way of Handling Missing Data in Multi-source Classification Based on Adaptive Imputation
Ikram Abdelkhalek, Afef Ben Brahim, Nadia Essoussi
MEDI3
2017 Learning Probabilistic Relational Models with (Partially Structured) Graph Databases
abstract
Probabilistic Relational Models (PRMs) such as Directed Acyclic Probabilistic Entity Relationship (DAPER) models are probabilistic models dealing with knowledge representation and relational data. Existing literature dealing with PRM and DAPER relies on well structured relational databases. In contrast, a large portion of real-world data is stored in Nosql databases specially graph databases that do not depend on a rigid schema. This paper builds on the recent work on DAPER models, and describes how to learn them from partially structured graph databases. Our contribution is twofold. First, we present how to extract the underlying ER model from a partially structured graph database. Then, we describe a method to compute sufficient statistics based on graph traversal techniques. Our objective is also twofold: we want to learn DAPERs with less structured data, and we want to accelerate the learning process by querying graph databases. Our experiments show that both objectives are completed, transforming the structure learning process into a more feasible task even when data are less structured than an usual relational database.
Marwa El Abri, Philippe Leray 0001, Nadia Essoussi
AICCSA3
2017 KP-S: A Spark-Based Design of the K-Prototypes Clustering for Big Data
abstract
Big data is often characterized by a huge volume and a mixed types of attributes namely, numeric and categorical. K-prototypes is one of the most well-known clustering methods to deal with mixed data. Several parallel alternatives based on MapReduce have been proposed to enable this method to handle large scale of mixed data. However, these solutions are not suitable when dealing with Big data, due to time and memory restrictions. To address this issue, we propose in this paper a new Spark-based k-prototypes clustering method which uses the reclustering technique. We take advantage of the in-memory operations of Spark to build grouping from large scale of mixed data. Experiments performed on simulated and real data sets show that the proposed method is scalable and improves the efficiency of the existing k-prototypes methods.
Mohamed Aymen Ben HajKacem, Chiheb-Eddine Ben N'cir, Nadia Essoussi
AICCSA3
2016 Parallel clustering method for non-disjoint partitioning of large-scale data based on spark framework
abstract
Clustering large scale data has become an important challenge which motivates several recent works. While the emphasis has been on the organization of massive data into disjoint groups, this work considers the identification of non-disjoint groups rather than the disjoint ones. In this setting, it is possible for data object to belong simultaneously to several groups since many real-world applications of clustering require non-disjoint partitioning to fit data structures. For this purpose, we propose the Parallel Overlapping k-means method (POKM) which is able to perform parallel clustering processes leading to non-disjoint partitioning of data. The proposed method is implemented within Spark framework to ensure the distribution of works over the different computation nodes. Experiments which we have performed on simulated and real-world multi-labeled datasets shows both faster execution times and high quality of clustering compared to existing methods.
Abir Zayani, Chiheb-Eddine Ben N'cir, Nadia Essoussi
IEEE BigData3
2016 A Hybrid Embedded-Filter Method for Improving Feature Selection Stability of Random Forests
Wassila Jerbi, Afef Ben Brahim, Nadia Essoussi
HIS3
2016 A Parallel Implementation of Relief Algorithm Using Mapreduce Paradigm
Jamila Yazidi, Bouaguel Waad, Nadia Essoussi
ICCCI (2)3
2016 Overlap regulation for additive overlapping clustering methods
abstract
Overlapping Clustering is an important technique in machine learning which aims to organize data into a set of non-disjoint groups rather than the disjoint one which is the case of conventional clustering methods. Several machine learning applications require that data object be assigned to one or several groups resulting in non-disjoint partitioning of data such as document clustering where each document can discuss one or many topics and then must be assigned to one or several groups. This paper presents a new partitional overlapping clustering method based on the additive model of overlaps. Compared to existing methods which build clusters with fixed size of overlaps, the proposed method gives users the ability to regulate this size. Experiments performed on simulated and real datasets show the performance of the proposed regulation principle to control the size of overlaps among groups.
Mohamed Ismail Maiza, Chiheb-Eddine Ben N'cir, Nadia Essoussi
RCIS3
2015 MapReduce-based k-prototypes clustering method for big data
abstract
Big data clustering is one of the recently challenging tasks that is used in many application domains. Traditional clustering methods are not able to deal with large-scale of data. Furthermore, Big data are often characterized by the mixed type of data, including numerical and categorical attributes. Thus, we propose in this paper the parallelization of k-prototypes clustering method (MR-KP) using MapReduce model to handle large-scale of mixed data. Experiments results show that MR-KP scales well with increasing data set sizes and achieves a close to linear speedup while maintaining the clustering accuracy.
Mohamed Aymen Ben HajKacem, Chiheb-Eddine Ben N'cir, Nadia Essoussi
DSAA3
2015 Parallel K-prototypes for Clustering Big Data
Mohamed Aymen Ben HajKacem, Chiheb-Eddine Ben N'cir, Nadia Essoussi
ICCCI (2)3
2015 Using Sequences of Words for Non-Disjoint Grouping of Documents
abstract
Grouping documents based on their textual content is an important application of clustering referred to as text clustering. This paper deals with two issues in text clustering which are the detection of non-disjoint groups and the representation of textual data. In fact, a text document can discuss several topics and then, it must belong to several groups. The learning algorithm must be able to produce non-disjoint clusters and assigns documents to several clusters. Given that text documents are considered as unstructured data, the application of a learning algorithm requires to prepare a set of documents for numerical analysis by using the vector space model (VSM). This representation of text avoids correlation between terms and does not give importance to the order of words in the text. Therefore, we present in this paper an unsupervised learning method, based on the word sequence kernel, where the correlation between adjacent words in text and the possibility of document to belong to more than one cluster are not ignored. In addition, to facilitate the use of this method in text-analytic practice, we present the "DocCO" software which is publicly available. Experiments performed on several text collections show that the proposed method outperforms existing overlapping methods using VSM representation in terms of clustering accuracy.
Chiheb-Eddine Ben N'cir, Nadia Essoussi
Int. J. Pattern Recognit. Artif. Intell.2
2014 Overlapping Clustering with Outliers Detection
abstract
Detecting overlapping groups is an important challenge in clustering offering relevant solutions for many applications domains. Recently, Parameterized R-OKM method was defined as an extension of OKM to control overlapping boundaries between clusters. However, the performance of both, OKM and Parameterized R-OKM is considerably reduced when data contain outliers. The presence of outliers affects the resulting clusters and yields to clusters which do not fit the true structure of data. In order to improve the existing methods, we propose a robust method able to detect relevant overlapping clusters with outliers identification.
Amira Rezgui, Chiheb-Eddine Ben N'cir, Nadia Essoussi
ICPRAM3
2014 Generalization of c-means for identifying non-disjoint clusters with overlap regulation
Chiheb-Eddine Ben N'cir, Guillaume Cleuziou, Nadia Essoussi
Pattern Recognit. Lett.3
2013 A Model-driven Process for Data Transformation of Heterogeneous Data
Haïfa Nakouri, Nadia Essoussi
MODELSWARD2