Mohamed Aymen Ben HajKacem

dblp:145/5865 · DBLP profile ↗
← Back
16ranked-venue papers
7as first author
5since 2021 · last 2025
0000-0002-0161-6646ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 5 first-author · 2 since 2021Databases, data management, data science and information retrieval · 8 · 4 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2025 Balancing Explainability and Accuracy in Credit Risk Classification Using Neuro-Fuzzy Model
abstract
Credit risk classification using financial and transactional data improves decision-making and enhances the accuracy of risk assessment. While deep learning models have shown strong predictive capabilities in credit risk classification, their black-box nature often limits interpretability and trust, especially in the banking sector. This paper investigates how to develop credit risk classification models that balance between accuracy and explainability, promoting transparency and confidence for financial decision makers. To achieve this, we propose a twophase explainable credit classification approach. The first phase uses neuro-fuzzy modeling to learn predictive models from data in the form of IF-THEN rules. The second phase introduces a pruning technique to reduce the number of generated rules by removing redundant or less important ones. Experiments conducted on two real credit risk datasets demonstrate that the proposed method maintains high predictive accuracy while enhancing explainability.
Sirine Ben Ghozzi, Mohamed Aymen Ben HajKacem, Nadia Essoussi
AICCSA2
2024 Explainable Ensemble Machine Learning Method for Credit Risk Classification
abstract
Credit risk classification (CRC) is a crucial task for banks to determine the financial position of the client for credit. Several machine learning models were proposed to deal with credit risk classification. However, conventional methods operate as black-box models that provide only the classification of clients without providing further explanations. To address this issue, we propose an explainable ensemble machine learning method for credit risk classification named EEML. The proposed method is based on combining five different machine learning models into an aggregate-learner, creating a standalone ensemble model. The interpretability and explainability of EEML outputs are enhanced by leveraging the capabilities of Shapley Additive exPlanations (SHAP). Experiments conducted on two real credit risk datasets have shown the performance of EEML compared to existing explainable credit risk classification methods. The EEML method gives high accuracy in classifying clients while providing explainability.
Sirine Ben Ghozzi, Mohamed Aymen Ben HajKacem, Nadia Essoussi
INISTA2
2024 Multi-view subspace text clustering
Maha Fraj, Mohamed Aymen Ben HajKacem, Nadia Essoussi
J. Intell. Inf. Syst.2
2022 Detection of Hot Topics Using Multi-view Text Clustering
Maha Fraj, Mohamed Aymen Ben HajKacem, Nadia Essoussi
iiWAS2
2021 Spark Based Text Clustering Method Using Hashing
Mohamed Aymen Ben HajKacem, Chiheb-Eddine Ben N'cir, Nadia Essoussi
DaWaK1
2020 Self-Organizing Map for Multi-view Text Clustering
Maha Fraj, Mohamed Aymen Ben HajKacem, Nadia Essoussi
DaWaK2
2020 Parallel K-Prototypes Clustering with High Efficiency and Accuracy
Hiba Jridi, Mohamed Aymen Ben HajKacem, Nadia Essoussi
DaWaK2
2019 Ensemble Method for Multi-view Text Clustering
Maha Fraj, Mohamed Aymen Ben HajKacem, Nadia Essoussi
ICCCI (1)2
2019 STiMR k-Means: An Efficient Clustering Method for Big Data
abstract
Big Data clustering has become an important challenge in data analysis since several applications require scalable clustering methods to organize such data into groups of similar objects. Given the computational cost of most of the existing clustering methods, we propose in this paper a new clustering method, referred to as STiMR [Formula: see text]-means, able to provide good tradeoff between scalability and clustering quality. The proposed method is based on the combination of three acceleration techniques: sampling, triangle inequality and MapReduce. Sampling is used to reduce the number of data points when building cluster prototypes, triangle inequality is used to reduce the number of comparisons when looking for nearest clusters and MapReduce is used to configure a parallel framework for running the proposed method. Experiments performed on simulated and real datasets have shown the effectiveness of the proposed method, with the existing ones, in terms of running time, scalability and internal validity measures.
Mohamed Aymen Ben HajKacem, Chiheb-Eddine Ben N'cir, Nadia Essoussi
Int. J. Pattern Recognit. Artif. Intell.1
2019 One-pass MapReduce-based clustering method for mixed large scale data
abstract
Big data is often characterized by a huge volume and a mixed types of attributes namely, numeric and categorical. K-prototypes has been fitted into MapReduce framework and hence it has become a solution for clustering mixed large scale data. However, k-prototypes requires computing all distances between each of the cluster centers and the data points. Many of these distance computations are redundant, because data points usually stay in the same cluster after first few iterations. Also, k-prototypes is not suitable for running within MapReduce framework: the iterative nature of k-prototypes cannot be modeled through MapReduce since at each iteration of k-prototypes, the whole data set must be read and written to disks and this results a high input/output (I/O) operations. To deal with these issues, we propose a new one-pass accelerated MapReduce-based k-prototypes clustering method for mixed large scale data. The proposed method reads and writes data only once which reduces largely the I/O operations compared to existing MapReduce implementation of k-prototypes. Furthermore, the proposed method is based on a pruning strategy to accelerate the clustering process by reducing the redundant distance computations between cluster centers and data points. Experiments performed on simulated and real data sets show that the proposed method is scalable and improves the efficiency of the existing k-prototypes methods.
Mohamed Aymen Ben HajKacem, Chiheb-Eddine Ben N'cir, Nadia Essoussi
J. Intell. Inf. Syst.1
2018 A Novel Tweets Clustering Method using Word Embeddings
abstract
Twitter users share a variety of information discussing multiple topics. Clustering-based methods have become an effective solution to group together tweets related to the same topics. However, due to the lack of context, short texts are challenging to model. Most of the existing clustering methods use the Vector Space Model (VSM) to transform tweets into a structured form. However, this representation do not consider the semantic relationships between words and suffers from high dimensionality and sparsity. To deal with these issues, we propose a new approach that aims to group tweets into topically coherent clusters by preserving the semantic links between words using word embeddings and low dimensional vector representations. The experimental results show the superior performance of the proposed method compared to existing ones.
Maha Fraj, Mohamed Aymen Ben HajKacem, Nadia Essoussi
AICCSA2
2018 Scalable Random Sampling K-Prototypes Using Spark
Mohamed Aymen Ben HajKacem, Chiheb-Eddine Ben N'cir, Nadia Essoussi
DaWaK1
2017 KP-S: A Spark-Based Design of the K-Prototypes Clustering for Big Data
abstract
Big data is often characterized by a huge volume and a mixed types of attributes namely, numeric and categorical. K-prototypes is one of the most well-known clustering methods to deal with mixed data. Several parallel alternatives based on MapReduce have been proposed to enable this method to handle large scale of mixed data. However, these solutions are not suitable when dealing with Big data, due to time and memory restrictions. To address this issue, we propose in this paper a new Spark-based k-prototypes clustering method which uses the reclustering technique. We take advantage of the in-memory operations of Spark to build grouping from large scale of mixed data. Experiments performed on simulated and real data sets show that the proposed method is scalable and improves the efficiency of the existing k-prototypes methods.
Mohamed Aymen Ben HajKacem, Chiheb-Eddine Ben N'cir, Nadia Essoussi
AICCSA1
2015 MapReduce-based k-prototypes clustering method for big data
abstract
Big data clustering is one of the recently challenging tasks that is used in many application domains. Traditional clustering methods are not able to deal with large-scale of data. Furthermore, Big data are often characterized by the mixed type of data, including numerical and categorical attributes. Thus, we propose in this paper the parallelization of k-prototypes clustering method (MR-KP) using MapReduce model to handle large-scale of mixed data. Experiments results show that MR-KP scales well with increasing data set sizes and achieves a close to linear speedup while maintaining the clustering accuracy.
Mohamed Aymen Ben HajKacem, Chiheb-Eddine Ben N'cir, Nadia Essoussi
DSAA1
2015 Parallel K-prototypes for Clustering Big Data
Mohamed Aymen Ben HajKacem, Chiheb-Eddine Ben N'cir, Nadia Essoussi
ICCCI (2)1
2014 A Three Stages to Implement Barriers in Bayesian-Based Bow Tie Diagram
Ahmed Badreddine, Mohamed Aymen Ben HajKacem, Nahla Ben Amor
IEA/AIE (2)2