Maha Elarbi

dblp:180/8339 · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
12since 2021 · last 2025
0000-0002-7220-9474ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 4 first-author · 9 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Cross-Project Code Smell Detection as a Dynamic Optimization Problem: An Evolutionary Memetic Approach
abstract
Code smells signal poor software design that can prevent maintainability and scalability. Identifying code smells is difficult because of the large volume of code, considerable detection expenses, and the substantial effort needed for manual tagging. Although current techniques perform well in within-project situations, they frequently struggle to adapt to cross-project environments that have varying data distributions. In this paper, we introduce CLADES (Cross-project Learning and Adaptation for Detection of Code Smells), a hybrid evolutionary approach consisting of three main modules: Initialization, Evolution, and Adaptation. The first module generates an initial population of decision tree detectors using labeled within-project data and evaluates their quality through fitness functions based on structural code metrics. The evolution module applies genetic operators (selection, crossover, and mutation) to create new offspring solutions. To handle cross-project scenarios, the adaptation module employs a clustering-based instance selection technique that identifies representative instances from new projects, which are added to the dataset and used to repair the decision trees through simulated annealing. These locally refined decision trees are then evolved using a genetic algorithm, thus enabling continuous adaptation to new project instances. The resulting optimized decision tree detectors are then employed to predict labels for the new unlabeled project instances. We assess CLADES across five open-source projects and we show that it has a better performance with respect to baseline techniques in terms of weighted F1-score and AUC-PR metrics. These results emphasize its capacity to effectively adjust to different project environments, facilitating precise and scalable detection of code smells while minimizing the need for manual review, contributing to more robust and maintainable software systems.
Sofien Boutaib, Maha Elarbi, Slim Bechikh, Carlos A. Coello Coello, Lamjed Ben Said
CEC2
2025 Adaptive Normal-Boundary Intersection Directions for Evolutionary Many-Objective Optimization with Complex Pareto Fronts
Maha Elarbi, Slim Bechikh, Carlos A. Coello Coello
EMO (1)1
2025 Bi-level Evolutionary Model Tree Chain Induction for Multi-output Regression
Safa Mahouachi, Maha Elarbi, Slim Bechikh
Neurocomputing2
2024 A Bi-Level Evolutionary Model Tree Induction Approach for Regression
abstract
Supervised machine learning techniques include classification and regression. In regression, the objective is to map a real-valued output to a set of input features. The main challenge that existing methods for regression encounter is how to maintain an accuracy-simplicity balance. Since Regression Trees (RTs) are simple to interpret, many existing works have focused on proposing RT and Model Tree (MT) induction algorithms. MTs are RTs with a linear function at the leaf nodes rather than a numerical value are able to describe the relationship between the inputs and the output. Traditional RT induction algorithms are based on a top-down strategy which often leads to a local optimal solution. Other global approaches based on Evolutionary Algorithms (EAs) have been proposed to induce RTs but they can require an important calculation time which may affect the convergence of the algorithm to the solution. In this paper, we introduce a novel approach called Bi-level Evolutionary Model Tree Induction algorithm for regression, that we call BEMTI, and which is able to induce an MT in a bi-level design using an EA. The upper-level evolves a set of MTs using genetic operators while the lower-level optimizes the Linear Models (LMs) at the leaf nodes of each MT in order to fairly and precisely compute their fitness and obtain the optimal MT. The experimental study confirms the outperformance of our BEMTI compared to six existing tree induction algorithms on nineteen datasets.
Safa Mahouachi, Maha Elarbi, Khaled Sethom, Slim Bechikh, Carlos A. Coello Coello
CEC2
2023 Discretization-Based Feature Selection as a Bilevel Optimization Problem
abstract
Discretization-based feature selection (DBFS) approaches have shown interesting results when using several metaheuristic algorithms, such as particle swarm optimization (PSO), genetic algorithm (GA), ant colony optimization (ACO), etc. However, these methods share the same shortcoming which consists in encoding the problem solution as a sequence of cut-points. From this cut-points vector, the decision of deleting or selecting any feature is induced. Indeed, the number of generated cut-points varies from one feature to another. Thus, the higher the number of cut-points, the higher the probability of selecting the considered feature; and vice versa. This fact leads to the deletion of possibly important features having a single or a low number of cut-points, such as the infection rate, the glycemia level, and the blood pressure. In order to solve the issue of the dependency relation between the feature selection (or removal) event and the number of its generated potential cut-points, we propose to model the DBFS task as a bilevel optimization problem and then solve it using an improved version of an existing co-evolutionary algorithm, named I-CEMBA. The latter ensures the variation of the number of features during the migration process in order to deal with the multimodality aspect. The resulting algorithm, termed bilevel discretization-based feature selection (Bi-DFS), performs selection at the upper level while discretization is done at the lower level. The experimental results on several high-dimensional datasets show that Bi-DFS outperforms relevant state-of-the-art methods in terms of classification accuracy, generalization ability, and feature selection bias.
Rihab Said, Maha Elarbi, Slim Bechikh, Carlos A. Coello Coello, Lamjed Ben Said
IEEE Trans. Evol. Comput.2
2022 Interval-based Cost-sensitive Classification Tree Induction as a Bi-level Optimization Problem
abstract
Cost-sensitive learning is one of the most adopted approaches to deal with data imbalance in classification. Unfortunately, the manual definition of misclassification costs is still a very complicated task, especially with the lack of domain knowledge. To deal with the issue of costs' uncertainty, some researchers proposed the use of intervals instead of scalar values. This way, each cost would be delimited by two bounds. Nevertheless, the definition of these bounds remains as a very complicated and challenging task. Recently, some researches proposed the use of genetic programming to simultaneously build classification trees and search for optimal costs' bounds. As for any classification tree there is a whole search space of costs' bounds, we propose in this paper a bi-level evolutionary approach for interval-based cost-sensitive classification tree induction where the trees are constructed at the upper level while misclassification costs intervals bounds are optimized at the lower level. This ensures not only a precise evaluation of each tree but also an effective approximation of optimal costs intervals bounds. The performance and merits of our proposal are shown through a detailed comparative experimental study on commonly used imbalanced benchmark data sets with respect to several existing works.
Rihab Said, Maha Elarbi, Slim Bechikh, Carlos A. Coello Coello, Lamjed Ben Said
CEC2
2022 Handling uncertainty in SBSE: a possibilistic evolutionary approach for code smells detection
Sofien Boutaib, Maha Elarbi, Slim Bechikh, Fabio Palomba, Lamjed Ben Said
Empir. Softw. Eng.2
2021 Software Anti-patterns Detection Under Uncertainty Using a Possibilistic Evolutionary Approach
Sofien Boutaib, Maha Elarbi, Slim Bechikh, Chih-Cheng Hung, Lamjed Ben Said
EuroGP2
2021 Dealing with Label Uncertainty in Web Service Anti-patterns Detection using a Possibilistic Evolutionary Approach
abstract
Like the case of any software, Web Services (WSs) developers could introduce anti-patterns due to the lack of experience and badly-planned changes. During the last decade, search-based approaches have shown their outperformance over other approaches mainly thanks to their global search ability. Unfortunately, these approaches do not consider the uncertainty of class labels. In fact, two experts could be uncertain about the smelliness of a particular WS interface but also about the smell type. Currently, existing works reject uncertain data that correspond to WSs interfaces with doubtful labels. Motivated by this observation and the good performance of the possibilistic K-NN classifier in handling uncertain data, we propose a new evolutionary detection approach, named Web Services Anti-patterns Detection and Identification using Possibilistic Optimized K-NNs (WS-ADIPOK), which can cope with the uncertainty based on the Possibility Theory. The obtained experimental results reveal the merits of our proposal regarding four relevant state-of-the-art approaches.
Sofien Boutaib, Maha Elarbi, Slim Bechikh, Mohamed Makhlouf, Lamjed Ben Said
ICWS2
2021 A Possibilistic Evolutionary Approach to Handle the Uncertainty of Software Metrics Thresholds in Code Smells Detection
abstract
A code smells detection rule is a combination of metrics with their corresponding crisp thresholds and labels. The goal of this paper is to deal with metrics' thresholds uncertainty; as usually such thresholds could not be exactly determined to judge the smelliness of a particular software class. To deal with this issue, we first propose to encode each metric value into a binary possibility distribution with respect to a threshold computed from a discretization technique; using the Possibilistic C-means classifier. Then, we propose ADIPOK-UMT as an evolutionary algorithm that evolves a population of PK-NN classifiers for the detection of smells under thresholds' uncertainty. The experimental results reveal that the possibility distribution-based encoding allows the implicit weighting of software metrics (features) with respect to their computed discretization thresholds. Moreover, ADIPOK-UMT is shown to outperform four relevant state-of-art approaches on a set of commonly adopted benchmark software systems.
Sofien Boutaib, Maha Elarbi, Slim Bechikh, Fabio Palomba, Lamjed Ben Said
QRS2
2021 Code smell detection and identification in imbalanced environments
Sofien Boutaib, Slim Bechikh, Fabio Palomba, Maha Elarbi, Mohamed Makhlouf, Lamjed Ben Said
Expert Syst. Appl.4
2021 On the importance of isolated infeasible solutions in the many-objective constrained NSGA-III
Maha Elarbi, Slim Bechikh, Lamjed Ben Said
Knowl. Based Syst.1
2020 Approximating Complex Pareto Fronts With Predefined Normal-Boundary Intersection Directions
abstract
Decomposition-based evolutionary algorithms using predefined reference points have shown good performance in many-objective optimization. Unfortunately, almost all experimental studies have focused on problems having regular Pareto fronts (PFs). Recently, it has been shown that the performance of such algorithms is deteriorated when facing irregular PFs, such as degenerate, discontinuous, inverted, strongly convex, and/or strongly concave fronts. The main issue is that the predefined reference points may not all intersect with the PF. Therefore, many researchers have proposed to update the reference points with the aim of adapting them to the discovered Pareto shape. Unfortunately, the adaptive update does not really solve the issue for two main reasons. On the one hand, there is a considerable difficulty to set the time and the frequency of updates. On the other hand, it is not easy to define how to update the search directions for an unknown PF shape. This article proposes to approximate irregular PFs using a set of predefined normal-boundary intersection (NBI) directions. The main motivation behind this article is that when using a set of well-distributed NBI directions, all these directions intersect with the PF regardless of its shape, except for the case of discontinuous and/or degenerate fronts. To handle the latter cases, a simple interaction mechanism between the decision maker (DM) and the algorithm is used. In fact, the DM is asked if the number of NBI directions needs to be increased in some stages of the evolutionary process. If so, the resolution of the NBI directions that intersect the PF is increased to properly cover discontinuous and/or degenerate PFs. Our experimental results on benchmark problems with regular and irregular PFs, having up to fifteen objectives, show the merits of our algorithm when compared to eight of the most representative state-of-the-art algorithms.
Maha Elarbi, Slim Bechikh, Carlos A. Coello Coello, Mohamed Makhlouf, Lamjed Ben Said
IEEE Trans. Evol. Comput.1
2019 A Hybrid Evolutionary Algorithm with Heuristic Mutation for Multi-objective Bi-clustering
abstract
Bi-clustering is one of the main tasks in data mining with several application domains. It consists in partitioning a data set based on both rows and columns simultaneously. One of the main difficulties in bi-clustering is the issue of finding the number of bi-clusters, which is usually a user-specified parameter. Recently, in 2017, a new multi-objective evolutionary clustering algorithm, called MOCK-II, has shown its effectiveness in data clustering while automatically determining the number of clusters. Motivated by the promising results of MOCK-II, we propose in this paper a hybrid extension of this algorithm for the case of bi-clustering. Our new algorithm, called MOBICK, uses an efficient solution encoding, an effective crossover operator, and a heuristic mutation strategy. Similarly to MOCK-II, MOBICK is able to find automatically the number of bi-clusters. The outperformance of our algorithm is shown on a set of real gene expression data sets against several existing state-of-the-art works. Moreover, to be able to compare MOBICK to MOCK-I and MOCK-II, we have designed two basic extensions of MOCK-I and MOCK-II for the case of bi-clustering that we named B-MOCK-I and B-MOCK-II. Again, the experimental results confirm the merits of our proposal.
Slim Bechikh, Maha Elarbi, Chih-Cheng Hung, Sabrine Hamdi, Lamjed Ben Said
CEC2
2018 A New Decomposition-Based NSGA-II for Many-Objective Optimization
abstract
Multiobjective evolutionary algorithms (MOEAs) have proven their effectiveness and efficiency in solving problems with two or three objectives. However, recent studies show that MOEAs face many difficulties when tackling problems involving a larger number of objectives as their behavior becomes similar to a random walk in the search space since most individuals are nondominated with respect to each other. Motivated by the interesting results of decomposition-based approaches and preference-based ones, we propose in this paper a new decomposition-based dominance relation to deal with many-objective optimization problems and a new diversity factor based on the penalty-based boundary intersection method. Our reference point-based dominance (RP-dominance), has the ability to create a strict partial order on the set of nondominated solutions using a set of well-distributed reference points. The RP-dominance is subsequently used to substitute the Pareto dominance in nondominated sorting genetic algorithm-II (NSGA-II). The augmented MOEA, labeled as RP-dominance-based NSGA-II, has been statistically demonstrated to provide competitive and oftentimes better results when compared against four recently proposed decomposition-based MOEAs on commonly-used benchmark problems involving up to 20 objectives. In addition, the efficacy of the algorithm on a realistic water management problem is showcased.
Maha Elarbi, Slim Bechikh, Abhishek Gupta 0001, Lamjed Ben Said, Yew-Soon Ong
IEEE Trans. Syst. Man Cybern. Syst.1
2017 On the importance of isolated solutions in constrained decomposition-based many-objective optimization
abstract
During the few past years, decomposition has shown a high performance in solving Multi-objective Optimization Problems (MOPs) involving more than three objectives, called as Many-objective Optimization Problems (MaOPs). The performance of most of the existing decomposition-based algorithms has been assessed on the widely used DTLZ and WFG unconstrained test problems. However, the number of works that have been devoted to tackle the problematic of constrained many-objective optimization is relatively very small when compared to the number of works handling the unconstrained case. Recently there has been some interest to exploit infeasible isolated solutions when solving Constrained MaOPs (CMaOPs). Motivated by this observation, we firstly propose an IS-update procedure (Isolated Solution-based update procedure) that has the ability to: (1) handle CMaOPs characterized by various types of difficulties and (2) favor the selection of not only infeasible solutions associated to isolated sub-regions but also infeasible solutions with smaller Constraint Violation (CV) values. The IS-update procedure is subsequently embedded within the Multi-Objective Evolutionary Algorithm-based on Decomposition (MOEA/D). The new obtained algorithm, named ISC-MOEA/D (Isolated Solution-based Constrained MOEA/D), has been shown to provide competitive and better results when compared against three recent works on the CDTLZ benchmark problems.
Maha Elarbi, Slim Bechikh, Lamjed Ben Said
GECCO1