Mouna Baklouti

dblp:71/7914 · DBLP profile ↗
← Back
26ranked-venue papers
3as first author
10since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 5 since 2021Systems, architecture and hardware · 9 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Predicting Digital Literacy Gains in Rural Contexts Using Multidimensional Data: A Machine Learning Approach
abstract
In low-resource educational contexts, the early identification of at-risk learners remains critical to preventing dropout and ensuring equitable skill acquisition. This study develops a predictive modeling framework to classify learner progression in digital literacy training programs targeting rural and semi-rural populations, across three domains: computer use, Internet navigation, and mobile literacy. Using a dataset of 1000 learners, six supervised classification algorithms were implemented via cross-validation, with Support Vector Machines (SVM) achieving the highest accuracy ($97 \%$) with strong recall for both no-risk ($97 \%$) and at-risk learners ($96 \%$). Feature analysis revealed that initial skills are the strongest predictor of progression, while contextual factors such as geographical disparities and household income level also significantly influence learning trajectories. These models enable the early identification of at-risk learners and the establishment of personalized support systems to reduce dropout and advance digital equity in marginalized communities.
Sameher Ajili, Rym Chéour, Mariem Abid, Mouna Baklouti, Richard Hotte
AICCSA4
2025 Hierarchical Clustering Strategy for Enhanced HPC System Configuration
abstract
This paper introduces the AI-Enhanced Cluster Configurator, a tool that uses artificial intelligence to streamline the configuration of high-performance computing (HPC) clusters. This tool offers a sophisticated recommendation system that assists users in selecting optimal server and storage solutions. It provides two operational modes: an Assisted Mode, which simplifies the configuration process for novice users through AI-generated recommendations, and an Expert Mode, which offers in-depth customization options for experienced users. This study delves into the AI methodologies employed, underscoring how they substantially enhance and simplify the design of HPC clusters for a broader range of users, thereby supporting the solution architecture of essential technology for cutting-edge research and development across various domains.
Nessrine Aloulou, Melek Kchaou, Baya Mezghanni, Mouna Baklouti, Saber Feki
HPCC4
2025 RaceHPC: A Cloud-Agnostic Automated Platform for HPC Hackathons
abstract
Managing High-Performance Computing (HPC) workloads in the cloud requires balancing performance, cost efficiency, and automation. Traditional HPC deployment on cloud is manual, complex, and error-prone leading to delays, misconfigurations. These workloads demand powerful and expensive compute, network, and storage resources that remain billable even during setup, further increasing financial overhead. The need for automated, scalable, and cost-optimized HPC environments has never been greater. To address these challenges, we introduce RaceHPC, a cloud-agnostic, fully automated platform designed to host HPC competitions, training programs, and benchmarking initiatives on Cloud with minimal manual effort. By leveraging DevOps, Continuous Integration / Continuous Delivery (CI/CD), and Kubernetes orchestration, it eliminates manual intervention by automatically provisioning all required resources and configurations using Infrastructure as Code (IaC) and containerized environments integrating idempotent and self-scaling deployment mechanisms and ensuring that each competitor receives a secure, isolated, and fully preloaded HPC workspace with the necessary dependencies and tools accessible via a web browser. Beyond automation, RaceHPC prioritizes cost efficiency by incorporating FinOps strategies, automatically scaling resources based on demand and decommissioning infrastructure at competition deadlines to eliminate unnecessary expenses.
Imen Châari, Mouna Baklouti, Saber Feki
HPCC2
2024 Optimization of Weighted Random Walk Tip Selection Algorithm on the IOTA Tangle
Imen Ahmed, Mariem Turki, Mouna Baklouti
AICCSA3
2024 Vision Transformers with Efficient LoRA Finetuning for Breast Cancer Classification
abstract
The diagnosis of breast cancer remains a very interesting area of research because of its high mortality rate among women worldwide. During the last decade, several studies have focused on the development of early detection systems for breast cancer based on new advances in deep learning models. In this context, we presented a new system of breast cancer detection using Vision Transformers (ViT) architecture. Due to the large number of trainable parameters in ViT, we have incorporated the Low-Rank Adaptation (LoRA) mechanism as an efficient fine-tuning method to significantly reduce the training calculation costs while maintaining the model performance. This reduction in the number of trainable parameters results in a significant reduction in memory consumption and training time. The experiments were conducted on two different histopathological breast images databases, and demonstrate the effectiveness of the proposed methodology in terms of accuracy, memory consumption and training process.
Tawfik Ezat Mousa, Ramzi Zouari, Mouna Baklouti
AICCSA3
2023 Breast Cancer Detection Based DenseNet with Attention Model in Mammogram Images
Tawfik Ezat Mousa, Ramzi Zouari, Mouna Baklouti
MEDI3
2022 Design of Multiprocessor Architecture for Watermarking and Tracing Images Using QR Code
Jalel Baaouni, Hedi Choura, Faten Chaabane, Tarek Frikha, Mouna Baklouti
KES-IDT5
2022 Blockchain for IoT-Based Healthcare using secure and privacy-preserving watermark
abstract
Patient medical data is the key data of an e-health application. Medical data frequently play an important role in disease analysis and are implicated in motivating values for educating diagnostic methods and modifying diagnostic results. Confidentiality of patient health documents is more important, and therefore researchers have considered many security measures, including access control, confidentiality, and integrity of these documents. To solve the security problem, this paper presents an advanced solution based on converting medical data into compact size, in the form of QR code before being deployed on our Blockchain. In this way, performing complex operations will be safely performed off-chain, such as collecting and preprocessing patient data, while submission of this data is done on an hourly basis. Chain. We recommend a data flow architecture that combines the Internet of Things (IoT) with blockchain (especially the Ethereum architecture), using smart contracts to access and store data while keeping Strong against many attacks. Our proposed system is effective in sharing and managing patient electronic health data. The tests involved implementing Ethereum chain configurations on PCs and two embedded IoT platforms, namely Raspberry Pi and PYNQ cards. Performance is evaluated in terms of security, execution time, memory, power and CPU consumption of various patient data scenarios.
Hedi Choura, Faten Chaabane, Mouna Baklouti, Tarek Frikha
SIN3
2021 Defensive approximation: securing CNNs using approximate computing
abstract
In the past few years, an increasing number of machine-learning and deep learning structures, such as Convolutional Neural Networks (CNNs), have been applied to solving a wide range of real-life problems. However, these architectures are vulnerable to adversarial attacks: inputs crafted carefully to force the system output to a wrong label. Since machine-learning is being deployed in safety-critical and security-sensitive domains, such attacks may have catastrophic security and safety consequences. In this paper, we propose for the first time to use hardware-supported approximate computing to improve the robustness of machine learning classifiers. We show that our approximate computing implementation achieves robustness across a wide range of attack scenarios. Specifically, we show that successful adversarial attacks against the exact classifier have poor transferability to the approximate implementation. The transferability is even poorer for the black-box attack scenarios, where adversarial attacks are generated using a proxy model. Surprisingly, the robustness advantages also apply to white-box attacks where the attacker has unrestricted access to the approximate classifier implementation: in this case, we show that substantially higher levels of adversarial noise are needed to produce adversarial examples. Furthermore, our approximate computing model maintains the same level in terms of classification accuracy, does not require retraining, and reduces resource utilization and energy consumption of the CNN. We conducted extensive experiments on a set of strong adversarial attacks; We empirically show that the proposed implementation increases the robustness of a LeNet-5 and an Alexnet CNNs by up to 99% and 87%, respectively for strong transferability-based attacks along with up to 50% saving in energy consumption due to the simpler nature of the approximate logic. We also show that a white-box attack requires a remarkably higher noise budget to fool the approximate classifier, causing an average of 4 dB degradation of the PSNR of the input image relative to the images that succeed in fooling the exact classifier.
Amira Guesmi, Ihsen Alouani, Khaled N. Khasawneh, Mouna Baklouti, Tarek Frikha, Mohamed Abid, Nael B. Abu-Ghazaleh
ASPLOS4
2021 Semantic Representation Driven by a Musculoskeletal Ontology for Bone Tumors Diagnosis
Mayssa Bensalah, Atef Boujelben, Yosr Hentati, Mouna Baklouti, Mohamed Abid
ISDA4
2020 A Novel On-Wrist Fall Detection System Using Supervised Dictionary Learning Technique
abstract
Wrist-based fall detection system provides a very comfortable and multi-modal healthcare solution, especially for elderly risking falls. However, the wrist location presents a very challenging and unstable spot to distinguish falls among other daily activities. In this paper, we propose a Supervised Dictionary Learning approach for wrist-based fall detection. Three Dictionary learning algorithms for classification are invoked in this study, namely SRC, FDDL, and LRSDL. To extract the best descriptive representation of the signal data we followed different preprocessing scenarios based on accelerometer, gyroscope, and magnetometer. A considerable overall performance was obtained by the SRC algorithms reaching respectively 99.8%, 100%, and 96.6% of accuracy, sensitivity, and specificity using raw data provided by a triaxial accelerometer, accordingly overthrowing previously proposed methods for wrist placement.
Farah Othmen, Mouna Baklouti, André Eugênio Lazzaretti, Marwa Jmal, Mohamed Abid
ICOST2
2019 HEAP: A Heterogeneous Approximate Floating-Point Multiplier for Error Tolerant Applications
abstract
Floating point arithmetic is one of the most commonly used units in nowadays computing systems and is deployed for a wide range of domains and applications. While floating point operators offer high precision calculations, a plethora of applications such as multimedia processing and machine learning tolerate errors and computation imprecision. In a context of limited power budget embedded systems, saving resources and energy with an acceptable precision loss is a challenging design task. Approximate computing is an emerging systems design paradigm that offers promising balance between accuracy on the one hand and power consumption and resource utilization on the other hand. While state of the art approximate techniques offer a wide design space at the operator level, few are the works that consider exploring different techniques to build a heterogeneous comprehensive approximate design. In this paper, we propose HEAP: a heterogeneous approximate floating point multiplier. Based on a design space exploration process, we present an approximation at the transistor level that reduces energy consumption of up to 68%. Experimental study on a set of machine learning applications shows promising results with comparable accuracy to exact multiplier based systems.
Amira Guesmi, Ihsen Alouani, Mouna Baklouti, Tarek Frikha, Mohamed Abid, Atika Rivenq
RSP3
2017 Performance Exploration of AMBA AXI4 Bus Protocols for Wireless Sensor Networks
abstract
Modern System-on-Chip (SoC) designs are faced with many challenges among which efficient communication managing is one of the most important. On-chip communication architectures can have a strong impact on the performance of SoC designs. The traditional SoC interconnects, cannot keep up with the high demands of today's SoC. To address this problem, SoC makers propose new protocols to implement high performance data transfer. AMBA AXI4, is one of the widely used protocols as on-chip bus in recent Wireless Sensor Network SoCs. It includes three distinct interconnect protocols: stream, burst and lite. In this paper, we highlight the various parameters that must be taken into consideration to select the adequate interface for a given application. We analyze and compare the on-chip interfaces for hardware/software (HW/SW) communication when various code transformations are applied on the Zynq-7000 Xilinx platform. Many experiments have been conducted to evaluate the communication between the main processor and the reconfigurable hardware accelerators.
Mariem Makni, Mouna Baklouti, Smaïl Niar, Mohamed Abid
AICCSA2
2017 A Rapid Data Communication Exploration Tool for Hybrid CPU-FPGA Architectures
abstract
Modern System-on-Chip (SoC) designs face many challenges. Choosing the best communication protocol among the different processing nodes is one of the most important design decisions. On-chip communication architectures can have a significant impact on the performance of SoC designs. However, in most of the existing design tools, only the computation cost is accurately estimated. To address this challenge, we present a high-level analytical tool to estimate the data communication cost for hybrid CPU-FPGA architectures. The proposed model allows to estimate, rapidly and accurately, both computation and communication cost of applications containing multiple nested loops. This paper also explores the benefits of applying various optimization pragmas including dataflow and loop pipelining, at the compilation phase. Experimental results show that the proposed model provides accurate data communication estimation for hybrid CPUFPGA architectures.
Mariem Makni, Smaïl Niar, Mouna Baklouti, Guanwen Zhong, Tulika Mitra, Mohamed Abid
PDP3
2016 Design and implementation of a fall detection system on a Zynq board
abstract
Population aging has become a worldwide problem. Fall accidents are considered as one of the major health risks, especially for elderly people. Fall detection devices are the key to distinguish a fall from daily activities, automatically alert when a fall occurred and significantly decrease the time of rescue when the monitored patient falls down. The proposed prototype is composed of a tri-axial accelerometer communicating to a Zybo board through Inter-Integrated Circuit (I2C) interface. A threshold-based algorithm has been implemented based on peaks acceleration detection and inactivity posture recognition after falling and executed as a standalone application on a Zynq Z-7010 Field Programmable Gate Array (FPGA). This first implementation using Vivado and Xilinx SDK showed prominent results in terms of power consumption and time of execution.
Sahar Abdelhedi, Mouna Baklouti, Riad Bourguiba, Jaouhar Mouine
AICCSA2
2016 On Exploiting Energy-Aware Scheduling Algorithms for MDE-Based Design Space Exploration of MP2SoC
abstract
Massively Parallel Multi-Processors System-on-Chip (MP2SoC) architectures have been widely deployed to run challenging high-performance computations. However, the ever greater demand for energy efficiency fosters energy budgeting in MP2SoC systems. Nowadays, having the appropriate Electronic Design Automation (EDA) tools for power estimation is mandatory. The major challenge for the design of such tools is to reach a better tradeoff between accuracy and time-to-market. This paper presents a Model Driven Engineering (MDE)-based energy-aware Design Space Exploration (DSE) approach allowing the designer to take the power consumption criterion into account early in the design flow. The originality of this approach is that it integrates the Energy-Aware Duplication (EAD) algorithm that strives to balance schedule lengths and energy savings by considering the most important sources of energy consumption in MP2SoC: the massive number of processing elements (PE) and the high-speed Network-on-Chip (NoC). To demonstrate the effectiveness of the proposed approach, we conducted experiments using the H.263 encoder application. The obtained results demonstrated that EAD can effectively save energy in MP2SoC systems. They also showed that our MDE approach is capable of accelerating the DSE process to make early energy-efficient design decisions.
Manel Ammar, Mouna Baklouti, Maxime Pelcat, Karol Desnos, Mohamed Abid
PDP2
2016 SCAC-Net: Reconfigurable Interconnection Network in SCAC Massively Parallel SoC
abstract
Parallel communication plays a critical role in massively parallel systems, especially in distributed memory systems executing parallel programs on shared data. Therefore, integrating an interconnection network in these systems becomes essential to ensure data inter-nodes exchange. Choose the most effective communication structure must meet certain criteria: speed, size and power consumption. Indeed, the communication phase should be as fast as possible to avoid compromising parallel computing, using small and low power consumption modules to facilitate the interconnection network extensibility in a scalable system. To meet these criteria and based on a module reuse methodology, we chose to integrate a reconfigurable SCAC-Net interconnection network to communicate data in SCAC Massively parallel SoC. This paper presents the detailed hardware implementation and discusses the performance evaluation of the proposed reconfigurable SCAC-Net network.
Hana Krichene, Mouna Baklouti, Philippe Marquet, Jean-Luc Dekeyser, Mohamed Abid
PDP2
2016 Development of a two-threshold-based fall detection algorithm for elderly health monitoring
abstract
Population aging has become a worldwide problem. Falls are considered as the first source of disabilities among elderly people. Fall detection algorithms are the key to distinguish a fall from daily activities, automatically alert when a fall occurred and significantly decrease the time of rescue when the monitored patient falls down. The algorithm presented in this paper uses tri-axial accelerometer outputs to discriminate between falls and daily activities. It is mainly based on a two-thresholds approach and inactivity posture recognition after falling. The algorithm showed prominent results compared to existing works and will be improved and implemented on a Zynq board for future applications.
Sahar Abdelhedi, Riad Bourguiba, Jaouhar Mouine, Mouna Baklouti
RCIS4
2016 Off-Line DVFS Integration in MDE-Based Design Space Exploration Framework for MP2SoC Systems
abstract
As the speed metric of Massively Parallel Multi-Processors System-on-Chip (MP2SoC) systems has increased over time, another metric has become more important: power consumption. Finding a tradeoff between power consumption and performance early in the design flow of MP2SoC systems in order to satisfy time-to-market is the design challenge of Electronic Design Automation (EDA) tools. This paper presents a Design Space Exploration (DSE) framework, named Energy-Aware Rapid Design of MP2SoC (EWARDS), aiming at exploring the performance and power capabilities of modern homogenous MP2SoC systems at design time using Model-Driven Engineering (MDE) techniques. The proposed framework extends the Modeling and Analysis of Real-Time and Embedded systems (MARTE)profile with power aspects of MP2SoC systems providing a high-level design entry. In addition, EWARDS integrates an energy-aware scheduler that strives to balance performance and energy savings by combining clustering scheduling algorithm with off-line Dynamic Voltage and Frequency Scaling (DVFS) power management techniques.
Manel Ammar, Mouna Baklouti, Maxime Pelcat, Karol Desnos, Mohamed Abid
WETICE2
2014 Design Space Exploration for Customized Asymmetric Heterogeneous MPSoC
abstract
Modern FPGA allows the design of very complex System-on-Chips (SoC). To fulfil modern application requirements, in terms of performance/energy consumption ratio, Heterogeneous Multiprocessor System-on-Chip (Ht- MPSoC) architectures represent a promising solution. In such systems, the processor instruction set is enhanced by application-specific custom instructions implemented on reconfigurable fabrics, namely FPGA. To increase area utilization and guarantee application constraint respect, we propose a new Ht-MPSoC architecture where hardware accelerators (HW accelerators) are shared among different processors in an intelligent manner. In this paper, we extend existing Ht-MPSoC architectures by considering asymmetric (AHt-MPSoC). In these architectures, cores have different resources that may share in different manners. Depending on the running applications and their needs in processing, private and shared HW accelerators are attached to the different cores. On a 8-core AHt-MPSoC we obtained a speed of 2.6 with a reduced number of HW accelerators for our benchmarks.
Bouthaina Damak, Rachid Benmansour, Mouna Baklouti, Smaïl Niar, Mohamed Abid
DSD3
2014 A mixed integer linear programming approach for design space exploration in FPGA-based MPSoC
abstract
Heterogeneous Multiprocessor System-on-Chip (Ht-MPSoC) architectures represent a promising approach as they allow a higher performance/energy consumption trade-off. In such systems, the processor instruction set is enhanced by application-specific custom instructions implemented on reconfigurable fabrics, namely FPGA. To increase area utilization and guarantee application constraint respect, we propose a new architecture where Ht-MPSoC hardware accelerators are shared among different processors in an intelligent manner. In this paper, a Mixed Integer Linear Programming (MILP) model is proposed to systematically explore the complex design space of the different configurations.
Bouthaina Damak, Rachid Benmansour, Smaïl Niar, Mouna Baklouti, Mohamed Abid
FPL4
2013 Master-Slave Control Structure for Massively Parallel System on Chip
abstract
The performance of massively parallel processing system depends mostly on the control configuration that is inherently part of the system. In particular, centralized control configuration is rigid and limits system scalability, and distributed control configuration is difficult to control in processing elements (PEs) interaction. Maintaining a flexible autonomous computation coupled with regular synchronous communication can assure a efficient parallel processing. The master-slave control structure is specified in such a way that previous features of the massively parallel System-on-Chip (mpSoC) are preserved and performance is improved. In this paper, we define the prototyping of a master-slave control structure for mpSoC in a FPGA-based platform. The structure implementation and related experiments using the vhdl language running on virtex6 ml605 of Xilinx board are described.
Hana Krichene, Mouna Baklouti, Mohamed Abid, Philippe Marquet, Jean-Luc Dekeyser
DSD2
2012 Extending MARTE to Support the Specification and the Generation of Data Intensive Applications for Massively Parallel SoC
abstract
The presented work proposes a MARTE (Modeling and Analysis of Real-Time and Embedded systems) extension for the specification of data-parallel applications designed to be executed on mppSoC, a massively parallel System-on-Chip. These applications can be clearly specified and generated using our transformation chain, which is automated and is a combination of contributions in different domains such as Model-Driven Engineering, MARTE modeling and automatic code generation. The modeling methodology as well as the generation process are validated by an image processing application example.
Manel Ammar, Mouna Baklouti, Mohamed Abid
DSD2
2010 IP Based Configurable SIMD Massively Parallel SoC
abstract
Significant advances in the field of configurable computing have enabled parallel processing within a single Field-Programmable Gate Array (FPGA) chip. This paper presents the implementation of a flexible and programmable Single Instruction Multiple Data (SIMD) processing system on FPGA that can be adapted to the application. Its implementation is based on an IP (Intellectual Property) assembling approach making its design fast and easy. A generation tool is also developed to generate the SIMD configuration depending on the application requirements. The proposed parallel processing system on chip is portable, scalable and flexible since it can be customized to match the needs of a data parallel application. Based on FPGA, different SIMD configurations have been evaluated in terms of performance and area trade-offs. The proposed parametric system shows good results executing some signal processing applications such as parallel matrices multiplication, FIR filter and RGB to YIQ image color conversion.
Mouna Baklouti, Mohamed Abid, Philippe Marquet, Jean-Luc Dekeyser
FPL1
2010 Scalable mpNoC for massively parallel systems - Design and implementation on FPGA
Mouna Baklouti, Yassine Aydi, Philippe Marquet, Jean-Luc Dekeyser, Mohamed Abid
J. Syst. Archit.1
2009 Study and integration of a parametric neighbouring interconnection network in a massively parallel architecture on FPGA
abstract
Single instruction multiple data processors are increasingly used in embedded systems for multimedia applications because of their area and energy-efficiency. Neighboring communications between the processing elements are a key issue in SIMD processors. They are present in most data parallel applications. However, the lack of flexibility in major parallel architectures is its main shortcoming. In order to improve the performances of a massively parallel architecture, especially in term of neighboring communication we need a flexible and parametric communication network. This paper focuses on the problems with the design of a parametric nearest neighborhood interconnection network in a SIMD architecture in system on chip (SoC). This network can be configured in multiple topologies making it flexible and parametric in order to suit different application needs. The proposed architecture is evaluated in terms of area (cost) and performance (execution time), which are deduced respectively from synthesis and simulation results. Experiments are performed on different architectures with various topologies. In order to evaluate the performance of the proposed architecture, a FIR application is finally implemented.
Mouna Baklouti, Mohamed Abid, Philippe Marquet, Jean-Luc Dekeyser
AICCSA1