EDBT 2026 Demo / reviewers in the wild / expert
Sidi Ahmed Mahmoudi
dblp:60/10646
· DBLP profile ↗
12ranked-venue papers
2as first author
6since 2021 · last 2026
0000-0002-1530-9524ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Real-time AI-powered monitoring for energy-efficient scheduling in multi-node heterogeneous systemsabstractLoad balancing is critical for maintaining computing systems’ stability and achieving optimal performance. Its significance is widely recognized across different computing fields, particularly in the context of heterogeneous systems. These systems comprise computing devices with varying computational capabilities and architectures, each optimized for specific workloads. This heterogeneity introduces dynamic resource constraints, architectural mismatches, and unpredictable task-device affinity, which aggravates the challenges of load balancing. This paper presents an AI-driven load balancing solution for real-time distributed heterogeneous systems. Our approach continuously monitors the system state, capturing key factors that influence performance, such as task and device characteristics. Leveraging AI-based models, it computes a dynamic load index for each device based on the collected data. Using these load estimations, the method predicts potential imbalances through a novel imbalance metric and proactively schedules incoming applications to the most suitable devices, ensuring system-wide balance. To validate our approach, we first evaluated the prediction models by comparing a variety of machine learning algorithms with device-specific deep learning models, with the latter achieving superior accuracy. We then compared our method against widely used scheduling techniques across diverse workloads. The results show that our approach achieves more balanced workload distribution, faster execution, higher throughput, improved resource utilization, and reduced energy consumption across all scenarios, showcasing its adaptability to dynamic conditions and its applicability in real-world settings. Taha Abdelazziz Rahmani, Ghalem Belalem, Sidi Ahmed Mahmoudi, Omar Rafik Merad Boudia |
Future Gener. Comput. Syst. | 3 |
| 2024 | Equalizer: Energy-efficient machine learning-based heterogeneous cluster load balancerabstractSummary Heterogeneous systems deliver high computing performance when effectively utilized. It is crucial to execute each application on the most suitable device while maintaining system balance. However, achieving equal distribution of the computing load is challenging due to variations in computing power and device architectures within the system. Moreover, scheduling applications at real‐time further complicates this task, as prior information about the submitted applications is absent. In this context, we introduce “Equalizer,” a real‐time load balancer for heterogeneous systems. “Equalizer” leverages machine learning to continuously monitor the system's state, predicting optimal devices for application execution at runtime. It assigns applications to devices that minimize system imbalance. To quantify system imbalance, we propose a novel metric that reflects the disparity in computing loads across the system's devices. This metric is calculated using predicted execution times of applications. To validate the performance of “Equalizer,” we conducted a comparative study against widely adopted approaches, namely Round Robin and Device Suitability. The experiments were performed on a heterogeneous cluster comprising a master host and three slave servers, equipped with a total of 4 central processing units (CPUs) and 4 graphics processing units (GPUs). All approaches were deployed on the cluster and evaluated using three distinct workloads categorized by their computing intensity: medium intensity, heavy intensity, and a combination of heavy and medium intensity, simulating real‐world scenarios. Each workload consisted of a set of 80 OpenCL applications with varying input data sizes. The experimental results demonstrate that “Equalizer” effectively minimized the system's imbalance, reduced the idle time of devices, and eliminated overloads. Moreover, “Equalizer” exhibited significant improvements in workload execution time, resource utilization, throughput, and energy consumption. Across all tested scenarios, “Equalizer” consistently outperformed alternative approaches, showcasing its robustness, adaptability to dynamic environments, and applicability in real‐world practice. Taha Abdelazziz Rahmani, Ghalem Belalem, Sidi Ahmed Mahmoudi, Omar Rafik Merad Boudia |
Concurr. Comput. Pract. Exp. | 3 |
| 2023 | Multimodal Approach for Harmonized System Code PredictionabstractThe rapid growth of e-commerce has placed considerable pressure on customs representatives, prompting advanced methods.In tackling this, Artificial intelligence (AI) systems have emerged as a promising approach to minimize the risks faced.Given that the Harmonized System (HS) code is a crucial element for an accurate customs declaration, we propose a novel multimodal HS code prediction approach using deep learning models exploiting both image and text features obtained through the customs declaration combined with e-commerce platform information.We evaluated two early fusion methods and introduced our MultConcat fusion method.To the best of our knowledge, few studies analyze the featurelevel combination of text and image in the state-of-the-art for HS code prediction, which heightens interest in our paper and its findings.The experimental results prove the effectiveness of our approach and fusion method with a top-3 and top-5 accuracy of 93.5% and 98.2% respectively.* These authors contributed equally to this work † The authors thank the support of Infortech institute and the E-origin project funded by the Walloon Region within the pole of logistics in Wallonia.1 Example of the Belgian governmental free access database of HS code nomenclature: https://eservices.minfin.fgov.be/extTariffBrowser/browseNomen.xhtml?suffix=80&lang=EN Otmane Amel, Sédrick Stassin, Sidi Ahmed Mahmoudi, Xavier Siebert |
ESANN | 3 |
| 2023 | Similarity versus Supervision: Best Approaches for HS Code PredictionabstractWith growing e-commerce flows and new legislative rules, customs representatives confront serious liabilities when completing customs declarations for their clients.In the latter, the Harmonized System (HS) code is a crucial component using 10 digits (HS10) to classify products and define national tax rates.In this paper, we first compare the performance of sentence embedding models using semantic similarity, and second, we assess the effectiveness of supervised models, both aimed at predicting up to the HS10 code.To the best of our knowledge, there is currently little research being conducted on this topic.We demonstrate the differences and respective strengths of each approach.Our results show the outstanding performance of the semantic similarity approach with a top-3 and top-5 accuracy of 89% and 94.8% respectively for HS10 prediction.* These authors contributed equally to this work.† The authors thank the support of the Infortech institute and the E-origin Sédrick Stassin, Otmane Amel, Sidi Ahmed Mahmoudi, Xavier Siebert |
ESANN | 3 |
| 2023 | Feedback Driven Multi Stereo Vision System for Real-Time Event Analysisabstract2D cameras are often used in interactive systems. Other systems like gaming consoles provide more powerful 3D cameras for short range depth sensing. Overall, these cameras are not reliable in large, complex environments. In this work, we propose a 3D stereo vision based pipeline for interactive systems, that is able to handle both ordinary and sensitive applications, through robust scene understanding. We explore the fusion of multiple 3D cameras to do full scene reconstruction, which allows for preforming a wide range of tasks, like event recognition, subject tracking, and notification. Using possible feedback approaches, the system can receive data from the subjects present in the environment, to learn to make better decisions, or to adapt to completely new environments. Throughout the paper, we introduce the pipeline and explain our preliminary experimentation and results. Finally, we draw the roadmap for the next steps that need to be taken, in order to get this pipeline into production. Mohamed Benkedadra, Matei Mancas, Sidi Ahmed Mahmoudi |
IMX | 3 |
| 2023 | Single node deep learning frameworks: Comparative study and CPU/GPU performance analysisabstractAbstract Deep learning presents an efficient set of methods that allow learning from massive volumes of data using complex deep neural networks. To facilitate the design and implementation of algorithms, deep learning frameworks provide a high‐level programming interface. Based on these frameworks, new models, and applications are able to make better and better predictions. One type of deep learning application is the Internet of Things that can gather a continuous flow of data, which causes an explosion of the amount of data. Therefore, to handle this data management issue, computation technologies can offer new perspectives to analyze more data with more complex models. In this context, a cluster of computers can operate to quickly deliver a model or to enable the design of a complex neural network spread among computers. An alternative is to distribute a deep learning task with HPC cloud computing resources and to scale cluster in order to quickly and efficiently train a neural network. As a first step to design an infrastructure aware framework which is able to scale the computing nodes, this work aims to review and analyze the state‐of‐the‐art frameworks by collecting device utilization data during the training task. We gather information about the CPU, RAM and the GPU utilization on deep learning algorithms with and without multi‐threading. The behavior of each framework is discussed and analyzed in order to shed light on the strengths and weaknesses of the different deep learning frameworks. Jean-Sébastien Lerat, Sidi Ahmed Mahmoudi, Saïd Mahmoudi |
Concurr. Comput. Pract. Exp. | 2 |
| 2020 | WALLeSMART: Cloud Platform for Smart FarmingabstractToday, agricultural practices are supported by bio-informatics and emerging technologies such as remote sensing, cloud computing and the Internet of Things (IoT), which leads to the concept of “Smart Farming”. Smart farming is a cycle of intelligent detection and monitoring, analysis and planning, as well as control of agricultural operations using a cloud-based event management system. In this paper, we propose WALLeSMART, a cloud-based framework built to capitalize the efforts invested in building smart farming management systems, applied to the Wallonia region of Belgium. The framework proposes an architecture to address the challenges of acquisition, processing, and visualization of massive amounts of data, in both batch and real-time basis. An initial prototype has been developed and tested with various farms and shows prominent results. Amine Roukh, Fabrice Nolack Fote, Sidi Ahmed Mahmoudi, Saïd Mahmoudi |
SSDBM | 3 |
| 2020 | Cloud architecture for plant phenotyping researchabstractSummary Digital phenotyping is an emergent science mainly based on imagery techniques. The tremendous amount of data generated needs important cloud computing for their processing. The coupling of recent advance of distributed databases and cloud computing offers new possibilities of big data management and data sharing for the scientific research. In this paper, we present a solution combining a lambda architecture built around Apache Druid and a hosting platform leaning on Apache Mesos. Lambda architecture has already proved its performance and robustness. However, the capacity of ingesting and requesting of the database is essential and can constitute a bottleneck for the architecture, in particular, for in terms of availability and response time of data. We focused our experimentation on the response time of different databases to choose the most adapted for our phenotyping architecture. Apache Druid has shown its ability to respond to typical queries of phenotyping applications in times generally inferior to the second. Olivier Debauche, Sidi Ahmed Mahmoudi, Nicolas De Cock, Saïd Mahmoudi, Pierre Manneback, Frédéric Lebeau |
Concurr. Comput. Pract. Exp. | 2 |
| 2020 | Multimedia processing using deep learning technologies, high-performance computing cloud resources, and Big Data volumesabstractSummary The last few years have been marked by the presence of very large sets of images and videos in our everyday lives. These multimedia objects have a very fast frequency of creation and sharing since images and videos can come from different devices such as smartphones, satellites, cameras, or drones. They are generally used to illustrate objects in different situations (public areas, train stations, hospitals, political and sport events and competitions, etc). As consequence, image and video processing algorithms have got increasing importance for several computer vision applications that should be adapted for managing large‐scale volumes and exploiting high performance computing resources (local or cloud). In this work, we propose a cloud‐based toolbox (platform) for computer vision applications. This platform integrates a toolbox of image and video processing algorithms that can (i) exploit high performance computing cloud resources, (ii) execute applications in real time, and (iii) manage large‐scale database using Big Data technologies. The related libraries and hardware drivers are automatically integrated and configured in order to offer to users an access to the different applications without the need to download, install, and configure software or hardware. Experiments were conducted using three kinds of applications: (i) image and video processing applications, (ii) deep learning techniques for images classification and multiobject localization, and (iii) images indexation and retrieval. These experiments demonstrated the interest of our platform for sharing, in an efficient way, our scientific contributions and annotated databases in order to improve the quality and performance of computer vision applications. Sidi Ahmed Mahmoudi, Mohammed Amin Belarbi, Saïd Mahmoudi, Ghalem Belalem, Pierre Manneback |
Concurr. Comput. Pract. Exp. | 1 |
| 2018 | Towards a smart selection of resources in the cloud for low-energy multimedia processingabstractSummary Nowadays, image and video processing applications have become widely used in many domains related to computer vision. Indeed, they can come from cameras, smartphones, social networks, or from medical devices. Generally, these images and videos are used for illustrating people or objects (cars, trains, planes, etc) in many situations such as airports, train stations, public areas, sport events, and hospitals. Thus, image and video processing algorithms have got increasing importance, they are required from various computer vision applications such as motion tracking, real time event detection, database (images and videos) indexation, and medical computer‐aided diagnosis methods. The main inconvenient of image and video processing applications is the high intensity of computation and the complex configuration and installation of the related materials and libraries. In this paper, we propose a new framework that allows users to select in a smart and efficient way the computing units (CPU or/and GPU) in a cloud‐based platform, in case of processing one image (or one video in real time) or many images (or videos). This framework enables to affect the local or remote computing units for calculation after analyzing the type of media and the algorithm complexity. The framework disposes of a set of selected CPU and GPU‐based computer vision methods, such as image denoising, histogram computation, features descriptors (SIFT, SURF), points of interest extraction, edges detection, silhouette extraction, and sparse and dense optical flow estimation. These primitive functions are exploited in various applications such as medical image segmentation, videos indexation, real time motion analysis, and left ventricle segmentation and tracking from 2D echocardiography. Experimental results showed a global speedup ranging from 5× to 273×(compared to CPU versions) as result of the application of our framework for the above‐mentioned methods. In addition to these performances, the parallel and heterogeneous implementations offered lower power consumption as result of the fast treatment. Sidi Ahmed Mahmoudi, Mohammed Amin Belarbi, Saïd Mahmoudi, Ghalem Belalem |
Concurr. Comput. Pract. Exp. | 1 |
| 2014 | A Multi-Resolution FPGA-Based Architecture for Real-Time Edge and Corner DetectionabstractThis work presents a new flexible parameterizable architecture for image and video processing with reduced latency and memory requirements, supporting a variable input resolution. The proposed architecture is optimized for feature detection, more specifically, the Canny edge detector and the Harris corner detector. The architecture contains neighborhood extractors and threshold operators that can be parameterized at runtime. Also, algorithm simplifications are employed to reduce mathematical complexity, memory requirements, and latency without losing reliability. Furthermore, we present the proposed architecture implementation on an FPGA-based platform and its analogous optimized implementation on a GPU-based architecture for comparison. A performance analysis of the FPGA and the GPU implementations, and an extra CPU reference implementation, shows the competitive throughput of the proposed architecture even at a much lower clock frequency than those of the GPU and the CPU. Also, the results show a clear advantage of the proposed architecture in terms of power consumption and maintain a reliable performance with noisy images, low latency and memory requirements. Paulo Ricardo Possa, Sidi Ahmed Mahmoudi, Naim Harb, Carlos Valderrama 0001, Pierre Manneback |
IEEE Trans. Computers | 2 |
| 2012 | A new self-adapting architecture for feature detectionabstractIn this paper, we present a FPGA based flexible self-adapting architecture for two features detectors, the Canny edge detector and the Harris corner detector, with reduced latency and memory requirements, and supporting variable resolution images. The new architecture uses neighbourhood extractors that can self-adapt its parameters on-the-fly and algorithm simplifications to reduce mathematical complexity, memory requirements and latency without losing reliability. Paulo Da Cunha Possa, Sidi Ahmed Mahmoudi, Naim Harb, Carlos Valderrama 0001 |
FPL | 2 |