EDBT 2026 Demo / reviewers in the wild / expert
Mahmoud Méribout
dblp:96/3652
· DBLP profile ↗
18ranked-venue papers
11as first author
5since 2021 · last 2026
0000-0002-0058-1090ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 5 first-author · 1 since 2021Computer networks · 5 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-authorArtificial intelligence and machine learning · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Electronic design automation · 44% Reconfigurable computing and FPGAs · 42% Embedded and real-time systems · 11% | |
| Computer graphics and multimedia
2 papers |
Image and video processing · 73% Computational photography and imaging · 19% Multimedia analysis and retrieval · 8% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Medical and health informatics · 100% |
Topics — the 13 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Reconfigurable computing and FPGAs
FPGA accelerator |
0.2 | 1 | 2016 | A New Parallel VLSI Architecture for Real-Time Electrical Capacitance Tomography · IEEE Trans. Computers 2016 |
Electronic design automation
hardware/software co-design |
0.2 | 1 | 2016 | A New Parallel VLSI Architecture for Real-Time Electrical Capacitance Tomography · IEEE Trans. Computers 2016 |
Medical and health informatics
medical imaging |
0.2 | 1 | 2015 | A Multimodal Image Reconstruction Method Using Ultrasonic Waves and Electrical Resistance Tomography · IEEE Trans. Image Process. 2015 |
Image and video processing
image reconstruction |
0.2 | 1 | 2015 | A Multimodal Image Reconstruction Method Using Ultrasonic Waves and Electrical Resistance Tomography · IEEE Trans. Image Process. 2015 |
Embedded and real-time systems
timing constraints |
0.1 | 1 | 2016 | A New Parallel VLSI Architecture for Real-Time Electrical Capacitance Tomography · IEEE Trans. Computers 2016 |
Computational photography and imaging › acoustic imaging
ultrasound imaging |
0.1 | 1 | 2015 | A Multimodal Image Reconstruction Method Using Ultrasonic Waves and Electrical Resistance Tomography · IEEE Trans. Image Process. 2015 |
Reconfigurable computing and FPGAs
dynamic reconfiguration |
0.0 | 1 | 2004 | A Combined Approach to High-Level Synthesis for Dynamically Reconfigurable Systems · IEEE Trans. Computers 2004 |
Electronic design automation
high-level synthesis |
0.0 | 1 | 2004 | A Combined Approach to High-Level Synthesis for Dynamically Reconfigurable Systems · IEEE Trans. Computers 2004 |
Image and video processing › feature detection
hough transform |
0.0 | 1 | 2000 | On using the CAM concept for parametric curve extraction · IEEE Trans. Image Process. 2000 |
Multimedia analysis and retrieval › image analysis
shape extraction |
0.0 | 1 | 2000 | On using the CAM concept for parametric curve extraction · IEEE Trans. Image Process. 2000 |
Electronic design automation
logic synthesis |
0.0 | 1 | 2004 | A Combined Approach to High-Level Synthesis for Dynamically Reconfigurable Systems · IEEE Trans. Computers 2004 |
Distributed systems
resource sharing |
0.0 | 1 | 2004 | A Combined Approach to High-Level Synthesis for Dynamically Reconfigurable Systems · IEEE Trans. Computers 2004 |
Memory systems
content-addressable memory |
0.0 | 1 | 2000 | On using the CAM concept for parametric curve extraction · IEEE Trans. Image Process. 2000 |
Methods — techniques the papers use, named apart from their topics
ultrasonic sensing · 0.4inversion with edge constraints · 0.4electrical resistance tomography · 0.4linear back projection · 0.2landweber algorithm · 0.2capacitance to digital conversion · 0.2weighted voting · 0.1hough transform · 0.1force-directed scheduling · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Edge-GPU-Based Face Tracking for Face Detection and Recognition AccelerationabstractCost-effective machine vision systems dedicated to real-time and accurate face detection and recognition in public places are crucial for many modern applications. However, despite their high performance, which could be reached using specialized edge or cloud AI hardware accelerators, there is still room for improvement in throughput and power consumption. This paper aims to suggest a combined hardware-software approach that optimizes face detection and recognition systems on NVIDIA Jetson AGX Orin for IoT applications. First, it leverages the simultaneous usage of all its hardware engines to improve processing time. This offers an improvement over previous works where these tasks were mainly allocated automatically and exclusively to the CPU or, to a higher extent, to the GPU core. Additionally, the paper suggests integrating a face tracker module to avoid redundantly running the face recognition algorithm on every frame but only when a new face appears in the scene. The results of extended experiments suggest that simultaneous usage of all the hardware engines that are available in the Orin GPU and tracker integration into the pipeline yield an impressive throughput of 290 FPS (frames per second) on 1920 x 1080 input size frames containing in average of 6 faces/frame. Additionally, a substantial saving of power consumption of around 800 mW was achieved when compared to running the task on the CPU/GPU engines only and without integrating a tracker into the Orin GPU’s pipeline. This hardware-codesign approach can pave the way to design high-performance machine vision systems at the edge, critically needed in video monitoring in public places where several nearby cameras are usually deployed for a same scene. Asma Baobaid, Mahmoud Méribout |
IEEE Internet Things J. | 2 |
| 2026 | A GPU-Enabled Multiframe Reconstruction for 3-D Electrical Tomography Using Tensor CoresabstractElectrical tomography (ET) has emerged as a sustainable and non-invasive imaging technique, offering nonionizing and non-radioactive solutions for various applications. Advances in hardware systems have enabled high-speed data acquisition and image reconstruction, but the computational demands of 3D ET systems with large electrode arrays and complex finite element models present significant challenges for real-time deployment, particularly in resource-constrained edge environments. This paper proposes a GPU-enabled Tensor Core-based reconstruction strategy for designing high-speed ET systems suitable for Internet-of-Things and embedded platforms. By leveraging Tensor Core units and employing a multi-frame reconstruction technique, the proposed approach demonstrates substantial performance improvements over traditional methods. Five different algorithms, encompassing both non-iterative and iterative approaches, were implemented to validate the benefits of this strategy. Experimental results on an embedded GPU achieved a speed gain of over 13 times compared to a general-purpose computer, and over four times when using Tensor Cores instead of CUDA cores alone. The system achieved frame rates of up to 3048 frames per second for a 32-electrode setup with 2 million mesh elements, highlighting its potential to meet the computational demands of sustainable, efficient, real-time, and portable ET systems. Varun Kumar Tiwari, Mahmoud Méribout |
IEEE Internet Things J. | 2 |
| 2026 | Scheduling Techniques of AI Models on Modern Heterogeneous Edge GPU - A Critical ReviewabstractIn recent years, the development of specialized edge computing devices has significantly increased, driven by the growing demand for artificial intelligence (AI) models. These devices, such as the NVIDIA Jetson series, must efficiently handle increased data processing and storage requirements. However, despite these advancements, there remains a lack of frameworks that automate the optimal execution of the deep neural network (DNN). Therefore, efforts have been made to create schedulers that can manage complex data processing needs while ensuring the efficient utilization of all available accelerators within these devices, including the CPU, GPU, deep learning accelerator (DLA), programmable vision accelerator (PVA), and video image compositor (VIC). Such schedulers would maximize the performance of edge computing systems, which is crucial in resource-constrained environments. This article aims to comprehensively review the various DNN schedulers implemented on NVIDIA Jetson devices. It examines their methodologies, performance, and effectiveness in addressing the demands of modern AI workloads. By analyzing these schedulers, this review highlights the current state of the research in the field. It identifies future research and development areas, further enhancing edge computing devices’ capabilities. Ashiyana Abdul Majeed, Mahmoud Méribout, Safa Mohammed Sali |
IEEE Trans. Ind. Informatics | 2 |
| 2025 | A Load Invariant and High-Speed IoT-Based Electrical Impedance Tomography SystemabstractElectrical impedance tomography (EIT) systems are widely used in various IoT-related fields as they yield real-time performance for two-dimensional (2D) and three-dimensional (3D) image reconstruction, are low-cost, and are not invasive. One of their most critical components is the current source that must deliver a constant amplitude AC electric current within the typical range of 10 kHz to 1 MHz, regardless of the load impedance. A slight fluctuation of the electric current would cause a deviation from the simulated Jacobian matrix, thereby yielding inaccurate image reconstruction. Thus, in practice, EIT is limited to applications where the process medium is highly conductive. This would neglect the load impedance, mitigating the fluctuations in the amplitude of the generated electric current. Consequently, many EIT systems were applied for processes dominantly comprising water, the conductivity of which ranges from 5 S/m for seawater to 5.5 x 10-6 and 5 x 10-3 S/m for pure and drinking water, respectively. Thus, even in the case of high water-cut fluid, unless the water is very salty, the usage of EIT may not be adequate. Indeed, even if the conductivity of the medium is known and high, the gross conductivity formed between a given pair of measuring electrodes depends heavily on the phases distribution pattern between them and can be excessively low, causing high electric current fluctuations. Changes in the contact impedance between the electrodes and the process are another major source of current amplitude fluctuations. The paper has two contributions. First, it experimentally assesses the effect of the electrical current fluctuations on the accuracy of the EIT image reconstruction using three different cutting-edge and most widely current sources designs. To the authors best knowledge, this study is the first of its kind, as all other prior works assume that the electrical current is constant in the formulation of the EIT forward and inverse problems. The paper also suggests a new cost-effective measurement circuit design that overcomes fluctuations while providing high-speed data acquisition throughput of up to 2,800 frames/s (fps). Extensive system assessment was conducted experimentally, and the associated results show the higher accuracy of the suggested design when using the Gauss-Newton (GN) method in terms of mean-squared error, which was decreased by 75%. Mohamed Elkhalil, Mahmoud Méribout, Varun Kumar Tiwari, Mohamed L. Seghier |
IEEE Internet Things J. | 2 |
| 2024 | A High-Speed ANN-Based Data Acquisition Hardware Accelerator Targeting Electrical Impedance TomographyabstractThis paper presents a novel design for a high-speed Data Acquisition (DAQ) system tailored for Electrical Impedance Tomography (EIT). Our proposed solution leverages a high-speed Analog-to-Digital Converter (ADC) to digitize analog signals from multiple electrode pairs within a single cycle, employing a time multiplexed approach. The resulting samples are then fed into an Artificial Neural Network (ANN) for accurate estimation of peak amplitudes across all channels, subsequently used for image reconstruction. To optimize the performance, we explored various ANN models with customized loss functions and devised an effective model selection approach using the grid search technique. In contrast to other multi-frequency techniques, our proposed approach eliminates the need for a multi-frequency current source, thereby simplifying the DAQ system. Additionally, it obviates the requirement for high-quality narrow-band pass band filters designed for different frequencies. By employing our approach, EIT systems can achieve remarkable throughput rates exceeding 2,800 Frames Per Second (fps) for a 50 kHz excitation signal, even with 32 or more electrodes. Extensive experimental testing demonstrated peak estimation accuracy surpassing 98%, even in scenarios with signals exhibiting 40 dB Signal-to-Noise Ratio (SNR). Consequently, our suggested approach exhibits tremendous potential for EIT applications that demand high SNR and rapid DAQ. Varun Kumar Tiwari, Mahmoud Méribout, Idowu Adeyemi, Mohamed Elkhalil |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2017 | A Pipelined Parallel Hardware Architecture for 2-D Real-Time Electrical Capacitance Tomography Imaging Using Interframe CorrelationabstractThis paper presents a new hardware algorithm for real-time and nonlinear 2-D electrical capacitance tomography (ECT) imaging, along with its parallel hardware architecture. A potential application of this system is to reconstruct in real time cross-sectional image of a two-phase fluid with different dielectric constants when it passes through a given section of a pipeline. The proposed hardware algorithm explores the spatial correlation that may occur between one or several consecutive frames. It uses this property to reformulate the classical regularized forward and inverse problems that are indeed very time consuming, preventing them to be used for fast ECT applications. As a result, the proposed algorithm features one-single-step iteration with a substantial reduction of the size of the Jacobian matrix. In addition, the intrinsically parallel feature of the algorithm makes it suitable for parallel hardware architecture. This architecture is based on a pipeline multiprocessor architecture using advanced features of field-programmable gate array technology. It explores the variable bit-width and floating-point multipliers array available in the digital signal processor blocks, to cooperatively perform the partial matrix product with associated arithmetic and logic units, and distributed memory. The experimental results obtained on a two-phase flow loop demonstrate the capability of the system to build in real time and with good accuracy (e.g., less than 3% error) the cross-sectional image of the fluid passing through the pipeline. Around 560 frames of 4096 moving boundary pixels can be reconstructed in 1 s using IEEE754 floating-point data representation and a clock frequency of 400 MHz, for a total power consumption of less than 33 W. Mahmoud Méribout, Samir Teniou |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2016 | A new systolic multiprocessor architecture for real-time soft tomography algorithms
Mahmoud Méribout, Ahmad Firadus |
Parallel Comput. | 1 |
| 2016 | A New Parallel VLSI Architecture for Real-Time Electrical Capacitance TomographyabstractThis paper presents a fixed-point reconfigurable parallel VLSI hardware architecture for real-time Electrical Capacitance Tomography (ECT). It is modular and consists of a front-end module which performs precise capacitance measurements in a time multiplexed manner using Capacitance to Digital Converter (CDC) technique. Another FPGA module performs the inverse steps of the tomography algorithm. A dual port built-in memory banks store the sensitivity matrix, the actual value of the capacitances, and the actual image. A two dimensional (2D) core multi-processing elements (PE) engine inter-communicates with these memory banks via parallel buses. A Hardware-software codesign methodology was conducted using commercially available tools in order to concurrently tune the algorithms and hardware parameters. Hence, the hardware was designed down to the bit-level in order to reduce both the hardware cost and power consumption, while satisfying real-time constraint. Quantization errors were assessed against the image quality and bit-level simulations demonstrate the correctness of the design. Further simulations indicate that the proposed architecture achieves a speed-up of up to three orders of magnitude over the software version when the reconstruction algorithm runs on 2.53 GHz-based Pentium processor or DSP Ti's Delphino TMS320F32837 processor. More specifically, a throughput of 17.241 Kframes/sec for both the Linear-Back Projection (LBP) and modified Landweber algorithms and 8.475 Kframes/sec for the Landweber algorithm with 200 iterations could be achieved. This performance was achieved using an array of [2 x 2] x [2 x 2] processing units. This satisfies the real-time constraint of many industrial applications. To the best of the authors' knowledge, this is the first embedded system which explores the intrinsic parallelism which is available in modern FPGA for ECT tomography. Ahmad Fajar Firdaus, Mahmoud Méribout |
IEEE Trans. Computers | 2 |
| 2015 | A Multimodal Image Reconstruction Method Using Ultrasonic Waves and Electrical Resistance TomographyabstractIn this paper, a new system that improves the image obtained by an array of ultrasonic sensors using electrical resistance tomography (ERT) is presented. One of its target applications can be in automatic exploration of soft tissues, where different organs and eventual anomalies exhibit simultaneously different electrical conductivities and different acoustic impedances. The exclusive usage of the ERT technique usually leads to some significant uncertainties around the regions' boundaries and usually generates images with relatively low resolutions. The proposed method shows that by properly combining this technique with an ultrasonic-based method, which can provide good localization of some edge points, the accuracy of the shape of individual cells can be improved, if these edge points are used as constraints during the inversion procedure. The performance of the proposed reconstruction method was assessed by conducting extensive tests on some simulated phantoms which mimic soft tissues. The obtained results clearly show the outperformance of this method over single modalities techniques that use either ultrasound or ERT imaging. Samir Teniou, Mahmoud Méribout |
IEEE Trans. Image Process. | 2 |
| 2004 | A Combined Approach to High-Level Synthesis for Dynamically Reconfigurable SystemsabstractThe increase in complexity of programmable hardware platforms results in the need to develop efficient high-level synthesis tools since that allows more efficient exploration of the design space while predicting the effects of technology specific tools on the design space. Much of the previous work, however, neglects the delay of interconnects (e.g. multiplexers) which can heavily influence the overall performance of the design. In addition, in the case of dynamic reconfigurable logic circuits, unless an appropriate design methodology is followed, an unnecessarily large number of configurable logic blocks may end up being used for communication between contexts, rather than for implementing function units. The aim of this paper is to present a new technique to perform interconnect-sensitive synthesis, targeting dynamic reconfigurable circuits. Further, the proposed technique exploits multiple hardware contexts to achieve efficient designs. Experimental results on several benchmarks, which have been done on our DRL LSI circuit [M. Meribout et al. [200]], [M. Meribout et al. (1997)], demonstrate that, by jointly optimizing the interconnect communication, and function unit cost, we can achieve higher quality designs than is possible with such previous techniques as Force-Directed-Scheduling. Mahmoud Méribout, Masato Motomura |
IEEE Trans. Computers | 1 |
| 2004 | Efficient metrics and high-level synthesis for dynamically reconfigurable logicabstractThe increase in complexity of programmable hardware platforms results in the need to develop efficient high-level synthesis (HLS) tools since it allows more efficient exploration of the design space while predicting the effects of technology specific tools on the design space. Much of the previous works however neglect the delay of interconnects (e.g. multiplexer) which can indeed contribute heavily on the overall performance of the design. In addition, in the case of dynamic reconfigurable logic (DRL) circuits, unless an appropriate design methodology is followed, large number of configurable logic blocks (CLBs) could be used for communication between contexts, rather than for implementing functional units (FUs). The aim of this paper is to present a new technique to perform interconnect-sensitive synthesis, targeting dynamic reconfigurable circuits. Further, the proposed technique exploits multiple hardware contexts to achieve efficient designs. Experimental results on several benchmarks, which have been done on our DRL LSI circuit (Meribout, 2000 and Motomura, 1997), demonstrate that by jointly optimizing the interconnect, communication, and function-unit cost, higher quality designs than other previous techniques (e.g. force-directed scheduling) can be achieved. Mahmoud Méribout, Masato Motomura |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2003 | On using a new dynamic reconfigurable logic (DRL) VLSI circuit for very high speed routingabstractRecent efforts to add new services to the Internet have increased the interest in designing flexible routers that are easy to extend and evolve. This paper describes a new hardware architecture based on dynamic reconfigurable logic (DRL) for high throughput networking applications. It mainly focuses on content-based router and on how to schedule efficiently its computation time. This scheduling task is difficult because of the various features of the underlying hardware such as multicontext, control-data path architecture and memory interface. Experimental results show some improvements over most recent network processors as well as a better hardware synthesis methodology. Mahmoud Méribout |
ICC | 1 |
| 2003 | On Using a new Dynamic Reconfigurable Logic (DRL) VLSI Circuit for Very High Speed RoutingabstractRecent efforts to add new services to the Internet have increased the interest in designing flexible routers that are easy to extend and evolve. This paper describes a new hardware architecture based on dynamic reconfigurable logic (DRL) for high throughput networking applications. It mainly focuses on content-based router and on how to schedule efficiently its computation time. This scheduling task is difficult because of the various features of the underlying hardware such as multicontext, control-data path architecture and memory interface. Experimental results show improvements over most recent network processors as well as a better hardware synthesis methodology. Mahmoud Méribout |
ISCC | 1 |
| 2003 | New design methodology with efficient prediction of quality metrics for logic level design towards dynamic reconfigurable logic
Mahmoud Méribout, Masato Motomura |
J. Syst. Archit. | 1 |
| 2002 | A parallel algorithm for real-time object recognition
Mahmoud Méribout, Mamoru Nakanishi, Takeshi Ogura |
Pattern Recognit. | 1 |
| 2000 | Hough Transform Algorithm for Three-Dimensional Segment Extraction and its Parallel Hardware Implementation
Mahmoud Méribout, Mamoru Nakanishi, Eiichi Hosoya, Takeshi Ogura |
Comput. Vis. Image Underst. | 1 |
| 2000 | On using the CAM concept for parametric curve extractionabstractIn this correspondence, a parallel algorithm, in order to extract parametric curves from a two-dimensional (2-D) image space, is proposed. It is based on the Hough transform (HT) and uses the content addressable memory (CAM) as the main processor. A set of simulated results for circular shape extraction are presented in order to demonstrate its merit. Hence, voting, thresholding, and three-dimensional (3-D) peak extraction are efficiently performed within the CAM. In addition, and in order to reduce the quantization errors, a weighted AT algorithm (WHT), which uses a weighted voting is proposed. Experimental results indicate that a real-time shape extraction for an image 256/spl times/256 can be achieved within a small amount of hardware. Therefore, CAM-based HT can be considered as a promising attraction for next generation pattern recognition platforms. Mahmoud Méribout, Takeshi Ogura, Mamoru Nakanishi |
IEEE Trans. Image Process. | 1 |
| 1993 | Real-time reprogrammable low-level image processing: edge detection and edge tracking acceleratorabstractCurrently, in image processing, segmentation algorithms comprise between real time video rate processing and accurate results. In this paper, we present an efficient and not recursive algorithm filter originated from Deriche filter. This algorithm is implemented in hardware by using FPGA technology. Thus, it permits video rate edge detection. In addition, the FPGA board is used as an edge tracking accelerator, it allows us to greatly reduce execution time by avoiding scanning the whole image. We also present the architecture of our vision system dedicated to build 3D scene every 200 ms. Mahmoud Méribout, Kun Mean Hou |
VCIP | 1 |