EDBT 2026 Demo / reviewers in the wild / expert
Mahmoud Masadeh
dblp:144/4689
· DBLP profile ↗
9ranked-venue papers
4as first author
5since 2021 · last 2025
0000-0001-7447-1276ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 4 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ML-based Load Value Approximator for Efficient Multimedia ProcessingabstractApproximate computing (AC) has gained traction as an alternative computing method for energy-efficient processing. This article proposes the exploitation of AC to address the memory wall. The proposed model predicts the memory load value using machine learning (ML). Subsequently, the ML model is a load value approximator (LVA) where the generated value is accepted as-is. The proposed LVA was tested under various approximate conditions, where 50% to 95% of the load instructions were approximated using a set of multimedia applications. The memory access operation using the proposed LVA was more than \(6\times\) faster in multiple cases. Additionally, the applications tested ran on average \(1.83\times\) faster. The peak signal-to-noise ratio (PSNR) exceeded 37 dB in several scenarios. The average normalized mean absolute error (NMAE) was 4.54%. Alain Aoun, Mahmoud Masadeh, Sofiène Tahar |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | Advancements in condition monitoring and fault diagnosis of rotating machinery: A comprehensive review of image-based intelligent techniques for induction motors
Omar AlShorman, Muhammad Irfan 0008, Ra'Ed Bani-Abdelrahman, Mahmoud Masadeh, Ahmad Alshorman, Muhammad Aman Sheikh, Nordin Bin Saad |
Eng. Appl. Artif. Intell. | 4 |
| 2023 | A Machine Learning Based Load Value Approximator Guided by the Tightened Value LocalityabstractThis paper addresses two essential memory bottlenecks: 1) memory wall, and 2) bandwidth wall. To accomplish this objective, we propose a machine learning (ML) based model that estimates the values to be loaded from the memory by a wide range of error-resilient applications. The proposed model exploits the feature of tightened value locality, which consists of a periodic load of few unique values. The proposed ML-based load value approximator (LVA) requires minimal overhead as it relies on a hash that encodes the history of events, e.g., history of accessed addresses, and values that can be extracted from the load instruction to be approximated. The proposed LVA completely eliminates memory accesses, i.e., 100% of accesses, in runtime and thus addresses the issue of memory wall and bandwidth wall. Compared to related work, our LVA delivers a maximum accuracy of 95.16% while offering a higher reduction in memory accesses. Alain Aoun, Mahmoud Masadeh, Sofiène Tahar |
ACM Great Lakes Symposium on VLSI | 2 |
| 2021 | A Quality-assured Approximate Hardware Accelerators-based on Machine Learning and Dynamic Partial ReconfigurationabstractMachine learning is widely used these days to extract meaningful information out of the Zettabytes of sensors data collected daily. All applications require analyzing and understanding the data to identify trends, e.g., surveillance, exhibit some error tolerance. Approximate computing has emerged as an energy-efficient design paradigm aiming to take advantage of the intrinsic error resilience in a wide set of error-tolerant applications. Thus, inexact results could reduce power consumption, delay, area, and execution time. To increase the energy-efficiency of machine learning on FPGA, we consider approximation at the hardware level, e.g., approximate multipliers. However, errors in approximate computing heavily depend on the application, the applied inputs, and user preferences. However, dynamic partial reconfiguration has been introduced, as a key differentiating capability in recent FPGAs, to significantly reduce design area, power consumption, and reconfiguration time by adaptively changing a selective part of the FPGA design without interrupting the remaining system. Thus, integrating “Dynamic Partial Reconfiguration” (DPR) with “Approximate Computing” (AC) will significantly ameliorate the efficiency of FPGA-based design approximation. In this article, we propose hardware-efficient quality-controlled approximate accelerators, which are suitable to be implemented in FPGA-based machine learning algorithms as well as any error-resilient applications. Experimental results using three case studies of image blending, audio blending, and image filtering applications demonstrate that the proposed adaptive approximate accelerator satisfies the required quality with an accuracy of 81.82%, 80.4%, and 89.4%, respectively. On average, the partial bitstream was found to be 28.6 smaller than the full bitstream . Mahmoud Masadeh, Yassmeen Elderhalli, Osman Hasan, Sofiène Tahar |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2021 | Machine-Learning-Based Self-Tunable Design of Approximate ComputingabstractApproximate computing (AC) is an emerging computing paradigm suitable for intrinsic error-tolerant applications to reduce energy consumption and execution time. Different approximate techniques and designs, at both hardware and software levels, have been proposed and demonstrated the effectiveness of relaxing the average output quality constraint. However, the output quality of AC is highly input-dependent, i.e., for some input data, the output errors may reach unacceptable levels. Therefore, there is a dire need for an input-dependent tunable approximate design. With this motivation, in this article, we propose a lightweight and efficient machine-learning-based approach to build an input-aware design selector, i.e., quality controller, to adapt the approximate design in order to meet the target output quality (TOQ). For illustration purposes, we use a library of 8-bit and 16-bit energy-efficient approximate array multipliers with 20 different settings, which are commonly used in image and audio processing applications. The simulation results, based on two sets of images, including an 8 Scene Categories Dataset, which is a benchmark of images data set, demonstrate the effectiveness of the lightweight selector where the proposed tunable design achieves a significant reduction in quality loss with relatively low overhead. Mahmoud Masadeh, Osman Hasan, Sofiène Tahar |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2019 | Using Machine Learning for Quality Configurable Approximate ComputingabstractApproximate computing (AC) is a nascent energy-efficient computing paradigm for error-resilient applications. However, the quality control of AC is quite challenging due to its input-dependent nature. Existing solutions fail to address fine-grained input-dependent controlled approximation. In this paper, we propose an input-aware machine learning based approach for the quality control of AC. For illustration purposes, we use 20 configurations of 8-bit approximate multipliers. We evaluate these designs for all combinations of possible input data. Then, we use machine learning algorithms to efficiently make predictive decisions for the quality control of the target approximate application, based on experimentally collected training data. The key benefits of the proposed approach include: (1) fine-grained input-dependent approximation, (2) no missed approximation opportunities, (3) no rollback recovery overhead, (4) applicable to any approximate computation with error-tolerant components, and (5) flexibility in adapting various error metrics. Mahmoud Masadeh, Osman Hasan, Sofiène Tahar |
DATE | 1 |
| 2018 | Comparative Study of Approximate MultipliersabstractApproximate multipliers are widely being advocated for energy-efficient computing in applications that exhibit an inherent tolerance to inaccuracy. In this paper, we identify three decisions for design and evaluation of approximate multiplier circuits: (1) the type of approximate full adder (FA) used to construct the multiplier, (2) the architecture, i.e., array or tree, of the multiplier and (3) the placement of sub-modules of approximate and exact multipliers in the target multiplier module. Based on FA cells implemented at the transistor level (TSMC65nm), we developed several approximate building blocks of 8x8 multipliers, as well as various implementations of higher order multipliers. These designs are evaluated based on their power, area, delay and error and the best designs are identified. We validate these designs on an image blending application using MATLAB, and compare them to related work. Mahmoud Masadeh, Osman Hasan, Sofiène Tahar |
ACM Great Lakes Symposium on VLSI | 1 |
| 2015 | Post-Bond Interconnect Test and Diagnosis for 3-D Memory Stacked on Logicabstract3-D stacked integrated circuit (IC) technology based on through-silicon vias (TSVs) provides numerous advantages as compared to traditional 2-D-ICs. A potential application is memory stacked on logic, providing enhanced throughput, and reduced latency and power consumption. However, testing the TSV interconnects between the two dies is challenging as both memory and logic dies might come from different providers. Currently, no standard exists and the proposed solutions fail to address dynamic and time-critical faults (at speed testing). In addition, memory vendors have not been in favor to put additional design-for-testability structures such as Joint Test Action Group for interconnect testing on their memory devices. This paper proposes a new memory-based interconnect test (MBIT) approach for 3-D memories stacked on logic (e.g., CPUs). A structural approach is used to develop fault models, their detection conditions, and test and diagnosis patterns. The test patterns are applied by read and write instructions to the memory and are validated by a case study where a 3-D memory is assumed to be stacked on a MIPS64 processor. The main benefits of the MBIT approach are: 1) zero area overhead; 2) the ability to detect both static and dynamic faults and perform at speed testing; 3) flexibility in applying any test pattern, as this can be executed by the CPU on the logic die; 4) extreme short test execution time; and 5) the ability to perform interconnect diagnosis. Mottaqiallah Taouil, Mahmoud Masadeh, Said Hamdioui, Erik Jan Marinissen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2014 | Interconnect test for 3D stacked memory-on-logicabstractThree-dimensional stacked IC (3D-SIC) technology based on Through-Silicon Vias (TSVs) provides numerous advantages as compared to traditional 2D-ICs. A potential application is memory stacked on logic, providing enhanced throughput, and reduced latency and power consumption. However, testing the TSV interconnects between the two dies is challenging, as both the memory and the logic die might come from different manufacturers. Currently, no standard exists and the proposed solutions fail to address dynamic and time-critical faults (at speed testing). In addition, memory vendors have not been in favor to put additional DfT structures such as JTAG for interconnect testing on their memory devices. This paper proposes a new Memory Based Interconnect Test (MBIT) approach for 3D stacked memories. Our test patterns are applied by read and write instructions to the memory and are validated by a case study where a 3D memory is assumed to be stacked on a MIPS64 processor. The main benefits of the MBIT approach are: (1) zero area overhead, (2) the ability to detect both static and dynamic faults and perform at speed testing, (3) flexibility in applying any test pattern, as this can be executed by the CPU on the logic die and (4) extreme short test execution time. Mottaqiallah Taouil, Mahmoud Masadeh, Said Hamdioui, Erik Jan Marinissen |
DATE | 2 |