Mohamed E. Fouda

dblp:141/0978 · also Mohammed E. Fouda · DBLP profile ↗
← Back
33ranked-venue papers
1as first author
28since 2021 · last 2026
0000-0001-7139-3428ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 24 · 1 first-author · 20 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Computer networks · 2 · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 A Synthesizable Mixed-Precision DCIM Macro with Parallel Write and Compute
Jinane Bazzi, Mohamed E. Fouda, Ahmed M. Eltawil
ISCAS2
2026 Sparse neural sampling mixers
Ahmed Elsheikh, Mohamed E. Fouda, Ahmed M. Eltawil
Neurocomputing2
2025 SoftmAP: Software-Hardware Co-Design for Integer-Only Softmax on Associative Processors
abstract
Recent research efforts focus on reducing the computational and memory overheads of Large Language Models (LLMs) to make them feasible on resource-constrained devices. Despite advancements in compression techniques, nonlinear operators like Softmax and Layernorm remain bottlenecks due to their sensitivity to quantization. We propose SoftmAP, a software-hardware co-design methodology that implements an integer-only low-precision Softmax using In-Memory Compute (IMC) hardware. Our method achieves up to three orders of magnitude improvement in the energy-delay product compared to A100 and RTX3090 GPUs, making LLMs more deployable without compromising performance.
Mariam Rakka, Jinhao Li 0006, Guohao Dai 0001, Ahmed M. Eltawil, Mohamed E. Fouda, Fadi J. Kurdahi
DATE5
2025 On Jailbreaking Quantized Language Models Through Fault Injection Attacks
abstract
The safety alignment of Language Models (LMs) is a critical concern, yet their integrity can be challenged by direct parameter manipulation attacks, such as those potentially induced by fault injection. As LMs are increasingly deployed using low-precision quantization for efficiency, this paper investigates the efficacy of such attacks for jailbreaking aligned LMs across different quantization schemes. We propose gradient-guided attacks, including a tailored progressive bit-level search algorithm introduced herein and a comparative word-level (single weight update) attack. Our evaluation on Llama-3.2-3B, Phi-4-mini, and Llama-3-8B across FP16 (baseline), and weight-only quantization (FP8, INT8, INT4) reveals that quantization significantly influences attack success. While attacks readily achieve high success (>80% Attack Success Rate, ASR) on FP16 models, within an attack budget of 25 perturbations, FP8 and INT8 models exhibit ASRs below 20% and 50%, respectively. Increasing the perturbation budget up to 150 bit-flips, FP8 models maintained ASR below 65%, demonstrating some resilience compared to INT8 and INT4 models that have high ASR. In addition, analysis of perturbation locations revealed differing architectural targets across quantization schemes, with (FP16, INT4) and (INT8, FP8) showing similar characteristics. Besides, jailbreaks induced in FP16 models were highly transferable to subsequent FP8/INT8 quantization (<5% ASR difference), though INT4 significantly reduced transferred ASR (avg. 35% drop). These findings highlight that while common quantization schemes, particularly FP8, increase the difficulty of direct parameter manipulation jailbreaks, vulnerabilities can still persist, especially through post-attack quantization.
Noureldin Zahran, Ahmad Tahmasivand, Ihsen Alouani, Khaled N. Khasawneh, Mohamed E. Fouda
ACM Great Lakes Symposium on VLSI5
2025 LM-Fix: Lightweight Bit-Flip Detection and Rapid Recovery Framework for Language Models
abstract
Bit-flip attacks threaten the reliability and security of Language Models (LMs) by altering internal parameters and compromising output integrity. Recent studies show that flipping only a few bits in model parameters can bypass safety mechanisms and jailbreak the model. Existing detection approaches for DNNs and CNNs are not suitable for LMs, as the massive number of parameters significantly increases timing and memory overhead for software-based methods and chip area overhead for hardware-based methods. In this work, we present LM-Fix, a lightweight LM-driven detection and recovery framework that leverages the model's own capabilities to identify and recover faults. Our method detects bit-flips by generating a single output token from a predefined test vector and auditing the output tensor of a target layer against stored reference data. The same mechanism enables rapid recovery without reloading the entire model. Experiments across various models show that LM-Fix detects more than 94% of single-bit flips and nearly 100% of multi-bit flips, with very low computational overhead <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">$(\approx 1 \%- 7.7 {\%}$</tex> at TVL <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">$=200$</tex> across models). Recovery achieves more than <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">$100 \times$</tex> speedup compared to full-model reload, which is critical in edge devices. LM-Fix can handle bit-flips affecting any part of the model's computation, including memory, cache, and arithmetic operations. Evaluation against recent LM-specific bit-flip attacks confirms its robustness and practical value for real-world deployment.
Ahmad Tahmasivand, Noureldin Zahran, Saba Al-Sayouri, Mohamed E. Fouda, Khaled N. Khasawneh
ICCD4
2025 Reconfigurable Precision INT4-8/FP8 Digital Compute-in-Memory Macro for AI Acceleration
abstract
Compute-in-memory (CIM) technology has emerged as a promising solution to address the computational demands of deep neural network (DNN) models, which require substantial multiply-accumulate (MAC) operations. However, there is a growing need for reconfigurable CIM architectures that can support both integer (INT) and floating-point (FP) operations within a single design. This flexibility is crucial for optimizing efficiency and resource utilization, especially in DNN applications involving mixed-precision computations. In this work, we propose a reconfigurable precision digital macro design that accelerates MAC computations while supporting INT4, INT8, and FP8 configurations within the same architecture. Both signed and unsigned operations are supported in INT mode. To enhance performance, the proposed design uses a parallel-input approach and a mantissa parallel-alignment technique in FP mode. The macro is implemented in 40nm CMOS technology. It achieves a peak throughput of 7123.48 GOPS in INT4 mode and 1187.25 GFLOPS in FP8 mode, with peak energy efficiencies of 367.45 TOPS/W and 23.14 TFLOPS/W, respectively.
Jinane Bazzi, Mohamed E. Fouda, Ahmed M. Eltawil
ISCAS2
2025 Towards Efficient IMC Accelerator Design Through Joint Hardware-Workload Co-optimization
abstract
Designing generalized in-memory computing (IMC) hardware that efficiently supports a variety of workloads requires extensive design space exploration, which is infeasible to perform manually. Optimizing hardware individually for each workload or solely for the largest workload often fails to yield the most efficient generalized solutions. To address this, we propose a joint hardware-workload optimization framework that identifies optimised IMC chip architecture parameters, enabling more efficient, workload-flexible hardware. We show that joint optimization achieves 36%, 36%, 20%, and 69% better energy-latency-area scores for VGG16, ResNet18, AlexNet, and MobileNetV3, respectively, compared to the separate architecture parameters search optimizing for a single largest workload. Additionally, we quantify the performance trade-offs and losses of the resulting generalized IMC hardware compared to workload-specific IMC designs.
Olga Krestinskaya, Mohamed E. Fouda, Ahmed M. Eltawil, Khaled N. Salama
ISCAS2
2025 Mitigating the Impact of ReRAM I-V Nonlinearity and IR Drop via Fast Offline Network Training
abstract
ReRAM crossbar arrays (RCAs) have the potential to provide extremely high efficiency for accelerating deep neural networks (DNNs). However, one crucial challenge for RCA-based DNN accelerators is functional inaccuracy due to nonidealities present in RCA hardware. While nonideality-aware training (NAT) could be used to mitigate the effect of nonidealities, with currently available methods it would take months to train even a medium size convolutional neural network (CNN). In this article we propose a nonideality prediction method that enables very fast training of RCA-based neural networks, and show its feasibility through NAT of DNNs. Our key ideas include 1) weight-centric nonideality modeling and 2) data-dependence elimination by tailored input randomization. Our experimental results using a multilayer perceptron and CNNs demonstrate that our method is very fast ($100\sim 15$$000\times $faster training speed) while achieving much better-crossbar-level accuracy ($2 \sim 90\times $lower-RMS error) and post-retraining validated accuracy than previous methods.
Sugil Lee, Mohamed E. Fouda, Chenghao Quan, Jongeun Lee, Ahmed M. Eltawil, Fadi J. Kurdahi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2025 Single-Cycle Independent Component Analysis Processor for In-Band Full Duplex Systems
abstract
This paper presents a single-cycle independent component analysis (ICA) algorithm and architecture for self-interference cancellation (SIC) in In-band Full-duplex (IBFD) communication systems. The proposed algorithm, AICA-EBM, incorporates the adaptive momentum (ADAM) approach with the entropy-bound estimation (EBM) to achieve rapid convergence. The simulation results show that the proposed AICA-EBM algorithm leads to a constant single-cycle processing time to achieve satisfactory SIC optimality in IBFD systems. The architecture of the ICA processor based on the proposed AICA-EBM algorithm is designed. Novel circuit structures are proposed to reduce the complexity of the ICA processor. The AICA-EBM processor is implemented by following an application-specific integrated circuit (ASIC) flow with the TSMC 90 nm process. The post-layout estimations show that Compared to previous ICA processors, the proposed design achieves a 2x improvement in processing throughput and leads to the best hardware efficiency.
Hao-Lun Weng, Chung-An Shen, Mohamed E. Fouda, Ahmed M. Eltawil
IEEE Trans. Circuits Syst. I Regul. Pap.3
2024 Battery Modeling with Mittag-Leffler Function
abstract
In various areas of life, rechargeable lithium-ion batteries are the technology of choice. Equivalent circuit models are utilized extensively in characterizing and modeling energy storage systems. In real-time applications, several generic-based battery models are created to simulate the battery’s charging and discharging behavior more accurately. In this work, we present two generic battery models based on Mittag-Leffler function using a generic Standard battery model as a reference. These models are intended to fit the continuous discharging cycles of lithium-ion, Nickel-cadmium, and Nickel-metal hydride batteries, as well as one set from the NASA randomized battery usage dataset. We formulate the parameter identification as an optimization problem, solved with Marine Predator Algorithm. The optimized models show very good matching against the measured data.
Shahenda M. Abdelhafiz, Mohamed E. Fouda, Ahmed Gomaa Radwan
ISCAS2
2024 Reconfigurable Precision SRAM-based Analog In-memory-compute Macro Design
abstract
In-memory computing (IMC) is a promising approach for accelerating multiply and accumulate (MAC) operations, which are the primary calculations used in artificial intelligence (AI). The demand for flexible architectures supporting different bit precisions in MAC computations becomes evident. This flexibility balances adapting to specific model requirements and optimizing design performance efficiency. As such, in this paper, we propose a reconfigurable IMC macro design, utilizing 8T static random-access memory (SRAM) bit-cells in 65nm technology, to efficiently perform MAC operations while supporting three bit precisions: 2, 3, and 4 bits for each of the input, weight, and output. The proposed 64×180 macro achieves a normalized peak throughput of 13.82 TOPS, a normalized peak energy efficiency of 291.66 TOPS/W, and a normalized peak area efficiency of 165.98 TOPS/mm2.
Jinane Bazzi, Rachid Jamil, Dana El Hajj, Rouwaida Kanj, Mohamed E. Fouda, Ahmed M. Eltawil
ISCAS5
2024 Chaotic neural network quantization and its robustness against adversarial attacks
Alaa Osama, Samar I. Gadallah, Lobna A. Said, Ahmed Gomaa Radwan, Mohamed E. Fouda
Knowl. Based Syst.5
2024 A Review of State-of-the-art Mixed-Precision Neural Network Frameworks
abstract
Mixed-precision Deep Neural Networks (DNNs) provide an efficient solution for hardware deployment, especially under resource constraints, while maintaining model accuracy. Identifying the ideal bit precision for each layer, however, remains a challenge given the vast array of models, datasets, and quantization schemes, leading to an expansive search space. Recent literature has addressed this challenge, resulting in several promising frameworks. This paper offers a comprehensive overview of the standard quantization classifications prevalent in existing studies. A detailed survey of current mixed-precision frameworks is provided, with an in-depth comparative analysis highlighting their respective merits and limitations. The paper concludes with insights into potential avenues for future research in this domain.
Mariam Rakka, Mohamed E. Fouda, Pramod P. Khargonekar, Fadi J. Kurdahi
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 High-Density FeFET-based CAM Cell Design Via Multi-Dimensional Encoding
abstract
Content addressable memory is one of the most frequently used technologies in Data-centric applications due to its exceptional search parallelism capability. SRAM cells were initially used to implement CAM designs. Recent innovations proposed using compact nonvolatile memories instead. FeFETs emerged as a multi-level NVM device with promising potential and 2T FeFET CAM designs were studied. In this paper, a new potential is discussed for increasing the density efficiency of FeFET CAM architectures by adapting higher-dimensional encoding using 3T and 4T CAM designs. We propose a scalable greedy search algorithm for maximizing encoding capabilities. We compare the density, latency, accuracy, and energy consumption of our designs to standard 2T architecture demonstrating a 4x and 8x decrease in fail probability with up to 16% and 26.5% increase in memory density (bits/unit-area) in the 3T and 4T designs respectively.
Hadi Noureddine, Omar Bekdache, Mohamad Al Tawil, Rouwaida Kanj, Ali Chehab, Mohamed E. Fouda, Ahmed M. Eltawil
ACM Great Lakes Symposium on VLSI6
2023 Low Precision Quantization-aware Training in Spiking Neural Networks with Differentiable Quantization Function
abstract
Deep neural networks have been proven to be highly effective tools in various domains, yet their computational and memory costs restrict them from being widely deployed on portable devices. The recent rapid increase of edge computing devices has led to an active search for techniques to address the abovementioned limitations of machine learning frameworks. The quantization of artificial neural networks (ANNs), which converts the full-precision synaptic weights into low-bit versions, emerged as one of the solutions. At the same time, spiking neural networks (SNNs) have become an attractive alternative to conventional ANNs due to their temporal information processing capability, energy efficiency, and high biological plausibility. Despite being driven by the same motivation, the simultaneous utilization of both concepts has yet to be thoroughly studied. Therefore, this work aims to bridge the gap between recent progress in quantized neural networks and SNNs. It presents an extensive study on the performance of the quantization function, represented as a linear combination of sigmoid functions, exploited in low-bit weight quantization in SNNs. The presented quantization function demonstrates the state-of-the-art performance on four popular benchmarks, CIFAR10-DVS, DVS128 Gesture, N-Caltech101, and N-MNIST, for binary networks (64.05%, 95.45%, 68.71%, and 99.43% respectively) with small accuracy drops and up to 31 × memory savings, which outperforms existing methods.
Ayan Shymyrbay, Mohamed E. Fouda, Ahmed M. Eltawil
IJCNN2
2023 Hardware Acceleration of DNA Pattern Matching with Binary Memristors
abstract
DNA pattern matching is a key technique applied in many bioinformatics applications. Recently, this technique has become very popular and is widely used for genetic disease diagnosis, where finding the number of consecutive repeats of a specific DNA pattern indicates the type and intensity of the patient's disorder. However, the remarkable growth of DNA data exacerbates the latency and power consumption required to perform DNA pattern matching. In this work, we propose a hardware accelerator design to detect the presence of different diseases efficiently using DNA pattern matching. We propose a novel CAM cell using binary memristors for reliable and robust data encoding. The proposed architecture consists of two main building blocks the Content-addressable memory (CAM) and pattern detector circuits in addition to the needed peripheral circuits for CAM read, write and match operation. CMOS PTM 45nm technology was used to design and simulate the full architecture. The evaluation of the proposed design shows$\sim 2\times$improvement in energy-delay-area product compared to the state-of-art work in the literature, in addition to robustness against noise and process variations.
Jinane Bazzi, Mohamed E. Fouda, Rouwaida Kanj, Ahmed M. Eltawil
ISCAS2
2023 Scalable Complementary FeFET CAM Design
abstract
CAMs are frequently employed for data-centric applications. They offer excellent parallelism. Traditionally, they were implemented using the area-consuming SRAM. Recent advancements suggest using compact nonvolatile memories (NVMs) to create CAM cells to reduce area. The ferroelectric field effect transistor (FeFET) has therefore emerged as an NVM device showing great potential in these memory architectures. In this work, we propose a novel multi-bit CAM architecture that utilizes p-type FeFETs – a topic yet to be explored in the literature – and we compare the latency, accuracy, and energy consumption of our design to other FeFET-based architectures demonstrating a 3-30× reduction in fail probability.
Omar Bekdache, Hadi Noureddine, Mohamad Al Tawil, Rouwaida Kanj, Mohamed E. Fouda, Ahmed M. Eltawil
ISCAS5
2023 High-Throughput Independent Component Analysis Processor for Full Duplex Systems
abstract
This paper presents the algorithm and very-large-scale integration (VLSI) architecture of a high-throughput and highly efficient independent component analysis (ICA) processor for self-interference cancellation (SIC) in in-band full-duplex (IBFD) systems. This is the first VLSI architecture reported in the literature based on the state-of-the-art entropy bound minimization (EBM) approach. A novel ICA algorithm is presented in this paper with momentum gradient descent optimization. Simulation results show that the number of iterations for the proposed algorithm is significantly reduced compared to the conventional ICA algorithms. Furthermore, a novel early-distribution estimation scheme is proposed in the designed ICA processor to compute multiple distribution functions with low latency and low complexity. The processing flow and the efficiency for the hardware utilization are specifically designed so that the processing speed is maximized with minimum employment of hardware components. The proposed ICA processor is designed and implemented based on the application-specific-integrated circuit (ASIC) flow. The post-layout estimations show that compared with the conventional EBM-based scheme, the proposed design improves the throughput and efficiency by 30x. In addition, compared to prior designs shown in the literature, the proposed ICA processor also demonstrates a significant enhancement in terms of throughput and efficiency.
Jen-Hao Cheng, Tien-Min Chang, Chung-An Shen, Mohamed E. Fouda, Ahmed M. Eltawil
IEEE J. Sel. Areas Commun.4
2023 Resistive Neural Hardware Accelerators
abstract
Deep neural networks (DNNs), as a subset of machine learning (ML) techniques, entail that real-world data can be learned, and decisions can be made in real time. However, their wide adoption is hindered by a number of software and hardware limitations. The existing general-purpose hardware platforms used to accelerate DNNs are facing new challenges associated with the growing amount of data and are exponentially increasing the complexity of computations. Emerging nonvolatile memory (NVM) devices and the compute-in-memory (CIM) paradigm are creating a new hardware architecture generation with increased computing and storage capabilities. In particular, the shift toward resistive random access memory (ReRAM)-based in-memory computing has great potential in the implementation of area- and power-efficient inference and in training large-scale neural network architectures. These can accelerate the process of IoT-enabled AI technologies entering our daily lives. In this survey, we review the state-of-the-art ReRAM-based DNN many-core accelerators, and their superiority compared to CMOS counterparts was shown. The review covers different aspects of hardware and software realization of DNN accelerators, their present limitations, and prospects. In particular, a comparison of the accelerators shows the need for the introduction of new performance metrics and benchmarking standards. In addition, the major concerns regarding the efficient design of accelerators include a lack of accuracy in simulation tools for software and hardware codesign.
Kamilya Smagulova, Mohamed E. Fouda, Fadi J. Kurdahi, Khaled N. Salama, Ahmed M. Eltawil
Proc. IEEE2
2023 Offline Training-Based Mitigation of IR Drop for ReRAM-Based Deep Neural Network Accelerators
abstract
Recently, resistive RAM (ReRAM)-based hardware accelerators showed unprecedented performance compared the digital accelerators. Technology scaling causes an inevitable increase in interconnect wire resistance, which leads to IR drops that could limit the performance of ReRAM-based accelerators. These IR drops deteriorate the signal integrity and quality, especially in the crossbar structures which are used to build high-density ReRAMs. Hence, finding a software solution, which can predict the effect of IR drop without involving expensive hardware or SPICE simulations, is very desirable. In this article, we propose two neural networks models to predict the impact of the IR drop problem. These models are used to evaluate the performance of the different deep neural network (DNN) models including binary and quantized neural networks showing similar performance (i.e., recognition accuracy) to the golden validation (i.e., SPICE-based DNN validation). In addition, these predication models are incorporated into the DNN training framework to efficiently retrain the DNN models and bridge the accuracy drop. To further enhance the validation accuracy, we propose incremental training methods. The DNN validation results, done through SPICE simulations, show very high improvement in performance close to the baseline performance, which demonstrates the efficacy of the proposed method even with challenging datasets, such as CIFAR10 and SVHN.
Sugil Lee, Mohamed E. Fouda, Jongeun Lee, Ahmed M. Eltawil, Fadi J. Kurdahi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 Training-Free Stuck-At Fault Mitigation for ReRAM-Based Deep Learning Accelerators
abstract
Although Resistive RAMs can support highly efficient matrix–vector multiplication, which is very useful for machine learning and other applications, the nonideal behavior of hardware, such as stuck-at fault (SAF) and IR drop is an important concern in making ReRAM crossbar array-based deep learning accelerators. Previous work has addressed the nonideality problem through either redundancy in hardware, which requires a permanent increase of hardware cost, or software retraining, which may be even more costly or unacceptable due to its need for a training dataset as well as high computation overhead. In this article, we propose a very lightweight method that can be applied on top of existing hardware or software solutions. Our method, called forward-parameter tuning (FPT), takes advantage of a certain statistical property existing in the activation data of neural network layers, and can mitigate the impact of mild nonidealities in ReRAM crossbar arrays (RCAs) for deep learning applications without using any hardware, a dataset, or gradient-based training. Our experimental results using MNIST, CIFAR-10, and CIFAR-100, and ImageNet datasets in binary and multibit networks demonstrate that our technique is very effective, both alone and together with previous methods, up to 20% fault rate, which is higher than even some of the previous remapping methods. We also evaluate our method in the presence of other nonidealities, such as variability and IR drop. Furthermore, we provide an analysis based on the concept of the effective fault rate (EFR), which not only demonstrates that EFR can be a useful tool to predict the accuracy of faulty RCA-based neural networks but also explains why mitigating the SAF problem is more difficult with multibit neural networks.
Chenghao Quan, Mohamed E. Fouda, Sugil Lee, Giju Jung, Jongeun Lee, Ahmed M. Eltawil, Fadi J. Kurdahi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2022 Wearable Vital Signal Monitoring Prototype Based on Capacitive Body Channel Communication
abstract
Wireless body area network (WBAN) provides a means for seamless individual health monitoring without imposing restrictive limitations on normal daily routines. To date, Radio Frequency (RF) transceivers have been the technology of choice, however, drawbacks such as vulnerability to body shadowing effects, higher power consumption due to omnidirectional radiation and security concerns, have prompted the adoption of transceivers that use the human body channel for communication. In this paper, a vital signal monitoring transceiver prototype based on the human body channel communication (HBC), using commercially available chipsets is presented. RF and HBC communications are briefly reviewed and compared, and different schemes of HBC are introduced. A circuit model that represents the human body channel is then discussed and simulations are presented to illustrate the influence of the return path capacitance and receiver terminations on the path loss. The architecture of the transceiver prototype is then introduced where it is designed at a 21 MHz IEEE 802.15.6 standard-compliant carrier frequency. Finally, the performance of the transceiver, including the bit error rate (BER) and power efficiency, are characterized. Path loss is measured for two different scenarios, where variations of up to 5 dB were observed due to environmental effects. Energy efficiency measured at a maximum data-rate of 1.3 Mbps was found to be 8.3 nJ/b.
Qi Huang 0002, Waseem Alkhayer, Mohamed E. Fouda, Abdulkadir Celik, Ahmed M. Eltawil
BSN3
2022 Accurate Prediction of ReRAM Crossbar Performance Under I-V Nonlinearity and IR Drop
abstract
Despite the promise of extremely efficient matrix-vector multiplication (MVM) by ReRAM crossbar arrays (RCAs), maintaining high accuracy has been challenging due to nonidealities such as wire resistance (also known as IR drop) and I-V nonlinearity (i.e., voltage-dependent conductance). For system architects, a fast method to accurately predict the MVM output of an RCA under nonidealities is highly desirable. While IR drop alone without I-V nonlinearity can be efficiently predicted, the existence of I-V nonlinearity makes the problem much harder. In this paper we propose a novel algorithm based on iterative refinement, which can predict with high accuracy the outcome of an MVM operation on an RCA in the presence of both I-V nonlinearity and IR drop. Our experiments using binary RCAs of different sizes demonstrate that our proposed method is order-of-magnitude more accurate than previous methods in terms of RMS error. We also present case studies predicting hardware-realistic accuracy of binarized neural networks on RCAs as well as nonideality-aware retraining, demonstrating the efficacy of our method for early design space exploration of ReRAM-based accelerators.
Sugil Lee, Mohamed E. Fouda, Jongeun Lee, Ahmed M. Eltawil, Fadi J. Kurdahi
ICCD2
2022 Toward the Optimal Design and FPGA Implementation of Spiking Neural Networks
abstract
The performance of a biologically plausible spiking neural network (SNN) largely depends on the model parameters and neural dynamics. This article proposes a parameter optimization scheme for improving the performance of a biologically plausible SNN and a parallel on-field-programmable gate array (FPGA) online learning neuromorphic platform for the digital implementation based on two numerical methods, namely, the Euler and third-order Runge-Kutta (RK3) methods. The optimization scheme explores the impact of biological time constants on information transmission in the SNN and improves the convergence rate of the SNN on digit recognition with a suitable choice of the time constants. The parallel digital implementation leads to a significant speedup over software simulation on a general-purpose CPU. The parallel implementation with the Euler method enables around 180× ( 20× ) training (inference) speedup over a Pytorch-based SNN simulation on CPU. Moreover, compared with previous work, our parallel implementation shows more than 300× ( 240× ) improvement on speed and 180× ( 250× ) reduction in energy consumption for training (inference). In addition, due to the high-order accuracy, the RK3 method is demonstrated to gain 2× training speedup over the Euler method, which makes it suitable for online training in real-time applications.
Wenzhe Guo, Hasan Erdem Yantir, Mohamed E. Fouda, Ahmed M. Eltawil, Khaled N. Salama
IEEE Trans. Neural Networks Learn. Syst.3
2022 Efficient Neuromorphic Hardware Through Spiking Temporal Online Local Learning
abstract
Local learning schemes have shown promising performance in spiking neural networks (SNNs) training and are considered a step toward more biologically plausible learning. Despite many efforts to design high-performance neuromorphic systems, a fast and efficient on-chip training algorithm is still missing, which limits the deployment of neuromorphic systems in many real-time applications. This work proposes a scalable, fast, and efficient spiking neuromorphic hardware system with on-chip local learning capability. We introduce an effective hardware-friendly local training algorithm compatible with sparse temporal input coding and binary random classification weights. The algorithm is demonstrated to deliver competitive accuracy in different tasks. The proposed digital system explores spike sparsity in communication, parallelism in vector–matrix operations and process-level dataflow, and locality of training errors, which leads to low cost and fast training speed. The system is optimized under various performance metrics. Taking into consideration energy, speed, resources, and accuracy, the proposed method shows around$10\times $efficiency over a recent work with a direct feedback alignment (DFA) method and$4.5\times $efficiency over the spike-timing-dependent plasticity (STDP) method. Moreover, our hardware architecture can easily scale up with the network size at a linear rate. Thus, our method has demonstrated great potential for use in various applications, especially those demanding low latency.
Wenzhe Guo, Mohamed E. Fouda, Ahmed M. Eltawil, Khaled N. Salama
IEEE Trans. Very Large Scale Integr. Syst.2
2022 Configurable Independent Component Analysis Preprocessing Accelerator
abstract
An independent component analysis (ICA) has been used in many applications, including self-interference cancellation (SIC) for in-band full-duplex (IBFD) wireless systems and anomaly detection in industrial Internet of Things (IoT). This article presents a high-throughput and highly efficient configurable preprocessing accelerator for the ICA algorithm. The proposed ICA accelerator has three major blocks that perform data centering, covariance matrix for computation, and eigenvalue decomposition (EVD). Specifically, the proposed accelerator is based on a high-performance matrix multiplication array (MMA). The proposed MMA architecture uses time-multiplexed processing, so that the efficiency of hardware utilization is greatly enhanced. Furthermore, the processing flow utilizes parallel processing, such that the centering, the calculation of the covariance matrix, and the EVD are conducted simultaneously and are individually pipelined to maximize throughput. This article presents the architecture, circuit design, and performance estimates based on post-layout extraction of the proposed preprocessing ICA accelerator. The proposed design achieves a throughput of 40.7 kMatrices/s at a complexity of 73.3 kGE.
Hsi-Hung Lu, Chung-An Shen, Mohamed E. Fouda, Ahmed M. Eltawil
IEEE Trans. Very Large Scale Integr. Syst.3
2021 Cost- and Dataset-free Stuck-at Fault Mitigation for ReRAM-based Deep Learning Accelerators
abstract
Resistive RAMs can implement extremely efficient matrix vector multiplication, drawing much attention for deep learning accelerator research. However, high fault rate is one of the fundamental challenges of ReRAM crossbar array-based deep learning accelerators. In this paper we propose a dataset-free, cost-free method to mitigate the impact of stuck-at faults in ReRAM crossbar arrays for deep learning applications. Our technique exploits the statistical properties of deep learning applications, hence complementary to previous hardware or algorithmic methods. Our experimental results using MNIST and CIFAR-10 datasets in binary networks demonstrate that our technique is very effective, both alone and together with previous methods, up to 20 % fault rate, which is higher than the previous remapping methods. We also evaluate our method in the presence of other non-idealities such as variability and IR drop.
Giju Jung, Mohamed E. Fouda, Sugil Lee, Jongeun Lee, Ahmed M. Eltawil, Fadi J. Kurdahi
DATE2
2021 Fast and Low-Cost Mitigation of ReRAM Variability for Deep Learning Applications
abstract
To overcome the programming variability (PV) of ReRAM crossbar arrays (RCAs), the most common method is program-verify, which, however, has high energy and latency overhead. In this paper we propose a very fast and low-cost method to mitigate the effect of PV and other variability for RCA-based DNN (Deep Neural Network) accelerators. Leveraging the statistical properties of DNN output, our method called Online Batch-Norm Correction (OBNC) can compensate for the effect of programming and other variability on RCA output without using on-chip training or an iterative procedure, and is thus very fast. Also our method does not require a nonideality model or a training dataset, hence very easy to apply. Our experimental results using ternary neural networks with binary and 4-bit activations demonstrate that our OBNC can recover the baseline performance in many variability settings and that our method outperforms a previously known method (VCAM) by large margins when input distribution is asymmetric or activation is multi-bit.
Sugil Lee, Mohamed E. Fouda, Jongeun Lee, Ahmed M. Eltawil, Fadi J. Kurdahi
ICCD2
2020 Learning to Predict IR Drop with Effective Training for ReRAM-based Neural Network Hardware
abstract
Due to the inevitability of the IR drop problem in passive ReRAM crossbar arrays, finding a software solution that can predict the effect of IR drop without the need of expensive SPICE simulations, is very desirable. In this paper, two simple neural networks are proposed as software solution to predict the effect of IR drop. These networks can be easily integrated in any deep neural network framework to incorporate the IR drop problem during training. As an example, the proposed solution is integrated in BinaryNet framework and the test validation results, done through SPICE simulations, show very high improvement in performance close to the baseline performance, which demonstrates the efficacy of the proposed method. In addition, the proposed solution outperforms the prior work on challenging datasets such as CIFAR10 and SVHN.
Sugil Lee, Giju Jung, Mohamed E. Fouda, Jongeun Lee, Ahmed M. Eltawil, Fadi J. Kurdahi
DAC3
2019 Non-Stationary Polar Codes for Resistive Memories
abstract
Resistive memories are considered a promising memory technology enabling high storage densities. However, the readout reliability of resistive memories is impaired due to the inevitable existence of wire resistance, resulting in the sneak path problem. Motivated by this problem, we study polar coding over channels with different reliability levels, termed non-stationary polar codes, and we propose a technique improving the bit error rate (BER) performance. We then apply the framework of non-stationary polar codes to the crossbar array and evaluate its BER performance under two modeling approaches, namely binary symmetric channels and binary asymmetric channels. Finally, we propose a technique for biasing the proportion of high-resistance states in the crossbar array and show its advantage in reducing further the BER. Several simulations are carried out using a SPICE-like simulator, exhibiting significant reduction in BER.
Marwen Zorgui, Mohamed E. Fouda, Zhiying Wang 0001, Ahmed M. Eltawil, Fadi J. Kurdahi
GLOBECOM2
2019 Simple MOS Transistor-Based Realization of Fractional-Order Capacitors
abstract
A new second-order MOS transistor based circuit block approximating the behavior of a fractional-order capacitor is proposed. The circuit is modular and therefore the order of the approximation can be increased by more stages of the same circuit in cascade or in parallel. Simulation results using a TSMC 65nm CMOS technology are provided and show less than 2° of phase error in two decades around the center frequency of the approximation. Experimental results of realized fractional-order capacitors and of a fractional-order relaxation oscillator are also shown.
Mohamed E. Fouda, Ahmed AboBakr, Ahmed S. Elwakil, Ahmed Gomaa Radwan, Ahmed M. Eltawil
ISCAS1
2018 Conditions and Emulation of Double Pinch-off Points in Fractional-order Memristor
abstract
Recently, double pinch-off points have been discovered in some memristive devices where the I-V hysteresis curve intersects in two points generating triple lobes. This paper investigates a fractional-order flux-controlled mathematical model which is able to develop the multiple pinch-off points or multiple lobes. The conditions for observing double pinch-off points (triple lobes) are derived in addition to the locations of the pinch-off points which do not appear in the integer domain. Also, expressions for maximum and minimum conductance are derived. Finally, a floating fractional flux-controlled memristor emulator circuit to generate the triple lobes is introduced and discussed. The PSICE results and the mathematical model results are matched.
Esraa M. Hamed, Mohamed E. Fouda, Ahmed Gomaa Radwan
ISCAS2
2016 Process variations-aware resistive associative processor design
abstract
Recent breakthroughs in memristive devices have demonstrated the potential of using resistive content addressable memories for associative processing. These architectures enable ultra-high density integrated circuits along with low-power computation. However, the reliability of memristive elements is limiting the widespread adoption of these architectures. In this study, we address the reliability issues that arise in high density, resistive associative processor architectures. We propose methods to design process variation immune resistive content addressable memories and minimize the error probabilities. According to SPICE-based circuit simulations, the reliability of the circuit increases significantly and thus positively influences the accuracy of arithmetic operations as well.
Hasan Erdem Yantir, Mohamed E. Fouda, Ahmed M. Eltawil, Fadi J. Kurdahi
ICCD2