Mohammed E. Elbtity

dblp:215/7276 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
6since 2021 · last 2023
0000-0002-3282-0076ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 3 first-author · 6 since 2021
YearPublicationVenuePosition
2023 IMAC-Sim: : A Circuit-level Simulator For In-Memory Analog Computing Architectures
abstract
With the increased attention to memristive-based in-memory analog computing (IMAC) architectures as an alternative for energy-hungry computer systems for machine learning applications, a tool that enables exploring their device- and circuit-level design space can significantly boost the research and development in this area. Thus, in this paper, we develop IMAC-Sim, a circuit-level simulator for the design space exploration of IMAC architectures. IMAC-Sim is a Python-based simulation framework, which creates the SPICE netlist of the IMAC circuit based on various device- and circuit-level hyperparameters selected by the user, and automatically evaluates the accuracy, power consumption, and latency of the developed circuit using a user-specified dataset. Moreover, IMAC-Sim simulates the interconnect parasitic resistance and capacitance in the IMAC architectures and is also equipped with horizontal and vertical partitioning techniques to surmount these reliability challenges. IMAC-Sim is a flexible tool that supports a broad range of device- and circuit-level hyperparameters. In this paper, we perform controlled experiments to exhibit some of the important capabilities of the IMAC-Sim, while the entirety of its features is available for researchers via an open-source tool at https://github.com/iCAS-Lab/IMAC-Sim.
Md Hasibul Amin, Mohammed E. Elbtity, Ramtin Zand
ACM Great Lakes Symposium on VLSI2
2023 Heterogeneous Integration of In-Memory Analog Computing Architectures with Tensor Processing Units
abstract
Tensor processing units (TPUs), specialized hardware accelerators for machine learning tasks, have shown significant performance improvements when executing convolutional layers in convolutional neural networks (CNNs). However, they struggle to maintain the same efficiency in fully connected (FC) layers, leading to suboptimal hardware utilization. In-memory analog computing (IMAC) architectures, on the other hand, have demonstrated notable speedup in executing FC layers. This paper introduces a novel, heterogeneous, mixed-signal, and mixed-precision architecture that integrates an IMAC unit with an edge TPU to enhance mobile CNN performance. To leverage the strengths of TPUs for convolutional layers and IMAC circuits for dense layers, we propose a unified learning algorithm that incorporates mixed-precision training techniques to mitigate potential accuracy drops when deploying models on the TPU-IMAC architecture. The simulations demonstrate that the TPU-IMAC configuration achieves up to 2.59× performance improvements, and 88% memory reductions compared to conventional TPU architectures for various CNN models while maintaining comparable accuracy. The TPU-IMAC architecture shows potential for various applications where energy efficiency and high performance are essential, such as edge computing and real-time processing in mobile devices. The unified training algorithm and the integration of IMAC and TPU architectures contribute to the potential impact of this research on the broader machine learning landscape.
Mohammed E. Elbtity, Brendan Reidy, Md Hasibul Amin, Ramtin Zand
ACM Great Lakes Symposium on VLSI1
2023 Work in Progress: Real-time Transformer Inference on Edge AI Accelerators
abstract
Transformer models have become a dominant architecture in the world of machine learning. From natural language processing to more recent computer vision applications, Transformers have shown remarkable results and established a new state-of-the-art in many domains. However, this increase in performance has come at the cost of ever-increasing model sizes requiring more resources to deploy. Machine learning (ML) models are used in many real-world systems, such as robotics, mobile devices, and internet of things (IoT) devices, that require fast inference with low energy consumption. For batterypowered devices, lower energy consumption directly translates into longer battery life. To address these issues, several edge AI accelerators have been developed. Among these, the Coral Edge TPU has shown promising results for image classification while maintaining very low energy consumption. Many of these devices, including the Coral TPU, were originally designed to accelerate convolutional neural networks, making deployment of Transformers challenging. Here, we propose a methodology to deploy Transformers on Edge TPU. We provide extensive latency, power, and energy comparisons among the leading edge devices and show that our methodology allows for real-time inference of Transformers while maintaining the lowest power and energy consumption of other edge devices on the market.
Brendan Reidy, Mohammadreza Mohammadi, Mohammed E. Elbtity, Heath Smith, Ramtin Zand
RTAS3
2022 MRAM-based Analog Sigmoid Function for In-memory Computing
abstract
We propose an analog implementation of the transcendental activation function leveraging two spin-orbit torque magnetoresistive random-access memory (SOT-MRAM) devices and a CMOS inverter. The proposed analog neuron circuit consumes 1.8-27x less power, and occupies 2.5-4931x smaller area, compared to the state-of-the-art analog and digital implementations. Moreover, the developed neuron can be readily integrated with memristive crossbars without requiring any intermediate signal conversion units. The architecture-level analyses show that a fully-analog in-memory computing (IMC) circuit that use our SOT-MRAM neuron along with an SOT-MRAM based crossbar can achieve more than 1.1x, 12x, and 13.3x reduction in power, latency, and energy, respectively, compared to a mixed-signal implementation with analog memristive crossbars and digital neurons. Finally, through cross-layer analyses, we provide a guide on how varying the device-level parameters in our neuron can affect the accuracy of multilayer perceptron (MLP) for MNIST classification.
Md Hasibul Amin, Mohammed E. Elbtity, Mohammadreza Mohammadi, Ramtin Zand
ACM Great Lakes Symposium on VLSI2
2022 Interconnect Parasitics and Partitioning in Fully-Analog In-Memory Computing Architectures
abstract
Fully-analog in-memory computing (IMC) architectures that implement both matrix-vector multiplication and nonlinear vector operations within the same memory array have shown promising performance benefits over conventional IMC systems due to the removal of energy-hungry signal conversion units. However, maintaining the computation in the analog domain for the entire deep neural network (DNN) comes with potential sensitivity to interconnect parasitics. Thus, in this paper, we investigate the effect of wire parasitic resistance and capacitance on the accuracy of DNN models deployed on fully-analog IMC architectures. Moreover, we propose a partitioning mechanism to alleviate the impact of the parasitic while keeping the computation in the analog domain through dividing large results for a $400 \times 120 \times 84 \times 10$ DNN model deployed on a results for a $400 \times 120 \times 84 \times 10$ DNN model deployed on a fully-analog IMC circuit show that a 94.84 % accuracy could be achieved for MNIST classification application with 16,8, and 8 horizontal partitions, as well as 8,8, and 1 vertical partitions for first, second, and third layers of the DNN, respectively, which is comparable to the $\sim 97$ % accuracy realized by digital implementation on CPU. It is shown that accuracy benefits are extra circuitry required for handling partitioning.
Md Hasibul Amin, Mohammed E. Elbtity, Ramtin Zand
ISCAS2
2022 APTPU: Approximate Computing Based Tensor Processing Unit
abstract
We propose an approximate tensor processing unit (APTPU), which includes two main components: (1) approximate processing elements (APEs) consisting of a low-precision multiplier and an approximate adder, and (2) pre-approximate units (PAUs) which are shared among the APEs in the APTPU’s systolic array, functioning as the steering logic to pre-process the operands and feed them to the APEs. We conduct extensive experiments to evaluate the performance of the APTPU across various configurations and various workloads. The results show that the APTPU’s systolic array achieves up to$5.2\times \textit {TOPS}/mm^{2}$and$4.4\times \textit {TOPS}/W$improvements compared to that of a conventional systolic array design. The comparison between the proposed APTPU and in-house TPU designs shows that we can achieve approximately$2.5\times $and$1.2\times $area and power reduction, respectively, while realizing comparable accuracy. Finally, a comparison with the state-of-the-art approximate systolic arrays shows that the APTPU can realize up to$1.58\times $,$2\times $, and$1.78\times $, reduction in delay, power, and area, respectively, while using similar design specifications and synthesis constraints.
Mohammed E. Elbtity, Peyton Chandarana, Brendan Reidy, Jason Kamran Eshraghian, Ramtin Zand
IEEE Trans. Circuits Syst. I Regul. Pap.1
2018 Small Area and Low Power Hybrid CMOS-Memristor Based FIFO for NoC
abstract
Area and power consumption are the main challenges in Network on Chip (NoC). Indeed, First Input First Output (FIFO) memory is the key element in NoC. Increasing the FIFO depth, produces an increas in the performance of NoC but at the cost of area and power consumption. This paper proposes a new hybrid CMOS-Memristor based FIFO architecture that consumes low power and has a small size compared to the conventional CMOS-based FIFOs. The predicted area is approximately equal to the half of that wasted in conventional FIFOs. The implementation of FIFO controller module is implemented using HDL. Moreover, the functionality test and the simulation results of the proposed architecture are presented. Simulation is done using ISF Xilinix and Cadence tools.
Mohammed E. Elbtity, Ahmed Gomaa Radwan
ISCAS1