Shady O. Agwa

dblp:87/9861 · also Shady Agwa · DBLP profile ↗
← Back
15ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0002-6678-6283ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 3 first-author · 10 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DiP: A Scalable, Energy-Efficient Systolic Array for Matrix Multiplication Acceleration
abstract
Transformers are gaining increasing attention across Natural Language Processing (NLP) application domains due to their outstanding accuracy. However, these data-intensive models add significant performance demands to the existing computing architectures. Systolic array architectures, adopted by commercial AI computing platforms like Google TPUs, offer energy-efficient data reuse but face throughput and energy penalties due to input-output synchronization via First-In-First-Out (FIFO) buffers. This paper proposes a novel scalable systolic array architecture featuring Diagonal-Input and Permutated weight stationary (DiP) dataflow for matrix multiplication acceleration. The proposed architecture eliminates the synchronization FIFOs required by state-of-the-art weight stationary systolic arrays. Beyond the area, power, and energy savings achieved by eliminating these FIFOs, DiP architecture maximizes the computational resource utilization, achieving up to 50% throughput improvement over conventional weight stationary architectures. Analytical models are developed for both weight stationary and DiP architectures, including latency, throughput, time to full PEs utilization (TFPU), and FIFOs overhead. A comprehensive hardware design space exploration using 22nm commercial technology demonstrates DiP’s scalability advantages, achieving up to a$2.02\times $improvement in energy efficiency per area. Furthermore, DiP outperforms TPU-like architectures on transformer workloads from widely-used models, delivering energy improvement up to$1.81\times $and latency improvement up to$1.49\times $. At a$64\times 64$size with 4096 PEs, DiP achieves a peak throughput of 8.192 TOPS with energy efficiency 9.548 TOPS/W.
Shady O. Agwa, Themistoklis Prodromakis
IEEE Trans. Circuits Syst. I Regul. Pap.2
2026 A Hybrid Edge Classifier: Combining TinyML-Optimised CNN With RRAM-CMOS ACAM for Energy-Efficient Inference
abstract
In recent years, the development of smart edge computing systems to process information locally is on the rise. Many near-sensor machine learning (ML) approaches have been implemented to introduce accurate and energy efficient template matching operations in resource-constrained edge sensing systems, such as wearables. To introduce novel solutions that can be viable for extreme edge cases, hybrid solutions combining conventional and emerging technologies have started to be proposed. Deep Neural Networks (DNN) optimised for edge application alongside new approaches of computing (both device and architecture -wise) could be a strong candidate in implementing edge ML solutions that aim at competitive accuracy classification while using a fraction of the power of conventional ML solutions. In this work, we are proposing a hybrid software-hardware edge classifier aimed at the extreme edge near-sensor systems. The classifier consists of two parts: (i) an optimised digital tinyML network, working as a front-end feature extractor, and (ii) a back-end RRAM-CMOS analogue content addressable memory (ACAM), working as a final stage template matching system. The combined hybrid system exhibits a competitive trade-off in accuracy versus energy metric with$E_{front-end}$=$96.23 nJ$and$E_{back-end}$=$1.45 nJ$for each classification operation compared with 78.06 μJ for the original teacher model, representing a 792-fold reduction, making it a viable solution for extreme edge applications.
Kieran Woodward, Eiman Kanjo, Georgios Papandroulidakis, Shady O. Agwa, Themistoklis Prodromakis
IEEE Trans. Knowl. Data Eng.4
2025 RRAM-Based Analogue Artificial Neuron for Gaussian Activation Function Edge Classifier
abstract
The computational bottlenecks encountered for memory access and data transfer in modern computing systems, especially for edge computing, require innovations both at device level and architectural level. Emerging technologies, like Resistive RAM (RRAM), can be used to develop novel analogue circuits and systems used to perform cornerstone machine learning operations in the analogue domain, without requiring data conversion in edge applications. In this work, we present RRAM-based artificial neurons used to implement Gaussian Activation Functions (GAFs) aimed at energy efficient analogue Radial Basis Function NNs (RBFNNs) implementation. We are show-casing a configurable RRAM-CMOS circuit emulating GAF-based neuron designed using a commercially available 180nm CMOS technology and in-house RRAM model, operated at 3.3V and 100MHz while dissipating 140fJ per cell per operation. We aim at using this cell as building block for designing a proof-of-concept memory-centric edge classifier. Through the capability of processing analogue information, the circuit can be used to process analogue information in near-sensor computing paradigms.
Georgios Papandroulidakis, Shady O. Agwa, Themistoklis Prodromakis
ISCAS2
2025 Live Demonstration: Hardware/Software Co-Design to Exploit RRAM Programmability for Emerging Edge Classification Using ArC TWO
abstract
In this demonstration, we present a hardware/software co-design methodology for Convolutional Neural Networks, where the classification section is managed through Resistive RAMs (RRAMs). To this aim, RRAM arrays are mounted onto the ArC TWO instrumentation board, which is interfaced to a laptop. A software Python front-end executes convolutional layers for feature extraction, generates stimuli for RRAMs, and controls the instrumentation board. As a proof of concept, handwritten digits classification is exhibited.
Cristian Sestito, Georgios Papandroulidakis, Patrick Foster, Spyros Stathopoulos, Shady O. Agwa, Themistoklis Prodromakis
ISCAS5
2025 A 9T4R RRAM-Based ACAM for Analogue Template Matching at the Edge
abstract
The continuous shift of computational bottlenecks to the memory access and data transfer, especially for AI applications, poses the urgent needs of re-engineering the computer architecture fundamentals. Many edge computing applications, like wearable and implantable medical devices, introduce increasingly more challenges to conventional computing systems due to the strict requirements of area and power at the edge. Emerging technologies, like Resistive RAM (RRAM), have shown a promising momentum in developing neuro-inspired analogue computing paradigms capable of achieving high classification capabilities alongside high energy efficiency. In this work, we present a novel RRAM-based Analogue Content Addressable Memory (ACAM) for on-line analogue template matching applications. This ACAM-based template matching architecture aims to achieve energy-efficient classification where low energy is of utmost importance. We are showcasing a highly tuneable novel RRAM-based ACAM pixel implemented using a commercial 180 nm CMOS technology and in-house RRAM technology and exhibiting low energy dissipation of approximately 0.036 pJ and 0.16 pJ for mismatch and match, respectively, at 66 MHz with 3.3 V voltage supply. A proof-of-concept system-level implementation based on this novel pixel design is also implemented in 180 nm.
Georgios Papandroulidakis, Shady O. Agwa, Ahmet Cirakoglu, Themistoklis Prodromakis
IEEE Trans. Circuits Syst. I Regul. Pap.2
2025 TrIM, Triangular Input Movement Systolic Array for Convolutional Neural Networks: Architecture and Hardware Implementation
abstract
Modern hardware architectures for Convolutional Neural Networks (CNNs), other than targeting high performance, aim at dissipating limited energy. Reducing the data movement cost between the computing cores and the memory is a way to mitigate the energy consumption. Systolic arrays are suitable architectures to achieve this objective: they use multiple processing elements that communicate each other to maximize data utilization, based on proper dataflows like the weight stationary and row stationary. Motivated by this, we have proposed TrIM, an innovative dataflow based on a triangular movement of inputs, and capable to reduce the number of memory accesses by one order of magnitude when compared to state-of-the-art systolic arrays. In this paper, we present a TrIM-based hardware architecture for CNNs. As a showcase, the accelerator is implemented onto a Field Programmable Gate Array (FPGA) to execute the VGG-16 and AlexNet CNNs. The architecture achieves a peak throughput of 453.6 Giga Operations per Second, outperforming a state-of-the-art row stationary systolic array up to$\sim 3 \times$in terms of memory accesses, and being up to$ \sim 11.9 \times$more energy-efficient than other FPGA accelerators.
Cristian Sestito, Shady O. Agwa, Themistoklis Prodromakis
IEEE Trans. Circuits Syst. I Regul. Pap.2
2024 An Energy-Efficient Capacitive-RRAM Content Addressable Memory
abstract
Content addressable memory is popular in intelligent computing systems as it allows parallel content-searching in memory. Emerging CAMs show a promising increase in bitcell density and a decrease in power consumption than pure CMOS solutions. This article introduced an energy-efficient 3T1R1C TCAM cooperating with capacitor dividers and RRAM devices. The RRAM as a storage element also acts as a switch to the capacitor divider while searching for content. CAM cells benefit from working parallel in an array structure. We implemented a$64\times 64$array and digital controllers to perform with an internal built-in clock frequency of 875MHz. Both data searches and reads take three clock cycles. Its worst average energy for data match is reported to be 1.71fJ/bit-search and the worst average energy for data miss is found at 4.69fJ/bit-search. The prototype is simulated and fabricated in 0.18um technology with in-lab RRAM post-processing. Such memory explores the charge domain searching mechanism and can be applied to data centers that are power-hungry.
Yihan Pan 0003, Adrian Wheeldon, Mohammed Mughal, Shady O. Agwa, Themistoklis Prodromakis, Alexander Serb
IEEE Trans. Circuits Syst. I Regul. Pap.4
2023 EVE: Ephemeral Vector Engines
abstract
There has been a resurgence of interest in vector architectures evident by recent adoption of vector extensions in mainstream instruction set architectures. Traditionally, vector engines leverage this abstraction by exploiting its inherent regularity to increase performance and efficiency. Recent work on SRAM-based compute-in-memory has shown promise in reducing the area overhead of these engines. In this work, we propose ephemeral vector engines (EVE) where we leverage SRAM-based compute-in-memory techniquesas well as bit-peripheral computations to facilitate efficient vector execution. EVE uses a novel approach of bit-hybrid execution, striking a balance between throughput and latency. Evaluated on the Rodinia and RiVEC benchmark suites, EVE achieves almost 8× speed-up compared to an out-of-order processor and 4.59× compared to an integrated vector unit. EVE achieves speed-ups comparable to an aggressive decoupled vector unit and increases the area-normalized performance by over 2 ×. By repurposing SRAM arrays in the L2 cache to create ephemeral vector execution units, EVE is able to efficiently achieve high performance while incurring as little as 11.7% area overhead.
Khalid Al-Hawaj, Tuan Ta, Nick Cebry, Shady O. Agwa, Olalekan Afuye, Eric Hall, Courtney Golden, Alyssa B. Apsel, Christopher Batten
HPCA4
2023 A 1T1R+2T Analog Content-Addressable Memory Pixel for Online Template Matching
abstract
The template matching approach has a promising momentum to build energy-efficient edge classifiers for var-ious implantable and wearable medical devices. To mitigate the analog/digital cross-domain interfacing complexity, analog content-addressable memories can be used efficiently to form the back-end classifiers by receiving the analog inputs and generating the digital classification outputs. This paper presents a novel memristor-based analog content-addressable memory pixel 1TIR+2T with 2.0x smaller footprint than its counterparts in the literature. The new compact pixel utilizes only one RRAM device through a 1T1R voltage divider circuit while exploiting the complementary behavior of the nMOS and pMOS transistors to determine the lower and upper bounds of the matching voltage range. The simulation results show that the 1T1R+2T pixel has a promising tunability with matching windows range from 50 mV to 200 mV according to the RRAM resistance value of the 1T1R voltage divider.
Shady O. Agwa, Georgios Papandroulidakis, Themistoklis Prodromakis
ISCAS1
2022 High-Density Digital RRAM-based Memory with Bit-line Compute Capability
abstract
The AI revolution shows the ever increasing performance demands of AI applications like Deep Neural Networks DNNs which consist of tens of layers and do computations on tens of millions of data weights [1]. Conventional Von Neumann architectures are currently struggling to meet these emerging performance demands with deep memory hierarchies to bridge the processor-memory performance gap [2]. Emerging technologies (like RRAMs) have meanwhile shown a real promise to address the increasing challenges of the conventional computing technology. While the main direction of research is focussing on exploiting the analogue memory attributes of RRAMs specially for analogue computing crossbars [3], this paper focuses on a different perspective of building high-density and digital-friendly RRAM-based memory that is a good alternative to the SRAM-based Last-Level Caches LLCs. This digital RRAM-based memory with conventional 1T1R bit-cells is proposed to be an on-chip gigantic data reservoir, with much higher density than SRAMs, to bridge the memory gap. The paper also shows that the digital RRAM-based memory is capable of doing robust bit-line compute which opens the door for digital in-memory computing architectures that can mitigate the Von Neumann bottleneck while adopting RRAM’s high-density promise. Unlike analogue RRAM crossbars, RRAMs’ digital in-memory computing capability should inherit the large scalability and the fast time-to-market of the digital domain with less engineering effort for optimisation as there is no need any more to build DACs and ADCs.
Shady O. Agwa, Yihan Pan 0003, Thomas Abbey, Alexander Serb, Themistoklis Prodromakis
ISCAS1
2022 A CMOS-based Characterisation Platform for Emerging RRAM Technologies
abstract
Mass characterisation of emerging memory devices is an essential step in modelling their behaviour for integration within a standard design flow for existing integrated circuit designers. This work develops a novel characterisation platform for emerging resistive devices with a capacity of up to 1 million devices on-chip. Split into four independent sub-arrays, it contains on-chip column-parallel DACs for fast voltage programming of the DUT. On-chip readout circuits with ADCs are also available for fast read operations covering 5-decades of input current (20nA to 2mA). This allows a device’s resistance range to be between 1k$\Omega$ and 10M$\Omega$ with a minimum voltage range of ±1.5V on the device.
Andrea Mifsud, Peilong Feng, Lijie Xie, Chaohan Wang, Yihan Pan 0003, Sachin Maheshwari, Shady O. Agwa, Spyros Stathopoulos, Shiwei Wang 0001, Alexander Serb, Christos Papavassiliou, Themistoklis Prodromakis, Timothy G. Constandinou
ISCAS8
2020 Towards a Reconfigurable Bit-Serial/Bit-Parallel Vector Accelerator using In-Situ Processing-In-SRAM
abstract
Vector accelerators can efficiently execute regular data-parallel workloads, but they require expensive multi-ported register files to feed large vector ALUs. Recent work on in-situ processing-in-SRAM shows promise in enabling area-efficient vector acceleration. This work explores two different approaches to leveraging in-situ processing-in-SRAM: BS-VRAM, which uses bit-serial execution, and BP-VRAM, which uses bit-parallel execution. The two approaches have very different latency vs. throughput trade-offs. BS-VRAM requires more cycles per operation, but is able to execute thousands of operations in parallel, while BP-VRAM requires fewer cycles per operation, but can only execute hundreds of operations in parallel. This paper is the first work to perform a rigorous evaluation of bit-serial vs. bit-parallel in-situ processing-in-SRAM. Our results show that both approaches have similar area overheads. For 32-bit arithmetic operations, BS-VRAM improves throughput by 1.3-5.0× compared to BP-VRAM, while BP-VRAM improves latency by 3.0-23.0× compared to BS-VRAM.
Khalid Al-Hawaj, Olalekan Afuye, Shady O. Agwa, Alyssa B. Apsel, Christopher Batten
ISCAS3
2020 Implementing Low-Diameter On-Chip Networks for Manycore Processors Using a Tiled Physical Design Methodology
abstract
Manycore processors are now integrating up to 1000 simple cores into a single die, yet these processors still rely on high-diameter mesh on-chip networks (OCNs) without complex flow-control nor custom circuits due to three reasons: (1) manycores require simple, low-area routers; (2) manycores usually use standard-cell-based design; and (3) manycores use a tiled physical design methodology. In this paper, we explore mesh and torus topologies with internal concentration and/or ruche channels that require low area overhead and can be implemented using a traditional standard-cell-based tiled physical design methodology. We use a combination of analytical and RTL modeling along with layout-level results for both hard macros and a 3×3mm 256-terminal OCN in a 14-nm technology for twelve topologies. Critically, the networks we study use a tiled physical design methodology meaning they: (1) tile a homogeneous hard macro across the chip; (2) implement chip top-level routing between hard macros via short wires to neighboring macros; and (3) use timing closure for the hard macro to quickly close timing at the chip top-level. Our results suggest that a concentration factor of four and a ruche factor of two in a 2D-mesh topology can reduce latency by over 2× at similar area and bisection bandwidth for both small and large messages compared to a 2D-mesh baseline.
Yanghui Ou, Shady O. Agwa, Christopher Batten
NOCS2
2019 PyOCN: A Unified Framework for Modeling, Testing, and Evaluating On-Chip Networks
abstract
There is a growing interest in the open-source hardware movement to amortize non-recurring engineering costs by using plug-and-play system-on-chip (SoC) designs, where the communication among different components is provided by an on-chip interconnection network. Unfortunately, building an on-chip network (OCN) that is suitable for a specific SoC design requires the exploration of a large number of design options and involves diverse research methodologies to evaluate performance, area, energy, and timing. In this paper, we propose PyOCN, a unified framework that vertically integrates multiple research methodologies to enable productively exploring the OCN design space. PyOCN is the first comprehensive framework for modeling (e.g., functional-level, cycle-level, and register-transfer-level), testing (e.g., unit testing, integration testing, and property-based random testing), and evaluating (e.g., simulating, generating, and characterizing) on-chip interconnection networks. We use a case study based on a 64-terminal butterfly network to illustrate the key features of PyOCN and to demonstrate the framework's potential in productively modeling, testing, and evaluating OCNs.
Cheng Tan 0002, Yanghui Ou, Shunning Jiang, Peitian Pan, Christopher Torng, Shady O. Agwa, Christopher Batten
ICCD6
2017 Power efficient AES core for IoT constrained devices implemented in 130nm CMOS
abstract
The Internet of Things (IoT) constrained devices show the urgent need for low power data security hardware cores. This paper presents a power efficient AES Core fabricated in UMC 130 nm CMOS technology by using Faraday standard cells library. The maximum throughput of the proposed AES Core is up to 2.6 Gb/s consuming about 0.2148 mW/MHz at 1.2V. The Dynamic Voltage and Frequency Scaling (DVFS) technique is applied to reduce the power consumption of the AES Core. The experimental measurements show about 3x reduction in power consumption, consuming about 0.0697 mW/MHz by scaling the supply voltage from 1.2 V to 0.7 V.
Shady O. Agwa, Eslam Yahya, Yehea I. Ismail
ISCAS1