Kashif Inayat

dblp:232/9403 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0001-5504-6274ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 SAPER-AI accelerator: a systolic array-based power-efficient reconfigurable AI accelerator
abstract
Deep learning (DL) accelerators are critical for handling the growing computational demands of modern neural networks. Systolic array (SA)-based accelerators consist of a 2D mesh of processing elements (PEs) working cooperatively to accelerate matrix multiplication. The power efficiency of such accelerators is of primary importance, especially considering the edge AI regime. This work presents the SAPER-AI accelerator, an SA accelerator with power intent specified via a unified power format representation in a simplified manner with negligible microarchitectural optimization effort. Our proposed accelerator switches off rows and columns of PEs in a coarse-grained manner, thus leading to SA microarchitecture complying with the varying computational requirements of modern DL workloads. Our analysis demonstrates enhanced power efficiency ranging between 10% and 25% for the best case 32×32 and 64×64 SA designs, respectively. Additionally, the power delay product (PDP) exhibits a progressive improvement of around 6% for larger SA sizes. Moreover, a performance comparison between the MobileNet and ResNet50 models indicates generally better SA performance for the ResNet50 workload. This is due to the more regular convolutions portrayed by ResNet50 that are more favored by SAs, with the performance gap widening as the SA size increases.
Fahad Bin Muslim, Kashif Inayat, Muhammad Zain Siddiqi, Safiullah Khan, Tayyeb Mahmood, Ihtesham Ul Islam
Frontiers Inf. Technol. Electron. Eng.2
2024 FPGA-assisted Design Space Exploration of Parameterized AI Accelerators: A Quickloop Approach
Kashif Inayat, Fahad Bin Muslim, Tayyeb Mahmood, Jaeyong Chung
J. Syst. Archit.1
2024 Factored Systolic Arrays Based on Radix-8 Multiplication for Machine Learning Acceleration
abstract
Systolic arrays (SAs) are re-gaining the attention as the heart to accelerate machine learning workloads. This article shows that a large design space exists at the logic level despite the simple structure of SAs and proposes two novel SAs based on factoring and radix-$8$multipliers: The first factored SA (FSA) extracts out the booth encoding and the hard-multiple generation which is common across all processing elements (PEs), reducing the delay and the area of the whole SA. This factoring is done at the cost of an increased number of registers; however, the reduced pipeline register requirement in radix-$8$offsets this effect. Our second proposed FSA compresses the interconnections further with two steps of hard-multiple addition. In the first part, carries are computed column-wise outside PEs, and in the second part, early-generated carries are used for hard-multiple final addition inside PEs. We called it hard-multiple carry portioned (HCP) FSA (HCP FSA). The first proposed factored$16$-bit multiplier achieves up to$15$%,$13$%, and$23$% better delay, area, and power, respectively, compared with the radix-$4$multipliers even if the register overhead is included. And first proposed FSA architecture improves delay, area, and power up to$11$%,$20$%, and$31$%, respectively, for different bitwidths when compared with the conventional radix-$4$SA. In addition, the second HCP FSA design eliminates the additional registered overhead associated with the first proposed FSA by reducing interconnections and shows further reductions in the area up to 11.7% and power up to 16.7% with little increase in delay for various sizes of SAs.
Kashif Inayat, Inayat Ullah, Jaeyong Chung
IEEE Trans. Very Large Scale Integr. Syst.1
2023 Quickloop: An Efficient, FPGA-Accelerated Exploration of Parameterized DNN Accelerators
abstract
Quickloop is a design-space exploration (DSE) framework of parameterized RTL generators, their software stack, and their simulation on FPGA. FPGAs are recently accelerating RTL simulations due to their rapid turnaround times (TAT), compared to ASIC. However, this TAT is still restrictive in DSE. We adopt a data-driven approach to optimize Quickloop's TAT and leverage this framework to extensively search the design space of an open source DNN accelerator. We show that our approach effectively slashes the TAT by above 30%, compared to conventional toolflow.
Tayyeb Mahmood, Kashif Inayat, Jaeyong Chung
PACT2
2022 Hybrid Accumulator Factored Systolic Array for Machine Learning Acceleration
abstract
Deep learning applications have become ubiquitous in today’s era and it has led to vast development in machine learning (ML) accelerators. Systolic arrays have been a primary part of ML accelerator architecture. To fully leverage the systolic arrays, it is required to explore the computer arithmetic data-path components and their tradeoffs in accelerators. We present a novel factored systolic array (FSA) architecture, in which the carry propagation adder (CPA) and carry-save adder (CSA) perform hybrid accumulation on least significant bit (LSB) bits and most significant bits (MSB) bits, respectively, inside each processing element. In addition, a small CPA to complete accumulation for MSB bits along with rounding logic for each column of the array is placed, which not only reduces the area, delay, and power but also balances the combinational and sequential area tradeoffs. We demonstrate the hybrid accumulator with partial CPA factoring in “Gemmini,” an open-source practical systolic array accelerator and factoring technique does not change the functionality of the base design. We implemented three baselines, original Gemmini and two variants of it, and show that the proposed approach leads to overall significant reduction in area within the range 12.8% – 50.2% and in power within the range 18.6% – 41% with improved or similar delay in comparison to the baselines.
Kashif Inayat, Jaeyong Chung
IEEE Trans. Very Large Scale Integr. Syst.1
2020 Factored Radix-8 Systolic Array for Tensor Processing
abstract
Systolic arrays are re-gaining the attention as the heart to accelerate machine learning workloads. This paper shows that a large design space exists at the logic level despite the simple structure of systolic arrays and proposes a novel systolic array based on factoring and radix-8 multipliers. The factored systolic array (FSA) extracts out the booth encoding and the hard-multiple generation which is common across all processing elements, reducing the delay and the area of the whole systolic array. This factoring is done at the cost of an increased number of registers, however, the reduced pipeline register requirement in radix-8 offsets this effect. The proposed factored 16-bit multiplier achieves up to 15%, 13%, and 23% better delay, area, and power, respectively, compared with the radix-4 multipliers even if the register overhead is included. The proposed FSA architecture improves delay, area, and power up to 11%, 20% and 31%, respectively, for different bitwidths when compared with the conventional radix-4 systolic array.
Inayat Ullah, Kashif Inayat, Joon-Sung Yang, Jaeyong Chung
DAC2