EDBT 2026 Demo / reviewers in the wild / expert
Aman Arora 0001
dblp:75/912-1
· DBLP profile ↗
37ranked-venue papers
10as first author
34since 2021 · last 2026
0000-0003-2547-4424ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 34 · 10 first-author · 31 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Accelerating Topology Optimization on AMD Versal AIE-ML EnginesabstractTopology optimization is a computational method used to determine the optimal material distribution within a prescribed design domain, aiming to minimize structural weight while satisfying load and boundary conditions. For critical infrastructure applications, such as real-time structural health monitoring of bridges and buildings, achieving real-time topology optimization is essential. Traditionally, topology optimization relies on finite element analysis (FEA), a computationally intensive process. Recent advances in deep neural networks (DNNs) have introduced data driven alternatives to FEA, substantially reducing computation time while maintaining solution quality. However, the inference latency of these neural network models, especially for complex designs, still remains a major obstacle to real-time deployment. To address this challenge, we propose a hardware accelerated implementation of a topology optimization neural network (CRONet) on the AMD Versal AI Engine-ML (AIE-ML) architecture. Our approach efficiently exploits the parallelism and memory hierarchy of AIE-ML engines to optimize the execution of various neural network operators. We develop custom, high-performance kernels for all network layers, including a modified version of the state-of-the-art GAMA framework for generalized matrix multiplication (GEMM) layers. Additionally, we offer integrated tools for performance simulation and functional verification prior to hardware deployment. Experimental results demonstrate that our implementation achieves 2× higher performance compared to an NVIDIA GTX 1050 Ti and up to 14× over a CPU baseline. These results highlight the potential of Versal AIE-ML based acceleration for enabling real-time topology optimization. Kaustubh Manohar, Vedant Tewari, Aditya Ray, Ridwan Olabiyi, Ashif Iquebal, Aman Arora 0001 |
FCCM | 7 |
| 2026 | MARU: An ML-Based Framework for Area Estimation from FPGA Resource UsageabstractFPGA design evaluation faces significant challenges due to heterogeneous resource reporting across vendors and architectures, which hinders performance comparisons and complicates design space exploration. We present MARU (Machine-learning for Area estimation from Resource Usage), a framework that enables FPGA and ASIC area estimation directly from HLS utilization reports. MARU achieves a mean absolute percentage error of 1.5% and an R² of 0.997 in cross-FPGA, multi-circuit predictions, while reducing estimation time from 2.5 days to just 5 minutes per design. The framework enables three key advancements: 1) unified, area-based estimation across diverse FPGA designs, 2) accelerated HLS design space exploration through real-time prediction, and 3) ASIC migration analysis by estimating equivalent area for predictive cost modeling. MARU bridges the FPGA/ASIC methodology gap by providing both a comparative framework for academic research and a practical tool for industry-scale co-design decisions—whether optimizing FPGA implementations or evaluating ASIC transition feasibility through predictive area modeling. Tarun Kholay, Anup Ashok Kedilaya, Aman Arora 0001, Jaydeep P. Kulkarni, Lizy Kurian John |
FPGA | 3 |
| 2026 | Closing the Loop on FPGA Verification: An Iterative Framework for Maximizing Routing Resource CoverageabstractPre-silicon verification of Field-Programmable Gate Array (FPGA) architectures and their supporting Computer-Aided Design (CAD) tools presents a significant challenge. While static benchmark suites fail to provide adequate coverage of the vast routing fabric, the application of formal verification is computationally not practical at the scale of modern FPGAs. To address this, we present an automated, iterative framework designed to systematically maximize routing resource graph (rr_graph) coverage. Our methodology employs a three-level, coarse-to-fine-grained heuristic engine that generates placement constraints and designs to strategically target all resources aiming for 100% coverage, while supporting user-defined exclusions. In the first level, the engine prioritizes coarse-grained logic resource coverage by targeting uncovered grid locations with diverse micro-benchmarks. In the second level, it maximizes interconnect utilization by forcing routing-intensive designs into identified routing coldspots. Finally, in the third level, it performs fine-grained targeting of the remaining hard-to-reach routing resources by generating specialized micro-designs with precise locking and stretching constraints. This approach not only provides a robust methodology for verifying FPGA fabrics and CAD tool correctness but also enables the quantitative analysis of FPGA architectures, systematically identifying unreachable or inefficiently designed resources and providing actionable feedback to architects before tape-out. Ruthwik Reddy Sunketa, Aman Arora 0001 |
FPGA | 2 |
| 2026 | XPNet: Cross-FPGA Power Prediction From High-Level Language CodeabstractMachine learning (ML) has been successfully employed to estimate power consumption for FPGAs using features derived from the results of High Level Synthesis (HLS). However, such models trained on one FPGA cannot be directly applied to another FPGA, even within the same FPGA series. Training a model for a new FPGA is time-consuming due to the significant effort required for dataset preparation. Researchers have to invest significant effort (weeks) in constructing a sufficient dataset with power value annotations, to train an accurate model for a new target FPGA. Another challenge is that existing model construction methods depend on many features extracted late in the HLS process, which are tool-specific and cannot be transferred between tools from different vendors. To address these challenges, we propose a novel cross-FPGA power modeling methodology called XPNet. With only frontend features from HLS, XPNet combines Transfer-Learning with innovative data selection techniques that enable efficient fine-tuning for a new target FPGA. With XPNet, models trained on one FPGA can be quickly adapted to a new target FPGA and used to efficiently predict the power on this new FPGA with high accuracy. Experiments with Polybench, Machsuite and CHStone demonstrate an average error of only 8.40% (10.34% if cross-vendor) when less than 1% of designs are used for the fine-tuning to the new target FPGA. In comparison to best prior model (with full training and 5.74% error), XPNet yields 232x speed up in dataset preparation and training, and 5x speed up in inference on a new FPGA. Zhigang Wei, Allison Seigler, Sean Lowe, Emily Shriver, Aman Arora 0001, Lizy Kurian John |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2026 | OpenFPGA-NoC: Automated Fabric and Bitstream Generation for NoC-based FPGAsabstractAs the demand for high-performance and flexible hardware accelerators increases, Network-on-Chip (NoC)-based Field Programmable Gate Arrays (FPGAs) offer a scalable solution for complex, data-intensive applications. While commercial FPGA vendors like Xilinx, Altera, and Achronix offer hardened NoCs in their flagship architectures, there are no academic or open source FPGAs with embedded NoCs. Although many open source soft NoC implementations exist, they pose challenges like high resource utilization and low frequencies, making them unsuitable for high-performance applications. Vendor-supplied FPGAs with fixed hard NoC topologies may not satisfy the requirements of new application domains, motivating the need to enable automatic design of customized NoC-based FPGAs. To address this need, we introduce OpenFPGA-NoC, an automated flow that generates fabric netlists and bitstreams for NoC-based FPGAs. Our work extends the OpenFPGA framework by adding an NoC-specific tag in the OpenFPGA architecture description, supporting custom configuration ports to handle address mapping of NoC routers, automating the generation of architecture files, and enabling a custom RTL-to-bitstream flow. OpenFPGA-NoC provides an easy-to-use interface that allows the FPGA architect to exploit the flexibility provided by the framework- providing NoC parameters like topology, number of routers, and key router parameters like data widths and buffer depths. By providing push-button flows, OpenFPGA-NoC significantly lowers the barrier to designing high-performance FPGA fabrics. Ruthwik Reddy Sunketa, Ganesh Gore, Allen Boston, Pierre-Emmanuel Gaillardon, Aman Arora 0001 |
ACM Trans. Reconfigurable Technol. Syst. | 6 |
| 2025 | Compute-In-Memory on FPGAs for Deep Learning: A ReviewabstractField-Programmable Gate Arrays (FPGAs) are increasingly recognized as an efficient platform for accelerating Deep Learning (DL) applications. This is due to their hardware-level configurability, which allows for tailored datapaths and low-precision inference capabilities. As data-intensive applications such as DL continue to grow, there is a renewed interest in Compute-In-Memory (CIM) architectures to address the bottle-neck caused by moving large volumes of data between memory blocks (off-chip or on-chip) and compute units. This paper delves deep into CIM architectures designed for FPGAs, focusing on making FPGAs better accelerators of DL applications. This paper provides an overview of CIM and FPGA architecture, motivating the need to design custom FPGA architectures with CIM capability. Detailed insights into CIM operation on FPGAs and the implementation of CIM on FPGAs are presented, focusing on modifications to the memory blocks on FPGAs. Existing FPGA CIM proposals are categorized and the nuances of architecting CIM blocks for FPGAs are explained. Lastly, it outlines the current challenges that must be addressed to facilitate broader adoption of CIM architectures on FPGAs. Aman Arora 0001 |
FCCM | 1 |
| 2025 | Analog In-Memory Computing Enhanced FPGA for High-Throughput and Energy-Efficient AccelerationabstractThe ever-growing demand for AI computing, coupled with slowing performance gains in chip manufacturing, has heightened the role of FPGA-based accelerators. FPGAs enable the implementation of application-customized parallel dataflows due to their reconfigurability, achieving high energy efficiency. However, the bit-level routing fabric on FPGAs often results in high overheads because large amounts of data must be shuttled between compute blocks and memory blocks on the FPGA. We propose enhancing FPGAs with in-memory computing macros, specifically analog Dot Product Engines based on non-volatile RRAM devices. Using the Verilog to Routing (VTR) framework, we simulate a novel 40 nm, 26.2 mm × 26.2 mm architecture and employ a custom event-driven simulator to evaluate its performance. Our design achieves 25.5 ×103TOPS/W, an average ×31.4 throughput improvement and an average ×9,380 energy efficiency improvement when compared to state-of-the-art FPGA implementations of AI models. Archit Gajjar, Omar Eldash, Aishwarya Natarajan, Xia Sheng, Giacomo Pedretti, Aman Arora 0001, Paolo Faraboschi, Jim Ignowski, Luca Buonanno |
FCCM | 7 |
| 2025 | High Throughput Low Latency Network Intrusion Detection on FPGAs: A Raw Packet Approach
Abid Rafique, Suhaib A. Fahmy, Aman Arora 0001 |
FPGA | 4 |
| 2025 | Performance Analysis of GEMM Workloads on the AMD Versal PlatformabstractAMD Versal is a new heterogeneous computing hardware architecture comprised of adaptive intelligence (AI) engines, programmable logic, and a processing system. General Matrix Multiplication (GEMM) is the fundamental building block of modern deep learning (DL) applications such as ChatGPT, and GEMM workloads can be mapped onto Versal in different ways, each with distinct trade-offs. This paper presents a thorough analysis of GEMM workloads of different shapes and sizes, showcasing performance artifacts associated with the AMD Versal architecture. Focusing on the unique aspects of the Versal architecture, multiple research questions related to performance scaling, sensitivity, and efficiency are explored. This paper aims to assist developers in the FPGA community looking to implement GEMM on AMD Versal by providing guidelines and insights for enhancing performance. Kaustubh Manohar, Venkata Guru Prasanth Mulleti, Curt John Bansil, Endri Taka, Aman Arora 0001 |
FPGA | 5 |
| 2025 | Systolic Sparse Tensor Slices: FPGA Building Blocks for Sparse and Dense AI AccelerationabstractFPGA architectures have recently been enhanced to meet the substantial computational demands of modern deep neural networks (DNNs). To this end, both FPGA vendors and academic researchers have proposed in-fabric blocks that perform efficient tensor computations. However, these blocks are primarily optimized for dense computation, while most DNNs exhibit sparsity. To address this limitation, we propose incorporating structured sparsity support into FPGA architectures. We architect 2D systolic in-fabric blocks, named systolic sparse tensor (SST) slices, that support multiple degrees of sparsity to efficiently accelerate a wide variety of DNNs. SSTs support dense operation, 2:4 (50%) and 1:4 (75%) sparsity, as well as a new 1:3 (66.7%) sparsity level to further increase flexibility. When demonstrating on general matrix multiplication (GEMM) accelerators, which are the heart of most current DNN accelerators, our sparse SST-based designs attain up to 5× higher FPGA frequency and 10.9× lower area, compared to traditional FPGAs. Moreover, evaluation of the proposed SSTs on state-of-the-art sparse ViT and CNN models exhibits up to 3.52× speedup with minimal area increase of up to 13.3%, compared to dense in-fabric acceleration. Endri Taka, Ning-Chi Huang, Chi-Chih Chang, Kai-Chiang Wu, Aman Arora 0001, Diana Marculescu |
FPGA | 5 |
| 2025 | GAMA: High-Performance GEMM Acceleration on AMD Versal ML-Optimized AI EnginesabstractGeneral matrix-matrix multiplication (GEMM) is a fundamental operation in machine learning (ML) applications. We present the first comprehensive performance acceleration of GEMM workloads on AMD's second-generation AIE-ML architecture, which is specifically optimized for ML applications. Compared to AI-Engine (AIE), AIE-ML offers increased compute throughput and larger on-chip memory capacity. We propose a novel design that maximizes AIE-ML memory utilization, incorporates custom buffer placement within the AIE-ML and staggered kernel placement across the AIE-ML array, significantly reducing performance bottlenecks such as memory stalls and routing congestion, resulting in improved performance and efficiency compared to the default AMD's compiler. We evaluate the performance benefits of our design at three levels: single AIEML, pack of AIE-ML's and the complete AIE-ML array. GAMA achieves state-of-the-art performance, delivering up to$\mathbf{1 6 5}$TOPS (85% of peak) for int8 precision and 83 TBFLOPS (86% of peak) for bfloat16 precision GEMM workloads. Our solution achieves$8.7 \%, 9 \%, 39 \%$and 53.6 % higher peak throughput efficiency compared to the state-of-the-art AIE frameworks AMA, MAXEVA, ARIES and CHARM, respectively. Kaustubh Manohar, Endri Taka, Aman Arora 0001 |
FPL | 3 |
| 2025 | ATAPP: Architecture and Technology Aware Power Predictor for Unseen FPGAS
Zhigang Wei, Aman Arora 0001, Emily Shriver, Lizy Kurian John |
FPL | 2 |
| 2025 | CarbonSet: A Dataset to Analyze Trends and Benchmark the Sustainability of CPUs and GPUs
Jiajun Hu, Chetan Choppali Sudarshan, Maxwell Clifford, Vidya A. Chhabria, Aman Arora 0001 |
ACM Great Lakes Symposium on VLSI | 5 |
| 2025 | Performance Analysis of GEMM Workloads on the AMD Versal PlatformabstractAMD Versal is a new heterogeneous computing hardware architecture comprised of adaptive intelligence (AI) engines, programmable logic, and a processing system. General Matrix Multiplication (GEMM) is the fundamental building block of modern deep learning (DL) applications such as ChatGPT, and GEMM workloads can be mapped onto Versal in different ways, each with distinct trade-offs. This paper presents a thorough analysis of GEMM workloads of different shapes and sizes, showcasing performance artifacts associated with the AMD Versal architecture. Focusing on the unique aspects of the Versal architecture, multiple research questions related to performance scaling, sensitivity, and efficiency are explored. This paper aims to assist FPGA developers looking to implement GEMM on AMD Versal by providing insights for enhancing performance. Kaustubh Manohar, Venkata Guru Prashanth Mulleti, Curt John Bansil, Endri Taka, Aman Arora 0001 |
ISPASS | 5 |
| 2025 | Field-Programmable Gate Array Architecture for Deep Learning: Survey and Future DirectionsabstractDeep learning (DL) is becoming the cornerstone of numerous applications both in large-scale datacenters and at the edge. Specialized hardware is often necessary to meet the performance requirements of state-of-the-art DL models, but the rapid pace of change in DL models and the wide variety of systems integrating DL make it impossible to create custom computer chips for all but the largest markets. Field-programmable gate arrays (FPGAs) present a unique blend of reprogrammability and direct hardware execution that make them suitable for accelerating DL inference. They offer the ability to customize processing pipelines and memory hierarchies to achieve lower latency and higher energy efficiency compared to general-purpose central processing units (CPUs) and graphics processing units (GPUs), at a fraction of the development time and cost of custom chips. Their diverse and high-speed inputs/outputs (IOs) also enable directly interfacing the FPGA to the network and/or a variety of external sensors, making them suitable for both datacenter and edge use cases. As DL has become an ever more important workload, FPGA architectures are evolving to enable higher DL performance. In this article, we survey both academic and industrialFPGA chip architectureenhancements for DL. First, we give a brief introduction on the basics of FPGA architecture and how its components lead to strengths and weaknesses for DL applications. Next, we discuss differentdesign stylesof DL inference accelerators implemented on FPGAs that achieve state-of-the-art performance and productive development flows, ranging from model-specific dataflow styles to software-programmable overlay styles. We survey DL-specific enhancements to traditional FPGA building blocks including the logic blocks (LBs), arithmetic circuitry, and on-chip memories, as well as new DL-specialized blocks that integrate into the FPGA fabric to accelerate tensor computations. Finally, we discuss hybrid devices that combine processors and coarse-grained accelerator blocks with FPGA-like interconnect and networks-on-chip (NoCs), and highlight promising future research directions. Andrew Boutros, Aman Arora 0001, Vaughn Betz |
Proc. IEEE | 2 |
| 2024 | GreenFPGA: Evaluating FPGAs as Environmentally Sustainable Computing SolutionsabstractGrowing global concerns about climate change highlight the need for environmentally sustainable computing. The ecological impact of computing, including operational and embodied, is crucial. Field Programmable Gate Arrays (FPGAs) stand out as promising sustainable computing platforms due to their reconfigurability across various applications. This paper introduces GreenFPGA, a tool estimating the total carbon footprint (CFP) of FPGAs over their lifespan, considering design, manufacturing, reconfigurability, operation, disposal, and recycling. Using GreenFPGA, the paper evaluates scenarios where the ecological benefits of FPGA reconfigurability outweigh operational and embodied carbon costs, positioning FPGAs as an environmentally sustainable choice for hardware acceleration compared to Application-specific integrated circuits (ASICs). Experimental results show that FPGAs have lower CFP than ASICs for multiple low-volume applications or short application lifespans. Chetan Choppali Sudarshan, Aman Arora 0001, Vidya A. Chhabria |
DAC | 2 |
| 2024 | Efficient Approaches for GEMM Acceleration on Leading AI-Optimized FPGAsabstractFPGAs are a promising platform for accelerating Deep Learning (DL) applications, due to their high performance, low power consumption, and reconfigurability. Recently, the leading FPGA vendors have enhanced their architectures to more efficiently support the computational demands of DL workloads. However, the two most prominent AI-optimized FPGAs, i.e., AMD/Xilinx Versal ACAP and Intel Stratix 10 NX, employ significantly different architectural approaches. This paper presents novel systematic frameworks to optimize the performance of General Matrix Multiplication (GEMM), a fundamental operation in DL workloads, by exploiting the unique and distinct architectural characteristics of each FPGA. Our evaluation on GEMM workloads for int8 precision shows up to 77 and 68 TOPs (int8) throughput, with up to 0.94 and 1.35 TOPs/W energy efficiency for Versal VC1902 and Stratix 10 NX, respectively. This work provides insights and guidelines for optimizing GEMM-based applications on both platforms, while also delving into their programmability trade-offs and associated challenges. Endri Taka, Dimitrios Gourounas, Andreas Gerstlauer, Diana Marculescu, Aman Arora 0001 |
FCCM | 5 |
| 2024 | Cross-FPGA Power Estimation from High Level Synthesis via Transfer-LearningabstractMachine learning (ML) has been successfully employed to estimate power consumption for FPGAs using features derived from post High Level Synthesis (HLS). As a result, the power evaluation of the design bypasses time-consuming logic synthesis and implementation. However, such models have noticeable drawbacks. Firstly, the dataset preparation is time-consuming since researchers invest significant effort in constructing a sufficient dataset to train an accurate model for a target FPGA. Secondly, the model trained on one FPGA cannot be directly applied to another. Without prior knowledge about the architecture of the second FPGA, the model's power estimation on this new FPGA is of unknown confidence. To address these challenges, we propose a novel cross-FPGA power modeling methodology called XPNet that combines Transfer-Learning with an innovative data selection technique that enables efficient fine-tuning. We start by applying Transfer-Learning with our data selection methodology to adapt a GNN-based power model to a second FPGA using only 20 data samples, resulting in 6.53% error. We then explore if our approach works for lighter-weight ML-based models, such as multi-layer perception (MLP), and show less than a 1% degradation in accuracy. Additionally, we explore the impact of using Meta-Learning algorithm on our model and show that with only 40 data samples from the target FPGA, the model still manages an error of 6%. Zhigang Wei, Aman Arora 0001, Emily Shriver, Lizy Kurian John |
FPGA | 2 |
| 2024 | PIMSAB: A Processing-In-Memory System with Spatially-Aware Communication and Bit-Serial-Aware ComputationabstractBit-serial Processing-In-Memory (PIM) is an attractive paradigm for accelerator architectures, for parallel workloads such as Deep Learning (DL), because of its capability to achieve massive data parallelism at a low area overhead and provide orders-of-magnitude data movement savings by moving computational resources closer to the data. While many PIM architectures have been proposed, improvements are needed in communicating intermediate results to consumer kernels, for communication between tiles at scale, for reduction operations, and for efficiently performing bit-serial operations with constants. We present PIMSAB, a scalable architecture that provides a spatially aware communication network for efficient intra-tile and inter-tile data movement and provides efficient computation support for generally inefficient bit-serial compute patterns. Our architecture consists of a massive hierarchical array of compute-enabled SRAMs (CRAMs), which is codesigned with a compiler to achieve high utilization. The key novelties of our architecture are (1) in providing efficient support for spatially aware communication by providing local H-tree network for reductions, by adding explicit hardware for shuffling operands, and by deploying systolic broadcasting, as well as (2) by taking advantage of the divisible nature of bit-serial computations through adaptive precision and efficient handling of constant operations. These innovations are integrated into a tensor expressions-based programming framework (including a compiler for easy programmability) that enables simple programmer control of optimizations for mapping programs into massively parallel binaries for millions of PIM processing elements. When compared against a similarly provisioned modern Tensor Core GPU (NVIDIA A100), across common DL kernels and end-to-end DL networks (Resnet18 and BERT), PIMSAB outperforms the GPU by 4.80×, and reduces energy by 3.76×. We compare PIMSAB with similarly provisioned state-of-the-art SRAM PIM (Duality Cache) and DRAM PIM (SIMDRAM), and observe a speedup of 3.7× and 3.88×, respectively. Kaustubh Manohar, Jian Weng 0002, Bagus Hanindhito, Zhengrong Wang, Tony Nowatzki, Lizy Kurian John, Aman Arora 0001 |
ACM Trans. Archit. Code Optim. | 8 |
| 2023 | COIN: Combinational Intelligent NetworksabstractWe introduce Combinational Intelligent Networks (COIN), a machine learning technique that targets edge inference using low-resourced FPGAs or ASICs. COIN is an improvement on LogicWiSARD, a recent weightless neural network that achieves low power, small area, and high throughput. We convert the LogicWiSARD model into a binary neural network, train it using backpropagation, and then convert it to a COIN model. As a result, COIN can achieve higher accuracy than LogicWiSARD or it can require significantly fewer hardware resources when comparing models with similar accuracies. In comparison to a BNN implementation, FINN, small and large COIN models are more energy efficient demonstrating up to 11.5x higher inferences/Joule at similar accuracy. Our tool executes the complete flow, from training to RTL. and is publicly available. Igor D. S. Miranda, Aman Arora 0001, Zachary Susskind, Josias S. A. Souza, Mugdha P. Jadhao, Luis A. Q. Villon, Diego Leonel Cadette Dutra, Priscila M. V. Lima, Felipe M. G. França, Maurício Breternitz, Lizy Kurian John |
ASAP | 2 |
| 2023 | HLSDataset: Open-Source Dataset for ML-Assisted FPGA Design using High Level SynthesisabstractMachine Learning (ML) has been widely adopted in design exploration using high level synthesis (HLS) for faster resource, timing and power estimation at very early stages for FPGA-based design. To perform prediction accurately, high-quality and large-volume datasets are required for training ML models. However, the current datasets used in this domain are proprietary or limited in use, and practitioners have to generate their own dataset to train HLS-related ML models. This paper presents a dataset for ML-assisted FPGA design using HLS, called HLSDataset. The dataset is generated from widely used HLS C benchmarks including Polybench, Machsuite, CHStone and Rossetta. The Verilog samples are generated with a variety of directives including loop unroll, loop pipeline, and array partition to make sure optimized and realistic designs are covered. The total number of generated Verilog samples is nearly 9,000 per FPGA type. The dataset repository includes CSV (comma separated values) files containing both HLS and implementation metrics which can be easily consumed by ML model. We also include original C source code with directives, Verilog designs, post-HLS reports, post-implementation reports for each sample in the dataset, so that any metrics not present in the CSV can be easily extracted. In order to extend the dataset for future benchmarks, generation and extraction scripts are also provided. To demonstrate the effectiveness of our dataset, we undertake case studies to perform power estimation and resource usage estimation with ML models trained with our dataset. All the code and dataset are public at our github page11https://github.com/UT-LCAIML4Accel-Dataset/tree/main/fpga_ml_dataset. We believe that HLSDataset can save valuable time for researchers by avoiding the tedious process of running tools, scripting and parsing files to generate the dataset, and enable them to spend more time where it counts, that is, in training ML models. Zhigang Wei, Aman Arora 0001, Ruihao Li 0002, Lizy Kurian John |
ASAP | 2 |
| 2023 | Infinity Stream: Portable and Programmer-Friendly In-/Near-Memory FusionabstractIn-memory computing with large last-level caches is promising to dramatically alleviate data movement bottlenecks and expose massive bitline-level parallelization opportunities. However, key challenges from its unique execution model remain unsolved: automated parallelization, transparently orchestrating data transposition/alignment/broadcast for bit-serial logic, and mixing in-/near-memory computing. Most importantly, the solution should be programmer friendly and portable across platforms. Zhengrong Wang, Christopher Liu, Aman Arora 0001, Lizy Kurian John, Tony Nowatzki |
ASPLOS (3) | 3 |
| 2023 | An FPGA-Based Weightless Neural Network for Edge Network Intrusion DetectionabstractAlgorithms for mobile networking are increasingly being moved from centralized servers towards the edge in order to decrease latency and improve the user experience. While much of this work is traditionally done using ASICs, 6G emphasizes the adaptability of algorithms for specific user scenarios, which motivates broader adoption of FPGAs. In this paper, we propose the FPGA-based Weightless Intrusion Warden (FWIW), a novel solution for detecting anomalous network traffic on edge devices. While prior work in this domain is based on conventional deep neural networks (DNNs), FWIW incorporates a weightless neural network (WNN), a table lookup-based model which learns sophisticated nonlinear behaviors. This allows FWIW to achieve accuracy far superior to prior FPGA-based work at a very small fraction of the model footprint, enabling deployment on small, low-cost devices. FWIW achieves a prediction accuracy of 98.5% on the UNSW-NB15 dataset with a total model parameter size of just 192 bytes, reducing error by 7.9x and model size by 262x vs. LogicNets, the best prior edge-optimized implementation. Implemented on a Xilinx Virtex UltraScale+ FPGA, FWIW demonstrates a 59x reduction in LUT usage with a 1.6x increase in throughput. The accuracy of FWIW comes within 0.6% of the best-reported result in literature (Edge-Detect), a model several orders of magnitude larger. Our results make it clear that WNNs are worth exploring in the emerging domain of edge networking, and suggest that FPGAs are capable of providing the extreme throughput needed. Zachary Susskind, Aman Arora 0001, Alan T. L. Bacellar, Diego Leonel Cadette Dutra, Igor D. S. Miranda, Maurício Breternitz, Priscila M. V. Lima, Felipe M. G. França, Lizy Kurian John |
FPGA | 2 |
| 2023 | ULEEN: A Novel Architecture for Ultra-low-energy Edge Neural Networksabstract‘‘Extreme edge” 1 devices, such as smart sensors, are a uniquely challenging environment for the deployment of machine learning. The tiny energy budgets of these devices lie beyond what is feasible for conventional deep neural networks, particularly in high-throughput scenarios, requiring us to rethink how we approach edge inference. In this work, we propose ULEEN, a model and FPGA-based accelerator architecture based on weightless neural networks (WNNs). WNNs eliminate energy-intensive arithmetic operations, instead using table lookups to perform computation, which makes them theoretically well-suited for edge inference. However, WNNs have historically suffered from poor accuracy and excessive memory usage. ULEEN incorporates algorithmic improvements and a novel training strategy inspired by binary neural networks (BNNs) to make significant strides in addressing these issues. We compare ULEEN against BNNs in software and hardware using the four MLPerf Tiny datasets and MNIST. Our FPGA implementations of ULEEN accomplish classification at 4.0–14.3 million inferences per second, improving area-normalized throughput by an average of 3.6× and steady-state energy efficiency by an average of 7.1× compared to the FPGA-based Xilinx FINN BNN inference platform. While ULEEN is not a universally applicable machine learning model, we demonstrate that it can be an excellent choice for certain applications in energy- and latency-critical edge environments. Zachary Susskind, Aman Arora 0001, Igor D. S. Miranda, Alan T. L. Bacellar, Luis A. Q. Villon, Rafael Fontella Katopodis, Leandro Santiago de Araújo, Diego Leonel Cadette Dutra, Priscila M. V. Lima, Felipe M. G. França, Maurício Breternitz, Lizy Kurian John |
ACM Trans. Archit. Code Optim. | 2 |
| 2023 | Koios 2.0: Open-Source Deep Learning Benchmarks for FPGA Architecture and CAD Researchabstractthe prevalence of deep learning (DL) in many applications, researchers are investigating different ways of optimizing field-programmable gate array (FPGA) architecture and CAD to achieve better quality-of-results (QoRs) on DL-based workloads. In this optimization process, benchmark circuits are an essential component; the QoR achieved on a set of benchmarks is the main driver for architecture and CAD design choices. However, current academic benchmark suites are inadequate, as they do not capture any designs from the DL domain. This work presents the second version of our suite of DL acceleration benchmark circuits for FPGA architecture and CAD research, called Koios. This suite of 40 circuits covers a wide variety of accelerated neural networks, design sizes, implementation styles, abstraction levels, and numerical precisions. These benchmarks include 32 DL designs and eight synthetic (proxy) benchmarks. The Koios benchmarks are larger, more data parallel, more heterogeneous, more deeply pipelined, and utilize more FPGA architectural features compared to existing open-source benchmarks. This enables researchers to pinpoint architectural inefficiencies for this class of workloads and optimize CAD tools on more representative benchmarks that stress the CAD algorithms in different ways. In this article, we describe the Koios designs, compare their characteristics to prior FPGA benchmark suites, and present results of running them through the verilog-to-routing (VTR) flow using a recent FPGA architecture model. Finally, we present case studies showing how exploration of DL-optimized FPGA architecture and CAD algorithms can be performed using our new benchmark suite. Aman Arora 0001, Andrew Boutros, Seyed Alireza Damghani, Karan Mathur, Vedant Mohanty, Tanmay Anand, Mohamed A. Elgammal, Kenneth B. Kent, Vaughn Betz, Lizy Kurian John |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | CoMeFa: Deploying Compute-in-Memory on FPGAs for Deep Learning AccelerationabstractBlock random access memories (BRAMs) are the storage houses of FPGAs, providing extensive on-chip memory bandwidth to the compute units implemented using logic blocks and digital signal processing slices. We propose modifying BRAMs to convert them to CoMeFa ( Co mpute-in- Me mory Blocks for F PG A s) random access memories (RAMs). These RAMs provide highly parallel compute-in-memory by combining computation and storage capabilities in one block. CoMeFa RAMs utilize the true dual-port nature of FPGA BRAMs and contain multiple configurable single-bit bit-serial processing elements. CoMeFa RAMs can be used to compute with any precision, which is extremely important for applications like deep learning (DL). Adding CoMeFa RAMs to FPGAs significantly increases their compute density while also reducing data movement. We explore and propose two architectures of these RAMs: CoMeFa-D (optimized for delay) and CoMeFa-A (optimized for area). Compared to existing proposals, CoMeFa RAMs do not require changing the underlying static RAM technology like simultaneously activating multiple wordlines on the same port, and are practical to implement. CoMeFa RAMs are especially suitable for parallel and compute-intensive applications like DL, but these versatile blocks find applications in diverse applications like signal processing and databases, among others. By augmenting an Intel Arria 10–like FPGA with CoMeFa-D (CoMeFa-A) RAMs at the cost of 3.8% (1.2%) area, and with algorithmic improvements and efficient mapping, we observe a geomean speedup of 2.55× (1.85×) across microbenchmarks from various applications and a geomean speedup of up to 2.5× across multiple deep neural networks. Replacing all or some BRAMs with CoMeFa RAMs in FPGAs can make them better accelerators of DL workloads. Aman Arora 0001, Atharva Bhamburkar, Aatman Borda, Tanmay Anand, Rishabh Sehgal, Bagus Hanindhito, Pierre-Emmanuel Gaillardon, Jaydeep P. Kulkarni, Lizy Kurian John |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2022 | Weightless Neural Networks for Efficient Edge InferenceabstractWeightless neural networks (WNNs) are a class of machine learning model which use table lookups to perform inference, rather than the multiply-accumulate operations typical of deep neural networks (DNNs). Individual weightless neurons are capable of learning non-linear functions of their inputs, a theoretical advantage over the linear neurons in DNNs, yet state-of-the-art WNN architectures still lag behind DNNs in accuracy on common classification tasks. Additionally, many existing WNN architectures suffer from high memory requirements, hindering implementation. In this paper, we propose a novel WNN architecture, BTHOWeN, with key algorithmic and architectural improvements over prior work, namely counting Bloom filters, hardware-friendly hashing, and Gaussian-based nonlinear thermometer encodings. These enhancements improve model accuracy while reducing size and energy per inference. BTHOWeN targets the large and growing edge computing sector by providing superior latency and energy efficiency to both prior WNNs and comparable quantized DNNs. Compared to state-of-the-art WNNs across nine classification datasets, BTHOWeN on average reduces error by more than 40% and model size by more than 50%. We demonstrate the viability of a hardware implementation of BTHOWeN by presenting an FPGA-based inference accelerator, and compare its latency and resource usage against similarly accurate quantized DNN inference accelerators, including multi-layer perceptron (MLP) and convolutional models. The proposed BTHOWeN models consume almost 80% less energy than the MLP models, with nearly 85% reduction in latency. In our quest for efficient ML on the edge, WNNs are clearly deserving of additional attention. Zachary Susskind, Aman Arora 0001, Igor D. S. Miranda, Luis A. Q. Villon, Rafael Fontella Katopodis, Leandro Santiago de Araújo, Diego Leonel Cadette Dutra, Priscila M. V. Lima, Felipe M. G. França, Maurício Breternitz, Lizy Kurian John |
PACT | 2 |
| 2022 | LogicWiSARD: Memoryless Synthesis of Weightless Neural NetworksabstractWeightless neural networks (WNNs) are an alternative pattern recognition technique where RAM nodes function as neurons. As both training and inference require mostly table lookups, few additions, and no multiplications, WNNs are suitable for high-performance and low-power embedded applications. This work introduces a novel approach to implement WiSARD, the leading WNN state-of-the-art architecture, completely eliminating memories and arithmetic circuits and utilizing only logic functions. The approach creates compressed minimized implementations by converting trained WNN nodes from lookup tables to logic functions. The proposed LogicWiSARD is implemented in FPGA and ASIC technologies to illustrate its suitability for edge inference. Experimental results show more than 80% reduction in energy consumption when the proposed LogicWiSARD model is compared with a multilayer perceptron network (MLP) of equivalent accuracy. Compared to previous work on FPGA implementations for WNNs, convolutional neural networks, and binary neural networks, the energy savings of LogicWiSARD range between 32.2% and 99.6%. Igor D. S. Miranda, Aman Arora 0001, Zachary Susskind, Luis A. Q. Villon, Rafael Fontella Katopodis, Diego Leonel Cadette Dutra, Leandro Santiago de Araújo, Priscila M. V. Lima, Felipe M. G. França, Lizy Kurian John, Maurício Breternitz |
ASAP | 2 |
| 2022 | Pruning Weightless Neural NetworksabstractWeightless neural networks (WNNs) are a type of machine learning model which perform prediction using lookup tables (LUTs) instead of arithmetic operations.Recent advancements in WNNs have reduced model sizes and improved accuracies, reducing the gap in accuracy with deep neural networks (DNNs).Modern DNNs leverage "pruning" techniques to reduce model size, but this has not previously been explored for WNNs.We propose a WNN pruning strategy based on identifying and culling the LUTs which contribute least to overall model accuracy.We demonstrate an average 40% reduction in model size with at most 1% reduction in accuracy. Zachary Susskind, Alan T. L. Bacellar, Aman Arora 0001, Luis A. Q. Villon, Renan Mendanha, Leandro Santiago de Araújo, Diego Leonel Cadette Dutra, Priscila M. V. Lima, Felipe M. G. França, Igor D. S. Miranda, Maurício Breternitz, Lizy Kurian John |
ESANN | 3 |
| 2022 | CoMeFa: Compute-in-Memory Blocks for FPGAsabstractBlock RAMs (BRAMs) are the storage houses of FPGAs, providing extensive on-chip memory bandwidth to the compute units implemented using Logic Blocks (LBs) and Digital Signal Processing (DSP) slices. We propose modifying BRAMs to convert them to CoMeFa (Compute-In-Memory Blocks for FPGAs) RAMs. These RAMs provide highly-parallel compute-in-memory by combining computation and storage capabilities in one block. CoMeFa RAMs utilize the true dual port nature of FPGA BRAMs and contain multiple programmable single-bit bit-serial processing elements. CoMeFa RAMs can be used to compute in any precision, which is extremely important for evolving applications like Deep Learning. Adding CoMeFa RAMs to FPGAs significantly increases their compute density. We explore and propose two architectures of these RAMs: CoMeFa-D (optimized for delay) and CoMeFa-A (optimized for area). Compared to existing proposals, CoMeFa RAMs do not require changing the underlying SRAM technology like simultaneously activating multiple rows on the same port, and are practical to implement. CoMeFa RAMs are versatile blocks that find applications in numerous diverse parallel applications like Deep Learning, signal processing, databases, etc. By augmenting an Intel Arria-10-like FPGA with CoMeFa-D (CoMeFa-A) RAMs at the cost of 3.8% (1.2%) area, and with algorithmic improvements and efficient mapping, we observe a geomean speedup of 2.55x (1.85x), across several representative benchmarks. Replacing all or some BRAMs with CoMeFa RAMs in FPGAs can make them better accelerators of modern compute-intensive workloads. Aman Arora 0001, Tanmay Anand, Aatman Borda, Rishabh Sehgal, Bagus Hanindhito, Jaydeep P. Kulkarni, Lizy Kurian John |
FCCM | 1 |
| 2022 | MathRAMs: Configurable Fused Compute-Memory Blocks for FPGAsabstractBlock RAMs (BRAMs) are the storage houses of FPGAs. We propose modifying BRAMs into new blocks called MathRAMs, which provide highly-parallel processing-in-memory (PIM) by combining computation and storage capabilities in one block. MathRAMs have a bit-serial precision-agnostic compute architecture. Compared to similar existing proposals, MathRAMs do not require changing the underlying SRAM technology like activating multiple wordlines and are not targeted to a specific application like Deep Learning. MathRAMs utilize the true dual port nature of FPGA BRAMs and contain multiple programmable single-bit bit-serial processing elements that operate in parallel. The area overhead of making these enhancements is 26% at the block level, which translates to 5.7% increase in FPGA die area for our baseline Stratix 10-like FPGA. A frequency reduction of 25% is observed in Compute mode (no change in Memory mode). MathRAMs increase the compute density of FPGAs and reduce power consumption by reducing data movement. MathRAMs find applications in numerous diverse parallel applications like signal processing, databases, deep learning, etc. In our evaluation of various applications, we observe speedup of 95% in moving average filter, upto 36% in matrix-vector multiplication, and upto 220% in bitwise operations, when compared to the baseline. Replacing all or some BRAMs with MathRAMs in FPGAs can make them more efficient and performant accelerators of modern compute-intensive workloads. Aman Arora 0001, Aatman Borda, Tanmay Anand, Bagus Hanindhito, Lizy Kurian John |
FPGA | 1 |
| 2022 | Tensor Slices: FPGA Building Blocks For The Deep Learning EraabstractFPGAs are well-suited for accelerating deep learning (DL) applications owing to the rapidly changing algorithms, network architectures and computation requirements in this field. However, the generic building blocks available on traditional FPGAs limit the acceleration that can be achieved. Many modifications to FPGA architecture have been proposed and deployed including adding specialized artificial intelligence (AI) processing engines, adding support for smaller precision math like 8-bit fixed point and IEEE half-precision (fp16) in DSP slices, adding shadow multipliers in logic blocks, etc. In this paper, we describe replacing a portion of the FPGA’s programmable logic area with Tensor Slices. These slices have a systolic array of processing elements at their heart that support multiple tensor operations, multiple dynamically-selectable precisions and can be dynamically fractured into individual multipliers and MACs (multiply-and-accumulate). These slices have a local crossbar at the inputs that helps with easing the routing pressure caused by a large block on the FPGA. Adding these DL-specific coarse-grained hard blocks to FPGAs increases their compute density and makes them even better hardware accelerators for DL applications, while still keeping the vast majority of the real estate on the FPGA programmable at fine-grain. Aman Arora 0001, Moinak Ghosh, Samidh Mehta, Vaughn Betz, Lizy Kurian John |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2021 | Tensor Slices to the Rescue: Supercharging ML Acceleration on FPGAsabstractFPGAs are well-suited for accelerating deep learning (DL) applications owing to the rapidly changing algorithms, network architectures and computation requirements in this field. However, the generic building blocks available on traditional FPGAs limit the acceleration that can be achieved. Many modifications to FPGA architecture have been proposed and deployed including adding specialized artificial intelligence (AI) processing engines, adding support for IEEE half-precision (fp16) math in DSP slices, adding hard matrix multiplier blocks, etc. In this paper, we describe replacing a small percentage of the FPGA's programmable logic area with Tensor Slices. These slices are arrays of processing elements at their heart that support multiple tensor operations, multiple dynamically-selectable precisions and can be dynamically fractured into individual adders, multipliers and MACs (multiply-and-accumulate). These tiles have a local crossbar at the inputs that helps with easing the routing pressure caused by a large slice. By spending ~3% of FPGA's area on Tensor Slices, we observe an average frequency increase of 2.45x and average area reduction by 0.41x across several ML benchmarks, including a TPU-like design, compared to an Intel Agilex-like baseline FPGA. We also study the impact of spending area on Tensor slices on non-ML applications. We observe an average reduction of 1% in frequency and an average increase of 1% in routing wirelength compared to the baseline, across the non-ML benchmarks we studied. Adding these ML-specific coarse-grained hard blocks makes the proposed FPGA a much efficient hardware accelerator for ML applications, while still keeping the vast majority of the real estate on the FPGA programmable at fine-grain. Aman Arora 0001, Samidh Mehta, Vaughn Betz, Lizy Kurian John |
FPGA | 1 |
| 2021 | Koios: A Deep Learning Benchmark Suite for FPGA Architecture and CAD ResearchabstractWith the prevalence of deep learning (DL) in many applications, researchers are investigating different ways of optimizing FPGA architecture and CAD to achieve better quality-of-results (QoR) on DL-based workloads. In this optimization process, benchmark circuits are an essential component; the QoR achieved on a set of benchmarks is the main driver for architecture and CAD design choices. However, current academic benchmark suites are inadequate, as they do not capture any designs from the DL domain. This work presents a new suite of DL acceleration benchmark circuits for FPGA architecture and CAD research, called Koios. This suite of 19 circuits covers a wide variety of accelerated neural networks, design sizes, implementation styles, abstraction levels, and numerical precisions. These designs are larger, more data parallel, more heterogeneous, more deeply pipelined, and utilize more FPGA architectural features compared to existing open-source benchmarks. This enables researchers to pin-point architectural inefficiencies for this class of workloads and optimize CAD tools on more realistic benchmarks that stress the CAD algorithms in different ways. In this paper, we describe the designs in our benchmark suite, present results of running them through the Verilog-to-Routing (VTR) flow using a recent FPGA architecture model, and identify key insights from the resulting metrics. On average, our benchmarks have 3.7× more netlist primitives, 1.8× and 4.7× higher DSP and BRAM densities, and 1.7× higher frequency with 1.9× more near-critical paths compared to the widely-used VTR suite. Finally, we present two example case studies showing how architectural exploration for DL-optimized FPGAs can be performed using our new benchmark suite. Aman Arora 0001, Andrew Boutros, Daniel Rauch, Aishwarya Rajen, Aatman Borda, Seyed Alireza Damghani, Samidh Mehta, Sangram Kate, Pragnesh Patel, Kenneth B. Kent, Vaughn Betz, Lizy Kurian John |
FPL | 1 |
| 2020 | Hamamu: Specializing FPGAs for ML Applications by Adding Hard Matrix Multiplier BlocksabstractDesigning efficient hardware for accelerating artificial intelligence (AI) and machine learning (ML) applications is a major challenge. Rapidly changing algorithms and neural network architectures make FPGA based designs an attractive solution. But the generic building blocks available in current FPGAs (Logic Blocks (LBs), multipliers, DSP blocks) limit the acceleration that can be achieved. We propose Hamamu, a modification to the current FPGA architecture that makes FPGAs specialized for ML applications. Specifically, we propose adding hard matrix multiplier blocks (matmuls) into the FPGA fabric. These matmuls are implemented using systolic arrays of MACs (Multiply-And-Accumulate) and can be connected using programmable direct interconnect between neighboring matmuls to make larger systolic matrix multipliers. We explore various matmul sizes ($2\times 2\times 2$, $4\times 4\times 4$, $8\times 8\times 8$, $16\times 16\times 16$) and various strategies to place these blocks on the FPGA (Columnar, Surround, Hybrid). We find that providing $4\times 4\times 4$ hard matrix multiplier blocks in an FPGA speeds up neural networks from MLPerf benchmarks by up to $\sim 3.9x$, compared to a Stratix-10 like FPGA with equal number of MACs, same MAC architecture and high DSP:LB ratio. Although the flexibility of the FPGA will reduce for non-ML applications, an FPGA with hard matrix multipliers is a faster, and more area efficient hardware accelerator for ML applications, compared to current FPGAs. Aman Arora 0001, Zhigang Wei, Lizy Kurian John |
ASAP | 1 |
| 2020 | Design Space Exploration for Softmax ImplementationsabstractDeep Neural Networks (DNN) are crucial components of machine learning in the big data era. Significant effort has been put into the hardware acceleration of convolution and fully-connected layers of neural networks, while not too much attention has been put on the Softmax layer. Softmax is used in terminal classification layers in networks like ResNet, and is also used in intermediate layers in networks like the Transformer. As the speed for other DNN layers keeps improving, efficient and flexible designs for Softmax are required. With the existence of several ways to implement Softmax in hardware, we evaluate various softmax hardware designs and the trade-offs between them. In order to make the design space exploration more efficient, we also develop a parameterized generator which can produce softmax designs by varying multiple aspects of a base architecture. The aspects or knobs are parallelism, accuracy, storage and precision. The goal of the generator is to enable evaluation of tradeoffs between area, delay, power and accuracy in the architecture of a softmax unit. We simulate and synthesize the generated designs and present results comparing them with the existing state-of-the-art. Our exploration reveals that the design with parallelism of 16 can provide the best area-delay product among designs with parallelism ranging from 1 to 32. It is also observed that look-up table based approximate LOG and EXP units can be used to yield almost the same accuracy as the full LOG and EXP units, while providing area and energy benefits. Additionally, providing local registers for intermediate values is seen to provide energy savings. Zhigang Wei, Aman Arora 0001, Pragenesh Patel, Lizy Kurian John |
ASAP | 2 |
| 2020 | The Case for Hard Matrix Multiplier Blocks in an FPGAabstractDesigning efficient hardware for accelerating machine learning (ML) applications is a major challenge. Rapid changing algorithms and network architectures in this field make FPGA based designs an attractive solution. But the generic building blocks available in current FPGAs (ALMs/CLBs, DSP blocks) limit the acceleration that can be achieved. We propose a modification to the current FPGA architecture that makes FPGAs specialized for ML applications. Specifically, we propose adding hard matrix multiplier blocks (matmuls) into the FPGA fabric. These matmuls are implemented using systolic arrays of MACs (Multiply-And-Accumulate) and can be connected using programmable direct interconnect between neighboring matmuls to make larger systolic matrix multipliers. We explore various matmul sizes (4x4x4, 8x8x8, 16x16x16, 32x32x32) and various strategies to place these blocks on the FPGA (clustered, surround, columnar). We recommend 4x4x4 matmul blocks with columnar placement after studying tradeoffs between area, frequency, fragmentation and channel width. Experimental results and analytical evaluation reveal that providing matmuls in an FPGA speeds up state-of-the-art neural networks (Resnet50, GNMT, Transformer, Minigo) by ~2.5x on average, compared to a DSP-heavy FPGA with equal number of MACs. Therefore, FPGAs with hard matrix multipliers can be used to design faster, more area (and hence, power) efficient hardware accelerators for ML applications, compared to current FPGAs, at the cost of reducing the flexibility of the FPGA for other applications. A matmul-heavy FPGA fabric could be a part of bigger FPGA, the rest of which can have general programmable logic, or fully ML-specific FPGAs with matmuls could be created. Aman Arora 0001, Zhigang Wei, Lizy Kurian John |
FPGA | 1 |