EDBT 2026 Demo / reviewers in the wild / expert
Tanmay Anand
dblp:295/2702
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2023
0000-0002-5725-7048ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 4 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Koios 2.0: Open-Source Deep Learning Benchmarks for FPGA Architecture and CAD Researchabstractthe prevalence of deep learning (DL) in many applications, researchers are investigating different ways of optimizing field-programmable gate array (FPGA) architecture and CAD to achieve better quality-of-results (QoRs) on DL-based workloads. In this optimization process, benchmark circuits are an essential component; the QoR achieved on a set of benchmarks is the main driver for architecture and CAD design choices. However, current academic benchmark suites are inadequate, as they do not capture any designs from the DL domain. This work presents the second version of our suite of DL acceleration benchmark circuits for FPGA architecture and CAD research, called Koios. This suite of 40 circuits covers a wide variety of accelerated neural networks, design sizes, implementation styles, abstraction levels, and numerical precisions. These benchmarks include 32 DL designs and eight synthetic (proxy) benchmarks. The Koios benchmarks are larger, more data parallel, more heterogeneous, more deeply pipelined, and utilize more FPGA architectural features compared to existing open-source benchmarks. This enables researchers to pinpoint architectural inefficiencies for this class of workloads and optimize CAD tools on more representative benchmarks that stress the CAD algorithms in different ways. In this article, we describe the Koios designs, compare their characteristics to prior FPGA benchmark suites, and present results of running them through the verilog-to-routing (VTR) flow using a recent FPGA architecture model. Finally, we present case studies showing how exploration of DL-optimized FPGA architecture and CAD algorithms can be performed using our new benchmark suite. Aman Arora 0001, Andrew Boutros, Seyed Alireza Damghani, Karan Mathur, Vedant Mohanty, Tanmay Anand, Mohamed A. Elgammal, Kenneth B. Kent, Vaughn Betz, Lizy Kurian John |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2023 | CoMeFa: Deploying Compute-in-Memory on FPGAs for Deep Learning AccelerationabstractBlock random access memories (BRAMs) are the storage houses of FPGAs, providing extensive on-chip memory bandwidth to the compute units implemented using logic blocks and digital signal processing slices. We propose modifying BRAMs to convert them to CoMeFa ( Co mpute-in- Me mory Blocks for F PG A s) random access memories (RAMs). These RAMs provide highly parallel compute-in-memory by combining computation and storage capabilities in one block. CoMeFa RAMs utilize the true dual-port nature of FPGA BRAMs and contain multiple configurable single-bit bit-serial processing elements. CoMeFa RAMs can be used to compute with any precision, which is extremely important for applications like deep learning (DL). Adding CoMeFa RAMs to FPGAs significantly increases their compute density while also reducing data movement. We explore and propose two architectures of these RAMs: CoMeFa-D (optimized for delay) and CoMeFa-A (optimized for area). Compared to existing proposals, CoMeFa RAMs do not require changing the underlying static RAM technology like simultaneously activating multiple wordlines on the same port, and are practical to implement. CoMeFa RAMs are especially suitable for parallel and compute-intensive applications like DL, but these versatile blocks find applications in diverse applications like signal processing and databases, among others. By augmenting an Intel Arria 10–like FPGA with CoMeFa-D (CoMeFa-A) RAMs at the cost of 3.8% (1.2%) area, and with algorithmic improvements and efficient mapping, we observe a geomean speedup of 2.55× (1.85×) across microbenchmarks from various applications and a geomean speedup of up to 2.5× across multiple deep neural networks. Replacing all or some BRAMs with CoMeFa RAMs in FPGAs can make them better accelerators of DL workloads. Aman Arora 0001, Atharva Bhamburkar, Aatman Borda, Tanmay Anand, Rishabh Sehgal, Bagus Hanindhito, Pierre-Emmanuel Gaillardon, Jaydeep P. Kulkarni, Lizy Kurian John |
ACM Trans. Reconfigurable Technol. Syst. | 4 |
| 2022 | CoMeFa: Compute-in-Memory Blocks for FPGAsabstractBlock RAMs (BRAMs) are the storage houses of FPGAs, providing extensive on-chip memory bandwidth to the compute units implemented using Logic Blocks (LBs) and Digital Signal Processing (DSP) slices. We propose modifying BRAMs to convert them to CoMeFa (Compute-In-Memory Blocks for FPGAs) RAMs. These RAMs provide highly-parallel compute-in-memory by combining computation and storage capabilities in one block. CoMeFa RAMs utilize the true dual port nature of FPGA BRAMs and contain multiple programmable single-bit bit-serial processing elements. CoMeFa RAMs can be used to compute in any precision, which is extremely important for evolving applications like Deep Learning. Adding CoMeFa RAMs to FPGAs significantly increases their compute density. We explore and propose two architectures of these RAMs: CoMeFa-D (optimized for delay) and CoMeFa-A (optimized for area). Compared to existing proposals, CoMeFa RAMs do not require changing the underlying SRAM technology like simultaneously activating multiple rows on the same port, and are practical to implement. CoMeFa RAMs are versatile blocks that find applications in numerous diverse parallel applications like Deep Learning, signal processing, databases, etc. By augmenting an Intel Arria-10-like FPGA with CoMeFa-D (CoMeFa-A) RAMs at the cost of 3.8% (1.2%) area, and with algorithmic improvements and efficient mapping, we observe a geomean speedup of 2.55x (1.85x), across several representative benchmarks. Replacing all or some BRAMs with CoMeFa RAMs in FPGAs can make them better accelerators of modern compute-intensive workloads. Aman Arora 0001, Tanmay Anand, Aatman Borda, Rishabh Sehgal, Bagus Hanindhito, Jaydeep P. Kulkarni, Lizy Kurian John |
FCCM | 2 |
| 2022 | MathRAMs: Configurable Fused Compute-Memory Blocks for FPGAsabstractBlock RAMs (BRAMs) are the storage houses of FPGAs. We propose modifying BRAMs into new blocks called MathRAMs, which provide highly-parallel processing-in-memory (PIM) by combining computation and storage capabilities in one block. MathRAMs have a bit-serial precision-agnostic compute architecture. Compared to similar existing proposals, MathRAMs do not require changing the underlying SRAM technology like activating multiple wordlines and are not targeted to a specific application like Deep Learning. MathRAMs utilize the true dual port nature of FPGA BRAMs and contain multiple programmable single-bit bit-serial processing elements that operate in parallel. The area overhead of making these enhancements is 26% at the block level, which translates to 5.7% increase in FPGA die area for our baseline Stratix 10-like FPGA. A frequency reduction of 25% is observed in Compute mode (no change in Memory mode). MathRAMs increase the compute density of FPGAs and reduce power consumption by reducing data movement. MathRAMs find applications in numerous diverse parallel applications like signal processing, databases, deep learning, etc. In our evaluation of various applications, we observe speedup of 95% in moving average filter, upto 36% in matrix-vector multiplication, and upto 220% in bitwise operations, when compared to the baseline. Replacing all or some BRAMs with MathRAMs in FPGAs can make them more efficient and performant accelerators of modern compute-intensive workloads. Aman Arora 0001, Aatman Borda, Tanmay Anand, Bagus Hanindhito, Lizy Kurian John |
FPGA | 3 |
| 2021 | A blockchain and deep neural networks-based secure framework for enhanced crop protection
Vikas Hassija, Siddharth Batra, Vinay Chamola, Tanmay Anand, Poonam Goyal, Navneet Goyal, Mohsen Guizani |
Ad Hoc Networks | 4 |