EDBT 2026 Demo / reviewers in the wild / expert
Tendayi Kamucheka
dblp:301/0093
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0002-7853-5780ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DA-VinCi: A Deep-Learning Accelerator Overlay Using In-Memory ComputingabstractThe matrix operations that underpin today’s deep learning models are routinely implemented in Single Instruction Multiple Data (SIMD) domain specific accelerators. SIMD accelerators including GPUs and array processors can effectively leverage parallelism in models that are compute-bound, but their effectiveness can be diminished for models that are memory-bound. Processing-in-Memory (PIM) architectures are being explored to provide better energy efficiency and scalable performance for these memory-bound models. Modern Field Programmable Gate Arrays (FPGAs) feature hundreds of megabits of Static Random Access Memory (SRAM) distributed across the device as disaggregated memory resources. This makes FPGAs ideal programmable platforms for developing custom Processor In/Near Memory accelerators. Several PIM array-based accelerator designs have been proposed to leverage this substantial internal bandwidth. However, results reported to date show the FPGA based PIM architectures operating at system clock frequencies well below a chips Block-RAM (BRAM) Fmax clock frequency. Results also show that the compute densities of the designs do not scale linearly with BRAM densities. These results indicate that FPGA PIM architectures will never be competitive with their custom Application-Specific Integrated Circuit (ASIC) counterparts. In this article, we introduce DA-VinCi, a D eep-Learning A ccelerator O v erlay using In -Memory C omput i ng. DA-VinCi is the first scalable FPGA based PIM deep-learning accelerator overlay capable of clocking at the maximum frequency of a device’s BRAM. Further, the architecture of DA-VinCi allows the number of compute units to scale linearly up to the maximum capacity of a devices BRAM, and at the maximum clock frequency of the BRAM. The DA-VinCi overlay has a programmable Instruction Set Architecture (ISA) that allows the same synthesized design to provide low-latency inferencing of a range of memory-bound deep-learning models, including Multilayer Perceptrons, Recurrent Neural Network, Long Short-Term Memory, and Gated Recurrent Unit networks. The scalability and high clocking frequency of DA-VinCi is achieved through a new Processor In Memory (PIM) tile architecture and a highly scalable system-level framework. We present results showing DA-VinCi linearly scaling the number of Processing Elements (PEs) to 100% of the BRAM capacity (over 60K PEs) on an Alveo U55 clocking at 737 MHz, the chips BRAM Fmax. We provide comparative studies on inference latency across multiple deep-learning applications that show DA-VinCi achieves up to a 201 \(\times\) improvement over a state-of-the-art PIM overlay accelerator, up to 87 \(\times\) improvement over existing PIM-based FPGA accelerators, and up to 57 \(\times\) improvement over custom deep-learning accelerators on FPGAs. M. D. Arafat Kabir, Nathaniel Fredricks, Tendayi Kamucheka, Joel Mandebi, Miaoqing Huang, Jason D. Bakos, David Andrews 0001 |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2024 | The BRAM is the Limit: Shattering Myths, Shaping Standards, and Building Scalable PIM AcceleratorsabstractMany recent FPGA-based Processor-in-Memory (PIM) architectures have appeared with promises of impressive levels of parallelism but with performance that falls short of expectations due to reduced maximum clock frequencies, an inability to scale processing elements up to the maximum BRAM capacity, and minimal hardware support for large reduction operations. In this paper, we propose a “Standard” set of design objectives for PIM array-based FPGA designs. We then propose a PIM array-based GEMV accelerator architecture as a case study to show the proposed Standard can be realized in practice. The GEMV accelerator serves as existence proof that dispels several myths surrounding what is normally accepted as clocking and scaling FPGA performance limitations. Specifically, the proposed accelerator clocks at the maximum frequency of the BRAM and scales to 100% of the available BRAMs. Comparative analyses show execution speeds over existing PIM-based GEMV engines on FPGAs and achieving a 2.65Χ – 3.2Χ faster clock. An AMD Alveo U55 implementation achieves a system clock speed of 737 MHz, providing 64K bit serial multiply-accumulate (MAC) units for GEMV operation. M. D. Arafat Kabir, Tendayi Kamucheka, Nathaniel Fredricks, Joel Mandebi, Jason D. Bakos, Miaoqing Huang, David Andrews 0001 |
FCCM | 2 |
| 2024 | Ph.D. Project: A Compiler-Driven Approach to HW/SW Co-Design of Deep-Learning AcceleratorsabstractThis work introduces SPAR, a compiler framework based on MLIR, tailored for deep learning inference applications. Alongside SPAR, we present SPAR HL, an extensible Instruction Set Architecture (ISA) designed to seamlessly interface custom FPGA-based accelerators with the SPAR compiler. Historically, custom accelerators on FPGA have posed a challenge as a compiler target. Prior compiler initiatives have addressed this issue by generating application-specific hardware during compilation, necessitating expertise across both application and hardware domains. In response, we offer an alternative approach by proposing an ISA as a unified compiler target and hardware interface for custom accelerators. Furthermore, we introduce a compiler capable of translating high-level machine learning models encoded in ONNX into code compatible with our proposed ISA, thus enabling efficient deployment on FPGA-based custom accelerators. Tendayi Kamucheka, David Andrews 0001 |
FCCM | 1 |
| 2024 | IMAGine: An In-Memory Accelerated GEMV Engine OverlayabstractProcessor-in-Memory (PIM) overlays and alternative reconfigurable tile fabrics have been proposed to eliminate the von Neumann bottleneck and enable processing performance to scale with BRAM capacity. The performance of these FPGA-based PIM architectures has been limited due to a reduction of the BRAMs maximum clock frequencies and less than ideal scaling of processing elements with increased BRAM capacity. This paper presents IMAGine, an In-Memory Accelerated GEMV engine, a PIM-array accelerator that clocks at the maximum frequency of the BRAM and scales to 100% of the available BRAMs. Comparative analyses are presented showing execution speeds over existing PIM-based GEMV engines on FPGAs and achieving a $2.65 \times-3.2 \times$ faster clock. An AMD Alveo U55 implementation is presented that achieves a system clock speed of 737 MHz, providing 64 K bit-serial multiply-accumulate (MAC) units for GEMV operation. This establishes IMAGine as the fastest PIM-based GEMV overlay, outperforming even the custom PIM-based FPGA accelerators reported to date. Additionally, it surpasses TPU v1-v2 and Alibaba Hanguang 800 in clock speed while offering an equal or greater number of multiply-accumulate (MAC) units. M. D. Arafat Kabir, Tendayi Kamucheka, Nathaniel Fredricks, Joel Mandebi, Jason D. Bakos, Miaoqing Huang, David Andrews 0001 |
FPL | 2 |
| 2022 | A Masked Pure-Hardware Implementation of Kyber Cryptographic AlgorithmabstractQuantum computing-specifically Shor's algorithm [1]-presents an existential threat to some standard cryptographic algorithms. In preparation, post-quantum cryptography (PQC) algorithms have been in development and are nearing mathematical and cryptanalytic maturity. Standardization efforts through the National Institute of Standards and Technology (NIST) PQC standardization process have chosen one PKE/KEM algorithm (i.e., CRYSTALS-Kyber) and three digital signature algorithms (i.e., CRYSTALS-Dilithium, Falcon, and SPHINCS+). CRYSTALS-Kyber is a lattice-based, IND-CCA2-secure, key-encapsulation mechanism (KEM) based on the learning-with-errors problem over module lattices. This paper presents a masked hardware implementation of Kyber that is demonstrably secure against side-channel power analysis methods. Tendayi Kamucheka, Alexander Nelson 0001, David Andrews 0001, Miaoqing Huang |
FPT | 1 |