M. D. Arafat Kabir

dblp:261/7898 · DBLP profile ↗
← Back
9ranked-venue papers
8as first author
7since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 8 first-author · 7 since 2021
YearPublicationVenuePosition
2026 Adaptive Redistribution Layer Routing for Chiplet-Package Co-Design in 2.5D System
abstract
2.5D packaging has become a popular alternative to integrate advanced logic and memory chiplets for high-performance computing and artificial intelligence systems. In the conventional design flow, chiplets and packages are independently designed and then integrated at the assembly stage. To bridge the gap between chiplet designs and package designs, existing chiplet-package co-design methods iteratively optimize chiplet layouts to improve the performance of the entire system. However, Redistribution Layer (RDL) routing, which finishes the interconnections between chiplets at the package level and significantly affects the system performance, is neglected in the existing co-design flows. Therefore, this article proposes an effective chiplet-package co-design flow focusing on the RDL routing to optimize the package system performance dynamically. The proposed co-design flow can fill in the missing link, package-level co-optimization, of previous design flows. In the proposed co-design flow, we propose an efficient RDL routing algorithm to iteratively optimize the substrate layout based on the cross-boundary timing context extracted from both chiplets and the package. The proposed RDL routing algorithm has two critical techniques, including (1) a Maximal Independent Set-based (MIS-based) pin assignment method to dynamically optimize the pin positions of nets and (2) a network-flow-based router to generate routing layouts. Experimental results show that the proposed design flow can gradually improve the maximum frequency of a real design to the target performance, 400 MHz.
Zhen Zhuang, Weishiun Hung, M. D. Arafat Kabir, Yarui Peng, Tsung-Yi Ho
ACM Trans. Design Autom. Electr. Syst.3
2025 DA-VinCi: A Deep-Learning Accelerator Overlay Using In-Memory Computing
abstract
The matrix operations that underpin today’s deep learning models are routinely implemented in Single Instruction Multiple Data (SIMD) domain specific accelerators. SIMD accelerators including GPUs and array processors can effectively leverage parallelism in models that are compute-bound, but their effectiveness can be diminished for models that are memory-bound. Processing-in-Memory (PIM) architectures are being explored to provide better energy efficiency and scalable performance for these memory-bound models. Modern Field Programmable Gate Arrays (FPGAs) feature hundreds of megabits of Static Random Access Memory (SRAM) distributed across the device as disaggregated memory resources. This makes FPGAs ideal programmable platforms for developing custom Processor In/Near Memory accelerators. Several PIM array-based accelerator designs have been proposed to leverage this substantial internal bandwidth. However, results reported to date show the FPGA based PIM architectures operating at system clock frequencies well below a chips Block-RAM (BRAM) Fmax clock frequency. Results also show that the compute densities of the designs do not scale linearly with BRAM densities. These results indicate that FPGA PIM architectures will never be competitive with their custom Application-Specific Integrated Circuit (ASIC) counterparts. In this article, we introduce DA-VinCi, a D eep-Learning A ccelerator O v erlay using In -Memory C omput i ng. DA-VinCi is the first scalable FPGA based PIM deep-learning accelerator overlay capable of clocking at the maximum frequency of a device’s BRAM. Further, the architecture of DA-VinCi allows the number of compute units to scale linearly up to the maximum capacity of a devices BRAM, and at the maximum clock frequency of the BRAM. The DA-VinCi overlay has a programmable Instruction Set Architecture (ISA) that allows the same synthesized design to provide low-latency inferencing of a range of memory-bound deep-learning models, including Multilayer Perceptrons, Recurrent Neural Network, Long Short-Term Memory, and Gated Recurrent Unit networks. The scalability and high clocking frequency of DA-VinCi is achieved through a new Processor In Memory (PIM) tile architecture and a highly scalable system-level framework. We present results showing DA-VinCi linearly scaling the number of Processing Elements (PEs) to 100% of the BRAM capacity (over 60K PEs) on an Alveo U55 clocking at 737 MHz, the chips BRAM Fmax. We provide comparative studies on inference latency across multiple deep-learning applications that show DA-VinCi achieves up to a 201 \(\times\) improvement over a state-of-the-art PIM overlay accelerator, up to 87 \(\times\) improvement over existing PIM-based FPGA accelerators, and up to 57 \(\times\) improvement over custom deep-learning accelerators on FPGAs.
M. D. Arafat Kabir, Nathaniel Fredricks, Tendayi Kamucheka, Joel Mandebi, Miaoqing Huang, Jason D. Bakos, David Andrews 0001
ACM Trans. Reconfigurable Technol. Syst.1
2024 The BRAM is the Limit: Shattering Myths, Shaping Standards, and Building Scalable PIM Accelerators
abstract
Many recent FPGA-based Processor-in-Memory (PIM) architectures have appeared with promises of impressive levels of parallelism but with performance that falls short of expectations due to reduced maximum clock frequencies, an inability to scale processing elements up to the maximum BRAM capacity, and minimal hardware support for large reduction operations. In this paper, we propose a “Standard” set of design objectives for PIM array-based FPGA designs. We then propose a PIM array-based GEMV accelerator architecture as a case study to show the proposed Standard can be realized in practice. The GEMV accelerator serves as existence proof that dispels several myths surrounding what is normally accepted as clocking and scaling FPGA performance limitations. Specifically, the proposed accelerator clocks at the maximum frequency of the BRAM and scales to 100% of the available BRAMs. Comparative analyses show execution speeds over existing PIM-based GEMV engines on FPGAs and achieving a 2.65Χ – 3.2Χ faster clock. An AMD Alveo U55 implementation achieves a system clock speed of 737 MHz, providing 64K bit serial multiply-accumulate (MAC) units for GEMV operation.
M. D. Arafat Kabir, Tendayi Kamucheka, Nathaniel Fredricks, Joel Mandebi, Jason D. Bakos, Miaoqing Huang, David Andrews 0001
FCCM1
2024 IMAGine: An In-Memory Accelerated GEMV Engine Overlay
abstract
Processor-in-Memory (PIM) overlays and alternative reconfigurable tile fabrics have been proposed to eliminate the von Neumann bottleneck and enable processing performance to scale with BRAM capacity. The performance of these FPGA-based PIM architectures has been limited due to a reduction of the BRAMs maximum clock frequencies and less than ideal scaling of processing elements with increased BRAM capacity. This paper presents IMAGine, an In-Memory Accelerated GEMV engine, a PIM-array accelerator that clocks at the maximum frequency of the BRAM and scales to 100% of the available BRAMs. Comparative analyses are presented showing execution speeds over existing PIM-based GEMV engines on FPGAs and achieving a $2.65 \times-3.2 \times$ faster clock. An AMD Alveo U55 implementation is presented that achieves a system clock speed of 737 MHz, providing 64 K bit-serial multiply-accumulate (MAC) units for GEMV operation. This establishes IMAGine as the fastest PIM-based GEMV overlay, outperforming even the custom PIM-based FPGA accelerators reported to date. Additionally, it surpasses TPU v1-v2 and Alibaba Hanguang 800 in clock speed while offering an equal or greater number of multiply-accumulate (MAC) units.
M. D. Arafat Kabir, Tendayi Kamucheka, Nathaniel Fredricks, Joel Mandebi, Jason D. Bakos, Miaoqing Huang, David Andrews 0001
FPL1
2023 Making BRAMs Compute: Creating Scalable Computational Memory Fabric Overlays
abstract
The increasing density of distributed BRAMs diffused throughout modern Field Programmable Gate Arrays (FP-GAs) is ideal for forming processor in/near memory architectures. This breaks the traditional von Neumann memory bottleneck limiting concurrency and degrading energy efficiency. Ideally, processing density should scale linearly with BRAM capacity, and clock frequencies should be set by the read/write access times of the BRAM. In this paper, we present a PIM overlay that achieves these goals. We observe an improvement of performance by 2.25 x, logic resource utilization by 2 x, and accumulation delay by 17 x compared to prior published work.
M. D. Arafat Kabir, Joshua Hollis, Atiyehsadat Panahi, Jason D. Bakos, Miaoqing Huang, David Andrews 0001
FCCM1
2023 FPGA Processor In Memory Architectures (PIMs): Overlay or Overhaul ?
abstract
The dominance of machine learning and the ending of Moore's law have renewed interests in Processor in Memory (PIM) architectures. This interest has produced several recent proposals to modify an FPGA's BRAM architecture to form a next-generation PIM reconfigurable fabric [1], [2]. PIM architectures can also be realized within today's FPGAs as overlays without the need to modify the underlying FPGA architecture. To date, there has been no study to understand the comparative advantages of the two approaches. In this paper, we present a study that explores the comparative advantages between two proposed custom architectures and a PIM overlay running on a commodity FPGA. We created PiCaSO, a Processor in/near Memory Scalable and Fast Overlay architecture as a representative PIM overlay. The results of this study show that the PiCaSO overlay achieves up to 80% of the peak throughput of the custom designs with 2.56 x shorter latency and 25% - 43% better BRAM memory utilization efficiency. We then show how several key features of the PiCaSO overlay can be integrated into the custom PIM designs to further improve their throughput by 18%, latency by 19.5%, and memory efficiency by 6.2%.
M. D. Arafat Kabir, Ehsan Kabir, Joshua Hollis, Eli Levy-Mackay, Atiyehsadat Panahi, Jason D. Bakos, Miaoqing Huang, David Andrews 0001
FPL1
2021 Cross-Boundary Inductive Timing Optimization for 2.5D Chiplet-Package Co-Design
abstract
With the popularity of 2.5D integration, an increasing number of chiplets are integrated into advanced system-in-package designs. In such systems, redistribution layer (RDL) wires become longer and denser, with a growing impact on system performance. However, RDL inductive impacts in timing analysis are ignored by the traditional CAD tools. This paper presents our chiplet-package co-optimization flow, which can capture the RDL inductance impact on system performance and automatically adjust the IO drivers to compensate for the inductance overhead. We develop our extraction and timing analysis tool that models RDL wire inductive timing impact on 2.5D system performance within +/-1% error. Our study shows 35% signal paths through RDL violate the timing requirement because of the inductive impact, and remain undetected through only RC-based STA.
M. D. Arafat Kabir, Dusan Petranovic, Yarui Peng
ACM Great Lakes Symposium on VLSI1
2020 Chiplet-Package Co-Design For 2.5D Systems Using Standard ASIC CAD Tools
abstract
Chiplet integration using 2.5D packaging is gaining popularity nowadays which enables several interesting features like heterogeneous integration and drop-in design method. In the traditional die-by-die approach of designing a 2.5D system, each chiplet is designed independently without any knowledge of the package RDLs. In this paper, we propose a Chip-Package Co-Design flow for implementing 2.5D systems using existing commercial chip design tools. Our flow encompasses 2.5D-aware partitioning suitable for SoC design, Chip-Package Floorplanning, and post-design analysis and verification of the entire 2.5D system. We also designed our own package planners to route RDL layers on top of chiplet layers. We use an ARM Cortex-M0 SoC system to illustrate our flow and compare analysis results with a monolithic 2D implementation of the same system. We also compare two different 2.5D implementations of the same SoC system following the drop-in approach. Alongside the traditional die-by-die approach, our holistic flow enables design efficiency and flexibility with accurate cross-boundary parasitic extraction and design verification.
M. D. Arafat Kabir, Yarui Peng
ASP-DAC1
2020 Coupling Extraction and Optimization for Heterogeneous 2.5D Chiplet-Package Co-Design
abstract
In recent years, 2.5D chiplet package designs have gained popularity in system integration of heterogeneous technologies. Currently, there exists no standard CAD flow that can design, analyze, and optimize a complete heterogeneous 2.5D system. The traditional die-by-die design approach does not consider any package layers during extraction and optimization, and an accurate chiplet-package extraction can not be applied to heterogeneous designs without fundamental changes in standard CAD tools. In this paper, we present our Holistic and In-Context chiplet-package co-design flows for high-performance high-density 2.5D systems using standard ASIC CAD tools with zero overhead on IO pipeline depth. Our flow encompasses 2.5D-aware partitioning, chiplet-package co-planning, in-context extraction, iterative optimization, and post-design analysis and verification of the entire 2.5D system. We design our package planner with a routing and pin-planning strategy to minimize package routing congestion and timing overhead. An ARM Cortex-M0-based microcontroller system is designed as the benchmark. The performance gap to the reference 2D design reduces by 62.5% when chip-package interactions are taken into account in the holistic flow. Our in-context extraction achieves only 0.71% and 0.79% error on ground and coupling capacitance on a homogeneous system. Further, we implement a heterogeneous 2.5D system to demonstrate our novel in-context design and optimization methodology.
M. D. Arafat Kabir, Dusan Petranovic, Yarui Peng
ICCAD1