EDBT 2026 Demo / reviewers in the wild / expert
Akhil Raj Baranwal
dblp:295/3410
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2026
0000-0003-1024-9101ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HiSpMM: High Performance High Bandwidth Sparse-Dense Matrix Multiplication on HBM-equipped FPGAsabstractSparse Matrix-Dense Matrix Multiplication (SpMM) is a critical operation in scientific computing, machine learning, and graph analytics. However, accelerating SpMM on FPGAs presents major challenges due to irregular memory access patterns and imbalanced workload distribution. In this work, we address a fundamental bottleneck in SpMM acceleration on High Bandwidth Memory (HBM)-equipped FPGAs: workload imbalance among processing elements (PEs). Additionally, we mitigate a scalability barrier present in state-of-the-art designs—namely, the tight coupling between PEs and HBM channels for dense matrix access. Furthermore, we provide an automated design space exploration framework. We propose HiSpMM, a high-performance SpMM accelerator architecture that introduces Dense Row Sharing to mitigate PE under-utilization by distributing heavy-row computations, a decoupled HBM access mechanism to allow independent scaling of PEs and memory bandwidth, and an automation tool that optimizes design parameters according to matrix structure-specific properties and user-defined hardware constraints. Our design achieves a geomean of \(5.81\times\) speedup and \(5.75\times\) energy efficiency improvement for imbalanced matrices compared to state-of-the-art designs, while also maintaining competitive performance for balanced matrices on the AMD/Xilinx U280 HBM FPGA board. Our HiSpMM project will be open sourced in the near future at https://github.com/SFU-HiAccel/HiSpMM . Ahmad Sedigh Baroughi, Manoj B. Rajashekar, Akhil Raj Baranwal, Zhenman Fang |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2025 | PoCo: Extending Task-Parallel HLS Programming with Shared Multi-Producer Multi-Consumer Buffer SupportabstractAdvancements in High-Level Synthesis (HLS) tools have enabled task-level parallelism on FPGAs. However, prevailing frameworks predominantly employ Single-Producer-Single-Consumer (SPSC) models for task communication, thus limiting application scenarios. Analysis of designs becomes non-trivial with an increasing number of tasks in task-parallel systems. Adding features to existing designs often requires re-profiling of several task interfaces, redesign of the overall inter-task connectivity, and describing a new floorplan. This article proposes PoCo, a novel framework to design scalable Multi-Producer-Multi-Consumer (MPMC) models on task-parallel systems. PoCo introduces a shared-buffer abstraction that facilitates dynamic and high-bandwidth access to share on-chip memory resources, incorporates latency-insensitive communication, and implements placement-aware design strategies to mitigate routing congestion. The frontend provides convenient APIs to access the buffer memory, while the backend features an optimized and pipelined datapath. Empirical evaluations demonstrate that PoCo achieves up to 50% reduction in on-chip memory utilization on SPSC models without performance degradation. Additionally, three case studies on distinct real-world applications reveal up to 1.5 \(\times\) frequency improvements and simplified dataflow management in heterogeneous FPGA accelerator designs. Akhil Raj Baranwal, Zhenman Fang |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2025 | MAD-HiSpMV: Matrix Adaptive Design with Hybrid Row Distribution for Imbalanced SpMV Acceleration on FPGAsabstractSparse Matrix–Vector Multiplication (SpMV) is fundamental in numerous applications such as scientific computing, Machine Learning (ML), and graph analytics. While recent studies have made tremendous progress in accelerating SpMV on HBM-equipped FPGAs, there are still multiple remaining challenges to accelerate imbalanced SpMV where the distribution of nonzeros in the sparse matrix is imbalanced across different rows. These include (1) imbalanced workload distribution among the parallel Processing Elements (PEs), (2) long-distance dependency for floating-point accumulation on the output vector, (3) a new bottleneck due to the often-overlooked dense vectors’ off-chip access after the SpMV acceleration, and (4) sub-optimal performance of generic accelerators for various types of sparse matrices. (5) Additionally, ML workloads often consist of both SpMV and General Matrix–Vector Multiplication (GeMV), which suffer from kernel switching inefficiencies. To address those challenges, we propose MAD-HiSpMV to accelerate imbalanced SpMV on HBM-equipped FPGAs with the following novel solutions: (1) a hybrid row distribution network to enable both inter-row and intra-row distribution for better balance, (2) a fully pipelined floating-point accumulation on the output vector using a combination of an adder chain and register-based circular buffer, (3) matrix adaptive design configurations generated by our automation framework via Design Space Exploration (DSE) to maximize performance for the given matrix, and (4) a GeMV overlay built into the same kernel for efficient acceleration of mixed workloads. Experimental results demonstrate that the DSE-picked configuration of MAD-HiSpMV achieves a geomean speedup of 1.3× (up to 2.12×) for the SpMV benchmark matrices and achieves a geomean 1.15× (up to 1.54×) better performance per watt, when compared to state-of-the-art generic designs. For the SpMV benchmark matrices, compared to Intel MKL running on a 24-core Xeon Silver 4214 CPU, MAD-HiSpMV achieves a geomean speedup of 8.80×. Compared to cuSparse running on an Nvidia GTX 1080ti GPU, MAD-HiSpMV achieves a geomean of 2.57× better performance per watt. Additionally, a GeMV overlay built into MAD-HiSpMV achieves a peak throughput of 156.7 GFLOPS, which is 2.64× better than the Vitis L2 GeMV benchmark on U280, and performs 2.7× better for an end-to-end mixed workload, when compared to Intel MKL running on a 24-core Xeon Silver 4214 CPU. MAD-HiSpMV is available at https://github.com/SFU-HiAccel/HiSpMV . Manoj B. Rajashekar, Akhil Raj Baranwal, Xingyu Tian, Zhenman Fang |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2024 | PASTA: Programming and Automation Support for Scalable Task-Parallel HLS Programs on Modern Multi-Die FPGAsabstractIn recent years, the adoption of FPGAs in datacenters has increased, with a growing number of users choosing High-Level Synthesis (HLS) as their preferred programming method. While HLS simplifies FPGA programming, one notable challenge arises when scaling up designs for modern datacenter FPGAs that comprise multiple dies. The extra delays introduced due to die crossings and routing congestion can significantly degrade the frequency of large designs on these FPGA boards. Due to the gap between HLS design and physical design, it is challenging for HLS programmers to analyze and identify the root causes, and fix their HLS design to achieve better timing closure. Recent efforts have aimed to address these issues by employing coarse-grained floorplanning and pipelining strategies on task-parallel HLS designs where multiple tasks run concurrently and communicate through FIFO stream channels. However, many applications are not streaming friendly and many existing accelerator designs heavily rely on buffer channel based communication between tasks. In this work, we take a step further to support a task-parallel programming model where tasks can communicate via both FIFO stream channels and buffer channels. To achieve this goal, we design and implement the PASTA framework, which takes a large task-parallel HLS design as input and automatically generates a high-frequency FPGA accelerator via HLS and physical design co-optimization. Our framework introduces a latency-insensitive buffer channel design, which supports memory partitioning and ping-pong buffering while remaining compatible with vendor HLS tools. On the frontend, we provide an easy-to-use programming model for utilizing the proposed buffer channel; while on the backend, we implement efficient placement and pipelining strategies for the proposed buffer channel. To validate the effectiveness of our framework, we test it on four widely used Rodinia HLS benchmarks and two real-world accelerator designs and show an average frequency improvement of 25%, with peak improvements of up to 89% on AMD/Xilinx Alveo U280 boards compared to Vitis HLS baselines. Moazin Khatti, Xingyu Tian, Ahmad Sedigh Baroughi, Akhil Raj Baranwal, Yuze Chi, Licheng Guo, Jason Cong, Zhenman Fang |
ACM Trans. Reconfigurable Technol. Syst. | 4 |
| 2021 | MemOReL: A Memory-oriented Optimization Approach to Reinforcement Learning on FPGA-based Embedded SystemsabstractReinforcement Learning (RL) represents the machine learning method that has come closest to showing human-like learning. While Deep RL is becoming increasingly popular for complex applications such as AI-based gaming, it has a high implementation cost in terms of both power and latency. Q-Learning, on the other hand, is a much simpler method that makes it more feasible for implementation on resource-constrained embedded systems for control and navigation. However, the optimal policy search in Q-Learning is a compute-intensive and inherently sequential process and a software-only implementation may not be able to satisfy the latency and throughput constraints of such applications. To this end, we propose a novel accelerator design with multiple design trade-offs for implementing Q-Learning on FPGA-based SoCs. Specifically, we analyze the various stages of the Epsilon-Greedy algorithm for RL and propose a novel microarchitecture that reduces the latency by optimizing the memory access during each iteration. Consequently, we present multiple designs that provide varying trade-offs between performance, power dissipation, and resource utilization of the accelerator. With the proposed approach, we report considerable improvement in throughput with lower resource utilization over state-of-the-art design implementations. Siva Satyendra Sahoo, Akhil Raj Baranwal, Salim Ullah, Akash Kumar 0001 |
ACM Great Lakes Symposium on VLSI | 2 |
| 2021 | ReLAccS: A Multilevel Approach to Accelerator Design for Reinforcement Learning on FPGA-Based SystemsabstractReinforcement learning (RL), specifically Q-learning, with human-like learning abilities to learn from experience without any a priori data, is being increasingly used in embedded systems in the field of control and navigation. However, finding the optimal policy in this approach can be highly compute-intensive, and a software-only implementation may not satisfy the application's timing constraints. To this end, we propose optimization methods at multiple levels of accelerator design for RL. Specifically, at the architecture-level, we exploit the instruction-level parallelism and the spatial parallelism in FPGAs to improve the throughput over state-of-the-art designs by up to 34%. Further, we propose lookup table-level optimizations to reduce the resource utilization and power dissipation of the accelerator. Finally, we propose algorithm-level approximation that can be used for acceleration of Q-learning problems with more states and for reducing the peak power dissipation. We report up to 10× reduction in power dissipation with marginal degradation in quality of results. Akhil Raj Baranwal, Salim Ullah, Siva Satyendra Sahoo, Akash Kumar 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |