Ashwin Bhat

dblp:316/8565 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0002-6395-9345ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-author · 7 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A ML-based Robust Channel Estimation Enhancer Against Multi-Tone Jamming in OFDM systems
Chaofan Deng, Ashwin Bhat, Adou Sangbone Assoa, Katarina Vuckovic, Subhashish Chakravarty, Arijit Raychowdhury
ISCAS2
2025 MDS-DOA: Fusing Model-Based and Data-Driven Approaches for Modular, Distributed, and Scalable Direction-of-Arrival Estimation
abstract
Massive MIMO systems are promising for wireless communications beyond 5G, but scalable Direction-of-Arrival (DOA) estimation in these systems is challenging due to the increasing number of required antennas. Existing solutions, model-based or data-driven (typically using neural networks), face scalability issues with the growing antenna array size. To address this issue, we propose a hybrid system that makes the overall approach scalable. In the front-end, we employ a modular distributed approach namely, the method of sparse linear inverse to compute a proxy spectrum from the sampled covariance matrix of the antenna subarrays. The proxy drives a fixed lightweight back-end which consists of a 1-dimensional Convolution Neural Network (1D-CNN) and a simplified peak extraction. The input proxy dimension being independent of the antenna count makes the neural network input invariant of the array size, enabling it to handle multiple array sizes without requiring any modification of the neural network structure. To reduce the computation of the covariance matrix and proxy spectrum, we employ a system of subarrays with Nearest-Neighbor communication. The proposed approach was implemented on a Xilinx ZCU102 FPGA targeting 100 MHz frequency for 8 to 256-element arrays. We achieve below 1 ms processing time for an array of 256 antennas while requiring significantly less computation than both model-based and data-driven approaches for large antenna arrays.
Adou Sangbone Assoa, Ashwin Bhat, Sigang Ryu, Arijit Raychowdhury
IEEE Trans. Circuits Syst. I Regul. Pap.2
2023 M5: Multi-modal Multi-task Model Mapping on Multi-FPGA with Accelerator Configuration Search
abstract
Recent machine learning (ML) models have advanced from single-modality single-task to multi-modality multi-task (MMMT). MMMT models typically have multiple backbones of different sizes along with complicated connections, exposing great challenges for hardware deployment. For scalable and energy-efficient implementations, multi-FPGA systems are emerging as the ideal design choices. However, finding the optimal solutions for mapping MMMT models onto multiple FPGAs is non-trivial. Existing mapping algorithms focus on either streamlined linear deep neural network architectures or only the critical path of simple heterogeneous models. Direct extensions of these algorithms for MMMT models lead to sub-optimal solutions. To address these shortcomings, we propose M5, a novel MMMT Model Mapping framework for Multi-FPGA platforms. In addition to handling multiple modalities present in the models, M5 can flexibly explore accelerator configurations and possible resource sharing opportunities to significantly improve the system performance. For various computation-heavy MMMT models, experiment results demonstrate that M5 can remarkably outperform existing mapping methods and lead to an average reduction of 35%, 62%, and 70% in the number of low-end, mid-end, and high-end FPGAs required to achieve the same throughput, respectively. Code is publicly available1,
Akshay Karkal Kamath, Stefan Abi-Karam, Ashwin Bhat, Cong Hao
DATE3
2023 A Scalable Platform for Single-Snapshot Direction Of Arrival (DOA) Estimation in Massive MIMO Systems
abstract
With the development of Radio Frequency (RF) massive Multiple Inputs-Multiple Outputs (MIMO) array systems for Beyond 5G (B5G) applications, real-time DOA estimation has become challenging due to large antenna architectures producing a staggering amount of data. Traditional DOA estimation techniques are not scalable since they either require multiple snapshots of data or computationally expensive matrix operations hindering fast processing. To address these challenges, we propose a single-snapshot DOA processor based on the Alternating Direction Method of Multipliers (ADMM). The algorithm is modified to handle complex-valued measurements. We develop a High-Level Synthesis (HLS) based scalable FPGA design to handle multiple array sizes ranging from 8 to 512 elements. Our system implemented on a Xilinx Ultra96-V2 FPGA, operates at a frequency of 100 MHz with a sub-200μs processing time for a 512-antenna array, thereby meeting the millisecond-level processing time specifications of B5G applications.
Adou Sangbone Assoa, Ashwin Bhat, Sigang Ryu, Arijit Raychowdhury
ACM Great Lakes Symposium on VLSI2
2023 Non-Uniform Interpolation in Integrated Gradients for Low-Latency Explainable-AI
abstract
There has been a surge in Explainable-AI (XAI) methods that provide insights into the workings of Deep Neural Network (DNN) models. Integrated Gradients (IG) is a popular XAI algorithm that attributes relevance scores to input features commensurate with their contribution to the model's output. However, it requires multiple forward & backward passes through the model. Thus, compared to a single forward-pass inference, there is a significant computational overhead to generate the explanation which hinders real-time XAI. This work addresses the aforementioned issue by accelerating IG with a hardware-aware algorithm optimization. We propose a novel non-uniform interpolation scheme to compute the IG attribution scores which replaces the baseline uniform interpolation. Our algorithm significantly reduces the total interpolation steps required without adversely impacting convergence. Experiments on the ImageNet dataset using a pre-trained InceptionV3 model demonstrate 2.6-3.6×performance speedup on GPU systems for iso-convergence. This includes the minimal 0.2-3.2% latency overhead introduced by the pre-processing stage of computing the non-uniform interpolation step-sizes.
Ashwin Bhat, Arijit Raychowdhury
ISCAS1
2022 Exploration into the Explainability of Neural Network Models for Power Side-Channel Analysis
abstract
In this work, we present a comprehensive analysis of explainability of Neural Network (NN) models in the context of power Side-Channel Analysis (SCA), to gain insight into which features or Points of Interest (PoI) contribute the most to the classification decision. Although many existing works claim state-of-the-art accuracy in recovering secret key from cryptographic implementations, it remains to be seen whether the models actually learn representations from the leakage points. In this work, we evaluated the reasoning behind the success of a NN model, by validating the relevance scores of features derived from the network to the ones identified by traditional statistical PoI selection methods. Thus, utilizing the explainability techniques as a standard validation technique for NN models is justified.
Anupam Golder, Ashwin Bhat, Arijit Raychowdhury
ACM Great Lakes Symposium on VLSI2
2022 Gradient Backpropagation based Feature Attribution to Enable Explainable-AI on the Edge
abstract
There has been a recent surge in the field of Explainable AI (XAI) which tackles the problem of providing insights into the behavior of black-box machine learning models. Within this field, feature attribution encompasses methods which assign relevance scores to input features and visualize them as a heatmap. Designing flexible accelerators for multiple such algorithms is challenging since the hardware mapping of these algorithms has not been studied yet. In this work, we first analyze the dataflow of gradient backpropagation based feature attribution algorithms to determine the resource overhead required over inference. The gradient computation is optimized to minimize the memory overhead. Second, we develop a High-Level Synthesis (HLS) based configurable FPGA design that is targeted for edge devices and supports three feature attribution algorithms. Tile based computation is employed to maximally use on-chip resources while adhering to the resource constraints. Representative CNNs are trained on CIFAR-10 dataset and implemented on multiple Xilinx FPGAs using 16-bit fixed-point precision demonstrating flexibility of our library. Finally, through efficient reuse of allocated hardware resources, our design methodology demonstrates a pathway to repurpose inference accelerators to support feature attribution with minimal overhead, thereby enabling real-time XAI on the edge.
Ashwin Bhat, Adou Sangbone Assoa, Arijit Raychowdhury
VLSI-SoC1