Kaustubh Manohar

dblp:270/9225 · also Kaustubh Manohar Mhatre · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
6since 2021 · last 2026
0009-0006-1256-1528ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Accelerating Topology Optimization on AMD Versal AIE-ML Engines
abstract
Topology optimization is a computational method used to determine the optimal material distribution within a prescribed design domain, aiming to minimize structural weight while satisfying load and boundary conditions. For critical infrastructure applications, such as real-time structural health monitoring of bridges and buildings, achieving real-time topology optimization is essential. Traditionally, topology optimization relies on finite element analysis (FEA), a computationally intensive process. Recent advances in deep neural networks (DNNs) have introduced data driven alternatives to FEA, substantially reducing computation time while maintaining solution quality. However, the inference latency of these neural network models, especially for complex designs, still remains a major obstacle to real-time deployment. To address this challenge, we propose a hardware accelerated implementation of a topology optimization neural network (CRONet) on the AMD Versal AI Engine-ML (AIE-ML) architecture. Our approach efficiently exploits the parallelism and memory hierarchy of AIE-ML engines to optimize the execution of various neural network operators. We develop custom, high-performance kernels for all network layers, including a modified version of the state-of-the-art GAMA framework for generalized matrix multiplication (GEMM) layers. Additionally, we offer integrated tools for performance simulation and functional verification prior to hardware deployment. Experimental results demonstrate that our implementation achieves 2× higher performance compared to an NVIDIA GTX 1050 Ti and up to 14× over a CPU baseline. These results highlight the potential of Versal AIE-ML based acceleration for enabling real-time topology optimization.
Kaustubh Manohar, Vedant Tewari, Aditya Ray, Ridwan Olabiyi, Ashif Iquebal, Aman Arora 0001
FCCM1
2025 Performance Analysis of GEMM Workloads on the AMD Versal Platform
abstract
AMD Versal is a new heterogeneous computing hardware architecture comprised of adaptive intelligence (AI) engines, programmable logic, and a processing system. General Matrix Multiplication (GEMM) is the fundamental building block of modern deep learning (DL) applications such as ChatGPT, and GEMM workloads can be mapped onto Versal in different ways, each with distinct trade-offs. This paper presents a thorough analysis of GEMM workloads of different shapes and sizes, showcasing performance artifacts associated with the AMD Versal architecture. Focusing on the unique aspects of the Versal architecture, multiple research questions related to performance scaling, sensitivity, and efficiency are explored. This paper aims to assist developers in the FPGA community looking to implement GEMM on AMD Versal by providing guidelines and insights for enhancing performance.
Kaustubh Manohar, Venkata Guru Prasanth Mulleti, Curt John Bansil, Endri Taka, Aman Arora 0001
FPGA1
2025 GAMA: High-Performance GEMM Acceleration on AMD Versal ML-Optimized AI Engines
abstract
General matrix-matrix multiplication (GEMM) is a fundamental operation in machine learning (ML) applications. We present the first comprehensive performance acceleration of GEMM workloads on AMD's second-generation AIE-ML architecture, which is specifically optimized for ML applications. Compared to AI-Engine (AIE), AIE-ML offers increased compute throughput and larger on-chip memory capacity. We propose a novel design that maximizes AIE-ML memory utilization, incorporates custom buffer placement within the AIE-ML and staggered kernel placement across the AIE-ML array, significantly reducing performance bottlenecks such as memory stalls and routing congestion, resulting in improved performance and efficiency compared to the default AMD's compiler. We evaluate the performance benefits of our design at three levels: single AIEML, pack of AIE-ML's and the complete AIE-ML array. GAMA achieves state-of-the-art performance, delivering up to$\mathbf{1 6 5}$TOPS (85% of peak) for int8 precision and 83 TBFLOPS (86% of peak) for bfloat16 precision GEMM workloads. Our solution achieves$8.7 \%, 9 \%, 39 \%$and 53.6 % higher peak throughput efficiency compared to the state-of-the-art AIE frameworks AMA, MAXEVA, ARIES and CHARM, respectively.
Kaustubh Manohar, Endri Taka, Aman Arora 0001
FPL1
2025 Performance Analysis of GEMM Workloads on the AMD Versal Platform
abstract
AMD Versal is a new heterogeneous computing hardware architecture comprised of adaptive intelligence (AI) engines, programmable logic, and a processing system. General Matrix Multiplication (GEMM) is the fundamental building block of modern deep learning (DL) applications such as ChatGPT, and GEMM workloads can be mapped onto Versal in different ways, each with distinct trade-offs. This paper presents a thorough analysis of GEMM workloads of different shapes and sizes, showcasing performance artifacts associated with the AMD Versal architecture. Focusing on the unique aspects of the Versal architecture, multiple research questions related to performance scaling, sensitivity, and efficiency are explored. This paper aims to assist FPGA developers looking to implement GEMM on AMD Versal by providing insights for enhancing performance.
Kaustubh Manohar, Venkata Guru Prashanth Mulleti, Curt John Bansil, Endri Taka, Aman Arora 0001
ISPASS1
2024 PIMSAB: A Processing-In-Memory System with Spatially-Aware Communication and Bit-Serial-Aware Computation
abstract
Bit-serial Processing-In-Memory (PIM) is an attractive paradigm for accelerator architectures, for parallel workloads such as Deep Learning (DL), because of its capability to achieve massive data parallelism at a low area overhead and provide orders-of-magnitude data movement savings by moving computational resources closer to the data. While many PIM architectures have been proposed, improvements are needed in communicating intermediate results to consumer kernels, for communication between tiles at scale, for reduction operations, and for efficiently performing bit-serial operations with constants. We present PIMSAB, a scalable architecture that provides a spatially aware communication network for efficient intra-tile and inter-tile data movement and provides efficient computation support for generally inefficient bit-serial compute patterns. Our architecture consists of a massive hierarchical array of compute-enabled SRAMs (CRAMs), which is codesigned with a compiler to achieve high utilization. The key novelties of our architecture are (1) in providing efficient support for spatially aware communication by providing local H-tree network for reductions, by adding explicit hardware for shuffling operands, and by deploying systolic broadcasting, as well as (2) by taking advantage of the divisible nature of bit-serial computations through adaptive precision and efficient handling of constant operations. These innovations are integrated into a tensor expressions-based programming framework (including a compiler for easy programmability) that enables simple programmer control of optimizations for mapping programs into massively parallel binaries for millions of PIM processing elements. When compared against a similarly provisioned modern Tensor Core GPU (NVIDIA A100), across common DL kernels and end-to-end DL networks (Resnet18 and BERT), PIMSAB outperforms the GPU by 4.80×, and reduces energy by 3.76×. We compare PIMSAB with similarly provisioned state-of-the-art SRAM PIM (Duality Cache) and DRAM PIM (SIMDRAM), and observe a speedup of 3.7× and 3.88×, respectively.
Kaustubh Manohar, Jian Weng 0002, Bagus Hanindhito, Zhengrong Wang, Tony Nowatzki, Lizy Kurian John, Aman Arora 0001
ACM Trans. Archit. Code Optim.2
2021 DeepDive: An Integrative Algorithm/Architecture Co-Design for Deep Separable Convolutional Neural Networks
abstract
Deep Separable Convolutional Neural Network (DSCNN) has become the emerging paradigm by offering modular networks with structural sparsity to achieve higher accuracy with relatively lower operations and parameters. However, there is a lack of customized architectures that can provide flexible solutions that fit the sparsity of the DSCNNs. This paper introduces DeepDive, a fully-functional vertical co-design framework, for power-efficient implementation of DSCNNs on edge FPGAs. DeepDive's architecture supports crucial heterogeneous Compute Units (CUs) to fully support DSCNNs with various convolutional operators interconnected with structural sparsity. It offers FPGA-aware training and online quantization combined with modular synthesizable C++ CUs, customized for DSCNNs. The execution results on Xilinx's ZCU102 FPGA board demonstrate 47.4 and 233.3 FPS/Watt for MobileNet-V2 and a compact version of EfficientNet, respectively, as two state-of-the-art depthwise separable CNNs. These comparisons showcase how DeepDive improves FPS/Watt by 2.2× and 1.51× over Jetson Nano high and low power modes, respectively. It also enhances FPS/Watt by about 2.27× and 37.25× over two other FPGA implementations.
Mohammadreza Baharani, Ushma Sunil, Kaustubh Manohar, Steven Furgurson, Hamed Tabkhi
ACM Great Lakes Symposium on VLSI3