Akshat Ramachandran

dblp:337/1025 · DBLP profile ↗
← Back
9ranked-venue papers
8as first author
9since 2021 · last 2026
0009-0000-4763-3321ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 5 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Algorithm-Hardware Co-Design of Digital Compute-in-Memory Architecture Supporting Flexible and Temporal N:M Sparsity
abstract
Structured pruning with fixed N:M sparsity ratios in large language models (LLMs) significantly constrains model expressivity, often leading to suboptimal accuracy. While supporting multiple N:M configurations can enhance representational flexibility, such a capability typically introduces substantial hardware complexity and overhead. To overcome these limitations, we first present FLOW, a flexible, layer-wise, outlier-density-aware N:M sparsity selection framework. FLOW adaptively determines the optimal N and M values per layer within a specified range by jointly considering the magnitude and distribution of outliers, thereby improving sparsity allocation and model fidelity. Extending this idea, we introduce FLOW++, which generalizes flexible N:M sparsity to the temporal domain for reasoning LLMs. FLOW++ enables temporal pruning by dynamically adapting sparsity patterns across reasoning thought types. To support efficient deployment of models with such dynamically varying sparsity patterns, we propose FlexCiM, a flexible, low-overhead, digital compute-in-memory (DCiM) architecture. FlexCiM partitions the DCiM macro into smaller submacros, which are adaptively aggregated and disaggregated through distribution and merging mechanisms for different values of N and M. We conduct experiments across different LLM families and state-space models and conclusively demonstrate that the proposed algorithm-hardware co-design framework achieves up to 36% higher accuracy,$1.75\times $faster inference, and$1.5\times $lower energy consumption compared to existing alternatives, establishing an effective balance between flexibility and hardware efficiency in sparse LLM inference. Code is available at:https://github.com/FLOW-open-project/FLOW
Akshat Ramachandran, Souvik Kundu 0009, Arnab Raha, Shamik Kundu, Deepak Mathaikutty, Tushar Krishna
IEEE Trans. Very Large Scale Integr. Syst.1
2025 AIRCHITECT v2: Learning the Hardware Accelerator Design Space Through Unified Representations
abstract
Design space exploration (DSE) plays a crucial role in enabling custom hardware architectures, particularly for emerging applications like AI, where optimized and specialized designs are essential. With the growing complexity of deep neural networks (DNNs) and the introduction of advanced large language models (LLMs), the design space for DNN accelerators is expanding at an exponential rate. Additionally, this space is highly non-uniform and non-convex, making it increasingly difficult to navigate and optimize. Traditional DSE techniques rely on search-based methods, which involve iterative sampling of the design space to find the optimal solution. However, this process is both time-consuming and often fails to converge to the global optima for such design spaces. Recently, AIRCHITECT vl, the first attempt to address the limitations of search-based techniques, transformed DSE into a constant-time classification problem using recommendation networks. However, AIRCHITECT v1 lacked generalizability and had poor performance in complex design spaces. In this work, we propose AIRCHITECT v2, a more accurate and generalizable learning-based DSE technique applicable to large-scale design spaces that overcomes the shortcomings of earlier approaches. Specifically, we devise an encoder-decoder transformer model that ($a$) encodes the complex design space into a uniform intermediate representation using contrastive learning and (b) leverages a novel unified representation blending the advantages of classification and regression to effectively explore the large DSE space without sacrificing accuracy. Experimental results evaluated on 105real DNN workloads demonstrate that, on average, AIRCHITECT v2 outperforms existing techniques by 15% in identifying optimal design points. Furthermore, to demonstrate the generalizability of our method, we evaluate performance on unseen model workloads and attain a 1.7 x improvement in inference latency on the identified hardware architecture. Code and dataset are available at: https://github.com/maestro-project/AIrchitect-v2.
Jamin Seo, Akshat Ramachandran, Yu-Chuan Chuang, Anirudh Itagi, Tushar Krishna
DATE2
2025 Ouromamba: a Data-Free Quantization Framework for Vision Mamba
abstract
We present OuroMamba, the first data-free post-training quantization (DFQ) method for vision Mamba-based models (VMMs). We identify two key challenges in enabling DFQ for VMMs, (1) VMM's recurrent state transitions restricts capturing of long-range interactions and leads to semantically weak synthetic data, (2) VMM activations exhibit dynamic outlier variations across time-steps, rendering existing static PTQ techniques ineffective. To address these challenges, OuroMamba presents a two-stage framework: (1) OuroMamba-Gen to generate semantically rich and meaningful synthetic data. It applies contrastive learning on patch level VMM features generated through neighborhood interactions in the latent state space, (2) OuroMamba-Quant to employ mixed-precision quantization with lightweight dynamic outlier detection during inference. In specific, we present a thresholding based outlier channel selection strategy for activations that gets updated every time-step. Extensive experiments across vision and generative tasks show that our data-free OuroMamba surpasses existing data-driven PTQ techniques, achieving state-of-the-art performance across diverse quantization settings. Additionally, we implement efficient GPU kernels to achieve practical latency speedup of up to 2.36x. Code and synthetic dataset are available here: https://github.com/georgia-tech-synergy-lab/ICCV-OuroMamba
Akshat Ramachandran, Souvik Kundu 0009, Tushar Krishna
ICCV1
2025 MicroScopiQ: Accelerating Foundational Models through Outlier-Aware Microscaling Quantization
Akshat Ramachandran, Souvik Kundu 0009, Tushar Krishna
ISCA1
2025 Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator
abstract
Large language model (LLM) pruning with fixed N:M structured sparsity significantly limits the expressivity of the sparse model, yielding sub-optimal performance. On the contrary, support for more than one N:M pattern to provide sparse representational freedom yields a costly overhead in the hardware. To mitigate these challenges for LLMs, we first present a flexible layer-wise outlier-density-aware N:M sparsity (FLOW) selection method. FLOW enables the identification of optimal layer-wise N and M values (from a given range) by simultaneously accounting for the presence and distribution of outliers, allowing a higher degree of representational freedom. To deploy the sparse models with such N:M flexibility, we then present a flexible low overhead, digital computein-memory architecture (FlexCiM). FlexCiM enables support for diverse sparsity patterns by partitioning a digital CiM (DCiM) macro into smaller sub-macros which are adaptively aggregated and disaggregated through distribution and merging mechanisms for different values of N and M. Extensive experiments on both transformer-based and recurrence-based state space foundation models (SSMs) demonstrate FLOW to outperform existing alternatives with an accuracy improvement of up to 36%, while FlexCiM delivers up to 1.75× lower inference latency and 1.5× lower energy consumption compared to existing sparse accelerators. Code is available at: https://github.com/FLOW-open-project/FLOW
Akshat Ramachandran, Souvik Kundu 0009, Arnab Raha, Shamik Kundu, Deepak K. Mathaikutty, Tushar Krishna
ISLPED1
2024 Algorithm-Hardware Co-Design of Distribution-Aware Logarithmic-Posit Encodings for Efficient DNN Inference
abstract
Traditional Deep Neural Network (DNN) quantization methods using integer or floating-point data types struggle to capture diverse DNN parameter distributions and often require large silicon overhead and intensive quantization-aware training. In this study, we introduce Logarithmic Posits (LP), an adaptive, hardware-friendly data type inspired by posits that dynamically adapts to DNN weight/activation distributions by parameterizing LP bit fields. We also develop a novel genetic-algorithm based framework, LP Quantization (LPQ), to find optimal layer-wise LP parameters while reducing representational divergence between quantized and full-precision models through a novel global-local contrastive objective. Additionally, we design a LP accelerator (LPA) architecture comprising of mixed-precision LP processing elements (PEs). Our algorithmhardware co-design demonstrates on average <1% drop in top-1 accuracy across various CNN and ViT models. It also achieves ~ 2× improvements in performance per unit area and 2.2× gains in energy efficiency compared to state-of-the-art quantization accelerators using different data types.
Akshat Ramachandran, Zishen Wan, Geonhwa Jeong, John L. Gustafson, Tushar Krishna
DAC1
2024 CLAMP-ViT: Contrastive Data-Free Learning for Adaptive Post-training Quantization of ViTs
Akshat Ramachandran, Souvik Kundu 0002, Tushar Krishna
ECCV (67)1
2023 NTrans-Net: A Multi-Scale Neutrosophic-Uncertainty Guided Transformer Network for Indoor Depth Completion
abstract
Dense depth maps are important constituents in a variety of tasks and have wide ranging applications. However, depth maps captured by indoor depth sensors have an extensive range of missing depth values and are also sparse in nature. Predicting dense depth from sparse input has been widely studied and is solved either as a regression or classification problem. In this paper, we propose a novel representation, termed Unified Ordinal Vectors, to realise the combined advantages of regression and classification methods. To disinter the potential of this representation, we propose NTrans-Net, a novel multi-scale network that can extract hierarchical and complementary information with neutrosophic indeterminacy feature handling. We also propose a dual encoder-decoder transformer structure to handle these neutrosophic domain features with guided attention to better capture inter-modal dependencies for superior depth completion performance. NTrans-Net is designed to be flexible enough to adapt to the dynamic nature of spatial input contexts and be robust to sensor-dependent distributions. We conduct extensive experiments on NYUv2 and ToF18K datasets to demonstrate the superiority of the proposed method in multiple settings, especially in realistic indoor environments as captured by commodity depth sensors.
Akshat Ramachandran, Ankit Dhiman, Basavaraja S. Vandrotti
ICIP1
2022 PositIV: A Configurable Posit Processor Architecture for Image and Video Processing
abstract
Image processing is essential for applications such as robot vision, remote sensing, computational photography, augmented reality etc. In the design of dedicated hardware for such applications, IEEE Std 754™ floating point (float) arithmetic units have been widely used. While float-based architectures have achieved favorable results, their hardware is complicated and requires a large silicon footprint. In this paper we propose a Posit-based Image and Video processor (PositIV), a completely pipelined, configurable, image processor using posit arithmetic that guarantees lower power use and smaller silicon footprint than floats. PositIV is able to effectively overlap computation with memory access and supports multidimensional addressing, virtual border handling, prefetching and buffering. It is successfully able to integrate configurability, flexibility, and ease of development with real-time performance characteristics. The performance of PositIV is validated on several image processing algorithms for different configurations and compared against state-of-the-art implementations. Additionally, we empirically demonstrate the superiority of posits in processing images for several conventional algorithms, achieving at least 35–40% improvement in image quality over standard floats.
Akshat Ramachandran, John Gustafson, Anusua Roy, Rizwan Ahmed Ansari, Rohin D. Daruwala
DSD1