VLDB 2026 Research / reviewers in the wild / expert
Pratyush Dhingra
dblp:365/4551
· DBLP profile ↗
8ranked-venue papers
5as first author
8since 2021 · last 2026
0009-0000-0422-3365ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 5 first-author · 8 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Focus Session: Hardware/Software Co-Design to Accelerate Generative AI Workloads on Heterogeneous ArchitecturesabstractEffective acceleration of Generative AI workloads is constrained by heterogeneous computational and memory characteristics. Large Language Models (LLMs), for example, are composed of diverse kernels, ranging from memory-bound attention mechanisms to compute-intensive feed-forward networks. Such variability leads to poor resource utilization and renders homogeneous architectures inefficient. This paper outlines a methodology to address this challenge using two synergistic approaches. First, we discuss the design of a three-dimensional (3D) manycore architecture enabled by heterogeneous integration. Specifically, the architecture leverages emerging non-volatile memory alongside CMOS-based computing cores to create an optimized kernel-to-compute mapping. Second, we explore a hardware and software co-design framework to enable fine-tuning of unimodal and multimodal LLMs. This framework ensures thermally robust inference when leveraging emerging non-volatile memories within the heterogeneous architecture. Experimental results on multiple LLMs demonstrate the efficacy of the 3D heterogeneous architecture and the co-design framework. Pratyush Dhingra, Vibhanshu Sharma, Janardhan Rao Doppa, Partha Pratim Pande |
DATE | 1 |
| 2026 | ThRIve: Thermally Robust CNN Inference via Low-Rank Adaptation in Heterogeneous PIM ArchitecturesabstractProcessing-In-Memory (PIM) has emerged as a promising technology for accelerating machine learning (ML) workloads. Specifically, non-volatile memory-based PIM architectures have enabled effective ML acceleration due to their ability to perform energy-efficient matrix-vector multiplication operations. However, these devices suffer from non-idealities such as thermal noise. This noise alters the stored values in the memory cells which correspond to actual model weights, compromising the inference accuracy. In this work, we introduce ThRIve, a noise-aware training methodology that leverages low-rank adaptation to enable thermally robust inference on heterogeneous PIM architectures. ThRIve selectively stores these low-rank noise-aware parameters on a hardware that is less susceptible to thermal noise, enabling robustness against temperature-induced noise variations. ThRIve mitigates the effects of thermal-noise and prevent the drop in inference accuracy across the entire operating temperature range. Experimental results demonstrate that ThRIve-enabled architectures maintain consistent inference accuracy, with the mean accuracy staying within 2% of the ideal (i.e., noise-free) accuracy, and the variation in accuracy across the entire operating temperature range remaining within 2% of the mean. The proposed methodology achieves accuracy and robustness comparable to thermally-resilient Static Random-Access Memory (SRAM)-based PIM systems, while delivering up to 5.4 × reduction in energy-delay product (EDP) during CNN model inferencing. Vibhanshu Sharma, Pratyush Dhingra, Janardhan Rao Doppa, Partha Pratim Pande |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2025 | Atleus: Accelerating Transformers on the Edge Enabled by 3D Heterogeneous Manycore ArchitecturesabstractTransformer architectures have become the standard neural network model for various machine learning (ML) applications, including natural language processing and computer vision. However, the compute and memory requirements introduced by transformer models make them challenging to adopt for edge applications. Furthermore, fine-tuning pretrained transformers (e.g., foundation models) is a common task to enhance the model’s predictive performance on specific tasks/applications. Existing transformer accelerators are oblivious to complexities introduced by fine-tuning. In this article, we propose the design of a three-dimensional (3D) heterogeneous architecture referred to as Atleus that incorporates heterogeneous computing resources specifically optimized to accelerate transformer models for the dual purposes of fine-tuning and inference. Specifically, Atleus utilizes nonvolatile memory and systolic array for accelerating transformer computational kernels using an integrated 3D platform. Moreover, we design a suitable NoC to achieve high performance and energy efficiency. Finally, Atleus adopts an effective quantization scheme to support model compression. Experimental results demonstrate that Atleus outperforms existing state-of-the-art by up to$56\times $and$64.5\times $in terms of performance and energy efficiency, respectively. Pratyush Dhingra, Janardhan Rao Doppa, Partha Pratim Pande |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2025 | ERGo: Energy-Efficient Hybrid Graph Neural Network Training on Processing-in-Memory ArchitecturesabstractProcessing-in-memory (PIM) has been proposed as an alternative computing paradigm for training Deep Neural Networks, including Graph Neural Networks (GNNs). Despite these advancements, training GNN workloads on PIM devices necessitates off-chip memory access. This off-chip access is expensive in terms of performance and energy, thereby impacting the overall energy efficiency of the training process on PIM platforms. In this article, we propose a novel hybrid training framework called ERGo that automatically switches from a full-parameter phase to a parameter-efficient phase during the training process with negligible loss in predictive accuracy. The parameter-efficient phase employs low-rank representations of GNN weights, effectively reducing the number of trainable parameters. This helps in reducing the off-chip access during the end-to-end training process on PIM architectures, leading to notable improvements in both performance and energy efficiency. ERGo outperforms the conventional full-parameter GNN training on a PIM-based platform by up to 3.15x in speedup and enhances energy efficiency by 10.5x. Furthermore, ERGo offers additional advantages when utilized on heterogeneous architectures incorporating non-volatile memory (NVM) and static random-access memory (SRAM). Training on the heterogeneous NVM-SRAM architecture typically utilizes NVM for the forward-pass and SRAM for the backward-pass. However, NVM cells suffer from low device endurance and repeated write operations due to weight updates can lead to poor lifetime. ERGo improves the lifetime of NVM devices by freezing the weights and preventing weight updates in the parameter-efficient phase. Experimental results demonstrate that ERGo improves the lifetime of heterogeneous NVM-SRAM architecture by up to 33x in comparison to full-parameter implementation. Pratyush Dhingra, Chibuike E. Ugwu, Janardhan Rao Doppa, Partha Pratim Pande |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2025 | A Heterogeneous Chiplet Architecture for Accelerating End-to-End Transformer ModelsabstractTransformers have revolutionized deep learning and generative modeling, enabling advancements in natural language processing tasks. However, the size of transformer models is increasing continuously, driven by enhanced capabilities across various deep learning tasks. This trend of ever-increasing model size has given rise to new challenges in terms of memory and compute requirements. Conventional computing platforms, including GPUs, suffer from suboptimal performance due to the memory demands imposed by models with millions/billions of parameters. The emerging chiplet-based platforms provide a new avenue for compute- and data-intensive machine learning applications enabled by a Network-on-Interposer (NoI). However, designing suitable hardware accelerators for executing Transformer inference workloads is challenging due to a wide variety of complex computing kernels in the Transformer architecture. In this article, we leverage chiplet-based heterogeneous integration to design a high-performance and energy-efficient multichiplet platform to accelerate transformer workloads. We demonstrate that the proposed NoI architecture caters to the data access patterns inherent in a transformer model. The optimized placement of the chiplets and the associated NoI links and routers enable superior performance compared to the state-of-the-art hardware accelerators. The proposed NoI-based architecture demonstrates scalability across varying transformer models and improves latency and energy efficiency by up to 11.8× and 2.36×, respectively when compared with the existing state-of-the-art architecture HAIMA. Pratyush Dhingra, Janardhan Rao Doppa, Ümit Y. Ogras, Partha Pratim Pande |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2025 | HpT: Hybrid Acceleration of Spatio-Temporal Attention Model Training on Heterogeneous Manycore ArchitecturesabstractTransformer models have become widely popular in numerous applications, and especially for building foundation large language models (LLMs). Recently, there has been a surge in the exploration of transformer-based architectures in non-LLM applications. In particular, the self-attention mechanism within the transformer architecture offers a way to exploit any hidden relations within data, making it widely applicable for a variety of spatio-temporal tasks in scientific computing domains (e.g., weather, traffic, agriculture). Most of these efforts have primarily focused on accelerating the inference phase. However, the computational resources required to train these attention-based models for scientific applications remain a significant challenge to address. Emerging non-volatile memory (NVM)-based processing-in-memory (PIM) architectures can achieve higher performance and better energy efficiency than their GPU-based counterparts. However, the frequent weight updates during training would necessitate write operations to NVM cells, posing a significant barrier for considering stand-alone NVM-based PIM architectures. In this paper, we presentHpT, a new hybrid approach to accelerate the training of attention-based models for scientific applications. Our approach is hybrid at two different layers: at the software layer, our approach dynamically switches from a full-parameter training mode to a lower-parameter training mode by incorporating intrinsic dimensionality; and at the hardware layer, our approach harnesses the combined power of GPUs, resistive random-access memory (ReRAM)-based PIM devices, and systolic arrays. This software-hardware co-design approach is aimed at adaptively reducing both runtime and energy costs during the training phase, without compromising on quality. Experiments on four concrete real-world scientific applications demonstrate that our hybrid approach is able to significantly reduce training time (up to$11.9\times$) and energy consumption (up to$12.05\times$), compared to the corresponding full-parameter training executing on only GPUs. Our approach serves as an example for accelerating the training of attention-based models on heterogeneous platforms including ReRAMs. Saiman Dahal, Pratyush Dhingra, Krishu K. Thapa, Partha Pratim Pande, Anantharaman Kalyanaraman |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2024 | FARe: Fault-Aware GNN Training on ReRAM-Based PIM AcceleratorsabstractResistive random-access memory (ReRAM)-based processing-in-memory (PIM) architecture is an attractive solution for training Graph Neural Networks (GNNs) on edge platforms. However, the immature fabrication process and limited write endurance of ReRAMs make them prone to hardware faults, thereby limiting their widespread adoption for GNN training. Further, the existing fault-tolerant solutions prove inadequate for effectively training GNNs in the presence of faults. In this paper, we propose a fault-aware framework referred to as FARe that mitigates the effect of faults during GNN training. FARe outperforms existing approaches in terms of both accuracy and timing overhead. Experimental results demonstrate that FARe framework can restore GNN test accuracy by 47.6% on faulty ReRAM hardware with a -1 % timing overhead compared to the fault-free counterpart. Pratyush Dhingra, Chukwufumnanya Ogbogu, Biresh Kumar Joardar, Janardhan Rao Doppa, Anantharaman Kalyanaraman, Partha Pratim Pande |
DATE | 1 |
| 2024 | HeTraX: Energy Efficient 3D Heterogeneous Manycore Architecture for Transformer AccelerationabstractTransformers have revolutionized deep learning and generative modeling to enable unprecedented advancements in natural language processing tasks and beyond. However, designing hardware accelerators for executing transformer models is challenging due to the wide variety of computing kernels involved in the transformer architecture. Existing accelerators are either inadequate to accelerate end-to-end transformer models or suffer notable thermal limitations. In this paper, we propose the design of a three-dimensional heterogeneous architecture referred to as HeTraX specifically optimized to accelerate end-to-end transformer models. HeTraX employs hardware resources aligned with the computational kernels of transformers and optimizes both performance and energy. Experimental results show that HeTraX outperforms existing state-of-the-art by up to 5.6x in speedup and improves EDP by 14.5x while ensuring thermally feasibility. Pratyush Dhingra, Janardhan Rao Doppa, Partha Pratim Pande |
ISLPED | 1 |