Pragnya Sudershan Nalla

dblp:366/8553 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2026
0009-0006-3688-6941ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 1 first-author · 6 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Chiplet-NAS: Chiplet-aware Neural Architecture Search for Efficient AI Inference on 2.5D Integration
abstract
The co-design of neural network architectures and their target chiplet-based hardware systems presents a significant challenge due to the vast and combinatorial design space. Identifying solutions that are Pareto-optimal across competing objectives of task accuracy, system latency, and power consumption requires solutions beyond manual design and brute-force methods. This paper proposes a closed-loop chiplet-aware neural architecture search (Chiplet-NAS) framework to automate the exploration and discover hardware-optimized models for efficient AI inference on 2.5 D chiplet-based systems. The framework integrates a Tree-structured Parzen Estimator (TPE) for sampleefficient search with CLAIRE, a chiplet-based library and fast performance benchmarking tool, to provide direct hardware feedback on latency and energy consumption, along with accuracy optimization. The framework is evaluated by co-designing ResNet-based model architectures with chiplet based hardware systems. Compared to a baseline NAS that optimizes only on the task accuracy, our Chiplet-NAS achieves significant power and performance benefits at the iso-accuracy.
Pragnya Sudershan Nalla, Nikhil K. Cherukuri, Sachin S. Sapatnekar, Chaitali Chakrabarti, Yu Cao 0001, Jeff Zhang 0001
ASP-DAC2
2026 vFPGA: Towards Sub-µs Reconfiguration via 3D FPGA and Packaging Co-Design
Nikhil K. Cherukuri, Sharad Nag, Pragnya Sudershan Nalla, Ashish K. Kola, Chetan S. Gadireddi, Kevin Dai, Jae-sun Seo, Zhenman Fang, Jeff Zhang 0001, Yu Cao 0001
FPGA3
2025 Invited: EDA for Heterogeneous Integration
abstract
The advent of heterogeneous integration (HI) places new demands on EDA tooling. Building large systems requires (1) methods for chiplet disaggregation that map the system to smaller chiplets, working in conjunction with system-technology co-optimization to determine the right design decisions that optimize computation and communication, together with the choice of substrate and chiplet technologies; (2) multiphysics and multiscale analyses that incorporate thermomechanical aspects into performance analysis, ranging from fast machine-learningdriven analyses in early stages to signoff-quality multiphysics-based analysis; (3) physical design techniques for placing and routing chiplets and embedded active/passive elements on and within the substrate, including the design of thermal and power delivery solutions; and (4) underlying infrastructure required to facilitate HI-based design, including the design and characterization of chiplet libraries and the establishment of data formats and standards. This paper overviews these issues and lays out a set of EDA needs for HI designs.
Emad Haque, Pragnya Sudershan Nalla, Chetal Choppali Sudarshan, Divya Yogi, Chaitali Chakrabarti, Vidya A. Chhabria, Ramesh Harjani, Jeff Zhang 0001, Sachin S. Sapatnekar
DAC2
2025 CLAIRE: Composable Chiplet Libraries for AI Inference
abstract
Artificial intelligence has made a significant impact on fields like computer vision, Natural Language Processing (NLP), healthcare, and robotics. However, recent AI models, such as GPT-4 and LLaMAv3, demand significant number of computational resources, pushing monolithic chips to their technological and practical limits. 2.5D chiplet-based heterogeneous architectures have been proposed to address these technological and practical limits. While chiplet optimization for models like Convolutional Neural Networks (CNNs) is well-established, scaling this approach to accommodate diverse AI inference models with different computing primitives, data volumes, and different chiplet sizes is very challenging. A set of hardened IPs and chiplet libraries optimized for a broad range of AI applications is proposed in this work. We derive the set of chiplet configurations that are composable, scalable and reusable by employing an analytical framework trained on a diverse set of AI algorithms. Testing these set of library synthesized configurations on a different set of algorithms, we achieve a$1.99\times-3.99\times$improvement in non-recurring engineering (NRE) chiplet design costs, with minimal performance overhead compared to custom chiplet-based ASIC designs. Similar to soft IPs for SoC development, the library of chiplets improves flexibility, reusability, and efficiency for AI hardware designs.
Pragnya Sudershan Nalla, Emad Haque, Yaotian Liu, Sachin S. Sapatnekar, Jeff Zhang 0001, Chaitali Chakrabarti, Yu Cao 0001
DATE1
2025 MAHL: Multi-Agent LLM-Guided Hierarchical Chiplet Design with Adaptive Debugging
abstract
As program workloads (e.g., AI) increase in size and algorithmic complexity, the primary challenge lies in their high dimensionality, encompassing computing cores, array sizes, and memory hierarchies. To overcome these obstacles, innovative approaches are required. Agile chip design has already benefited from machine learning integration at various stages, including logic synthesis, placement, and routing. With Large Language Models (LLMs) recently demonstrating impressive proficiency in Hardware Description Language (HDL) generation, it is promising to extend their abilities to 2.5D integration, an advanced technique that saves area overhead and development costs. However, LLM-driven chiplet design faces challenges such as flatten design, high validation cost and imprecise parameter optimization, which limit its chiplet design capability. To address this, we propose MAHL, a hierarchical LLM-based chiplet design generation framework that features six agents which collaboratively enable AI algorithm-hardware mapping, including hierarchical description generation, retrieval-augmented code generation, diverseflow-based validation, and multi-granularity design space exploration. These components together enhance the efficient generation of chiplet design with optimized Power, Performance and Area (PPA). Experiments show that MAHL not only significantly improves the generation accuracy of simple RTL design, but also increases the generation accuracy of real-world chiplet design, evaluated by Pass@5, from 0 to 0.72 compared to conventional LLMs under the best-case scenario. Compared to state-of-the-art CLARIE (expert-based), MAHL achieves comparable or even superior PPA results under certain optimization objectives.
Jinwei Tang, Jiayin Qin, Nuo Xu 0013, Pragnya Sudershan Nalla, Yu Cao 0001, Yang Zhao 0013, Caiwen Ding
ICCAD4
2025 HISIM: Analytical Performance Modeling and Design Space Exploration of 2.5D/3D Integration for AI Computing
abstract
Monolithic designs face significant fabrication cost and data movement challenges, especially when executing complex and diverse AI models. Advanced 2.5D/3D packaging promises high bandwidth and connection density to overcome these challenges, yet it also introduces new electro-thermal constraints. This article develops a suite of analytical performance models to enable efficient benchmarking of a 2.5D/3D heterogeneous system for energy-efficient AI computing. These models encompass various performance metrics related to computing units, network-on-chip (NoC), and network-on-package (NoP). The results are summarized into a new tool, HISIM, which is$10^{4} \times $–$10^{6} \times $faster than state-of-the-art AI benchmark tools. Furthermore, HISIM integrates rapid thermal simulation for the 2.5D/3D system, helping shed light on both the potential and limitations of 2.5D/3D heterogeneous integration (HI) on representative AI algorithms. The code of HISIM is available athttps://github.com/mec-UMN/HISIM.
Zhenyu Wang 0016, Pragnya Sudershan Nalla, Jingbo Sun 0003, A. Alper Goksoy, Sumit K. Mandal, Jae-sun Seo, Vidya A. Chhabria, Jeff Zhang 0001, Chaitali Chakrabarti, Ümit Y. Ogras, Yu Cao 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2