VLDB 2026 Research / reviewers in the wild / expert
Kailash Prasad
dblp:282/9745
· DBLP profile ↗
9ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0002-4873-7728ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 2 first-author · 7 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | VeriBench: Benchmarking Large Language Models for Verilog Code Generation and Design SynthesisabstractIn the rapidly advancing field of hardware design, Electronic Design Automation (EDA) tools can be significantly improved using Machine Learning. This study evaluates the efficacy of various Large Language Models (LLMs) for automating Electronic Design Automation for Verilog design, testbench generation, and Formal Verification (FV) assertion synthesis by comparing 3 closed-source LLMs and 14 Open-Source LLM variants. In our setup of 33 Verilog designs, ChatGPT-4 generates 22 synthesizable Verilog designs in one-shot without feedback, while the Llama 3 (8B) model generates 20. Both models generate all testbenches correctly, 9 of which are given in our setup. For generating Formal Verification properties, ChatGPT-4 generates all properties correctly, whereas Llama 3 synthesizes 7 out of 9 properties correctly. Of the sample synthesized in Vivado, ChatGPT-4 codes result into power-efficient designs as compared to Llama-3, whereas in Genus there is no clear winner. These results underscore the efficacy of open-source models, which perform competitively despite having significantly fewer parameters (8 billion) compared to closed-source models such as ChatGPT-4. This study demonstrates the potential of parameter-efficient, open-source models for hardware design and verification tasks. Mihir Agarwal, Zaqi Momin, Kailash Prasad, Joycee Mekie |
ISCAS | 3 |
| 2023 | Hardware-Software Codesign of DNN Accelerators Using Approximate Posit MultipliersabstractEmerging data intensive AI/ML workloads encounter memory and power wall when run on general-purpose compute cores. This has led to the development of a myriad of techniques to deal with such workloads, among which DNN accelerator architectures have found a prominent place. In this work, we propose a hardware-software co-design approach to achieve system-level benefits. We propose a quantized data-aware POSIT number representation that leads to a highly optimized DNN accelerator. We demonstrate this work on SOTA SIMBA architecture, extendable to any other accelerator. Our proposal reduces the buffer/storage requirements within the architecture and reduces the data transfer cost between the main memory and the DNN accelerator. We have investigated the impact of using integer, IEEE floating point, and posit multipliers for LeNet, ResNet and VGG NNs trained and tested on MNIST, CIFAR10 and ImageNet datasets, respectively. Our system-level analysis shows that the proposed approximate-fixed-posit multiplier when implemented on SIMBA architecture, achieves on average ~2.2× speed up, consumes ~3.1× less energy and requires ~3.2× less area, respectively, against the baseline SOTA architecture, without loss of accuracy (~±1%) Tom Glint, Kailash Prasad, Jinay Dagli, Krishil Gandhi, Vrajesh Patel, Joycee Mekie |
ASP-DAC | 2 |
| 2023 | PVC-RAM:Process Variation Aware Charge Domain In-Memory Computing 6T-SRAM for DNNsabstractThis work introduces PVC-RAM, a process variation aware in-memory computing (IMC) static random-access memory (SRAM) macro designed for efficient convolutional neural network (CNN) inference. PVC-RAM is a charge-domain based compact IMC and is the first fully analog IMC for 4b-weights/4b-inputs MAC operation for deep neural networks in 6T-SRAM to the best of our knowledge. Further, PVC-RAM fully computes 4-bit MAC in the analog domain and requires fewer invocations of the ADCs. Implemented in 28nm technology, PVC-RAM achieves a bitwise throughput of 6964.48 TOPS, which is 1.5× higher than the SOTA, and a bitwise energy efficiency of 75.12 TOPS/W, which is 1.3× higher than the SOTA. Sai Shubham, Shubham Pandit, Kailash Prasad, Joycee Mekie |
DAC | 3 |
| 2023 | PIC-RAM: Process-Invariant Capacitive Multiplier Based Analog In Memory Computing in 6T SRAMabstractIn-Memory Computing (IMC) is a promising approach to enabling energy-efficient Deep Neural Network-based applications on edge devices. However, analog domain dot product and multiplication suffers accuracy loss due to process variations. Furthermore, wordline degradation limits its minimum pulsewidth, creating additional non-linearity and limiting IMC's dynamic range and precision. This work presents a complete end-to-end process invariant capacitive multiplier based IMC in 6T-SRAM (PIC-RAM). The proposed architecture employs the novel idea of two-step multiplication in column-major IMC to support 4-bit multiplication. The PIC-RAM uses an operational amplifier-based capacitive multiplier to reduce bitline discharge allowing good enough WL pulse width. Further, it employs process tracking voltage reference and fuse capacitor to tackle dynamic and post-fabrication process variations, respectively. Our design is compute-disturb free and provides a high dynamic range. To the best of our knowledge, PIC-RAM is the first analog SRAM IMC approach to tackle process variation with a focus on its practical implementation. PIC-RAM has a high energy efficiency of about 25.6 TOPS/W for$4-\text{bit}\times 4-\text{bit}$multiplication and has only 0.5% area overheads due to the use of the capacitance multiplier. We obtain 409 bit-wise TOPS/W, which is about 2× better than state-of-the-art. PIC-RAM shows the TOP-1 accuracy for ResNet-18 on CIFAR10 and MNIST is 89.54% and 98.80% for$4bit\times 4bit$multiplication. Kailash Prasad, Aditya Biswas, Arpita Kabra, Joycee Mekie |
DATE | 1 |
| 2023 | Process Variation Resilient Current-Domain Analog In Memory ComputingabstractIn-Memory Computing (IMC) has emerged as one of the energy-efficient solutions for data and compute-intensive machine learning applications. Analog IMC architectures have high throughput, but limited bit precision. Process variation further degrades the bit-precision. This work proposes an efficient way to track process variation and compensate for it to achieve high bit-resolution, which, to the best of our knowledge, is first such proposal. PV tracking is achieved by using an additional SRAM column and compensation by a non-conventional word-line driver. The proposed circuit can be augmented to any analog IMC architecture to make it resilient to process variations. To demonstrate the versatility of the proposal, we have implemented and analyzed 2-bit dot product operation in IMC architectures with six different SRAM cell configurations, and 2-bit, 4-bit, and 8-bit dot product on 6T SRAM IMC. For these, we report a reduction of$4\times$to$14\times$in the standard deviation of statistical variations in bit-line voltage for different SRAM cells, increase in the bit-resolution from 2 bits to 4 bits or 6 bits. Kailash Prasad, Sai Shubham, Aditya Biswas, Joycee Mekie |
DATE | 1 |
| 2023 | HyGain: High-performance, Energy-efficient Hybrid Gain Cell-based Cache HierarchyabstractIn this article, we propose a “full-stack” solution to designing high-apacity and low-latency on-chip cache hierarchies by starting at the circuit level of the hardware design stack. We propose a novel half V DD precharge 2T Gain Cell (GC) design for the cache hierarchy. The GC has several desirable characteristics, including ~50% higher storage density and ~50% lower dynamic energy as compared to the traditional 6T SRAM, even after accounting for peripheral circuit overheads. We also demonstrate data retention time of 350 us (~17.5× of eDRAM) at 28 nm technology with V DD = 0.9V and temperature = 27°C that, combined with optimizations like staggered refresh, makes it an ideal candidate to architect all levels of on-chip caches. We show that compared to 6T SRAM, for a given area budget, GC-based caches, on average, provide 30% and 36% increase in IPC for single- and multi-programmed workloads, respectively, on contemporary workloads, including SPEC CPU 2017. We also observe dynamic energy savings of 42% and 34% for single- and multi-programmed workloads, respectively. Finally, in a quest to utilize the best of all worlds, we combine GC with STT-RAM to create hybrid hierarchies. We show that a hybrid hierarchy with GC caches at L1 and L2 and an LLC split between GC and STT-RAM is able to provide a 46% benefit in energy-delay product (EDP) as compared to an all-SRAM design, and 13% as compared to an all-GC cache hierarchy, averaged across multi-programmed workloads. Sarabjeet Singh, Neelam Surana, Kailash Prasad, Pranjali Jain, Joycee Mekie, Manu Awasthi |
ACM Trans. Archit. Code Optim. | 3 |
| 2022 | An Automated Approach to Compare Bit Serial and Bit Parallel In-Memory Computing for DNNsabstractThis paper presents an exhaustive comparison of two different techniques for In-Memory Computing in SRAM: bit-serial arithmetic (BSA), and bit-parallel arithmetic (BPA). We have modeled both BSA and BPA and integrated them with CACTI for the CMOS 28nm technology node. The results are analyzed for both approaches with ten different sub-array configurations ranging from 128x128 to 2048x2048. We performed the convolution operation on the ImageNet dataset for comparison. The key observation is that the BPA begins to yield at least 25% better (lower) Energy Delay Product (EDP) as compared to that in BSA for large (2048x2048) sub-array sizes. The BSA yields $4 \times$ lower delay on an average, while the BPA yields $\sim 6 \times$ lower dynamic energy. Hence, the choice of the IMC architecture needs to be made depending on the application need (low energy/high performance). We have also simulated 12 different multi-bank IMC arrangements and show that just modifying the memory array structure can improve the EDP by up to $8 \times$. Alok Parmar, Kailash Prasad, Nanditha P. Rao, Joycee Mekie |
ISCAS | 2 |
| 2020 | FPAD: A Multistage Approximation Methodology for Designing Floating Point Approximate DividersabstractApproximate computing has emerged as a unique proposition for error-resilient applications such as image/video processing, neural networks, and the like, where both performance and power can be simultaneously reduced by trading off output quality. In this paper we propose a multistage approximation methodology for designing IEEE 754 floating point approximate dividers (FPADs). We propose a number of FPADs with varying upper bounds on error. For the same mean error, FPAD is 2.84× better in terms of power-delay product (PDP) as compared to state of the art approximate floating point divider. Further, when applied on applications, such as image enhancement, mean filtering and JPEG compression, FPAD outperforms the existing state-of-the-art approximate divider in terms of PDP by 25%. We also show that FPAD gives same PDP benefits in Alexnet convolutional neural network without noticeable drop in top-5 and top-1 accuracy. Chandan Kumar Jha 0001, Kailash Prasad, Vibhor Kumar Srivastava, Joycee Mekie |
ISCAS | 2 |
| 2020 | SEDAAF: FPGA Based Single Exact Dual Approximate Adders for Approximate ProcessorsabstractApproximate circuits for ASICs have gained immense traction in recent years due to the benefits obtained in both energy and performance with little or no loss in output quality. Approximation in FPGAs remain a challenge due to the higher level of granularity at which logic is implemented on FPGAs. The smallest configurable blocks in FPGAs used for implementing logic consists of the look up tables (LUTs). In this paper, we exploit the inherent structures available in the FPGAs to implement SEDAAF. SEDAAF is a runtime configurable approximate adder that can perform a one-bit exact addition or two-bit approximate addition using the same hardware. SEDAAF also has a maximum bounded error, i.e. for an n-bit adder if m-bits are approximated the maximum error is 2m- 1. SEDAAF consumes 25% lesser power and has a 17% lesser power delay product as compared to existing designs. SEDAAF outperforms the existing state of the art designs in terms of output quality for Sobel edge detection application and can be used in approximate processors for performing both exact and approximate additions. Chandan Kumar Jha 0001, Kailash Prasad, Arun Singh Tomar, Joycee Mekie |
ISCAS | 2 |