EDBT 2026 Demo / reviewers in the wild / expert
Prashanth H. C.
dblp:321/5744
· DBLP profile ↗
10ranked-venue papers
6as first author
10since 2021 · last 2024
0000-0002-9650-3731ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Performance Analysis of OFA-NAS ResNet Topologies Across Diverse Hardware Compute UnitsabstractNetwork architecture search (NAS) is a tedious process and hence a different approach to train a large over parameterized network, followed by a progressively shrinking algorithm towards targeting the best efficient models for hardware platforms is practiced, which is referred to as Once-For-All (OFA) network. This paper focuses on utilizing OFA-defined NAS runs on ResNet topologies for wide range of hardware platforms ranging from workstations CPU, GPU, mobile CPU, GPUs, VPU, and DPU run on Xilnx FPGA. The OFA extracted model runs on the VPU unit offers speed improvement of 86% and 46% over mobile CPU A76, and mobile GPU 630 system respectively. The latency improvement of 100% was achieved for 3-threaded execution over single-thread for the DPU unit. Roofline model of OFA defined NAS extracted ResNet topologies indicated that all models are compute-bound when operated on Threadripper 960X CPU workstation; with 16-threaded execution offers maximum performance for batch size of 16. For GPU workstation, a throughput improvement of 66.66% to 100% was achieved for the OFA NAS generated models when configured to run on 3-threads over single-thread. Roofline performance analysis for the OFA NAS extracted model running on a Xilnx FPGA ZCU102 showed that most of the models are compute-bound for single and multi-threaded execution. Prashanth H. C., Madhav Rao |
ACM Great Lakes Symposium on VLSI | 1 |
| 2023 | GCells: A Graph-Search Approach to Design Custom Cells for Computational SubsystemsabstractStandard cell design is a challenge considering its impact on the overall synthesis of the design. The standard cells are generally offered by the manufacturers and the cells are updated to match the advancement in the fabrication facility. However matching of the system designed by the non-manufacturing team with the standard cells may not always present the best results. The exploration of standard cell design is very limited and not much is disclosed due to intellectual protection (IPs). This paper proposes a graph-search algorithm to extract the popular set of connected nodes, where the graph represents the computational subsystem and the nodes reflect individual gates. The proposed graph-search approach was applied on MAC designs of different bit-widths to extract top four ranked custom cells of four inputs and further characterized to incorporate them in the CMOS implemented standard cell library. The extracted top four cells were optimized for transistor widths to three different targets including critical path delay, product-of-power-and-delay (PDP), and equal rising and falling resistances separately. The augmented library with extracted custom cells was applied for synthesizing MAC design of four different bit-widths using three different set of library cells. The custom cells in the range of 38.23% to 52.80% were mapped in the synthesized versions of MAC designs for all the three optimized versions of cells, which validates the approach and emphasizes the need to explore new cells for acquiring performance and power-efficient subsystem designs. The standard cell library incorporated with custom cells offered maximum performance improvement of 35.3% and power savings of 56% over the standard cell library when synthesized for MAC designs of different bit-widths. The custom library synthesized MAC designs when adopted as hardware accelerator units for LeNet, AlexNet, and VGG-16 network showcased a performance gain of 23% to 33.80%, over the MAC designs synthesized by the original standard cell library. The graph-search approach is a step towards automating the custom cells for the design under synthesis. The approach has the potential to realize complex functions with customized cells, applicable to modern day SoC design. Mayank Kabra, Shreyas V. S, Prashanth H. C., Kedar Deshpande, Madhav Rao |
DSD | 3 |
| 2023 | IMAC: : A Pre-Multiplier And Integrated Reduction Based Multiply-And-Accumulate UnitabstractMultiply-and-accumulate (MAC) units are primarily utilized for convolution operations targeted towards signal and image processing workload. The compressors are applied at the partial product reduction stages to extract the multiplier output bits, which are later accumulated with an extra adder unit. The paper proposes an integrated approach where the other operand of the MAC unit is directly fed to the partial-product-matrix (PPM) before the product bits are evaluated. This integrated Multiplier-and-Accumulate (IMAC) approach saves an additional adder unit and instead extends the compressor, which is already used to reduce partial-product bits of the multiplier design. Compressors employed exact and approximate IMAC architectures were designed and evaluated through ASIC and FPGA flow. Five versions of inexact IMAC design were independently compared with traditional one-level approximation and two-level approximation in MAC designs. The proposed work is found to be hardware efficient when compared with state-of-art MAC units. The error metrics were either comparable or better for IMAC design when compared with separately designed approximate multipliers followed by exact or approximate adder units. The image blending application was considered to measure the quality metrics. The proposed IMAC design files are made freely available for further usage by the research and development community. Bindu G. Gowda, Prashanth H. C., Madhav Rao |
ACM Great Lakes Symposium on VLSI | 2 |
| 2023 | EBASA: Error Balanced Approximate Systolic Array Architecture DesignabstractSystolic array (SA) is an architecture which is conceptually similar to an arithmetic pipeline and is created by uniformly connecting group of identical data processing elements (PE). Approximate computing benefits in hardware and performance, but incurs accuracy loss, thereby limiting it to error-resilient applications. Majority of inexact multipliers offer one-sided Error Distribution (ErD), and SA architecture with such multipliers results in large accumulated errors. This paper investigates SA architecture with various arrangement of approximate multipliers (AM) with dissimilar ErD for image smoothing and outline extracting applications. Among all the patterns, the Ring arrangement comprising of AMs with opposite-sided ErD placed in nested loops of the SA, was found to accelerate performance by 22.31%, and enhance image quality metrics by 18.15%. For FPGA implementation, alternate arrangement with equal number of AMs with opposite-sided ErD in the SA offered 12.14% LUT savings and comparable flip-flops usage when compared with one-sided AMs in the SA. Sai Karthik Nandigama, Bindu G. Gowda, Prashanth H. C., Madhav Rao |
ACM Great Lakes Symposium on VLSI | 3 |
| 2023 | ApproxCNN: Evaluation Of CNN With Approximated Layers Using In-Exact MultipliersabstractApproximate computing in the hardware design space has gained attention owing to the significant benefits achieved in power savings, performance improvement, and compact spacing. Most of the arithmetic operations including addition, multiplication, division, and even a few of the activation functions are realized using approximate computing techniques in the past. These are primarily applicable for error-resilient image processing applications where the visual perception of humans has limited abilities to distinguish between the original and approximated image. This has been extended to other signal-processing domains under similar constraints on human sensing modalities. Although most of the work intends to apply approximate computing on Artificial Intelligence (AI) workloads, but realizing approximate computing on the same has not been feasible. Hence, this has deprived to feel the impact of hardware benefits in conjunction with network accuracy compromises. This research work aims to establish the adoption of a specific set of approximate multipliers in the Convolutional Neural Network (CNN) and present hardware characteristics along with the validation accuracy of the network. Eight different approximate multipliers that are categorized along the positive and negative error distributions are applied to 6-layer, 3-layer, and 1-layer CNNs that are trained on benchmark datasets of CIFAR-10, MNIST, and F-MNIST respectively. The extensive design space exploration has allowed extracting the optimal sequence of approximate multipliers along different layers of CNNs reporting hardware gain and least accuracy drop. For the 6-layered CNN, the optimal hardware design with approximation applied for 1st, and 3rdlayer offered hardware gains of 16.5%, 10.2%, 2.4%, and 18.2% in power, footprint, delay, and PDP respectively, with comparable validation accuracy as that of exact CNN. Similarly, the best 3 layered approximated CNN model offered an improvement of 38.09%, 31.57%, 5.97%, and 43.84% in power, area, delay, and PDP parameters respectively over exact CNN model, without much drop in the validation accuracy. The 3 layered CNN demonstrated alternate arrangement of positive and negative error distributed type of multipliers along the layers to minimize the errors and maintain model accuracy close to the original exact multiplier adopted model. The correlation between five different error metrics for the sequence of approximate multipliers applied along the layers of CNNs and the overall accuracy loss in network inference were computed. The proposed work sets an example to leverage any approximate multipliers along the layers of CNN in the future and estimate the most optimal hardware design alongside retain the original network accuracy. The work is a step towards designing custom power-performance-area efficient neural network accelerators. Bindu G. Gowda, S. N. Raghava, Prashanth H. C., Pratyush Nandi, Madhav Rao |
ICCD | 3 |
| 2023 | Meta-Heuristic Optimization of Transistor Sizing in CMOS Digital Designs
Prashanth H. C., Madhav Rao |
IJCCI | 1 |
| 2022 | Evolutionary Standard Cell Synthesis of Unconventional DesignsabstractConventional synthesis algorithms transform the behavioral RTL design to a standard cell mapped gate level netlist, with support to customize optimization effort of few operators. HDL description standards and current synthesis methods lack support to generate netlist of custom functions for quick validation and characterization of the design. Additionally, synthesis does not cater directly to various mathematical functions, design efforts towards approximating the desired function is needed. Hence a synthesis method for realizing circuits applicable to not only arithmetic but also to non-linear functions will be highly valuable and appreciated among the VLSI design community. This work employs Cartesian Genetic Programming (CGP) algorithm, an evolutionary design methodology suitable to synthesize digital circuits. CGP benefits in accelerating the design process and offers the ease to realize complex functions with little to no design effort. Activation functions are difficult to realize as combinational circuits using traditional design methods, this work validates the synthesis results for 6 non-linear activation functions using both classical and standard cell synthesis oriented CGP. The ability to incorporate such unconventional designs to the traditional synthesis flow will be instrumental for implementing accelerators in hardware space, and eventually for efficient design of heterogeneous SoC systems. Prashanth H. C., Madhav Rao |
ACM Great Lakes Symposium on VLSI | 1 |
| 2022 | Design and Evaluation of In-Exact Compressor based Approximate MultipliersabstractVLSI implementation of arithmetic functions are of high demand considering the rise in hardware realization of image and digital signal processing modules for various autonomous applications. The hardware implementation offers faster results and desirable outcome, but expecting the same design metrics in the form of power, footprint and delay on a tiny decision-making edge devices with limited resources needs design improvisation. Approximate computing promises to support the required hardware metrics in error resilient applications where the inexact output is not deviated much from the expected one, and decision made remains unchanged. Multiplier design blocks are heavily used in the multimedia functional chip, and introducing approximation in these blocks effectively benefits design metrics and chip cost of the developed system-on-chip(SoC). The proposed work attempts to design and use various sizes of approximate AND-OR re-coded compressors in the multiple reduction stages, along with various fast adders in the final addition stage of multiplier design. Further, design metrics and resources utilized for different multiplier designs were characterized in ASIC and FPGA synthesis flows respectively, along with their error statistics. Designed approximate multipliers were employed in Gaussian smoothing application to evaluate the quality-hardware resource trade-off of approximation Prashanth H. C., Soujanya S. R, Bindu G. Gowda, Madhav Rao |
ACM Great Lakes Symposium on VLSI | 1 |
| 2022 | SOMALib: Library of Exact and Approximate Activation Functions for Hardware-efficient Neural Network AcceleratorsabstractApproximate computing along with quantized low-precision computing has gained significant interest in today’s neural network (NN) implementation. This paper proposes a library of VLSI implementations of different activation functions, aimed towards designing hardware-efficient NN accelerators. Cartesian genetic programming (CGP), an evolutionary algorithm was employed to generate gate-level designs of approximate and exact representations of activation functions. We open-source the hardware library of 9444 circuits containing a majority of the activation functions employed in NN architectures, including Sigmoid, Hyperbolic-Tangent, Gaussian, ReLU, GeLU, Softplus, and Binary-Step. The library also presents the error characteristics and hardware metrics of the designs which will aid in the usage of the library in future research. Additionally a hardware comparison of the proposed circuits against existing implementations including piecewise-linear (PWL), memory-based, hls4ml, DNNweaver implementations to realize activation functions on FPGA and ASIC flow is presented. The CGP evolved hardware library shows minimal silicon space requirement, least power consumption when investigated for ASIC flow, and the least LUT utilization’s in FPGA flow. Besides, SOMALib designs are purely combinatorial, allowing various synthesis stage optimizations towards the target Power-Performance-Area budget, which is not possible in standard memory block implementations. Prashanth H. C., Madhav Rao |
ICCD | 1 |
| 2022 | Improving Digital Circuit Synthesis of Complex Functions using Binary Weighted Fitness and Variable Mutation Rate in Cartesian Genetic Programming
Prashanth H. C., Madhav Rao |
IJCCI | 1 |