EDBT 2026 Demo / reviewers in the wild / expert
Bindu G. Gowda
dblp:321/5858
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2024
0000-0003-2797-2363ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Integrated MAC-based Systolic Arrays: Design and Performance EvaluationabstractIn the rapidly advancing landscape of computing, hardware accelerator designs are pivotal for satisfying high performance and low power demands. Systolic array (SA) architectures, tailored for general matrix multiplication (GEMM) operations, are ideal for image processing workloads. In this work, an integrated MAC (IMAC)-factored SA is proposed. Unlike prior focus on standalone multipliers and adders, IMAC optimizes multiplier-and-accumulator (MAC) units. The new IMAC approach was introduced to three category of Processing Elements (PE) that define SAs, and were further evaluated against four state-of-the-art (SOTA) SA designs. IMAC-SA reported noteworthy advantages: a design footprint reduction of 17.30% to 26.40%, power savings from 5.46% to 15.85%, and a maximum critical path delay improvement of 9.47% over other SOTA designs. Dantu Nandini Devi, Gandi Ajay Kumar, Bindu G. Gowda, Madhav Rao |
ACM Great Lakes Symposium on VLSI | 3 |
| 2023 | IMAC: : A Pre-Multiplier And Integrated Reduction Based Multiply-And-Accumulate UnitabstractMultiply-and-accumulate (MAC) units are primarily utilized for convolution operations targeted towards signal and image processing workload. The compressors are applied at the partial product reduction stages to extract the multiplier output bits, which are later accumulated with an extra adder unit. The paper proposes an integrated approach where the other operand of the MAC unit is directly fed to the partial-product-matrix (PPM) before the product bits are evaluated. This integrated Multiplier-and-Accumulate (IMAC) approach saves an additional adder unit and instead extends the compressor, which is already used to reduce partial-product bits of the multiplier design. Compressors employed exact and approximate IMAC architectures were designed and evaluated through ASIC and FPGA flow. Five versions of inexact IMAC design were independently compared with traditional one-level approximation and two-level approximation in MAC designs. The proposed work is found to be hardware efficient when compared with state-of-art MAC units. The error metrics were either comparable or better for IMAC design when compared with separately designed approximate multipliers followed by exact or approximate adder units. The image blending application was considered to measure the quality metrics. The proposed IMAC design files are made freely available for further usage by the research and development community. Bindu G. Gowda, Prashanth H. C., Madhav Rao |
ACM Great Lakes Symposium on VLSI | 1 |
| 2023 | EBASA: Error Balanced Approximate Systolic Array Architecture DesignabstractSystolic array (SA) is an architecture which is conceptually similar to an arithmetic pipeline and is created by uniformly connecting group of identical data processing elements (PE). Approximate computing benefits in hardware and performance, but incurs accuracy loss, thereby limiting it to error-resilient applications. Majority of inexact multipliers offer one-sided Error Distribution (ErD), and SA architecture with such multipliers results in large accumulated errors. This paper investigates SA architecture with various arrangement of approximate multipliers (AM) with dissimilar ErD for image smoothing and outline extracting applications. Among all the patterns, the Ring arrangement comprising of AMs with opposite-sided ErD placed in nested loops of the SA, was found to accelerate performance by 22.31%, and enhance image quality metrics by 18.15%. For FPGA implementation, alternate arrangement with equal number of AMs with opposite-sided ErD in the SA offered 12.14% LUT savings and comparable flip-flops usage when compared with one-sided AMs in the SA. Sai Karthik Nandigama, Bindu G. Gowda, Prashanth H. C., Madhav Rao |
ACM Great Lakes Symposium on VLSI | 2 |
| 2023 | ApproxCNN: Evaluation Of CNN With Approximated Layers Using In-Exact MultipliersabstractApproximate computing in the hardware design space has gained attention owing to the significant benefits achieved in power savings, performance improvement, and compact spacing. Most of the arithmetic operations including addition, multiplication, division, and even a few of the activation functions are realized using approximate computing techniques in the past. These are primarily applicable for error-resilient image processing applications where the visual perception of humans has limited abilities to distinguish between the original and approximated image. This has been extended to other signal-processing domains under similar constraints on human sensing modalities. Although most of the work intends to apply approximate computing on Artificial Intelligence (AI) workloads, but realizing approximate computing on the same has not been feasible. Hence, this has deprived to feel the impact of hardware benefits in conjunction with network accuracy compromises. This research work aims to establish the adoption of a specific set of approximate multipliers in the Convolutional Neural Network (CNN) and present hardware characteristics along with the validation accuracy of the network. Eight different approximate multipliers that are categorized along the positive and negative error distributions are applied to 6-layer, 3-layer, and 1-layer CNNs that are trained on benchmark datasets of CIFAR-10, MNIST, and F-MNIST respectively. The extensive design space exploration has allowed extracting the optimal sequence of approximate multipliers along different layers of CNNs reporting hardware gain and least accuracy drop. For the 6-layered CNN, the optimal hardware design with approximation applied for 1st, and 3rdlayer offered hardware gains of 16.5%, 10.2%, 2.4%, and 18.2% in power, footprint, delay, and PDP respectively, with comparable validation accuracy as that of exact CNN. Similarly, the best 3 layered approximated CNN model offered an improvement of 38.09%, 31.57%, 5.97%, and 43.84% in power, area, delay, and PDP parameters respectively over exact CNN model, without much drop in the validation accuracy. The 3 layered CNN demonstrated alternate arrangement of positive and negative error distributed type of multipliers along the layers to minimize the errors and maintain model accuracy close to the original exact multiplier adopted model. The correlation between five different error metrics for the sequence of approximate multipliers applied along the layers of CNNs and the overall accuracy loss in network inference were computed. The proposed work sets an example to leverage any approximate multipliers along the layers of CNN in the future and estimate the most optimal hardware design alongside retain the original network accuracy. The work is a step towards designing custom power-performance-area efficient neural network accelerators. Bindu G. Gowda, S. N. Raghava, Prashanth H. C., Pratyush Nandi, Madhav Rao |
ICCD | 1 |
| 2022 | Design and Evaluation of In-Exact Compressor based Approximate MultipliersabstractVLSI implementation of arithmetic functions are of high demand considering the rise in hardware realization of image and digital signal processing modules for various autonomous applications. The hardware implementation offers faster results and desirable outcome, but expecting the same design metrics in the form of power, footprint and delay on a tiny decision-making edge devices with limited resources needs design improvisation. Approximate computing promises to support the required hardware metrics in error resilient applications where the inexact output is not deviated much from the expected one, and decision made remains unchanged. Multiplier design blocks are heavily used in the multimedia functional chip, and introducing approximation in these blocks effectively benefits design metrics and chip cost of the developed system-on-chip(SoC). The proposed work attempts to design and use various sizes of approximate AND-OR re-coded compressors in the multiple reduction stages, along with various fast adders in the final addition stage of multiplier design. Further, design metrics and resources utilized for different multiplier designs were characterized in ASIC and FPGA synthesis flows respectively, along with their error statistics. Designed approximate multipliers were employed in Gaussian smoothing application to evaluate the quality-hardware resource trade-off of approximation Prashanth H. C., Soujanya S. R, Bindu G. Gowda, Madhav Rao |
ACM Great Lakes Symposium on VLSI | 3 |