EDBT 2026 Demo / reviewers in the wild / expert
Srinivasu Bodapati 0001
dblp:254/8984-1 · also B. Srinivasu 0001, Bodapati Srinivasu 0001
· DBLP profile ↗
7ranked-venue papers
1as first author
6since 2021 · last 2025
0000-0003-0974-8245ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 5 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | High-Performance Gemmini-Based Matrix Multiplication Accelerator for Deep Learning WorkloadsabstractTransformer models have recently gained considerable attention in computer vision applications due to their capacity to capture relations among features, resulting in enhanced performance. Moreover, deep neural networks (DNNs) have been extensively researched due to their superior performance and applicability in several domains, including image classification, detection, and recognition. Matrix multiplications are crucial operations in Transformers and DNN, and computations such as weight stationary (WS) and output stationary (OS) are considered to meet the dataflow constraints. This work presents a systolic array (SA)-based general matrix multiplication (GEMM) architecture for Gemmini accelerators, enabling the performance of convolutions in neural networks and self-attention in Transformers. First, this work proposes a multiplier-optimized SA (MOSA) by integrating a novel multiplier. The multiplier is designed using a single-stage stacking-based 3:2 and 4:2 compressors. The proposed MOSA achieves significant improvements, with area savings ranging from 12% to 92% for 8-bit designs and 87%–89% for 16-bit designs across$8 \times 8$,$16 \times 16$,$32 \times 32$, and$64 \times 64$SAs. The proposed multiplier is employed to present a high-performance Gemmini SA for WS and OS dataflow; further modifications to the processing element (PE) are presented to achieve better performance over the existing Gemmini accelerator. In particular, the proposed WS PE achieves power-delay product (PDP) savings of 56%–59% and area-delay product (ADP) reductions of 45%–51%, while the OS PE shows PDP savings of 10%–46% and ADP reductions of 6%–42%. The analysis is further extended to both convolutional neural networks (CNNs) and Transformers by integrating the proposed SA into complete inference pipelines, including AlexNet, MobileNetV2, and ResNet-50 for CNNs, and the self-attention mechanism of TinyBERT for Transformers. The MOSA-32 design delivers a GOPS improvement of 12%–91% across different CNN models and up to$5.24\times $for Transformer workloads. Similarly, MOSA-64 achieves up to 98% improvements across CNNs and up to$6.14\times $gains for Transformers. These improvements validate the effectiveness of the proposed energy-efficient PE design and its seamless integration into Gemmini SA, making the proposed architecture suitable for next-generation AI and Transformer accelerators for edge computing applications. L. Hemanth Krishna, Srinivasu Bodapati 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2024 | High-Speed Serial and Semi-Parallel IMPLY-based Approximate Adders through Memristors for In-Memory ComputingabstractThis paper introduces three memristive approximate adder designs using Material Implication (IMPLY) logic. Two designs follow a serial approach, while the third adopts a semi-parallel technique. The proposed design achieves the approximate output in just 2 steps, utilizing only 3 memristors, in contrast to the 8 steps and 5 memristors required by the best competing IMPLY-based approximate adder. Furthermore, when integrated into an 8-bit Ripple Carry Adder (RCA), the proposed designs enhance latency by 4% to 28% compared to the existing IMPLY-based approximate design. Compared to exact serial and semi-parallel adders in literature, latency is improved by 27% to 56.8%. Similarly, an improvement in the range of 4% to 42% is achieved for energy consumption. The performance of the proposed approximate adder designs is then evaluated in image processing applications, specifically Image Masking and RGB to Grayscale conversion. The results demonstrate better Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index Measure (SSIM) values compared to the existing approximate IMPLY design. Nandit Kaushik, L. Hemanth Krishna, Srinivasu Bodapati 0001 |
ISCAS | 3 |
| 2024 | Energy Efficient Accurate and Approximate Modified Adders for Ternary MultipliersabstractThis paper presents energy-efficient approximate multiplier designs using proposed modified adder designs. In this paper, we propose accurate and approximate designs of modified half adder and full adders, MHA, FA1and FA2designed using a combination of 3×1 and 2×1 multiplexers. The proposed designs were simulated using the Synopsys HSPICE using MOSFET-like CNFET. The proposed modified half-adder consumes 63.2% less energy than the conventional half-adder. The proposed modified full adders FA1and FA2, respectively have a savings of 24% and 44% energy than the conventional full adder. The overall approximate multiplier saves 26% of energy than the conventional full adder-based multiplier. The applications of the proposed multipliers explored in image processing applications, such as image multiplication and the proposed approximation lead to better Peak signal-to-noise ratio (PSNR) and structural similarity index measure (SSIM) over the best existing designs. L. Hemanth Krishna, Nandit Kaushik, Srinivasu Bodapati 0001 |
ISCAS | 3 |
| 2024 | PUF-based Lightweight Mutual Authentication Protocol for Internet of Things (IoT) DevicesabstractThe Internet of Things(IoT) can transfer data between the sensor node and the cloud server with the Internet’s help for various automated tasks like remote monitoring and controlling. IoT has various applications in different sections, including healthcare, also called the Internet of Medical Things(IoMT). IoT used in healthcare applications can collect patient’s biomedical data through medical sensors, which is sent to the cloud server with the help of the Internet. IoT/IoMT systems are having serious challenging concerns such as data security and IoT device’s security. Hence, this paper proposes a Physically Unclonable Function(PUF) based lightweight mutual authentication and key agreement protocol for IoT/IoMT devices. A lightweight AEAD cipher, ASCON, and a PUF are used for mutual authentication and key agreement between the IoT/IoMT device and the server. The agreed key is used for encrypting the sensor node data using ASCON cipher. The protocol is implemented in Artix-7 FPGA, and the formal verification is performed using the automated tool Proverif. The proposed protocol takes 912 bits of communication cost, which is 13% less compared to the best existing protocol. Further, the protocol requires a node storage cost of 128 bits, which is only 66% of the best existing protocol. Kamal Raj, Srinivasu Bodapati 0001, Anupam Chattopadhyay |
ISCAS | 2 |
| 2024 | Priority Arbiter PUF: Analysis
Meenakshi Kansal, Animesh Roy 0004, Dibyendu Roy 0001, Srinivasu Bodapati 0001, Anupam Chattopadhyay |
Discret. Appl. Math. | 4 |
| 2022 | PA-PUF: A Novel Priority Arbiter PUFabstractThis paper proposes a 3-input arbiter-based novel physically unclonable function (PUF) design. Firstly, a 3-input priority arbiter is designed using a simple arbiter, two multiplexers (2:1), and an XOR logic gate. The priority arbiter has an equal probability of 0’s and 1’s at the output, which results in excellent uniformity (49.45%) while retrieving the PUF response. Secondly, a new PUF design based on priority arbiter PUF (PA-PUF) is presented. The PA-PUF design is evaluated for uniqueness, non-linearity, and uniformity against the standard tests. The proposed PA-PUF design is configurable in challenge-response pairs through an arbitrary number of feed-forward priority arbiters introduced to the design. We demonstrate, through extensive experiments, reliability of 100% after performing the error correction techniques and uniqueness of 49.63%. Finally, the design is compared with the literature to evaluate its implementation efficiency, where it is clearly found to be superior compared to the state-of-the-art. Simranjeet Singh, Srinivasu Bodapati 0001, Sachin B. Patkar, Rainer Leupers, Anupam Chattopadhyay, Farhad Merchant |
VLSI-SoC | 2 |
| 2017 | A Transistor-Level Probabilistic Approach for Reliability Analysis of Arithmetic Circuits With Applications to Emerging TechnologiesabstractSeveral field-effect transistor (FET)-based device technologies are emerging as powerful alternatives to the classical metal oxide semiconductor FET (mosfet) for computing applications. The focus of this paper is on the analysis of reliability of combinational logic circuits at the transistor level with the goal of application to these technologies. To this end, we present an approach for the calculation of the output probabilities for basic logic primitives that comprise a combinational circuit. Using the output probabilities, a computationally efficient algorithm (based on Hadamard product) is presented to calculate the reliability. As an application of the proposed algorithm, we analyze the reliability of adder circuits composed of carbon nanotube FETs considering fabrication-level parameters. We then investigate adder characteristics that lead to high reliability independent of the technology. In particular, we propose a new multibit adder termed hybrid adder offering high reliability with low area requirement for various transistor-based emerging technologies. Detailed simulation results are presented to support the analysis. Srinivasu Bodapati 0001, K. Sridharan 0001 |
IEEE Trans. Reliab. | 1 |