Sanampudi Gopala Krishna Reddy

dblp:356/4823 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2025
0009-0003-5427-4285ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021
YearPublicationVenuePosition
2025 Design of Cascade and One-Shot Mixed-Mode Recursive Multipliers for GF(2N) Polynomials
abstract
Finite field polynomial multiplication for large operand sizes forms the building block for designing modern Cryptography Systems. This work presents two new approaches to realize three-operand multiplier architecture for Galois Field (2N) polynomial operations, referred to as Cascade and One-Shot Mixed mode configurations. A meta-heuristic approach enabled with a single-objective fitness run was employed to design optimal solutions in the form of non-homogeneous recursive sequences in Karatsuba multiplication, targeted for three-operand multiplication in Galois Field (GF). During the optimization runs, the candidate design solutions were hardware characterized through ASIC process using Cadence Genus tool with 45 nm technology library files. The proposed architectural designs offered improvement in compute-latency of 12.77% and footprint complexity of 61.97% with benefits in the area-delay product (ADP) of 70.72%, and power savings of 83.62% and 51.23% improvement in PPA compared to the existing state-of-the-art (SOTA) designs. All the hardware design files are made freely available for further usage to the designers and researchers community.
D. R. Vasanthi, Daksh Sharma, Sanampudi Gopala Krishna Reddy, Madhav Rao
ISCAS3
2024 FPGA-based Hardware Software Co-design to Accelerate Brain Tumour Segmentation
abstract
Brain tumors are a major concern, being the leading cause of cancer-related deaths. Computer-aided diagnosis significantly reduces the workload on physicians and improves cancer diagnosis and treatment. Brain tumor segmentation is a computationally intensive image-processing task. In this paper, we propose an FPGA-based Hardware-Software Co-design to accelerate this task using Watershed and Otsu thresholding algorithms. The FPGA handles parallel components, while the CPU manages sequential tasks in the same System-on-Chip (SoC). Using PolarFire Icicle FPGA platform, we process 20 MRI brain scan images (128x128) from the Kaggle dataset. Implementing both algorithms in parallel on the FPGA results in a 1.97× acceleration compared to a CPU-only implementation, mainly achieved by a 1973× reduction in latency when moving the Otsu algorithm from the CPU to the FPGA. This optimization employs DSP/MATH blocks, loop unrolling, and pipelining techniques.
Vinay Rayapati, Gogireddy Ravi Kiran Reddy, Gandi Ajay Kumar, Saketh Gajawada, Sanampudi Gopala Krishna Reddy, Nanditha Rao
ISCAS5
2024 HRM: M-Term Heterogeneous Hybrid Blend Recursive Multiplier for GF(2n) Polynomial
abstract
Hardware-efficient polynomial multipliers are desired to satisfy the ever-growing demands of computing within the finite field space toward developing a strong cryptosystems. This research meticulously explores polynomial multiplication from the context of algebraic structures by introducing a novel hetero-blend recursive multiplier that harnesses the strengths of the contemporary state-of-the-art (SOTA) designs. The heterogeneous-blend recursive multiplier (HRM) adeptly merges the footprint efficiency of the Karatsuba multiplier (KM) and the compute-latency benefits of the overlap-free KM (OKM) at higher stages, while at lower bounds, it capitalizes the optimal balance of footprint and compute-latency benefits of the schoolbook multiplier (SBM). To further enhance the performance, HRM integrates the heterogeneous term division throughout its stages which is a characteristic find taken from the prior work on$M$-term nonhomogeneous Karatsuba multiplier (MNHKA). Furthermore, a MATLAB framework has been devised to expedite the exploration process in the finite field design space resulting from the heterogeneous usage of the$M$terms across multiple stages. The presented HRM design undergoes comprehensive evaluation when benchmarked against contemporary SOTA designs including KM, OKM, their corresponding homogeneous$M$term variants referred to as$M$-term Karatsuba multiplier (MKM),$M$-term OKM (MOKM) alongside recent variants of composite$M$-term Karatsuba multipliers (CMKA), MNHKA, and equivalent overlap-free variant$M$-term nonhomogeneous overlap-free Karatsuba multiplier (MNHOKA). The field-programmable gate array (FPGA) synthesized results for the HRM designs on Zynq ZCU-104 board showcase a best-case of 17.288% lookup table (LUT) savings, 5.68% reduction in delay, and 20.88% gain in area-delay product (ADP) compared with the optimal SOTA design, while also revealing a 13.49% reduction in LUT usage, 5.45% decrease in delay, and 12.97% improvement in ADP when compared with the best among MNHKA and MNHOKA designs. Furthermore HRM designs synthesized on the Cadence GPDK45 library achieved a best-case footprint saving of 16.18%, a critical path delay improvement of 29.53%, a remarkable 45.66% gain in the ADP, a substantial 30.37% reduction in power consumption, and a noteworthy 38.63% improvement in power per area when compared with the optimal SOTA design. In comparison to the leading MNHKA and MNHOKA designs, the HRM designs exhibit a best-case footprint improvement of 5.77%, 8.31% reduction in delay, 16.76% enhancement in ADP, a significant 20.18% power savings, and a notable 17.71% improvement in power-per-unit-area (PPA). To catalyze ongoing research and innovation, hardware designs assessed in this article are made publicly available for further usage.
D. R. Vasanthi, Sanampudi Gopala Krishna Reddy, Madhav Rao
IEEE Trans. Very Large Scale Integr. Syst.2
2023 MNHOKA - PPA Efficient M-Term Non-Homogeneous Hybrid Overlap-free Karatsuba Multiplier for GF (2n) Polynomial Multiplier
abstract
In the constantly evolving field of multiplication architectures, the Karatsuba algorithm and its extensions have captivated the minds of researchers with their performance metrics. One such optimized design is the Overlap-free Karatsuba (OKA) algorithm which has emerged as an innovative architecture, specifically aimed at enhancing power, performance, and area (PPA) parameters. In this paper, we introduce a novel technique referred to as M-term Non-Homogeneous Hybrid Overlap-free Karatsuba polynomial multiplier (MNHOKA), which surpasses existing state-of-the-art (SOTA) designs, including Karatsuba multiplier (KA), M-Term Karatsuba-like multiplier (MKA), Composite M-term Karatsuba-like multiplier (CMKA), and Overlap-free Karatsuba multiplier (OKA), across various operand sizes. In this paper, a detailed analysis of the proposed MNHOKA and its corresponding M-Term Non-homogeneous Hybrid Karatsuba Algorithm (MNHKA) is presented, highlighting its performance improvements on both Cadence 45 nm process and the ZYNQ ZCU-104 FPGA board for popular bit widths. In ASIC implementations, MNHOKA achieves significant ADP improvements of 28.33%, 28.99%, 58.23%, and 11.95% for operand sizes of 128, 232, 282, and 750 bits, respectively, compared to the best-case SOTA design. Furthermore, our method yields lower power consumption. When comparing FPGA results of the proposed MNHKA design with the best-case SOTA works, ADP improvement of 22.72%, 16.10%, 2.52%, and 11.36% improvement was achieved for the respective bit-widths of 128, 232, 282, and 750 bits respectively. The advantages of the proposed MNHOKA, along with its equivalent MNHKA design variants, are evident in their superior hardware characteristics over existing SOTA designs. This research represents a significant step towards realizing efficient Cryptosystems in the immediate future. To foster further research and innovation, we have made the hardware design files freely available to the researchers and designer community.
Gogireddy Ravi Kiran Reddy, Sanampudi Gopala Krishna Reddy, D. R. Vasanthi, Madhav Rao
ICCD2