Vinay Rayapati

dblp:372/0762 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2024 POCO: Hardware Characterization of Activation Functions using POSIT-CORDIC Architecture
abstract
POSIT offers a wider dynamic range when compared to floating-point (FP) formats with lesser number of bits. Such data formats are required to address the need for low-bit high-precision hardware architectures for neural networks (NNs) on edge platforms. Activation functions (Af) which introduce non-linearity during the feature extraction process remain as a core component for realizing NN systems. CORDIC (COordinate Rotation Digital Computer) architecture is a hardware efficient technique to realize complex non-linear functions and is deemed suitable to implement Afs. Hence, this work aims to investigate POSIT data formatted CORDIC architecture to realize Afs (Tanh, Sigmoid and Softmax) in different architectural styles. A benchmark evaluation for the proposed POSIT data formatted Afs with the improved CORDIC architecture over SOTA (IEEE 754 FP formats) based designs are presented. The noticeable improvement in hardware design space and error metrics makes the CORDIC architecture-based POSIT formatted Afs stand out over other methods. All the design files are made publicly available for easy adoption and further usage to the designers’ and researchers’ community.
Mahati Basavaraju, Vinay Rayapati, Madhav Rao
ISCAS2
2024 VPU-CIM: A 130nm, 33.98 TOPS/W RRAM based Compute-In-Memory Vector Co-Processor
abstract
Deep Learning inference on edge devices requires reduced memory load/store latency and low bit-precision computations. To address these challenges, we present VPU-CIM: a novel RRAM-based Compute-In-Memory (CIM) variable bit-precision vector co-processor. We introduce vector extensions to the RISC-V ISA and implement it as an in-memory compute unit with a unique data mapping strategy. The design is implemented using open-source Skywater 130nm PDK, with area estimates provided for TSMC 28nm and ASAP 7nm PDKs. Our design achieves an energy efficiency of 33.98 TOPS/W for a 4,4 (I, W) precision configuration. The results demonstrate the potential of RRAM-based vector computations in memory.
J. Chithambara Moorthii, Vinay Rayapati, Nanditha Rao, Manan Suri
ISCAS2
2024 FPGA-based Hardware Software Co-design to Accelerate Brain Tumour Segmentation
abstract
Brain tumors are a major concern, being the leading cause of cancer-related deaths. Computer-aided diagnosis significantly reduces the workload on physicians and improves cancer diagnosis and treatment. Brain tumor segmentation is a computationally intensive image-processing task. In this paper, we propose an FPGA-based Hardware-Software Co-design to accelerate this task using Watershed and Otsu thresholding algorithms. The FPGA handles parallel components, while the CPU manages sequential tasks in the same System-on-Chip (SoC). Using PolarFire Icicle FPGA platform, we process 20 MRI brain scan images (128x128) from the Kaggle dataset. Implementing both algorithms in parallel on the FPGA results in a 1.97× acceleration compared to a CPU-only implementation, mainly achieved by a 1973× reduction in latency when moving the Otsu algorithm from the CPU to the FPGA. This optimization employs DSP/MATH blocks, loop unrolling, and pipelining techniques.
Vinay Rayapati, Gogireddy Ravi Kiran Reddy, Gandi Ajay Kumar, Saketh Gajawada, Sanampudi Gopala Krishna Reddy, Nanditha Rao
ISCAS1
2023 High Performance and Energy Efficient AMD and BWAD Pooling Schemes Characterised for CNN Accelerators
abstract
Convolution Neural Network (CNN) accelerator designs have a plethora of applications but they account for high computational complexity and demand huge resources. The challenge is to attain reduction in various hardware parameters in size-constrained edge inferencing systems. Multiple methods such as using approximate multipliers, or systolic arrays or quantization techniques, to name a few, were experimented on, previously, to achieve the same. But, one of the important layers in a CNN is a pooling layer where the feature map size is reduced based on a predefined scheme. These pooling schemes play a prominent role in determining the accuracy obtained and also affect the hardware usage in accelerator designs. This paper discusses the hardware design of two new pooling methods, Binary Weighted Absolute Deviation (BWAD) and Absolute Maximum Deviation (AMD) which showcase promising results in terms of area, delay, area-delay-product (ADP), and power-delay-product (PDP), and also maintain comparable accuracy with that of the existing state-of-the-art (SOTA) methods, Max and Absolute Average Deviation (AAD) pooling. The proposed pooling schemes make use of the best features in the existing methods such as calculating absolute deviation between consecutive pixels for maintaining fairly comparable accuracy and consist of simple designs, without the use of multipliers or dividers, when mapped on to hardware. The two pooling methods are validated by incorporating the corresponding layers in multiple CNN architectures and datasets for fair evaluation. The hardware efficiency of the pooling methods is investigated by synthesizing on Zynq 7000 series Zedboard (FPGA). ASIC flow synthesis results are obtained from Cadence Genus tool for 45 nm and 130 nm technology nodes. Hence, both the proposed pooling schemes are potential candidates for designing on-chip neural network hardware accelerators which exhibit a fine balance between network accuracy and hardware benefits.
Vinay Rayapati, Mahati Basavaraju, Madhav Rao
DSD1