Bishnu Prasad Das

dblp:71/5738 · DBLP profile ↗
← Back
14ranked-venue papers
1as first author
8since 2021 · last 2025
0000-0001-6993-2744ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Light-Weight Low-Latency Reconfigurable CORDIC Architecture With a New Non-Power-of-2 Angle Set of Microrotations
abstract
Trigonometric functions are widely used in various scientific and engineering applications, including digital signal processing. This article proposes a low-complexity low-latency reconfigurable CORDIC (COordinate Rotation Digital Computer) algorithm that computes trigonometric functions, which is based on a novel angle set of microrotations. The CORDIC algorithm is an iterative algorithm that reduces the performance due to the high latency of computation. The proposed pipelined scale-free CORDIC algorithm computes the functions in only four iterations with an optimized angle set of microrotations with high accuracy. It can be configured in two modes with different trajectories 1) rotation-mode; and 2) vectoring-mode for circular and hyperbolic trajectories. On-board FPGA implementation of the proposed architecture consumes 34.95% less register, 48.12% less occupied slices, 19.06% less LUTs, 66.51% less slice-latency product (SLP), 59.29% less LUT-latency product (LLP), and 71.23% less register-latency product (RLP) compared to the best existing state-of-the-art CORDIC architectures. Moreover, the latency of the proposed architecture is 72.34% less than the best of existing designs. Besides, the parasitic-extracted layout simulation results of the proposed design across different word lengths are reported in TSMC 65-nm CMOS technology for the validation of the proposed architecture.
Anu Verma, Bishnu Prasad Das
IEEE Trans. Circuits Syst. I Regul. Pap.2
2025 A Unified Approach to a Secure and Lightweight Mutual Authentication Protocol Using Pre-Characterized COTS SRAM ICs for IoT Applications
abstract
Traditional Physical Unclonable Function (PUF)-based authentication protocols are vulnerable to machine learning attacks and evolving cyber threats. Moreover, these protocols lack suitability for resource-constrained IoT devices due to the involvement of heavy cryptographic primitives, error correction modules, and significant computational overhead. This article proposes a mutual authentication protocol and session key agreement utilizing commercial-off-the-shelf (COTS) SRAM integrated circuits (ICs) to extract a secret key. We introduce a block-based lightweight fuzzy extractor to minimize the overhead associated with error correction modules on IoT devices. Our protocol relies only hash, XOR and masking functions for the identity verification for both parties and stores only one challenge-response pair (CRP) on the server, reducing memory overhead on the authentication server. In addition, we perform a rigorous informal security analysis against well-known attacks and formal security analysis using Verifpal tool considering an active attacker in the communication link. Furthermore, the performance evaluation and comparative analysis indicate that the proposed protocol significantly outperforms the state-of-the-art protocols in terms of communication, computational overhead, storage overhead, and energy consumption by up to 69%, 81%, 87.5%, and 76.6%, respectively. We have implemented the proposed authentication protocol on ESP32 and Raspberry Pi 3 boards to show its applicability and scalability in a real-world IoT framework.
Aranya Gupta, Amit Surpur, Bishnu Prasad Das, Sanjeev Manhas 0001
ACM Trans. Embed. Comput. Syst.3
2024 PVT-Insensitive Time-Domain-based In-Memory Computation with Improved Linearity for Binary Neural Networks
abstract
The in-memory computation technique has a strong potential to improve the speed and energy efficiency of data-intensive tasks used in Artificial intelligence (AI). In this work, a time domain-based in-memory computation (TD-IMC) architecture is proposed to perform XNOR-and-accumulation (XAC) operations using two types of delay cells. The first XNOR and delay cell (XDC) is the base cell, which contains an XNOR gate and a stack-inverter-based delay cell to perform XNOR- and-accumulation (XAC) operations. The second process-voltagetemperature (PVT)-insensitive XNOR-and-delay cell (PXDC) is a PVT-tolerant cell, which contains the base cell and an array of controllable load capacitors to compensate for PVT variation. To demonstrate the application of the proposed TD-IMC XAC operation, the MNIST image classification is performed using a binary neural network (BNN) algorithm. The post-layout simulation results of the proposed architecture in an industrial 65 nm CMOS process show a classification accuracy of 99.6% compared to the software results. Moreover, the proposed TD-IMC achieves 1057/3.01 GOPS and 673/13.12 TOPS/W energy efficiency for XDC/PXDC-based TD-IMC architectures, which are comparable to the state-of-the-art TD-IMC architectures available in the literature.
Bishnu Prasad Das
ISCAS2
2024 A Peak-detector-based Ultra Low Power ECG ASIC for Early Detection of Cardio-Vascular Diseases
abstract
Cardio-vascular diseases (CVDs) are growing rapidly these days, which motivates to design an ultra-low power wearable cardiac monitor for early detection. In this work, we proposed an ultra-low power application-specific integrated circuit (ASIC) that can detect the P, R, and T peaks of the Electrocardiogram (ECG) signal. These peaks are useful for early detection of CVDs. The proposed design consists of a low-power analog front end followed by a peak-detection circuit, which can detect the P, R, and T peaks from an incoming single-lead differential ECG signal. The time-varying amplitude difference between the adjacent ECG peaks are efficiently detected by the proposed design. The proposed peak detector circuit has a self-adjustable discharge rate, which depends on the amplitude of the peak. The circuit is designed at the sub-threshold region to reduce the power consumption of the overall design. The proposed design extracts the digital output with the temporal information from the peaks of the ECG signal without using an analog-to-digital converter (ADC), one of the most power-hungry blocks in conventional biomedical signal acquisition architectures. A test chip has been fabricated in an industrial 0.18 µm CMOS process and occupies an active area of only 0.848 mm2. Measured results show that the proposed design consumes only 299.45 nW of power with a supply voltage of 0.9 V at typical operating conditions, which represents an improvement of 27.07%, 34.57%, and 51.93% over the existing state-of-the-art design. As it consumes ultra-low power, occupies less silicon area, and reliably extracts the ECG peaks, the proposed design is highly suitable for wearable healthcare applications.
Sidharth Thomas, Jaskirat Singh Virdi, Anshul Verma, Bishnu Prasad Das, Kenichi Okada 0001, Pratap Narayan Singh
ISCAS4
2024 Process-Variation-Aware In-Memory Computation With Improved Linearity Using On-Chip Configurable Current-Steering Thermometric DAC
abstract
The in-memory computation (IMC) is a potential technique to improve the speed and energy efficiency of data-intensive designs. However, the scalability of IMC to large systems is hindered by the non-linearities of analog multiply-and-accumulate (MAC) operations and process variation, which impacts the precision of high bit-width MAC operations. In this paper, we present an IMC architecture that is capable of performing multi-bit MAC operations with improved speed, linearity, and computational accuracy. To improve the speed/linearity of the IMC-MAC operations, the image and weight data are applied by using the pulse amplitude modulation (PAM) and thermometric techniques, respectively. Although the PAM technique improves the speed of the IMC-MAC operations, it has linearity issues that need to be addressed. Based on the detailed linearity analysis of the IMC-MAC circuit, we proposed two approaches to improve the linearity and the signal margin (SM) of the IMC architecture. The proposed configurable current steering thermometric digital-to-analog converter (CST-DAC) array is employed to provide the PAM signals with various dynamic ranges and non-linear gaps that are required to improve the linearity/SM. The proposed combined PAM and thermometric IMC (PT-IMC) architecture is designed and fabricated in the TSMC 180-nm CMOS process. The post-silicon calibration of the design point mitigates the process-variation issues and provides the maximum SM (close to the simulation results). Furthermore, the proposed PT-IMC architecture performs MNIST/CIFAR-10 data set classification with an accuracy of 98%/88%. In addition, the PT-IMC architecture achieves a peak throughput of 12.41 GOPS, a normalized energy efficiency of 30.64 TOPS/W, a normalized figure-of-merit (FOM) of 3039, a loss in the SM of 8.3% with respect to the ideal SM, and a computational error of 0.41%.
Prasanna Kumar Saragada, Bishnu Prasad Das
IEEE Trans. Circuits Syst. I Regul. Pap.2
2023 An Efficient Scaling-Free Folded Hyperbolic CORDIC Design Using a Novel Low-Complexity Power-of-2 Taylor Series Approximation
abstract
Hyperbolic trigonometric functions are widely used in several engineering and scientific applications, including digital signal processing (DSP), communication systems, and many others. In this article, we propose a scaling-free hyperbolic coordinate rotation digital computer (CORDIC) algorithm and its architecture based on a novel power-of-2 coefficient low-complexity Taylor series approximation to implement sinh and cosh functions. CORDIC architectures are generally slow due to their high latency of computation. The proposed architecture reduces the latency and achieves the desired precision with only four iterations where an optimized angle set comprised of six CORDIC microrotations are mapped into a four-stage folded-pipeline structure leveraging mutually exclusive behavior of two pairs of microrotations. The proposed design is implemented on field-programmable gate arrays (FPGAs) Xilinx Zedboard using 65.38% less registers with ~63.63% less latency and 48.97% less power consumption compared with the best of the existing designs. The proposed design is synthesized by Synopsys Design Compiler and place and route (PnR) tool using Taiwan Semiconductor Manufacturing Company (TSMC) 65-nm CMOS process. It consumes ~76.31% less area, 68.75% less computational delay, and 68.92% less power consumption compared with the best of the existing designs. Moreover, the proposed architecture involves 46.89% less energy per output (EPO) than the best of the existing designs. The error–energy performance (EEP) and the error–area performance (EAP) of the proposed design are, respectively, ~1.25 times and ~2.8 times better than that of the best of the existing designs. Besides, the proposed architecture is also implemented and verified on a silicon chip in the TSMC 180-nm CMOS process for the validation of the algorithm and architecture.
Anu Verma, Khyati Kiyawat, Bishnu Prasad Das, Pramod Kumar Meher
IEEE Trans. Very Large Scale Integr. Syst.3
2022 RISC-V Core with Approximate Multiplier for Error-Tolerant Applications
abstract
RISC-V is an open-source instruction set architecture with customizable extensions to introduce operations like multiplication, division, atomic functions, and floating-point operations. In this paper, a new approximate multiplier is integrated with RI5CY (CV32E40P) processor, which can perform integer and floating-point multiplication for error-tolerant applications. The multiplication operation is required in various engineering and scientific applications, including image processing, digital signal processing, and many others. The proposed approximate multiplier is based on linear CORDIC (COordinate Rotation Digital Computer) algorithm and implemented by using only shift-add operations. It can perform multiplication and MAC (Multiply and accumulate) operations. The FPGA (Field programmable gate arrays) implementation results and ASIC (Application-specific integrated circuit) synthesis results for the proposed approximate multiplier along with RI5CY core are reported. The proposed design with RI5CY core is implemented on FPGA Xilinx Zedboard, which improves the performance by 20% and reduces power delay product (PDP) by 15.79% over the existing multipliers of the RI5CY core. Moreover, RI5CY core with the proposed approximate multiplier is synthesized using Industrial 130 nm standard cell library (ISCL) and Sub-threshold 130 nm standard cell library (STSCL) in Synopsys DC compiler. In the case of STSCL, RI5CY core with proposed approximate multiplier has 11.76% less power-consumption, 27.27% less delay, and 38.77% PDP compared to the existing multipliers of the RI5CY core.
Anu Verma, Priyamvada Sharma, Bishnu Prasad Das
DSD3
2022 In-Memory Computation With Improved Linearity Using Adaptive Sparsity-Based Compact Thermometric Code
abstract
The article presents an efficient static random access memory (SRAM)-based in-memory computation (IMC) architecture which is capable of performing image classification with improved linearity. In this work, we proposed a thermometric code-based IMC (TC-IMC) to perform multibit multiply-and-accumulate (MAC) operations with improved linearity. An input sparsity-aware compact thermometric code approach is proposed to reduce the number of SRAM bitcells compared to the thermometric code-based encoding without loss of accuracy in the MAC operation. We proposed an optimal sampling time to improve the linearity of the TC-IMC MAC operation with maximum signal margin (SM) based on the detailed nonlinearity analysis of 8T SRAM-based TC-IMC architecture. The test chip measurement results in 180 nm process show that the proposed TC-IMC has 72% better linearity than the traditional IMC. The measured results on 100 Modified National Institute of Standards and Technology (MNIST) and CIFAR-10 test images show an accuracy of 97% and 87%, respectively. In addition, the proposed TC-IMC architecture achieves MAC compute latency of 25 ns, GOPS/kb (normalized to 1 b$\times $1 b) of 29.8, and energy efficiency of 2.3 TOPS/W. The simulation result of the entire 10000 MNIST and CIFAR-10 images considering process variation effects shows an accuracy of 98.7% and 90.4%, respectively.
Prasanna Kumar Saragada, Bishnu Prasad Das
IEEE Trans. Very Large Scale Integr. Syst.2
2018 Metastability immune and area efficient error masking flip-flop for timing error resilient designs
Govinda Sannena, Bishnu Prasad Das
Integr.2
2018 Low Overhead Warning Flip-Flop Based on Charge Sharing for Timing Slack Monitoring
Govinda Sannena, Bishnu Prasad Das
IEEE Trans. Very Large Scale Integr. Syst.2
2016 A Metastability Immune Timing Error Masking Flip-Flop for Dynamic Variation Tolerance
abstract
In this paper, two timing error masking flip-flops have been proposed, which are immune to metastability. The proposed flip-flops exploit the concept of either delayed data or pulse based approach to detect timing errors. The timing violations are masked by passing direct data instead of master latch output to slave latch. Simulation results show that the proposed flip-flops such as type-A and type-B reduce the error masking latency up to 23% and 42% respectively in typical process corners and increase the effective timing error monitoring window compared to state of the art metastable immune flip-flops [14]. The proposed flip-flops can be used in dynamic voltage and frequency scaling (DVFS) applications. A 16-bit adder is implemented to evaluate the functionality of the proposed flip-flops in DVFS frame work and the simulation results show that the adder using the proposed flip-flop can reduce up to 48% power consumption or improve the performance up to 50% in typical process corners compared to conventional worst case design.
Govinda Sannena, Bishnu Prasad Das
ACM Great Lakes Symposium on VLSI2
2014 Detecting Reliability Attacks during Split Fabrication using Test-only BEOL Stack
abstract
Split fabrication, the process of splitting an IC into an untrusted and trusted tier, facilitates access to the most advanced semiconductor manufacturing capabilities available in the world without requiring disclosure of design intent. While obfuscation techniques have been proposed to prevent malicious circuit insertion or modifications in the untrusted tier, detecting a pernicious reliability attack induced in the offshore foundry is more elusive. We describe a methodology for exhaustive testing of components in the untrusted tier using a specialized test-only metal stack for selected sacrificial dies.
Kaushik Vaidyanathan, Bishnu Prasad Das, Lawrence T. Pileggi
DAC2
2014 Frequency-Independent Warning Detection Sequential for Dynamic Voltage and Frequency Scaling in ASICs
abstract
In this paper, a metastability immune warning flip-flop (FF) is proposed, which consists of an edge detector, a warning window generator, and a warning detector along with a traditional FF. The delayed data are monitored during the warning window to flag a warning signal before the data enter the erroneous zone. In this scheme, the warning window is independent of input clock frequency and hence is suitable for frequency scaling application. A 16-bit Kogge-stone adder is implemented in 65-nm technology, which uses warning FF for dynamic voltage and frequency scaling (DVFS). The warning FF-based DVFS allows elimination of safety margins and operates till the point of first warning of the adder without any erroneous results. The experiments were conducted with different supply voltages, phase-shifted clocks, and process conditions. The circuit is helpful to determine when to stop further reduction in supply voltage by producing the warning signal with predefined timing slacks in DVFS application. The test chip results demonstrate that the proposed circuit can track the critical path delay of 2.4-7.5 ns at warning voltage of 1.15-0.72 V, respectively. The measured results from 10 different chips show the effectiveness of the proposed concept across process variation.
Bishnu Prasad Das, Hidetoshi Onodera
IEEE Trans. Very Large Scale Integr. Syst.1
2010 Learning Collaboration Moderator Services: Supporting Knowledge Based Collaboration
Alok Kumar Choudhary, Jenny A. Harding, Rahul Swarnkar, Bishnu Prasad Das, Robert I. M. Young
PRO-VE4