Roy P. Paily

dblp:60/4004 · also Roy Paily, Roy Paily Palathinkal · DBLP profile ↗
← Back
12ranked-venue papers
0as first author
5since 2021 · last 2025
0000-0003-3004-9369ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5Computer networks · 3 · 1 since 2021
YearPublicationVenuePosition
2025 Non-Linear Cyclic Variable Clock Feistel Bridge-Inspired Countermeasure for Securing RISC-V Crypto-Core Against Power Attacks
abstract
With the increasing popularity of Internet of Things Edge (IoTe) devices, RISC-V emerges as the most suitable architecture for various applications. Since attackers have direct physical access to these IoTe devices, securing these RISC-V designs is of major concern. This research article first focuses on designing a low-power high-performance, RISC-V crypto-core for high security applications, which significantly enhances the efficiency of the AES encryption algorithm, but still susceptible to Power Analysis Attacks (PAA). A novel Non-Linear Cyclic Variable Clock Feistel Bridge-Inspired Countermeasure (NCVCFB) providing variation in both amplitude and temporal domains of signal was introduced to enhance AES security in RISC-V based IoTe devices, specially to address the vulnerability against PAA. The NCVCFB countermeasure employs the rolling architecture and will operate in tandem with each round of AES to obscure power consumption patterns. It will execute additional random number of rounds cyclically in the first and last round of AES, based on the randomly generated number. The newly designed, secured RISC-V achieved remarkably low area and power overheads of 1.05% and 1.23%, respectively, without affecting the maximum operating frequency. Nevertheless, the throughput was decreased due to variable clock count for different plaintext. The PAA was performed using the power traces captured from post-layout design on ASIC platform at UMC 65 nm technology node as well as on the experimental hardware setup employing a Side-channel Attack Security Evaluation Board (SASEBO). The resilience of the secured RISC-V architecture against PAA was tested by subjecting it to 2 Million traces, and none of the bytes got recovered. Thus, the NCVCFB secured RISC-V attained a Measurement To Disclose (MTD) >2M, Signal to Noise Ratio (SNR) <0.5, Mutual Information (MI) in the milli range, and Test Vector Leakage Assessment (TVLA) within +/−4.5 limits.
Titu Mary Ignatius, Roy P. Paily
IEEE Trans. Circuits Syst. I Regul. Pap.2
2025 Design and Analysis of Energy Efficient Approximate Multipliers for Image Processing and Deep Neural Network
abstract
Numerous obstacles in enhancing the performance of computing systems have spurred the emergence of approximate computing. Extensive studies have been reported on approximate computing to develop high-performance, energy-efficient hardware designs tailored to error-resilient applications. In this brief, we proposed 8-bit approximate multipliers with 15 levels of accuracy using three techniques: recursive, bit-wise, and hybrid approximation using partial bit OR (PBO). Compared to the existing multipliers, investigated designs have significantly improved the area, power, delay, Power Delay Product (PDP), and Power Area Delay Product (PADP) by 41.68%, 73.16%, 35.57%, 72.65%, and 75.42% respectively on average. On resemblance with the accurate multiplier, the area, power, delay, PDP, and PADP were enhanced by 54.41%, 57.57%, 25.73%, 60.14%, and 74.33% correspondingly on average. Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index Measure (SSIM) values surpassing (30 dB, 94%), (31 dB, 96%), and (26 dB, 95%) by applying them to benchmarks in image smoothing, edge detection, and image sharpening successively. Moreover, upon scrutinizing the efficacy of multipliers in hardware implementations of deep neural networks attaining the performance exceeding 95%. The obtained results confirm that suggested multipliers are well-suited for their widespread applications.
Anupam Kumari, Roy P. Paily
IEEE Trans. Circuits Syst. I Regul. Pap.2
2025 Continuous Flow 4096-Point FFT/IFFT Hardware Architecture for 5G Applications
abstract
This paper presents a high-throughput, low-latency 4096-point FFT/IFFT hardware design tailored for 5G applications. Using the radix-16 FFT algorithm, the architecture efficiently supports both FFT and IFFT operations with minimal modifications. It employs only three radix-16 butterfly units constructed from pipelined radix-4 structures to achieve lower operational complexity at higher radix levels. To accommodate real-time application requirements, a continuous flow design is proposed utilizing a Conflict-Free Memory Addressing (CFMA) scheme for data ordering to access various memory banks in parallel. The twiddle factors are generated using Canonical Signed Digit (CSD) and CORDIC architectures, optimizing area and power usage while eliminating the need for ROM units. The design verified both FFT and IFFT operations using MATLAB and synthesized with the Cadence Genus tool in a UMC 65nm process at 250 MHz. It achieves a latency of 2.43$\mu $s and a throughput of 4 GS/s, representing a 29.03% increase in throughput compared to the best results reported in the literature, with a Signal-to-Quantization Noise Ratio (SQNR) of 51.2 dB.
Aditi Paul, Shaik Rafi Ahamed, Roy P. Paily
IEEE Trans. Circuits Syst. I Regul. Pap.3
2024 Improvement in Resilience of AES Design With Reconfigured CFB Mode Against Power Attacks
abstract
Advanced encryption standard (AES) is used to secure the communication process on the Internet-of-Things (IoT) hardware. It is implementable in various 128-bit modes, such as electronic code book (ECB), cipher block chaining (CBC), cipher feedback (CFB), output feedback (OFB), and counter (CTR), to facilitate parallel processing of data. The noninvasive nature of power analysis attacks (PAAs) to retrieve secret information off a physical device renders such hardware to be unsafe from the adversaries. Also, the assessment of the aforementioned modes for security remains obscured, which is undertaken by this work as a novel attempt. In addition, this work proposes a novel 64-bit version of CFB mode, which provides the highest security with respect to other modes and several unprotected AES designs. PAAs are performed on ASIC platform utilizing UMC 65-nm technology node and a hardware experimental setup using side-channel attack security evaluation board (SASEBO), both at 16-MHz AES frequency and traces sampled at the rate of 1 GSa/s. The measurements to disclose (MTDs) of >1 000 000 provided by the proposed CFB-64 are significantly more than that provided by usual unprotected AES designs. It also offers the highest MTD, and least signal-to-noise ratio (SNR) and mutual information (MI) among other modes, indicating the highest security. The proposed CFB-64 acts as a countermeasure upon integration with an unprotected AES.
Thockchom Birjit Singha, Basa Sanjana, Titu Mary Ignatius, Roy P. Paily, Shaik Rafi Ahamed
IEEE Trans. Very Large Scale Integr. Syst.4
2023 Securing AES Designs Against Power Analysis Attacks: A Survey
abstract
With the advent of Internet of Things (IoT), the call for hardware security has been seriously demanding due to the risks of side-channel attacks from adversaries. Advanced encryption standard (AES) is the de facto security standard for such applications and needs to ensure a low power, low area, and moderate throughput design apart from providing high security to these devices. Substitution-box (S-box), being the core component of AES, has always drawn the attention of the cryptographic community. A chronological development of the S-box over a period of 20 years since the inception of AES is presented. This article provides the first comprehensive review of the state-of-the-art S-box design techniques, identifying current advancements and analyzing their impact on gate count, area, maximum frequency of operation, throughput, and power. The other goal of the survey is to study the countermeasures designed for AES to protect it against side-channel attacks. In particular, we consider the power analysis attacks (PAAs), and the countermeasures are investigated in terms of their security metrics and design overheads, such as area, power, and performance. The countermeasures are based on hiding or masking approaches depending on their design principle. Similar to the S-box survey, a chronological development of the countermeasures since the discovery of PAAs in 1999, is presented. Finally, we suggest some open research gaps and possible direction of research in terms of S-box and countermeasure designs.
Thockchom Birjit Singha, Roy P. Paily, Shaik Rafi Ahamed
IEEE Internet Things J.2
2019 Analysis of Electromagnetic Actuation System for Different Coil Topologies
abstract
Electromagnetic actuation systems are used to generate magnetic forces for manipulation of magnetic nanoparticles. In this paper, we present a comparative performance analysis of actuation systems, consisting of coils having the following five different topologies: single coil, Helmholtz coil, Maxwell coil, Helmholtz-Maxwell pair and differential current coil (DCC). The experiments are performed using PASCO EX-5540A setup while the numerical analyses are done using COMSOL Multiphysics software. It is observed that the experimental results conform to analytical results as well as simulation results. Furthermore, the results highlight that both Helmholtz-Maxwell pair and DCC topologies produce strong gradient magnetic flux density, thus leading to better manipulation of magnetic nanoparticles as compared to that of other topologies. Additionally, the DCC approach with two coils achieves better gradient flux density while having lesser geometrical complexity as compared to that of Helmholtz-Maxwell pair consisting of four coils.
Pralay Chakrabarty, Siddhanta Roy, Roy P. Paily
TENCON3
2019 Low Power 10T SRAM Cell with Improved Stability Solving Soft Error Issue
abstract
In this paper, we present a new 10T design for static random access memory (SRAM). The proposed design simultaneously aims to address stability, power and half select issues. The design metrics such as power, delay, and stability of the proposed design are compared with 7T, 11T, and 9T SRAM cells. It is observed that the proposed design reduces read and write power by 27.69% and 41.72% as compared to 11T and 9T SRAM cell. Furthermore, WSNM of the proposed SRAM is larger as compared to 7T, 9T and 11T. In addition, the RSNM is larger than 7T and 9T. The simulation result confirms that the proposed design is effective for low power application.
Rohit Lorenzo, Roy P. Paily
TENCON2
2019 Fabrication of Back to Back Schottky Micro-Diodes Using Silver Nanoparticle Film and Zinc Oxide Nanowire Mat for Biological Interactions
abstract
In this work, fabrication of silver nanoparticle film (AgNPF) and zinc oxide nanowire mat (ZnO NWM) based Schottky micro-diodes has been reported. AgNPF has been used as electrode material while ZnO NWM has been utilized as channel material to bridge the gap between the two electrodes. The films have been deposited over cleaned SiO2/Si surface by simply drop-casting dispersions of both the materials. The fabricated device structures have been electrically characterized for different voltage levels. The I-V characteristics establish the fabricated device structures as back to back Schottky diodes. Since the ZnO NWM channel dimensions are near 80 μm, the fabricated device structures have been referred as Schottky micro-diodes. To verify the repeatability of the proposed device structure, 16 such structures have been fabricated by drop-casting 0.5 μL of ZnO NW dispersion in low-surface tension de-ionised water (LST-DI) in four different concentrations (200 μL, 400 μL, 800 μL and 1200 μL) over SiO2/Si surface. Electrical characterization has been carried out for all the 16 devices which shows that I-V curves are in good agreement with the characteristics of back to back Schottky micro-diodes (BBSMD). Moreover, the fabricated device has been found to show interaction with collagen protein fibres based on the shift in the average electrical resistance. Such electro-biological interactions can be helpful in the development of efficient nanomaterials for healthcare and medicinal improvements.
Vimal Kumar Singh Yadav, Siddhanta Roy, Gayatri Natu, Roy P. Paily
TENCON4
2017 On-Chip Photovoltaic Power Harvesting System With Low-Overhead Adaptive MPPT for IoT Nodes
abstract
Extracting maximum power from a photovoltaic (PV) harvester with minimum power transfer loss is one of the primary design goals of an energy processing circuit. This paper presents a fully integrated PV power harvesting system with a low-overhead adaptive maximum power point tracking (MPPT) scheme for Internet-of-Things (IoT) nodes. The proposed scheme tracks the MPPs within 12 μs by utilizing an inherent negative feedback loop, within a tracking error of 0.6%. The tracking range has been improved by ~57% using a current-starved voltage-controlled oscillator (CS-VCO) instead of a polynomial VCO. The overhead area and power consumed by this tracking scheme are approximately 0.013% and 0.1%, respectively. Using commercially available solar cell of area 11.3 cm2, the proposed system can provide 833 μW of power with a light intensity of 600 lx. The proposed energy processing circuit has been designed using 0.18-μm CMOS technology node and the circuit simulations demonstrate that the proposed scheme can track maximum power point (MPP) under rapidly changing atmospheric conditions with a peak tracking efficiency of 99%.
Saroj Mondal, Roy P. Paily
IEEE Internet Things J.2
2017 Low-Power Digital Baseband Transceiver Design for UWB Physical Layer of IEEE 802.15.6 Standard
abstract
This paper presents the design and implementation of ultrawideband digital baseband transceiver for wireless body area networks (WBAN) applications. The power dissipation and area of the proposed architecture are minimized by combining algorithmic and architectural level modifications. A new algorithm for Bose-Chaudhuri- Hocquenghem (BCH) encoding is implemented and the coding gain of the BCH decoder is improved by 0.5 dB, with a cost of only one percent area overhead. An area efficient, low-complexity, and low-power BCH decoder is implemented and it has 42% lower area and 38% lower power dissipation compared to a conventional hard-decision decoder. Other salient features of this paper include a low-complexity, low-power packet detection unit, and a low-power module for removing the shortening bits. The baseband transceiver has been designed in 90 nm CMOS technology and it has an energy efficiency of 73 pJ/bit in transmitter mode and 225 pJ/bit in receiver mode.
Manchi Pavan Kumar, Roy P. Paily, Anup Kumar Gogoi
IEEE Trans. Ind. Informatics2
2013 Performance and throughput analysis of turbo decoder for the physical layer of digitalvideo-broadcasting-satellite-services-tohandhelds standard
abstract
In this study, coding performance of turbo decoder compliant to the physical layer of digital‐video‐broadcasting‐satellite‐services‐to‐handhelds (DVB‐SH) standard for additive‐white‐Gaussian‐noise (AWGN) and frequency selective fading channels are presented. The modulation of transmitted bits is carried out with orthogonal‐frequency‐division‐multiplexing (OFDM) technique, incorporating 1 K‐fast‐Fourier‐transform (1K‐FFI) where each subcarrier is modulated using quadrature‐phase‐shift‐keying (QPSG) or quadrature‐amplitude‐modulation (QAM) schemes. Performance analysis of turbo decoder for the decoding iterations of 3, 8, 14 and 18 as well as the sliding window sizes of 10, 20, 30 and 40 are investigated for both the channels. Discussion on the values of these design metrics to achieve optimum coding performance is also presented. The optimisation of system throughput for turbo decoder based on the decoding iteration and sliding window size for various processor speed ranging from 200 MHz to 1 GHz is carried out. Such an analysis is presented for non‐parallel radix‐2 as well as parallel radix‐4 configuration of turbo decoder to meet the system throughput specification of third‐generation wireless communication standard ranging from 100 to 300 Mbps. The coding performance of turbo decoder based on max‐log‐MAP, log‐MAP and Maclaurin series‐based algorithms are studied for both the channel conditions. Simultaneously, the running time for each of these algorithm in a 64 bit processor is also presented for comparison. Finally, the coding performance of turbo decoder for various code rates of 1/5, 2/9, 1/4, 2/7, 1/3, 2/5 and 1/2 are carried out.
Rahul Shrestha, Roy P. Paily
IET Commun.2
2008 A 1.8/2.4-ghz dualband cmos low noise amplifier using miller capacitance tuning
abstract
This paper describes an inductively source degenerated dual-band low noise amplifier (LNA) designed in a standard CMOS 0.18 µm TSMC process. The dual-band LNA can be tuned to 1.8-GHz and 2.4-GHz. The impedance matching is obtained at the required frequency bands using Miller-capacitance tuning. The designed LNA exhibits a gain of 18.2dB and 16.2dB and a noise figure of 5.0dB and 3.7dB at 1.8 and 2.4-GHz respectively. The LNA design is carried out using Mentor Graphics Eldo software.
Depak Balemarthy, Roy P. Paily
ISLPED2