Shaik Rafi Ahamed

dblp:128/8521 · also Rafi Ahamed Shaik, Rafiahamed Shaik · DBLP profile ↗
← Back
22ranked-venue papers
3as first author
13since 2021 · last 2026
0000-0003-1617-2299ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CTPNet: Achieving Real-Time Semantic Segmentation on Resource-Constrained Edge Devices for Autonomous Driving
Nadeem Atif, Saquib Mazhar, Mohammed Ameen, Shaik Rafi Ahamed, Manas Kamal Bhuyan
IEEE Trans. Circuits Syst. I Regul. Pap.4
2025 Multi-Scale Temporal Attention Convolutional Neural Network for Sleep Stage Classification Using Single-Channel EEG
abstract
Accurate and efficient classification of sleep stages is essential for advancing widespread sleep monitoring solutions. This study presents a novel multi-scale temporal attention Convolutional Neural Network (CNN) tailored for single-channel EEG (Fpz-Cz) sleep stage analysis. Our architecture innovatively combines temporal attention mechanisms across multiple scales within a convolutional neural network framework. By emphasizing temporal relationships in the EEG signal across varied receptive fields, the model effectively captures discriminative features without the computational overhead of channel-wise operations common in multi-channel systems. This multi-scale strategy enables the network to autonomously identify and prioritize significant patterns, ranging from short-term transients to longer rhythmic structures. The proposed design strikes an optimal balance between performance and computational efficiency compared to existing deep learning approaches. Evaluated on the Sleep-EDF-20 dataset, our model achieves a competitive accuracy of 76.46% and a Cohen's kappa$\kappa$of 0.69, with approximately 141.5 k parameters. This research highlights the potential of multi-scale temporal attention in CNNs for accurate and computationally efficient single-channel EEG sleep stage classification, paving the way for future advancements and practical applications.
Aikendrajit Ningthoujam, Shaik Rafi Ahamed
TENCON2
2025 Toward interpretable schizophrenia detection from EEG using autoencoder and EfficientNet
Umesh Kumar Naik Mudavath, Shaik Rafi Ahamed, KongFatt Wong-Lin
Pattern Recognit. Lett.2
2025 Continuous Flow 4096-Point FFT/IFFT Hardware Architecture for 5G Applications
abstract
This paper presents a high-throughput, low-latency 4096-point FFT/IFFT hardware design tailored for 5G applications. Using the radix-16 FFT algorithm, the architecture efficiently supports both FFT and IFFT operations with minimal modifications. It employs only three radix-16 butterfly units constructed from pipelined radix-4 structures to achieve lower operational complexity at higher radix levels. To accommodate real-time application requirements, a continuous flow design is proposed utilizing a Conflict-Free Memory Addressing (CFMA) scheme for data ordering to access various memory banks in parallel. The twiddle factors are generated using Canonical Signed Digit (CSD) and CORDIC architectures, optimizing area and power usage while eliminating the need for ROM units. The design verified both FFT and IFFT operations using MATLAB and synthesized with the Cadence Genus tool in a UMC 65nm process at 250 MHz. It achieves a latency of 2.43$\mu $s and a throughput of 4 GS/s, representing a 29.03% increase in throughput compared to the best results reported in the literature, with a Signal-to-Quantization Noise Ratio (SQNR) of 51.2 dB.
Aditi Paul, Shaik Rafi Ahamed, Roy P. Paily
IEEE Trans. Circuits Syst. I Regul. Pap.2
2024 EEG Data Augmentation Using Generative Adversarial Network for Improved Emotion Recognition
Raktim Acharjee, Shaik Rafi Ahamed
ICPR (11)2
2024 Improvement in Resilience of AES Design With Reconfigured CFB Mode Against Power Attacks
abstract
Advanced encryption standard (AES) is used to secure the communication process on the Internet-of-Things (IoT) hardware. It is implementable in various 128-bit modes, such as electronic code book (ECB), cipher block chaining (CBC), cipher feedback (CFB), output feedback (OFB), and counter (CTR), to facilitate parallel processing of data. The noninvasive nature of power analysis attacks (PAAs) to retrieve secret information off a physical device renders such hardware to be unsafe from the adversaries. Also, the assessment of the aforementioned modes for security remains obscured, which is undertaken by this work as a novel attempt. In addition, this work proposes a novel 64-bit version of CFB mode, which provides the highest security with respect to other modes and several unprotected AES designs. PAAs are performed on ASIC platform utilizing UMC 65-nm technology node and a hardware experimental setup using side-channel attack security evaluation board (SASEBO), both at 16-MHz AES frequency and traces sampled at the rate of 1 GSa/s. The measurements to disclose (MTDs) of >1 000 000 provided by the proposed CFB-64 are significantly more than that provided by usual unprotected AES designs. It also offers the highest MTD, and least signal-to-noise ratio (SNR) and mutual information (MI) among other modes, indicating the highest security. The proposed CFB-64 acts as a countermeasure upon integration with an unprotected AES.
Thockchom Birjit Singha, Basa Sanjana, Titu Mary Ignatius, Roy P. Paily, Shaik Rafi Ahamed
IEEE Trans. Very Large Scale Integr. Syst.5
2023 Block attention network: A lightweight deep network for real-time semantic segmentation of road scenes in resource-constrained devices
Saquib Mazhar, Nadeem Atif, Manas Kamal Bhuyan, Shaik Rafi Ahamed
Eng. Appl. Artif. Intell.4
2023 Securing AES Designs Against Power Analysis Attacks: A Survey
abstract
With the advent of Internet of Things (IoT), the call for hardware security has been seriously demanding due to the risks of side-channel attacks from adversaries. Advanced encryption standard (AES) is the de facto security standard for such applications and needs to ensure a low power, low area, and moderate throughput design apart from providing high security to these devices. Substitution-box (S-box), being the core component of AES, has always drawn the attention of the cryptographic community. A chronological development of the S-box over a period of 20 years since the inception of AES is presented. This article provides the first comprehensive review of the state-of-the-art S-box design techniques, identifying current advancements and analyzing their impact on gate count, area, maximum frequency of operation, throughput, and power. The other goal of the survey is to study the countermeasures designed for AES to protect it against side-channel attacks. In particular, we consider the power analysis attacks (PAAs), and the countermeasures are investigated in terms of their security metrics and design overheads, such as area, power, and performance. The countermeasures are based on hiding or masking approaches depending on their design principle. Similar to the S-box survey, a chronological development of the countermeasures since the discovery of PAAs in 1999, is presented. Finally, we suggest some open research gaps and possible direction of research in terms of S-box and countermeasure designs.
Thockchom Birjit Singha, Roy P. Paily, Shaik Rafi Ahamed
IEEE Internet Things J.3
2022 Hardware Implementation of Low Complexity High-speed Perceptron Block
abstract
Perceptron is the basic computation unit of neural network architectures. This work proposes a resource-efficient and fast hardware for perceptron and Multi-layer Perceptron (MLP) network. The inner product computation unit and activation function unit is designed using Offset Binary Coding (OBC) and Co-ordinate Rotation Digital Computer (CORDIC) respectively. The proposed hardware is implemented on Field Programmable Logic Array (FPGA) and synthesized on 65 nm Application Specific Integrated Chips (ASIC). It achieved a speed-up of at least $21 \times$ as compared to software. The total area and power consumed was $0.085662 mm^{2}$ and 2.28 mW respectively @200 MHz.
Rituparna Choudhury, Shaik Rafi Ahamed, Prithwijit Guha
ISCAS2
2022 A high throughput hardware architecture for deblocking filter in HEVC
P. Kopperundevi, Matcha Surya Prakash, Shaik Rafi Ahamed
Signal Process. Image Commun.3
2021 Fractional-Harris hawks optimization-based generative adversarial network for osteosarcoma detection using Renyi entropy-hybrid fusion
abstract
Osteosarcoma is the malignant bone sarcoma that is characterized by widespread genomic disruption and the inclination for metastatic spread. Early detection of osteosarcoma increases the survival rate. Various osteosarcoma detection methods are adopted to detect osteosarcoma at an early stage, but evaluating the slides under the microscope to find the degree of tumor necrosis and tumor result is a major challenge in the medical sector. Hence, an effective detection method is developed using the proposed Fractional-Harris Hawks Optimization-based Generative Adversarial Network (F-HHO-based GAN) for detecting osteosarcoma at an early stage. Here, the proposed F-HHO is designed by the integration of Fractional Calculus and HHO, respectively. Accordingly, the classification of viable tumor, nontumor, and the necrotic tumor is carried out by GAN using the histology image slides. GAN is used to perform osteosarcoma detection based on the features extracted from the image through the process of cell segmentation. The training process of GAN is done using the proposed F-HHO algorithm. However, the proposed F-HHO obtained better performance using the metrics, namely, accuracy, sensitivity, and specificity with the values of 98%, 98%, and 98% for training percentage and 96.282%, 97.552%, and 95.651% for K-fold, respectively.
Syed Jahangir Badashah, Shaik Shafiulla Basha, Shaik Rafi Ahamed, S. P. V. Subba Rao
Int. J. Intell. Syst.3
2021 A Distributed Arithmetic based realization of the Least Mean Square Adaptive Decision Feedback Equalizer with Offset Binary Coding scheme
Matcha Surya Prakash, Shaik Rafi Ahamed
Signal Process.2
2021 Training Accelerator for Two Means Decision Tree
abstract
Decision trees (DTs) are profusely used in machine learning (ML) applications on account of their fast execution and high interpretability. As DT training is time-consuming, in this brief, we proposed a hardware training accelerator to speedup the training process. The proposed training accelerator is implemented on the field-programmable gate array (FPGA) having a maximum operating frequency of 62 MHz. The proposed architecture uses a combination of parallel execution for training time reduction and pipelined execution to minimize resource consumption. For a given design, the proposed hardware implementation is found to be at least 14× faster than the C-based software implementation. Moreover, the proposed architecture can be easily retrained for the next set of data using a single RESET signal. This on-the-go training makes the hardware versatile for any kind of application.
Rituparna Choudhury, Shaik Rafi Ahamed, Prithwijit Guha
IEEE Trans. Very Large Scale Integr. Syst.2
2020 A Novel Time-Shared and LUT-Less Pipelined Architecture for LMS Adaptive Filter
abstract
This article presents a novel time-shared and lookup table (LUT)-less pipelined architecture for a least-mean-square (LMS) adaptive filter (ADF). The proposed approach first employs a time-shared architecture for the pipelined LMS ADF to compute each filter partial product and coefficient increment term using a single multiplier. Critical path analysis of this architecture is carried out to determine the pipeline requirements. Next, a novel LUT-less multiplier is suggested by exploiting the symmetries between the odd-multiples. Due to the symmetries between the odd-multiples, offset terms are added using an adder tree. For higher wordlength coefficients, only a few adders are required to generate the odd-multiples, and only one offset adder tree is required. Finally, a novel super-latch is developed to pipeline the LUT-less multiplier with adaptation delays of the pipelined LMS ADF. From the implementation results, it is found that the proposed design for the 32nd-order filter occupies 60.32% less area, consumes 61.93% less power, and utilizes 58.83% less sliced LUTs and 63.28% fewer flip-flops over the best existing design.
Ranendra Kumar Sarma, Mohd. Tasleem Khan, Shaik Rafi Ahamed, Jinti Hazarika
IEEE Trans. Very Large Scale Integr. Syst.3
2020 Low-Complexity Distributed-Arithmetic-Based Pipelined Architecture for an LSTM Network
abstract
Long short-term memory (LSTM) networks have addressed the shortcomings of recurrent neural networks, such as vanishing gradients and the lack of ability in developing connections across discontinuous parts of sequences. However, the implementations of state-of-the-art LSTM networks face the computational bottleneck of having multiple high-order matrix-vector multiplications (MVMs). This article presents a generalized approach to accelerate a circulant MVM (C-MVM), and hence, it is applicable to many neural networks. The proposed scheme presents a novel low-complexity distributed arithmetic (DA) architecture for optimizing C-MVMs. Unlike conventional offset binary coding-based DA (OBC-DA), it is based on separate generation and selection of partial products. Only one partial product generator (PPG) with several partial product selectors (PPSs) is required. The complexity of PPSs is reduced by sharing the minterms across Boolean expressions. Fine-grained pipelining is employed to achieve approximately one adder delay. From the implementation results, the proposed design with 512 × 512 LSTM layer occupies 74.54% less core area, consumes 68.66% less core power, offers 2.61 times more throughput, and 3.89 times more hardware efficiency over the best existing design.
Krishna Praveen Yalamarthy, Saurabh Dhall, Mohd. Tasleem Khan, Shaik Rafi Ahamed
IEEE Trans. Very Large Scale Integr. Syst.4
2018 Analysis and Implementation of Block Least Mean Square Adaptive Filter using Offset Binary Coding
abstract
In this paper, we present a new mathematical analysis and implementation of block-least-mean-square (BLMS) for finite-impulse-response (FIR) adaptive filter using offset binary coding (OBC). The proposed approach is based on distributed arithmetic (DA) in which OBC combination of input samples are stored in look-up-table (LUT). The filter output and weight adaptation terms are computed by successive shift-and-accumulation (SA) of LUT contents. The recursive use of OBC scheme have reduced the complexity of LUT. Also, we suggested new structure for LUT update unit which involve fewer adders. In addition, a new SA unit is proposed to include the offset term in OBC representation of weights and error signals. Application Specific Integrated Circuit (ASIC) synthesis results show that the proposed scheme utilizes 51.73% less area, consumes 48.32% less power, provides nearly 2.33 times higher throughput for 32ndorder filter with block length of 4 over the best existing scheme.
Mohd. Tasleem Khan, Shaik Rafi Ahamed
ISCAS2
2016 DA based approach for the implementation of block adaptive decision feedback equaliser
abstract
Adaptive decision feedback equalisers (ADFEs) are used in wireless transmission systems for mitigating the InterSymbol Interference (ISI) that occurs due to multipath propagation of the transmitted signal. In case of high speed communications which involve rapid‐varying channels, fast convergent and low complexity ADFEs are required. Frequency domain block processing of signals is an effective means of handling the increased complexities in such high speed systems. However, block ADFEs being inherently non‐causal, the feedback filter (FBF) suffers from the lack of unknown decisions in every block computation. In this study, the authors propose an efficient approach for the computation of these unknown decisions. For this, the authors tried to minimise a cost function based on two criteria mean‐absolute difference (MAD) and mean‐square difference between the ADFE output and the adder that sums up the feedforward filter (FFF) and FBF. A bank of registers which store all the symbols used in the modulation scheme is also utilised for this purpose. Using the proposed solution, the authors recast block ADFE equations in the frequency domain using distributed arithmetic (DA). The algorithm has a good convergence performance and utilises only few computations compared with the existing schemes.
Matcha Surya Prakash, Shaik Rafi Ahamed
IET Signal Process.2
2013 A block floating point treatment to finite precision realization of the adaptive decision feedback equalizer
Shaik Rafi Ahamed, Mrityunjoy Chakraborty
Signal Process.1
2011 Efficient sign based normalized adaptive filtering techniques for cancelation of artifacts in ECG signals: Application to wireless biotelemetry
Md. Zia Ur Rahman 0001, Shaik Rafi Ahamed, D. V. Rama Koti Reddy
Signal Process.2
2009 An Efficient Noise Cancellation Technique to Remove Noise from the ECG Signal Using Normalized Signed Regressor LMS Algorithm
abstract
In this paper, we present a simple and efficient normalized signed regressor LMS (NSRLMS) algorithm, that can be applied to ECG signal in order to remove various artifacts from them. This algorithm enjoys less computational complexity because of the sign present in the algorithm and good filtering capability because of the normalized term. As a result it is particularly suitable for applications requiring large signal to noise ratios with less computational complexity. Simulation studies shows that the proposed realization gives better performance compared to existing realizations in terms of signal to noise ratio.
Md. Zia Ur Rahman 0001, Shaik Rafi Ahamed, D. V. Rama Koti Reddy
BIBM2
2008 An efficient finite precision realization of the block adaptive decision feedback equalizer
abstract
Recently, a block based adaptive decision feedback equalizer (ADFE) is presented which first uses an iterative scheme to evaluate a block of unknown decisions. FFT based block processing is then used on the received input block and the decision block to carry out the block ADFE operation. A direct floating point (FP) based realization of this scheme, however, pushes up the cost and complexity of processing hugely, as each FP operation involves several additional steps not present in its fixed point (FxP) counterpart. To overcome this problem, a block floating point (BFP) based treatment is presented in this paper for realization of the block ADFE. The proposed scheme, while maintaining FP like high dynamic range, deploys mostly FxP operations and thus reduces the processing cost and complexity substantially.
Shaik Rafi Ahamed, Mrityunjoy Chakraborty, Santanu Chattopadhyay
ISCAS1
2007 An Efficient Finite Precision Realization of the Adaptive Decision Feedback Equalizer
abstract
A scheme for efficient finite precision realization of the adaptive decision feedback equalizer using block floating point (BFP) arithmetic is presented. The scheme adopts separate BFP formats for the feed forward and the feedback filter weights and works out separate update relations for their respective mantissas and exponents. Care is taken to prevent overflow in all computations by using a dynamic scaling of the data and a carefully chosen upper bound for the step sizeμ. Since no block processing of the feedback input is possible, an efficient scheme is presented for block formatting the data stored in the feedback filter memory, at each time index. The proposed scheme mostly employs simple fixed point operations and achieves considerable speed up over its floating point counterpart.
Shaik Rafi Ahamed, Mrityunjoy Chakraborty
ISCAS1