Md Tanvir Arafin

dblp:161/2445 · DBLP profile ↗
← Back
18ranked-venue papers
9as first author
10since 2021 · last 2026
0000-0002-5179-5216ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 17 · 9 first-author · 9 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 GPU Acceleration of the Sum-Check Protocol Over Towers of Binary Fields for Verifiable Computing
abstract
Emerging zero-knowledge proof protocols such as Binius and Binius-FRI operate over towers of binary fields, allowing for ultra-fast polynomial commitments over a base field. Sum-check, a key protocol in algebraic proof systems, is one of the key implementation bottlenecks for Binius and similar protocols. While sum-check is a massively parallel algorithm, GPU acceleration of sum-check has received little attention due to the lack of native GPU support for binary field multiplication. Hence, in this paper, we explore the key issues in existing GPU-based sum-check accelerators and present SumCATS - an efficient GPU implementation for sum-check acceleration. SumCATS leverages two fundamental improvements over the existing solutions. First, it adapts a CPU-based algorithmic improvement to sum-check proving and applies it to GPUs by recognizing the reduction pattern and shared memory optimizations. Secondly, SumCATS reduces the number of global memory accesses by precomputing products of random challenges and using base field operations to reconstruct extension field elements. When these optimizations are combined, SumCATS achieves a significant speedup (1.81× on NVIDIA RTX 3090 Ti, 1.62× on NVIDIA A100) over the baseline GPU implementation (Binius-GPU) for sum-check over binary tower fields. The code and research artifacts for SumCATS design are available at https://github.com/SPIRE-GMU/sum_cats.
Andrew Fan, Yanze Wu, Harry Han, Md Tanvir Arafin
DATE4
2025 Energy-Efficient Acceleration of Hash-Based Post-Quantum Cryptographic Schemes on Embedded Spatial Architectures
abstract
This work introduces AXIOS, a novel spatial architecture for accelerating hash-based post-quantum cryptography (PQC) primitives. AXIOS demonstrates that structural regularities in hash-based algorithms can be efficiently mapped to spatial accelerators that support FPGA-based programming for granular control and coarse-grained reconfigurable arrays (CGRA) for repeated tasks. AXIOS selects the key generation task of the eXtended Markle Signature Scheme (XMSS), which embodies critical implementation challenges in modern hash-based (PQC) algorithms. The AXIOS implementation on AMD’s VCK190 platform demonstrates an $8.54 \times$ improvement in runtime and a $71.65 \times$ improvement in energy efficiency compared to a benchmark implementation on Intel’s Core $\mathbf{i 9 - 1 4 9 0 0 K}$. AXIOS also breaks the current record of XMSS acceleration in terms of execution time on an embedded SoC or an FPGA platform. To our knowledge, this is the first efficient hardware implementation of compute-intensive hash-based PQC schemes in an embedded spatial architecture. Albeit complex, this FPGA+CGRA-based design is a promising step to support compute-intensive PQC applications at the edge. This work’s code and experimental artifacts are publicly available at https://github.com/SPIRE-GMU/AXIOS.
Yanze Wu, Md Tanvir Arafin
PACT2
2025 Thena: Torus Fully Homomorphic Encryption on Energy-Efficient Heterogeneous Architecture
abstract
Fully Homomorphic Encryption (FHE) enables privacy-preserving computations on encrypted data with strong security guarantees. Torus-based FHE (TFHE) emerges as a promising candidate among FHE variants due to its efficient Boolean logic operation and unlimited computational depth. However, it heavily relies on bootstrapping, a computationally intensive technique. Although there has been significant progress in improving the throughput and latency of the bootstrapping process, there exists a gap in the energy efficiency research of this process without compromising its speed. Also, energy-efficient implementation of TFHE is a key requirement for its application in energy-constrained systems. This work introduces THENA, an energy-efficient bootstrapping accelerator for TFHE built on a heterogeneous Versal adaptive system on chip (ASoC) platform to address this gap. THENA partitions the bootstrapping workload into different parts of ASoC: the serial operations are handled by the processing system (PS), the compute-intensive torus multiplications are mapped to the adaptive intelligent engine (AIE), and the memory and communication operations are allocated on the programming logic (PL). THENA derives a wavefront arraybased energy-efficient multiplier, achieving a higher$(2 \times)$improvement in throughput over a similar implementation (SaberNTT, TCAS '23). THENA uses this multiplier to deliver an end-to-end bootstrapping accelerator on the Versal VCK-190 platform. THENA delivers$7 \times$better energy efficiency for bootstrapping than GPU-based CuFHE (RTX 3090) and outperforms existing complete FPGA designs, such as YKP (HPEC '22) by demonstrating 17.5%, and 35.6% decrease in latency and energy consumption. To the best of our knowledge, this is the first PS+PL+AIE-based heterogeneous TFHE accelerator on Versal ASoCs. THENA's code and experimental artifacts are published at https://github.com/SPIRE-GMU/tfhe-aie/.
Yanze Wu, Md Tanvir Arafin
ICCD2
2025 DEMO: Radio Unit Activity Fingerprinting through Electromagnetic Side-Channel Analysis in O-RAN Networks
abstract
While the disaggregated architecture of the industry-driven Open Radio Access Network (O-RAN) promises to foster vendor competition, accelerate innovation, and reduce cost for 5G/6G cellular network deployments, it also exposes the cellular network to various new cybersecurity and privacy vulnerabilities. This demo paper highlights one such new potential cybersecurity vulnerability in the Radio Unit (RU) of O-RAN networks, where an adversary can infer RU activity by analyzing electromagnetic side-channel emissions. We present a custom-built, open-source cellular O-RAN testbed equipped with EM measurement capabilities that enables direct observation of the FPGA-based RU during operation. By capturing EM emissions from the RU, we extract side-channel traces that reveal the underlying RU activity. These traces are then analyzed using a Random Forest-based machine learning classifier, which accurately distinguishes between different RU activity patterns. Our preliminary findings demonstrate the feasibility of inferring RU-level operations via passive EM observation, highlighting a previously unexplored security threat in O-RAN systems. All code and experimental artifacts are made publicly available at https://github.com/SPIRE-GMU/NextGRadio_Sidechanel.
Sreenithya Somavarapu, Harshita Chaudhari, Nour El Houda Aidlaid, Nongnapat Adchariyavivit, Qais Dib, Moinul Hossain, Vijay Kumar Shah, Md Tanvir Arafin
WISEC8
2024 An RRAM-Based Computing-in-Memory Architecture and Its Application in Accelerating Transformer Inference
abstract
Deep neural network (DNN)-based transformer models have demonstrated remarkable performance in natural language processing (NLP) applications. Unfortunately, the unique scaled dot-product attention mechanism and intensive memory access pose a significant challenge during inference on power-constrained edge devices. One emerging solution to this challenge is computing-in-memory (CIM), which uses memory cells for logic computation to reduce data movement and overcome the memory wall. However, existing CIM designs do not support high-precision computations, such as floating-point operations, which are essential for NLP applications. Furthermore, CIM architectures require complex control modules and costly peripheral circuits to harness the full potential of in-memory computation. Hence, this article proposes a scalable RRAM-based in-memory floating-point computation architecture (RIME) that uses single-cycle NOR, NAND, and minority logic to implement in-memory floating-point operations. RIME features efficient parallel and pipeline capabilities with a centralized control module and a simplified peripheral circuit to eliminate data movement during computation. Furthermore, the article proposes pipelined implementations of matrix–matrix multiplication (MatMul) and softmax functions, enabling the construction of a transformer accelerator based on RIME. Extensive experimental results show that compared with GPU-based implementation, the RIME-based transformer accelerator improves timing efficiency by$2.3\times $and energy efficiency by$1.7\times $without compromising inference accuracy.
Zhaojun Lu, Md Tanvir Arafin, Haoxiang Yang, Zhenglin Liu, Jiliang Zhang 0002, Gang Qu 0001
IEEE Trans. Very Large Scale Integr. Syst.3
2023 ADLPT: Improving 3D NAND Flash Memory Reliability by Adaptive Lifetime Prediction Techniques
abstract
NAND flash memory has become increasingly popular in various computing systems. Although NAND flash memory offers attractive performance, it suffers limited operable programming and erasing cycles. To improve the reliability of flash-based systems, previous works introduce machine learning models to predict flash lifetime. These works generally focus on improving prediction accuracy but present little research about the resources required for flash lifetime prediction. In application scenarios, the overheads and the frequency of lifetime predictions are important for storage systems. Excessive prediction actions would lead to unnecessary resource consumption. For building an efficient storage system, resource requirements need to be taken into consideration when designing flash lifetime prediction schemes. In this paper, we propose adaptive lifetime prediction techniques (ADLPT) that minimize redundant prediction operations by exploiting reliability variation. To explore reliability variation, we investigate the error distribution of different 3D flash chips. Based on the investigation, a prediction judgment method is presented. The method identifies the necessary prediction by detecting the variation of erase duration and raw bit errors. Furthermore, we provide a method to improve the performance of the static model. The experimental result shows that our approach can reduce about 90% of redundant predictions with over 0.8 F1-Score.
Yuqian Pan, Zhaojun Lu, Haichun Zhang, Md Tanvir Arafin, Zhenglin Liu, Gang Qu 0001
IEEE Trans. Computers5
2022 Computation-in-Memory Accelerators for Secure Graph Database: Opportunities and Challenges
abstract
This work presents the challenges and opportunities for developing computing-in-memory (CIM) accelerators to support secure graph databases (GDB). First, we examine the database backend of common GDBs to understand the feasibility of CIM-based hardware architectures to speed up database queries. Then, we explore standard accelerator designs for graph computation. Next, we present the security issues of graph databases and survey how advanced cryptographic techniques such as homomorphic encryption and zero-knowledge protocols can execute privacy-preserving queries in a secure graph database. After that, we illustrate possible CIM architectures for integrating secure computation with GDB acceleration. Finally, we discuss the design overheads, useability, and potential challenges for building CIM-based accelerators for supporting data-centric calculations. Overall, we find that computing-in-memory primitives have the potential to play a crucial role in realizing the next generation of fast and secure graph databases.
Md Tanvir Arafin
ASP-DAC1
2022 Voltage Over-Scaling-Based Lightweight Authentication for IoT Security
abstract
It is a challenging task to deploy lightweight security protocols in resource-constrained IoT applications. A hardware-oriented lightweight authentication protocol based on device signature generated during voltage over-scaling (VOS) was recently proposed to address this issue. VOS-based authentication employs the computation unit such as adders to generate the process variation dependent error, which is combined with secret keys to create a two-factor authentication protocol. In this article, machine learning (ML)-based modeling attacks to break such authentication is presented. We also propose achallengeself-obfuscationstructure (CSoS) which employs previous challenges combined with keys or random numbers to obfuscate the current challenge for the VOS-based authentication to resist ML attacks. Experimental results show that ANN, RNN, and CMA-ES can clone the challenge-response behavior of VOS-based authentication with up to 99.65 percent prediction accuracy, while the prediction accuracy is less than 51.2 percent after deploying our proposed ML resilient technique. In addition, our proposed CSoS also shows good obfuscation ability for strong PUFs. Experimental results show that the modeling accuracy is below 54 percent when 106challenge-response pairs (CRPs) are collected to model the CSoS-based Arbiter PUF with ML attacks based on LR, SVM, ANN, RNN, and CMA-ES.
Jiliang Zhang 0002, Chaoqun Shen, Haihan Su, Md Tanvir Arafin, Gang Qu 0001
IEEE Trans. Computers4
2021 RIME: A Scalable and Energy-Efficient Processing-In-Memory Architecture for Floating-Point Operations
abstract
Processing in-memory (PIM) is an emerging technology poised to break the memory-wall in the conventional von Neumann architecture. PIM reduces data movement from the memory systems to the CPU by utilizing memory cells for logic computation. However, existing PIM designs do not support high precision computation (e.g., floating-point operations) essential for critical data-intensive applications. Furthermore, PIM architectures require complex control module and costly peripheral circuits to harness the full potential of in-memory computation. These peripherals and control modules usually suffer from scalability and efficiency issues.
Zhaojun Lu, Md Tanvir Arafin, Gang Qu 0001
ASP-DAC2
2021 Security of Neural Networks from Hardware Perspective: A Survey and Beyond
abstract
Recent advances in neural networks (NNs) and their applications in deep learning techniques have made the security aspects of NNs an important and timely topic for fundamental research. In this paper, we survey the security challenges and opportunities in the computing hardware used in implementing deep neural networks (DNN). First, we explore the hardware attack surfaces for DNN. Then, we report the current state-of-the-art hardware-based attacks on DNN with focus on hardware Trojan insertion, fault injection, and side-channel analysis. Next, we discuss the recent development on detecting these hardware-oriented attacks and the corresponding countermeasures. We also study the application of secure enclaves for the trusted execution of NN-based algorithms. Finally, we consider the emerging topic of intellectual property protection for deep learning systems. Based on our study, we find ample opportunities for hardware based research to secure the next generation of DNN-based artificial intelligence and machine learning platforms.
Qian Xu 0022, Md Tanvir Arafin, Gang Qu 0001
ASP-DAC2
2020 Security Challenges of Processing-In-Memory Systems
abstract
Emerging memory systems such as resistive random access memory (RRAM), phase-change memory (PCM), and spin-transfer torque magneto-resistive random access memory (STT-MRAM) offer unique physical properties useful in designing next-generation processing in-memory (PIM) circuits and systems. Modified dynamic random access memory (DRAM) designs are also demonstrating on-chip data processing and bulk data operation capabilities. However, in-memory computation can fundamentally change the security models and assumptions of existing systems due to several key factors, such as modified system architecture, disparate programming models, side-channel effects, device reliability, hardware Trojans, and malicious perturbations in data processing. Therefore, in this paper, we survey and examine fundamental vulnerabilities arising from processing-in-memory systems. We aim to present the PIM system architects and designers an overview of security issues that can jeopardize the future of in-memory computation.
Md Tanvir Arafin, Zhaojun Lu
ACM Great Lakes Symposium on VLSI1
2019 LPN-based Device Authentication Using Resistive Memory
abstract
Recent progress in the design and implementation of resistive memory components such as RRAMs and PCMs has introduced opportunities for developing novel hardware security solutions using unique physical properties of these devices. In this work, we utilize the faults in HfOx-based resistive RRAMs to design secure, lightweight device authentication protocols. To detail our design, first, we introduce the device breakdown problem due to high bias conditions in resistive memory and the physics behind non-recoverable resistive states. Then, using the concepts of learning with parity noise (LPN) based authentication protocols, we demonstrate that simple READ and WRITE operations on resistive memory cells with defects can perform necessary calculation required for LPN-based authentication schemes. Next, we design two simple authentication protocols using resistive memory based hardware and provide a detailed security analysis for these protocols. We find that these authentication mechanisms can offer significant improvement against its CMOS counterpart regarding the area and power budget. Finally, we provide detailed physical design requirements for the memory components. The resistive memory components that are capable of performing the proposed authentication protocols have also been designed and fabricated. From our analysis, we find that these memory dependent authentication protocols are lightweight, resistant to learning attacks from active and passive adversaries, and reliable under normal changes in operating conditions.
Md Tanvir Arafin, Haoting Shen, Mark Tehranipoor, Gang Qu 0001
ACM Great Lakes Symposium on VLSI1
2018 Memristors for Secret Sharing-Based Lightweight Authentication
abstract
User authentication is one of the most fundamental security problems that design effective ways of identifying single or multiple entities using shared information, signatures, or intrinsic properties of the user(s). Password-based authentication is standard in the computer systems; however, passwords usually have low entropy content, and therefore vulnerable to dictionary attacks. Furthermore, password storage and simultaneous multiparty authentication also pose security and privacy concerns. Secret sharing-based techniques during password enrollment are found to be helpful in securing key storage in the authentication server, and in assisting multiparty authentication without exposing individual identity. However, secret sharing techniques, such as Shamir's secret sharing, are computationally expensive; therefore, its implementation in power-constrained systems is elusive. To address these problems, we have demonstrated how secure and lightweight user authentication techniques can be designed using several well-known properties of memristive devices. For developing our secret sharing-based computationally lightweight user authentication protocols, first, we define essential utility functions, such as Read State, SET Pulse Count, Preconditioning, and so on, for controlling conductive filament formation in memristive devices. Next, we demonstrate the implementation of hardware-dependent simple authentication protocols that can ensure secure key storage using secret sharing protocol derived from Shamir and Naor's visual cryptographic constructs. Then, we lay out the required hardware design and discuss the potential attacks to these protocols and the corresponding countermeasures. We conclude that, under realistic attacking assumptions, the proposed protocols are secure. Finally, using PTM's 65-nm MOSFET models and Stanford's variation-aware memristor models, we perform HSPICE simulation of the secret-reconstruction and authentication units to demonstrate the reliability of the hardware designs against SET-RESET unbalance, noise, temperature fluctuations, and aging.
Md Tanvir Arafin, Gang Qu 0001
IEEE Trans. Very Large Scale Integr. Syst.1
2017 VOLtA: Voltage over-scaling based lightweight authentication for IoT applications
abstract
Incorporating security protocols in IoT components is challenging due to their extremely constrained resources. We address this challenge by proposing a hardware-oriented lightweight authentication protocol based on device signature generated during voltage over-scaling (VOS). First, we demonstrate that VOS-based computing leaves a process variation dependent error signature in its approximate results. This error can be methodically profiled to extract information about the underlying process variation in the computation unit. We then combine this error profile with security key based authentication schemes to create a two-factor authentication mechanism. To understand the effectiveness of this protocol, we perform detailed security analysis under various attack scenarios. Finally, we simulate the authentication hardware using a process variation aware 45nm design library in HSpice. Simulation results show that our VOS-based assumptions are valid, and this authentication mechanism can withstand basic environmental variations. Overall, our approach provides a unique approach for using hardware process variations as a key for authentication.
Md Tanvir Arafin, Gang Qu 0001
ASP-DAC1
2017 A Low-Cost GPS Spoofing Detector Design for Internet of Things (IoT) Applications
abstract
The civilian Global Positioning System (GPS) is widely used for precise positioning, timekeeping, and synchronization in embedded systems. As a result, emerging digital infrastructure such as the Internet of Things (IoT) are dependent on GPS to locate and synchronize \textit{Things} in the network. From a security perspective, civilian GPS signals are vulnerable to malintent because they are not encrypted and can easily be spoofed. Several countermeasures have been proposed to detect GPS spoofing attacks, but most of them require extensive signal processing capabilities and additional electronic components to capture and analyze RF signals. These add-ons may not be available to IoT devices, and if present, they will affect the device\textquoteright s power budget significantly. Therefore, new techniques for spoofing detection and survival are required before integrating GPS receivers with IoT devices and other critical infrastructures where energy and computation power are limited. In this work, we propose a novel GPS spoofing detection scheme based on hardware oscillators. Our design depends on measuring the frequency drift and offset of a free-running crystal oscillator with respect to the GPS signals. In our secure GPS spoofing detector design the trust is intrinsic, \textit{i.e.}, the receiver only trusts the on-board free running local oscillator. Intrinsic properties of these oscillators exhibit a strong correlation with the authentic GPS signals and any anomaly in this measurement will indicate potential attacks on the received GPS signals. This proposed design is cost-effective, secure, backward compatible with existing receivers, and does not require additional RF circuitry or network connection with other clocks for detecting attacks.
Md Tanvir Arafin, Dhananjay Anand, Gang Qu 0001
ACM Great Lakes Symposium on VLSI1
2017 Hardware-based anti-counterfeiting techniques for safeguarding supply chain integrity
abstract
Counterfeit integrated circuits (ICs) and systems have emerged as a menace to the supply chain of electronic goods and products. Simple physical inspection for counterfeit detection, basic intellectual property (IP) laws, and simple protection measures are becoming ineffective against advanced reverse engineering and counterfeiting practices. As a result, hardware security-based techniques have emerged as promising solutions for combating counterfeiting, reverse engineering, and IP theft. However, these solutions have their own merits and shortcomings, and therefore, these options must be carefully studied. In this work, we present a comparative overview of available hardware security solutions to fight against IC counterfeiting. We provide a detailed comparison of the techniques in terms of integration effort, deployability, and security matrices that would assist a system designer to adopt any one of these security measures for safeguarding the product supply chain against counterfeiting and IP theft.
Md Tanvir Arafin, Andrew Stanley, Praveen Sharma
ISCAS1
2016 Secret Sharing and Multi-user Authentication: From Visual Cryptography to RRAM Circuits
abstract
In this era of Internet of Things (IoT), connectivity exists everywhere, among everything (including people) at all times. Therefore, security, trust, and privacy become crucial to the design and implementation of IoT devices [12]. However, it is challenging to build security into IoT devices because most of them are constrained by extremely limited resources such as the battery, memory, and computation power etc. Inspired by the concept of visual cryptography [4] that requires the least amount of computation and a recent work on pure hardware-based single-user authentication [6], we present a novel solution to the secret sharing and multi-user authentication problem. Our solution is built on the observation that non-volatile resistive memories display nice monotonic and additive properties during resistive state transitions. We demonstrate how to design a hardware dependent multi-user authentication protocol using resistive random access memory (RRAM)-based hardware and provide the necessary circuits for the application. Finally, we simulate the proposed circuit to understand the nature of the operation and practical problems that these designs encounter during operation.
Md Tanvir Arafin, Gang Qu 0001
ACM Great Lakes Symposium on VLSI1
2015 RRAM Based Lightweight User Authentication
abstract
Resistance switching memories have emerged as a promising solution for low power and high density non-volatile storage. Unique electronic properties of resistive RAMs (and memristors) have attracted not only memory applications, but other applications such as neuromorphic computation and security as well. In this paper, we investigate how to take advantage of the availability of RRAM devices or components in the system to perform lightweight user authentication. Based on several well-known features of RRAM devices, we argue that the basic requirements for user authentication are met in RRAM devices. Then, we design three RRAM utility functions, namely Read State, Read Pulse Write State, and Copy State that are critical to develop RRAM based user authentication protocols. We propose two such protocol primitives to illustrate the concepts, layout the hardware design, and discuss the potential attacks to these protocols and the corresponding countermeasures. We conclude that under realistic attacking assumptions, the proposed protocols are secure. Finally, we use PTM's 65nm MOSFET models and perform HSPICE simulation of our proposed RRAM based hardware authentication units to demonstrate the reliability of our protocols against environmental variations such as temperature, noise, unbalanced set/reset, filament formation variation and device aging.
Md Tanvir Arafin, Gang Qu 0001
ICCAD1