EDBT 2026 Demo / reviewers in the wild / expert
Himanshu Thapliyal
dblp:t/HimanshuThapliyal
· DBLP profile ↗
45ranked-venue papers
16as first author
19since 2021 · last 2026
0000-0001-9157-4517ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 42 · 14 first-author · 19 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 first-authorComputer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hybrid Quantum-Classical Optimization for MRI-Based Detection of Alzheimer's Disease and Related DementiasabstractQuantum machine learning has emerged as a potential tool for neuroimaging-based detection of Alzheimer’s disease and related dementias (ADRD). In this paper, we present a hybrid quantum-classical training framework for binary dementia detection using sagittal MRI images from the OASIS-2 dataset. The proposed framework uses a ResNet-style convolutional neural network as a backbone with two prediction heads: a classical deep neural network head and a variational quantum head. We construct a λ -weighted objective function to examine how the classical and quantum heads behave during joint optimization. Under ideal simulation, the optimized joint-loss framework achieves strong performance, with both heads reaching 0.9913 accuracy and 0.9952 F1-score. However, in a noisy simulation, the quantum head achieves an accuracy of 0.9396, indicating that the quantum modules are sensitive to noise but still preserve strong discriminative behavior. The proposed framework also outperforms the pure QNN baseline trained on a reduced 16-dimensional feature space, which achieves 0.5909 accuracy, and remains competitive with a prior quantum transfer learning baseline. These results suggest that current quantum neural networks work better as supporting tools inside classical medical imaging systems rather than replacing MRI classifiers completely. Sounak Bhowmik, Himanshu Thapliyal |
ACM Great Lakes Symposium on VLSI | 2 |
| 2026 | SELSA: SDD-Based Efficient Logic Synthesis of Adiabatic CircuitsabstractWe propose SELSA (SDD-Based Efficient Logic Synthesis of Adiabatic Circuits), a novel design automation framework for multi-level adiabatic logic. Efficient design of multi-level adiabatic logic is challenging due to the pipelined nature of adiabatic logic gates. The proposed framework addresses this challenge by strategically collapsing regions of an And-Inverter Graph (AIG) into Sentential Decision Diagrams (SDDs), while optimizing the SDDs with a novel reordering heuristic which accounts for the area overhead of adiabatic gates and buffers. SDDs are mapped to adiabatic logic gates using a simple synthesis algorithm (similar to those used for Binary Decision Diagrams) which interprets SDD internal nodes with n branches as adiabatic gates with n rails. We evaluate the overall performance of the proposed framework on the 76 combinational benchmarks in the LGSynth’91 benchmark set and compare against another recently-proposed AIG-based tool. Our tool produces an average 26.0% reduction in transistor count compared to the AIG-based tool. Joseph Clark, Himanshu Thapliyal |
ACM Great Lakes Symposium on VLSI | 2 |
| 2026 | HPAL: High-Performance Adiabatic Logic in the GHz Regime for Energy-Efficient ComputingabstractThe increasing demand for high-performance computing in modern applications has intensified the need for energy efficient digital design at GHz frequencies. While adiabatic logics offer the potential for energy efficient operation, frequency range of energy saving for these logics is fundamentally bound to lower frequencies. However, in practice adiabatic logics lose their energy efficiency before reaching this fundamental frequency and their energy efficiency at higher frequency is practically limited by circuit level non-idealities such as race condition and short circuit conduction. Thus, in this paper, by structurally minimizing non-adiabatic energy dissipation caused by race condition and short circuit conduction, an energy efficient high performance adiabatic logic (HPAL) family is proposed. HPAL employs a novel dual amplitude power clocking scheme and non-interfering charge and discharge path to reduce the energy dissipation. Post-layout simulation results of different gates show the superiority of HPAL in energy efficiency over its CMOS and adiabatic counterparts. Furthermore, the simulation results of multiplier benchmark, including 64 four-bit multipliers and 192 buffers operating at 1.31GHz, show that even by including the energy dissipation of sinusoidal clock generator, HPAL consumes at least 60% lower energy per cycle and 30% lower energy per operation, compared to CMOS counterpart. These results indicate that HPAL overcomes key limitations of conventional adiabatic logic at high frequencies and provides a practical path toward energy efficient digital systems for compute-intensive applications at GHz regime. Milad Tanavardi Nasab, Himanshu Thapliyal |
ACM Great Lakes Symposium on VLSI | 2 |
| 2026 | Scalability Analysis of Quantum Models for Stress and Emotion DetectionabstractStress and emotion detection from high-dimensional physiological signals is a challenging task, particularly when aiming for accurate classification across diverse behavioral states. Quantum machine learning (QML) is promising for modeling such high-dimensional data, but scalability is limited by qubit resources and the exponential cost of classical statevector simulation. This work studies the scalability of quantum support vector machines (QSVMs) for binary stress detection and three-class emotion recognition (Negative/Neutral/Positive) under varying qubit counts and angle-encoding strategies. We also present a comparison study with one-feature-per-qubit (1:1) and two-features-per-qubit (2:1) mappings. Experiments are executed on HPC infrastructure using NVIDIA CUDA-Q to evaluate performance, variance, and class-dependent separability at higher-qubit setups. Results show that larger Hilbert spaces can improve peak accuracy but may increase instability. At the same time, dense 2:1 encoding yields more consistent stress detection performance. For emotion recognition, scaling improves discrimination for classes like Negative and Positive more than Neutral. We find that effective QML scaling is task-dependent and benefits more from encoding design than simply increasing qubit count. Md. Saif Hassan Onim, Travis S. Humble, Himanshu Thapliyal |
ACM Great Lakes Symposium on VLSI | 3 |
| 2026 | Advancing Quantum Workforce Development Through Hands-on Education in Circuit Design, Optimization, QML, and SecurityabstractThis educational paper presents a modular workshop framework for teaching quantum programming and hardware-oriented reasoning within microelectronics education. The framework integrates five connected instructional elements. They are Qiskit foundations laboratory on gates, measurement, and Bell-state entanglement; a noise-aware quantum arithmetic module featuring GHZ circuits, ripple-carry addition, and quantum carry-lookahead addition; a structured adder-analysis exercise focused on depth, ancilla, and reversible cleanup; a capstone project on noise-resilient and security-aware 4-bit quantum adders; and a bridge from classical machine learning workflows to quantum machine learning. The paper redefines these activities explicitly as learning modules with student learning outcomes, module learning outcomes, assessment options, and discussion prompts. The resulting sequence helps learners bridge from syntax and circuit building to architecture-level trade-offs, noise interpretation, approximate design, and research-oriented thinking, all of which are well-suited for the quantum workforce. Himanshu Thapliyal, Sounak Bhowmik, Rajnish Bajpai, Md. Saif Hassan Onim |
ACM Great Lakes Symposium on VLSI | 1 |
| 2025 | Late Breaking Results: Novel Design of MTJ-Based Unified LIF Spiking Neuron and PUFabstractDue to the higher energy and hardware efficiency of spiking neural networks (SNNs) compared to deep neural networks, they have attracted a lot of attention. However, their security must be investigated, given that they have access to private and confidential data. Physically unclonable functions (PUFs) are a class of circuits with security applications like device authentication, embedded licensing, device-specific cryptographic key generation, and anti-counterfeiting. Therefore, PUFs can be used to enhance the security of SNN. Accordingly, in this paper, an MTJ-based LIF Neuron/PUF has been proposed. The proposed design can function as both a LIF neuron and PUF. The results of the Monte Carlo simulation show that the proposed design has better uniqueness and uniformity values compared to its counterparts. These values for the proposed design are $50.07 \%$ and $49.66 \%$, which are close to their ideal value of $50 \%$. Also, the mean value of Shannon entropy for the 128-bit PUF response of the proposed design is 0.9974, which is close to its ideal value of 1. Milad Tanavardi Nasab, Himanshu Thapliyal |
DAC | 3 |
| 2025 | Quantum Transfer Learning to Boost Dementia Detection
Sounak Bhowmik, Talita Perciano, Himanshu Thapliyal |
ACM Great Lakes Symposium on VLSI | 3 |
| 2025 | Emotion Detection in Older Adults Using Physiological Signals from Wearable SensorsabstractEmotion detection in older adults is crucial for understanding their cognitive and emotional well-being, especially in hospital and assisted living environments. In this work, we investigate an edge-based, non-obtrusive approach to emotion identification that uses only physiological signals obtained via wearable sensors. Our dataset includes data from 40 older individuals. Emotional states were obtained using physiological signals from the Empatica E4 and Shimmer3 GSR+ wristband and facial expressions were recorded using camera-based emotion recognition with the iMotion's Facial Expression Analysis (FEA) module. The dataset also contains twelve emotion categories in terms of relative intensities. We aim to study how well emotion recognition can be accomplished using simply physiological sensor data, without the requirement for cameras or intrusive facial analysis. By leveraging classical machine learning models, we predict the intensity of emotional responses based on physiological signals. We achieved the highest 0.782 r2 score with the lowest 0.0006 MSE on the regression task. This method has significant implications for individuals with Alzheimer's Disease and Related Dementia (ADRD), as well as veterans coping with Post-Traumatic Stress Disorder (PTSD) or other cognitive impairments. Our results across multiple classical regression models validate the feasibility of this method, paving the way for privacy-preserving and efficient emotion recognition systems in real-world settings. Md. Saif Hassan Onim, Andrew Kiselica, Himanshu Thapliyal |
ACM Great Lakes Symposium on VLSI | 3 |
| 2025 | Crosstalk Attack Resilient RNS Quantum AdditionabstractAs quantum computers scale, the rise of multi-user and cloud-based quantum platforms can lead to new security challenges. Attacks within shared execution environments become increasingly feasible due to the crosstalk noise that, in combination with quantum computer’s hardware specifications, can be exploited in form of crosstalk attack. Our work pursues crosstalk attack implementation in ion-trap quantum computers. We propose three novel quantum crosstalk attacks designed for ion trap qubits: (i) Alternate CNOT attack (ii) Superposition Alternate CNOT (SAC) attack (iii) Alternate Phase Change (APC) attack. We demonstrate the effectiveness of proposed attacks by conducting noise-based simulations on a commercial 20-qubit ion-trap quantum computer. The proposed attacks achieve an impressive reduction of up to 42.2% in output probability for Quantum Full Adders (QFA) having 6 to 9-qubit output. Finally, we investigate the possibility of mitigating crosstalk attacks by using Residue Number System (RNS) based Parallel Quantum Addition (PQA). We determine that PQA achieves higher attack resilience against crosstalk attacks in the form of 24.3% to 133.5% improvement in output probability against existing Non Parallel Quantum Addition (NPQA). Through our systematic methodology, we demonstrate how quantum properties such as superposition and phase transition can lead to crosstalk attacks and how parallel quantum computing can provide security. Bhaskar Gaur, Himanshu Thapliyal |
ISCAS | 2 |
| 2025 | Detection of Physiological Data Tampering Attacks with Quantum Machine LearningabstractThe widespread use of cloud-based medical devices and wearable sensors has made physiological data susceptible to tampering. These attacks can compromise the reliability of healthcare systems which can be critical and life-threatening. Detection of such data tampering is of immediate need. Machine learning has been used to detect anomalies in datasets but the performance of Quantum Machine Learning (QML) is still yet to be evaluated for physiological sensor data. Thus, our study compares the effectiveness of QML for detecting physiological data tampering, focusing on two types of white-box attacks: data poisoning and adversarial perturbation. The results show that QML models are better at identifying label-flipping attacks, achieving accuracy rates of 75% − 95% depending on the data and attack severity. This superior performance is due to the ability of quantum algorithms to handle complex and high-dimensional data. However, both QML and classical models struggle to detect more sophisticated adversarial perturbation attacks, which subtly alter data without changing its statistical properties. Although QML performed poorly against this attack with around 45% − 65% accuracy, it still outperformed classical algorithms in some cases. Md. Saif Hassan Onim, Himanshu Thapliyal |
ISCAS | 2 |
| 2025 | MTJ/CMOS-Based CLB Design for Low-Power and CPA-Resistant Secure Nonvolatile FPGAabstractModern applications such as the Internet of Things (IoT) devices, AI, and automotive applications widely use field-programmable gate arrays (FPGAs). However, many of these applications have limited power resources. Also, the existing FPGAs are vulnerable to side-channel attacks (SCAs) such as correlation-based power analysis (CPA) attacks. Therefore, designing low-power, CPA-resistant, and secure-by-design FPGA is required. In this article, two low-power and CPA-resistant hybrid CMOS/magnetic tunnel junction (MTJ) logic-in-memory-based configurable logic blocks (CLBs) have been proposed and compared to a state-of-the-art counterpart. The first proposed design is single output, and the second one is multioutput. The simulation results show that compared to the state-of-the-art secure CLB counterpart [secured CLB (sCLB) by Zooker et al. (2020)], the proposed CLB designs have 42% and 33% lower delay, 85% and 18% lower power consumption, and 86% and 63% fewer equivalent transistors. To implement one round of the PRESENT algorithm, the first and second designs have 85% and 77% fewer transistors, 42% and 33% lower delay, and 86% and 50% lower power consumption compared to their silicon-proven secure counterpart. Also, to implement convolution layers of binarized neural network (BNN), compared to this counterpart, the first and second proposed designs have 85% and 90% fewer equivalent transistors, 42% and 33% lower delay, and 86% and 79% lower power consumption. Also, the resiliency of the proposed designs against power analysis attacks has been investigated by exhaustive simulations and performing CPA attacks on PRESENT and Advanced Encryption Standard (AES) SBOX. Also, this resiliency has been investigated for different tunnel magnetoresistance ratios (TMRs) and supply voltages. Milad Tanavardi Nasab, Himanshu Thapliyal, Garrett S. Rose |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2025 | Guest Editorial: Selected Papers From IEEE Computer Society Annual Symposium on VLSI (ISVLSI) 2024
Himanshu Thapliyal, Jürgen Becker 0001, Garrett S. Rose, Tosiron Adegbija, Selçuk Köse |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2024 | Novel Optimized Designs of Modulo 2n+1 Adder for Quantum ComputingabstractQuantum modular adders are one of the most fundamental yet versatile quantum computation operations. They help implement the functions of higher complexity, such as subtraction and multiplication, which are used in applications, such as quantum cryptanalysis, quantum image processing, and securing communication. To the best of our knowledge, there is no existing design of quantum modulo ($2^{n}+1$) adder (QMA). In this work, we propose four quantum adders targeted specifically for modulo ($2^{n}+1$) addition. These adders can provide both regular and modulo ($2^{n}+1$) sum concurrently, enhancing their application in residue number system-based arithmetic. Our first design, QMA1, is a novel quantum modulo ($2^{n}+1$) adder. The second proposed adder, QMA2, optimizes the utilization of quantum gates within the QMA1, resulting in 37.5% reduced CNOT gate count, 46.15% reduced CNOT depth, and 26.5% decrease in both Toffoli gates and depth. We propose a third adder QMA3 that uses zero resets, a dynamic circuits-based feature that reuses qubits, leading to 25% savings in qubit count. Our fourth design, QMA4, demonstrates the benefit of incorporating additional zero resets to achieve a purer$|0$$\rangle $state, reducing quantum state preparation errors. Notably, we conducted experiments using 5-qubit configurations of the proposed modulo ($2^{n}+1$) adders on the IBM Washington, a 127-qubit quantum computer based on the Eagle R1 architecture, to demonstrate a 28.8% reduction in QMA1’s error of which do the following: 1) 18.63% error reduction happens due to gate/depth reduction in QMA2; 2) 2.53% drop in error due to qubit reduction in QMA3; and 3) 7.64% error decreased due to application of additional zero resets in QMA4. Bhaskar Gaur, Himanshu Thapliyal |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2023 | TDAG: Tree-based Directed Acyclic Graph Partitioning for Quantum CircuitsabstractWe propose the Tree-based Directed Acyclic Graph (TDAG) partitioning for quantum circuits, a novel quantum circuit partitioning method which partitions circuits by viewing them as a series of binary trees and selecting the tree containing the most gates. TDAG produces results of comparable quality (number of partitions) to an existing method called ScanPartitioner (an exhaustive search algorithm) with an 95% average reduction in execution time. Furthermore, TDAG improves compared to a faster partitioning method called QuickPartitioner by 38% in terms of quality of the results with minimal overhead in execution time. Joseph Clark, Travis S. Humble, Himanshu Thapliyal |
ACM Great Lakes Symposium on VLSI | 3 |
| 2023 | Lightweight Hierarchical Root-of-Trust Framework for CAN-based 3D Printing SecurityabstractController Area Network (CAN) has been demonstrated to have excellent applications in 3D printer communication. However, single-bus CAN network designs are plagued with vulnerabilities, such as hijacking, denial-of-service, and eavesdropping. Exploitation of these issues can result in every node in the network being compromised. In response, we propose a hierarchical tree-based design focused on protecting the CAN bus and isolating critical systems in a novel approach against these threats. By organizing ASCON-encrypted network routing into smaller, authenticated CAN sub-nets, modular 3D printers can maintain integrity, confidentiality, and authenticity of all traffic. Our preliminary results demonstrate an 88-99% hardware efficiency between pre-existing client nodes and the total nodes after conversion to our framework. Tyler Cultice, Joseph Clark, Himanshu Thapliyal |
ACM Great Lakes Symposium on VLSI | 3 |
| 2023 | Noise-Resilient and Reduced Depth Approximate Adders for NISQ Quantum ComputingabstractThe "Noisy intermediate-scale quantum" NISQ machine era primarily focuses on mitigating noise, controlling errors, and executing high-fidelity operations, hence requiring shallow circuit depth and noise robustness. Approximate computing is a novel computing paradigm that produces imprecise results by relaxing the need for fully precise output for error-tolerant applications including multimedia, data mining, and image processing. We investigate how approximate computing can improve the noise resilience of quantum adder circuits in NISQ quantum computing. We propose five designs of approximate quantum adders to reduce depth while making them noise-resilient, in which three designs are with carryout, while two are without carryout. We have used novel design approaches that include approximating the Sum only from the inputs (pass-through designs) and having zero depth, as they need no quantum gates. The second design style uses a single CNOT gate to approximate the SUM with a constant depth of O(1). We performed our experimentation on IBM Qiskit on noise models including thermal, depolarizing, amplitude damping, phase damping, and bitflip: (i) Compared to exact quantum ripple carry adder without carryout the proposed approximate adders without carryout have improved fidelity ranging from 8.34% to 219.22%, and (ii) Compared to exact quantum ripple carry adder with carryout the proposed approximate adders with carryout have improved fidelity ranging from 8.23% to 371%. Further, the proposed approximate quantum adders are evaluated in terms of various error metrics. Bhaskar Gaur, Travis S. Humble, Himanshu Thapliyal |
ACM Great Lakes Symposium on VLSI | 3 |
| 2023 | A Logarithmic Depth Quantum Carry-Lookahead Modulo (2n - 1) AdderabstractQuantum Computing is making significant advancements toward creating machines capable of implementing quantum algorithms in various fields, such as quantum cryptography, quantum image processing, and optimization. The development of quantum arithmetic circuits for modulo addition is vital for implementing these quantum algorithms. While it is ideal to use quantum circuits based on fault-tolerant gates to overcome noise and decoherence errors, the current Noisy Intermediate Scale Quantum (NISQ) era quantum computers cannot handle the additional computational cost associated with fault-tolerant designs. Our research aims to minimize circuit depth, which can reduce noise and facilitate the implementation of quantum modulo addition circuits on NISQ machines. This work presents quantum carry-lookahead modulo (2n - 1) adder (QCLMA), which is designed to receive two n-bit numbers and perform their addition with an O(log n) depth. Compared to existing work of O(n) depth, our proposed QCLMA reduces the depth and helps increase the noise fidelity. In order to increase error resilience, we also focus on creating a tree structure based Carry path, unlike the chain based Carry path of the current work. We run experiments on Quantum Computer IBM Cairo to evaluate the performance of the proposed QCLMA against the existing work and define Quantum State Fidelity Ratio (QSFR) to quantify the closeness of the correct output to the top output. When compared against existing work, the proposed QCLMA achieves a 47.21% increase in QSFR for 4-qubit modulo addition showcasing its superior noise fidelity. Bhaskar Gaur, Edgard Muñoz-Coreas, Himanshu Thapliyal |
ACM Great Lakes Symposium on VLSI | 3 |
| 2023 | CASD-OA: Context-Aware Stress Detection for Older Adults with Machine Learning and Cortisol BiomarkerabstractStress can aggravate age-related diseases that can lead to significant clinical impairment and decrease the quality of life in older adults. To mitigate the harmful effects of stress and aging, it is important to monitor and manage stress. In this paper, we have developed context-aware stress detection for older adults with machine learning and cortisol biomarker. The Trier Social Stress Test (TSST), a well-known experimental protocol that consistently inflicts stress on people in a social context, was used as the stress protocol for this study. We have used salivary cortisol as a stress biomarker for ground truth estimation. The proposed machine learning model classifies stress into three different levels (no-stress, low-stress, and high-stress) based on data collected from Electro-Dermal Activity (EDA), Blood Volume Pressure (BVP), and Inter Beat Interval (IBI) sensors. To develop a context-aware machine learning model, we have used context features captured from the TSST protocol. Using sensor fusion, our proposed context-aware machine learning model achieved a macro-average F1-score of 0.937 and an accuracy of 92.48% in distinguishing among the three stress levels. We have also illustrated that using context improves the macro-average F1-score by 0.20 and accuracy by over 20% compared to the machine learning model without context. Md. Saif Hassan Onim, Himanshu Thapliyal |
ACM Great Lakes Symposium on VLSI | 2 |
| 2022 | Fortifying Vehicular Security through Low Overhead Physically Unclonable FunctionsabstractWithin vehicles, the Controller Area Network (CAN) allows efficient communication between the electronic control units (ECUs) responsible for controlling the various subsystems. The CAN protocol was not designed to include much support for secure communication. The fact that so many critical systems can be accessed through an insecure communication network presents a major security concern. Adding security features to CAN is difficult due to the limited resources available to the individual ECUs and the costs that would be associated with adding the necessary hardware to support any additional security operations without overly degrading the performance of standard communication. Replacing the protocol is another option, but it is subject to many of the same problems. The lack of security becomes even more concerning as vehicles continue to adopt smart features. Smart vehicles have a multitude of communication interfaces an attacker could exploit to gain access to the networks. In this work, we propose a security framework that is based on physically unclonable functions (PUFs) and lightweight cryptography (LWC). The framework does not require any modification to the standard CAN protocol while also minimizing the amount of additional message overhead required for its operation. The improvements in our proposed framework result in major reduction in the number of CAN frames that must be sent during operation. For a system with 20 ECUs, for example, our proposed framework only requires 6.5% of the number of CAN frames that is required by the existing approach to successfully authenticate every ECU. Carson Labrado, Himanshu Thapliyal, Saraju P. Mohanty |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2020 | Special Session: A Novel Low-Power and Energy-Efficient Adiabatic Logic-In-Memory Architecture Using CMOS/MTJabstractHybridization of Magnetic Tunnel Junction (MTJ) devices and CMOS transistors are being investigated to design Logic-In-Memory (LIM) architecture. However, hybrid CMOS/MTJ based LIM architecture suffers from significant power consumption due to charging and discharging of a capacitive output load. Adiabatic logic is one of the low-power design techniques to design energy-efficient hardware. Therefore, we apply the adiabatic logic in hybrid CMOS/MTJ based designs to propose the novel concept of Adiabatic Logic-In-Memory (ALIM) circuits. The proposed ALIM based CMOS/MTJ circuits have reduced dynamic power consumption as compared to the existing CMOS/MTJ circuits by using the energy recovery property. As a case study, we have designed an ALIM based magnetic full adder. Simulations are performed using 45nm CMOS technology with perpendicular anisotropy CoFeB/MgOMTJ model using Cadence Spectre simulator. From the simulation results, it is verified that the proposed ALIM based Magnetic Full Adder (MFA) saves 37% of energy and 43% of power as compared to the existing Pre-Charge Sense Amplifier (PCSA) based MFA. Further, the proposed ALIM based MFA also saves 38 % of the area as compared to the existing PCSA based MFA due to reduction in the number of transistors. The low-area, low-power and low-energy consumption makes the proposed ALIM architecture an the attractive choice to design low-power circuits. Himanshu Thapliyal, S. Dinesh Kumar |
ICCD | 1 |
| 2020 | Special Session: Quantum Carry Lookahead Adders for NISQ and Quantum Image ProcessingabstractProgress in quantum hardware design is progressing toward machines of sufficient size to begin realizing quantum algorithms in disciplines such as encryption and physics. Quantum circuits for addition are crucial to realize many quantum algorithms on these machines. Ideally, quantum circuits based on fault-tolerant gates and error-correcting codes should be used as they tolerant environmental noise. However, current machines called Noisy Intermediate Scale Quantum (NISQ) machines cannot support the overhead associated with fault-tolerant design. In response, low depth circuits such as quantum carry lookahead adders (QCLA)s have caught the attention of researchers. The risk for noise errors and decoherence increase as the number of gate layers (or depth) in the circuit increases. This work presents an out-of-place QCLA based on Clifford+T gates. The QCLAs optimized for T gate count and make use of a novel uncomputation gate to save T gates. We base our QCLAs on Clifford+T gates because they can eventually be made fault-tolerant with error-correcting codes once quantum hardware that can support fault-tolerant designs becomes available. We focus on T gate cost as the T gate is significantly more costly to make fault-tolerant than the other Clifford+T gates. The proposed QCLAs are compared and shown to be superior to existing works in terms of T-count and therefore the total number of quantum gates. Finally, we illustrate the application of the proposed QCLAs in quantum image processing by presenting quantum circuits for bilinear interpolation. Himanshu Thapliyal, Edgard Muñoz-Coreas, Vladislav Khalus |
ICCD | 1 |
| 2020 | Design of Adiabatic Logic-Based Energy-Efficient and Reliable PUF for IoT DevicesabstractInternet of Things (IoT) devices have stringent constraints on power and energy consumption. Adiabatic logic has been proposed as a novel computing platform to design energy-efficient IoT devices. Physically Unclonable Functions (PUFs) is a promising paradigm to solve security concerns such as Integrated Circuit (IC) piracy, IC counterfeiting, and the like. PUFs have shown great promise for generating the secret bits that can be used in the secure systems in an inexpensive way. However, designing a reliable PUF along with energy-efficiency is a big challenge. Therefore, for energy-efficient and reliable PUFs, we are proposing a novel energy-efficient adiabatic logic-based PUF structure. The proposed adiabatic PUF uses energy recovery concept to achieve high energy efficiency and uses the time ramp voltage to exhibit the reliable start-up behavior. The channel length of the transistors play a major role in controlling manufacturing variations. So, in this article, the circuit simulations are performed with 180nm and 45nm Complementary metal-oxide-semiconductor (CMOS) technology in a Cadence Spectre simulator to analyze the impact of channel length variations. The proposed adiabatic PUF has worst-case reliability of 96.84% and 99.6% with temperature variations at 180nm and 45nm CMOS technology, respectively. Further, the proposed adiabatic PUF consumes 1.071fJ/bit-per cycle at 180nm CMOS technology and 0.08fJ/bit-per cycle at 45nm CMOS technology. S. Dinesh Kumar, Himanshu Thapliyal |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2019 | Solving Energy and Cybersecurity Constraints in IoT Devices Using Energy Recovery ComputingabstractWith the growth of Internet-of-Things (IoT), the potential threat vectors for malicious cyber and hardware attacks are rapidly expanding. As the IoT paradigm emerges, there are challenging requirements to design energy-efficient and secure systems. To address these challenges, we illustrate energy recovery computing as a potential solution to design low-energy hardware security primitives for IoT devices. Energy Recovery (ER) is a circuit design technique in which circuits recycle the charge stored in the load capacitor. This overview provides example applications of ER computing in (i) low-energy and Differential Power Analysis (DPA) resistant design, (ii) low-energy Physically Unclonable Function (PUF), and (iii) hardware trojan detection. Himanshu Thapliyal, Zachary Kahleifeh |
ACM Great Lakes Symposium on VLSI | 1 |
| 2019 | Design of a Piezoelectric-Based Physically Unclonable Function for IoT SecurityabstractAccording to a report from McAfee and the center for strategic and international studies, worldwide financial loss due to cybercrime was estimated to be $600 billion in 2017. Researchers are currently exploring new methods for preventing cybercrime in Internet of Things (IoT) devices. Physically unclonable functions (PUFs) show promise as a device that could help in the fight against cybercrime. PUFs are a class of circuit that are unique and unclonable due to inherent variations caused by the device manufacturing process. We can take advantage of these PUF properties by using the outputs of PUFs to generate secret keys or pseudonyms that are similarly unique and unclonable. In recent years, energy harvesting devices, such as piezoelectric devices have been integrated with IoT devices for various purposes such as power generation and sensing applications. In this paper we propose a PUF design based on piezo sensors which are already commonly found in IoT devices. Our proposed PUF is tested in terms of reliability and uniformity. Carson Labrado, Himanshu Thapliyal |
IEEE Internet Things J. | 2 |
| 2019 | Quantum Circuit Design of a T-count Optimized Integer MultiplierabstractQuantum circuits of many qubits are extremely difficult to realize; thus, the number of qubits is an important metric in a quantum circuit design. Further, scalable and reliable quantum circuits are based on fault tolerant implementations of quantum gates such as Clifford+T gates. An efficient quantum circuit saves quantum hardware resources by reducing the number of T gates without substantially increasing the number of qubits. This work presents a T-count optimized quantum circuit for integer multiplication with only 4 · n + 1 qubits and no garbage outputs. The proposed quantum multiplier design reduces the T-count by using a novel quantum conditional adder circuit. Also, where one operand to the conditional adder is zero, the conditional adder is replaced with a Toffoli gate array to further save T gates. Average T-count savings of 46:12, 47:55, 62:71 and 26.30 percent are achieved compared to the recent works by Kotiyal et al., Babu, Lin et al., and Jayashree et al., respectively. Edgard Muñoz-Coreas, Himanshu Thapliyal |
IEEE Trans. Computers | 2 |
| 2018 | T-count and Qubit Optimized Quantum Circuit Design of the Non-Restoring Square Root AlgorithmabstractQuantum circuits for basic mathematical functions such as the square root are required to implement scientific computing algorithms on quantum computers. Quantum circuits that are based on Clifford+T gates can easily be made fault tolerant, but the T gate is very costly to implement. As a result, reducing T-count has become an important optimization goal. Further, quantum circuits with many qubits are difficult to realize, making designs that save qubits and produce no garbage outputs desirable. In this work, we present a T-count optimized quantum square root circuit with only 2 ṡ n + 1 qubits and no garbage output. To make a fair comparison against existing work, the Bennett’s garbage removal scheme is used to remove garbage output from existing works. We determined that our proposed design achieves an average T-count savings of 43.44%, 98.95%, 41.06%, and 20.28% as well as qubit savings of 85.46%, 95.16%, 90.59%, and 86.77% compared to existing works. Edgard Muñoz-Coreas, Himanshu Thapliyal |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2018 | FinSAL: FinFET-Based Secure Adiabatic Logic for Energy-Efficient and DPA Resistant IoT DevicesabstractWith the emergence of Internet of Things (IoT), there is an urgent need to design energy-efficient and secure IoT devices. For example, IoT devices such as radio frequency identification tags and wireless sensor nodes employ AES cryptographic module that are susceptible to differential power analysis (DPA) attacks. With the scaling of technology, leakage power in the cryptographic device increases, which increases their vulnerability to DPA attack. This paper presents a novel FinFET-based secure adiabatic logic (FinSAL), that is energy-efficient and DPA-immune. The proposed adiabatic FinSAL is used to design logic gates such as buffers, XOR, and NAND. Further, the logic gates based on adiabatic FinSAL are used to implement a positive polarity Reed Muller architecture-based S-box circuit. SPICE simulations at 12.5 MHz show that adiabatic FinSAL (20-nm FinFET technology) S-box circuit saves up to 81% of energy per cycle as compared to the conventional S-box circuit implemented using FinFET (20-nm FinFET technology). Further, the security of adiabatic FinSAL S-box circuit has been evaluated by performing the DPA attack through SPICE simulations. We proved that the FinSAL S-box circuit is resistant to a DPA attack through a developed DPA attack flow applicable to SPICE simulations. Further, the impact of FinSAL on hardware security at different technology nodes of FinFETs (7, 10, 14, and 16 nm) are evaluated. From the simulation results, FinSAL gates at 14-nm FinFET offer superior security with optimum power consumption, therefore is the best candidate to design low-power secure IoT devices. S. Dinesh Kumar, Himanshu Thapliyal, Azhar Mohammad |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2017 | Design of majority logic based approximate arithmetic circuitsabstractThe increasing amount of circuit density possible in CMOS technology has the consequence of also increasing the power consumption of circuits using the technology. One possible method of offsetting these increased power demands is to use approximate computing designs in circuits where complete accuracy is not a strict requirement. These circuits use fewer logic gates which reduces power consumption at the cost of accuracy. Another possible method for reducing power consumption is to use an emerging nanotechnology which is already low power in nature. Combining approximate computing with an emerging nanotechnology has the potential to further cut power consumption. Unfortunately, existing approximate computing circuits were designed using standard logic gates found in CMOS technology which in turn can limit their effectiveness when implemented with the majority based logic used by some emerging nanotechnologies. For that reason, we propose designs of approximate arithmetic units which are specifically designed for use in majority logic based technologies. Carson Labrado, Himanshu Thapliyal, Fabrizio Lombardi |
ISCAS | 2 |
| 2017 | Energy-efficient magnetic circuits based on nanoelectronic devicesabstractAs CMOS technology scales down to the nanoscale, high leakage power consumption becomes the main problem and challenge of electronic circuits. To overcome this challenge, nano-emerging technologies and logic-in-memory structure are being studied. Magnetic tunnel junction (MTJ) is an emerging technology which has many advantages when used in logic in memory structures in conjunction with CMOS. In this paper, we present novel designs of hybrid MTJ/CMOS circuits; AND, XOR and 1-bit full adder. The proposed MTJ/CMOS full adder design has 71% lower Power-delay-product (PDP) compared to the previous MTJ/CMOS full adder. To further improve the energy efficiency we investigated the use of nanoelectronic devices (CNFET, FinFET) in the proposed circuits and compared them with the CMOS based designs. The hybrid MTJ/CNFET and MJT/FinFET full adders have about 18 and 11 times lower PDP, respectively, when compared to the MJT/CMOS design. Also, the MTJ/CNFET based full adder has 66% lower PDP than the MTJ/FinFET based design. Fazel Sharifi, Himanshu Thapliyal |
ISCAS | 2 |
| 2017 | Design exploration of a Symmetric Pass Gate Adiabatic Logic for energy-efficient and secure hardware
S. Dinesh Kumar, Himanshu Thapliyal, Azhar Mohammad, Kalyan S. Perumalla |
Integr. | 2 |
| 2017 | Automatic synthesis of quaternary quantum circuits
Mozammel H. A. Khan, Himanshu Thapliyal, Edgard Muñoz-Coreas |
J. Supercomput. | 2 |
| 2016 | Ancilla-input and garbage-output optimized design of a reversible quantum integer multiplier
H. V. Jayashree, Himanshu Thapliyal, Hamid R. Arabnia, Vinod Kumar Agrawal |
J. Supercomput. | 2 |
| 2016 | Design procedures and NML cost analysis of reversible barrel shifters optimizing garbage and ancilla lines
Himanshu Thapliyal, Carson Labrado |
J. Supercomput. | 1 |
| 2016 | Erratum to: Design procedures and NML cost analysis of reversible barrel shifters optimizing garbage and ancilla lines
Himanshu Thapliyal, Carson Labrado |
J. Supercomput. | 1 |
| 2015 | Reversible logic based multiplication computing unit using binary tree data structure
Saurabh Kotiyal, Himanshu Thapliyal, N. Ranganathan |
J. Supercomput. | 2 |
| 2013 | Design of efficient reversible logic-based binary and BCD adder circuitsabstractReversible logic is gaining significance in the context of emerging technologies such as quantum computing since reversible circuits do not lose information during computation and there is one-to-one mapping between the inputs and outputs. In this work, we present a class of new designs for reversible binary and BCD adder circuits. The proposed designs are primarily optimized for the number of ancilla inputs and the number of garbage outputs and are designed for possible best values for the quantum cost and delay. In reversible circuits, in addition to the primary inputs, some constant input bits are used to realize different logic functions which are referred to as ancilla inputs and are overheads that need to be reduced. Further, the garbage outputs which do not contribute to any useful computations but are needed to maintain reversibility are also overheads that need to be reduced in reversible designs. First, we propose two new designs for the reversible ripple carry adder: (i) one with no input carry c 0 and no ancilla input bits, and (ii) one with input carry c 0 and no ancilla input bits. The proposed reversible ripple carry adder designs with no ancilla input bits have less quantum cost and logic depth (delay) compared to their existing counterparts in the literature. In these designs, the quantum cost and delay are reduced by deriving designs based on the reversible Peres gate and the TR gate. Next, four new designs for the reversible BCD adder are presented based on the following two approaches: (i) the addition is performed in binary mode and correction is applied to convert to BCD when required through detection and correction, and (ii) the addition is performed in binary mode and the result is always converted using a binary to BCD converter. The proposed reversible binary and BCD adders can be applied in a wide variety of digital signal processing applications and constitute important design components of reversible computing. Himanshu Thapliyal, N. Ranganathan |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2013 | Design of Testable Reversible Sequential CircuitsabstractIn this paper, we propose the design of two vectors testable sequential circuits based on conservative logic gates. The proposed sequential circuits based on conservative logic gates outperform the sequential circuits implemented in classical gates in terms of testability. Any sequential circuit based on conservative logic gates can be tested for classical unidirectional stuck-at faults using only two test vectors. The two test vectors are all 1's, and all 0's. The designs of two vectors testable latches, master-slave flip-flops and double edge triggered (DET) flip-flops are presented. The importance of the proposed work lies in the fact that it provides the design of reversible sequential circuits completely testable for any stuck-at fault by only two test vectors, thereby eliminating the need for any type of scan-path access to internal memory cells. The reversible design of the DET flip-flop is proposed for the first time in the literature. We also showed the application of the proposed approach toward 100% fault coverage for single missing/additional cell defect in the quantum-dot cellular automata (QCA) layout of the Fredkin gate. We are also presenting a new conservative logic gate called multiplexer conservative QCA gate (MX-cqca) that is not reversible in nature but has similar properties as the Fredkin gate of working as 2:1 multiplexer. The proposed MX-cqca gate surpasses the Fredkin gate in terms of complexity (the number of majority voters), speed, and area. Himanshu Thapliyal, N. Ranganathan, Saurabh Kotiyal |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2012 | Mach-Zehnder interferometer based design of all optical reversible binary adderabstractIn recent years reversible logic has emerged as a promising computing model for applications in dissipation less optical computing, low power CMOS, quantum computing, etc. In reversible circuits there exist a one-to-one mapping between the inputs and the outputs resulting in no loss of information. Researchers have implemented reversible logic gates in optical computing domain as it can provide high speed and low energy requirement along with easy fabrication at the chip level [1]. The all optical implementation of reversible gates are based on semiconductor optical amplifier (SOA) based Mach-Zehnder interferometer (MZI) due to its significant advantages such as high speed, low power, fast switching time and ease in fabrication. In this work we present the all optical implementation of an n bit reversible ripple carry adder for the first time in literature. The all optical reversible adder design is based on two new optical reversible gates referred as optical reversible gate I (ORG-I) and optical reversible gate II (ORG-II) and the existing all optical Feynman gate. The two new reversible gates ORG-I and ORGI-I are proposed as they can implement a reversible adder with reduced optical cost which is the measure of number of MZIs switches and the propagation delay, and with zero overhead in terms of number of ancilla inputs and the garbage outputs. The proposed all optical reversible adder design based on the ORG-I and ORG-II reversible gates are compared and shown to be better than the other existing designs of reversible adder proposed in non-optical domain in terms of number of MZIs, delay, number of ancilla inputs and the garbage outputs. The proposed all optical reversible ripple carry adder will be a key component of an all optical reversible ALU that can be applied in a wide variety of optical signal processing applications. Saurabh Kotiyal, Himanshu Thapliyal, N. Ranganathan |
DATE | 2 |
| 2011 | A new reversible design of BCD adderabstractReversible logic is one of the emerging technologies having promising applications in quantum computing. In this work, we present new design of the reversible BCD adder that has been primarily optimized for the number of ancilla input bits and the number of garbage outputs. The number of ancilla input bits and the garbage outputs is primarily considered as an optimization criteria as it is extremely difficult to realize a quantum computer with many qubits. As the optimization of ancilla input bits and the garbage outputs may degrade the design in terms of the quantum cost and the delay, thus the quantum cost and the delay parameters are also considered for optimization with primary focus towards the optimization of the number of ancilla input bits and the garbage outputs. Firstly, we propose a new design of the reversible ripple carry adder having the input carry Co and is designed with no ancilla input bits. The proposed reversible ripple carry adder design with no ancilla input bits has less quantum cost and the logic depth (delay) compared to its existing counterparts. The existing reversible Peres gate and a new reversible gate called the TR gate is efficiently utilized to improve the quantum cost and the delay of the reversible ripple carry adder. The improved quantum design of the TR gate is also illustrated. Finally, the reversible design of the BCD adder is presented which is based on a 4 bit reversible binary adder to add the BCD number, and finally the conversion of the binary result to the BCD format using a reversible binary to BCD converter. Himanshu Thapliyal, N. Ranganathan |
DATE | 1 |
| 2010 | Design of reversible sequential circuits optimizing quantum cost, delay, and garbage outputsabstractReversible logic has shown potential to have extensive applications in emerging technologies such as quantum computing, optical computing, quantum dot cellular automata as well as ultra low power VLSI circuits. Recently, several researchers have focused their efforts on the design and synthesis of efficient reversible logic circuits. In these works, the primary design focus has been on optimizing the number of reversible gates and the garbage outputs. The number of reversible gates is not a good metric of optimization as each reversible gate is of different type and computational complexity, and thus will have a different quantum cost and delay. The computational complexity of a reversible gate can be represented by its quantum cost. Further, delay constitutes an important metric, which has not been addressed in prior works on reversible sequential circuits as a design metric to be optimized. In this work, we present novel designs of reversible sequential circuits that are optimized in terms of quantum cost, delay and the garbage outputs. The optimized designs of several reversible sequential circuits are presented including the D Latch, the JK latch, the T latch and the SR latch, and their corresponding reversible master-slave flip-flop designs. The proposed master-slave flip-flop designs have the special property that they don't require the inversion of the clock for use in the slave latch. Further, we introduce a novel strategy of cascading a Fredkin gate at the outputs of a reversible latch to realize the designs of the Fredkin gate based asynchronous set/reset D latch and the master-slave D flip-flop. Finally, as an example of complex reversible sequential circuits, the reversible logic design of the universal shift register is introduced. The proposed reversible sequential designs were verified through simulations using Verilog HDL and simulation results are presented. Himanshu Thapliyal, N. Ranganathan |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2009 | Concurrently Testable FPGA Design for Molecular QCA using Conservative Reversible Logic GateabstractReversible logic is attracting the researchers attention for fault susceptible nanotechnologies including molecular QCA. In this paper, we propose concurrently testable FPGA design for molecular QCA using conservative reversible Fredkin gate. Fredkin gate is conservative reversible in nature, in which there would be an equal number of 1s in the outputs as there would be on the inputs, in addition to one-to-one mapping. Fault patterns in Fredkin gate are analyzed using HDLQ tool due to a single missing/additional cell defect in molecular QCA. Exhaustive simulation shows that if there is a fault in molecular QCA implementation of Fredkin gate, there is a parity mismatch between the inputs and the outputs; otherwise the inputs parity is same as outputs parity. Thus, any permanent and transient fault in molecular QCA that results in parity mismatch can be concurrently detected. The logic block and the routing fabric (both are programmable) are the two key components of an FPGA. Thus, we have shown the Fredkin gate based concurrently testable designs of the configurable logic block (CLB) and the routing switch of a molecular QCA-based FPGA. Analysis of power dissipation in the proposed FPGA is also shown. Himanshu Thapliyal, N. Ranganathan |
ISCAS | 1 |
| 2007 | Design of Reversible Sequential Elements With Feasibility of Transistor ImplementationabstractThis paper presents the novel designs of reversible sequential circuits (latches and flip flops). The proposed reversible latches and flip flops are designed from reversible Fredkin, Feynman and Toffoli gates. Two new reversible gates called modified Fredkin gate (MFG) and modified Toffoli gate (MTG) are also proposed to design the optimized implementations. The proposed designs are better than the recently proposed ones in terms of number of reversible gates and garbage outputs. In order to reach towards the goal of transistor implementations of proposed reversible sequential circuits, transistor implementation of the existing Feynman gate, Fredkin gate, Toffoli gates as well as the proposed MTG and MFG are also proposed. The proposed transistor implementations are completely reversible in nature, i.e., suitable for both the forward and backward computation. Himanshu Thapliyal, A. Prasad Vinod 0001 |
ISCAS | 1 |
| 2007 | Designing Efficient Online Testable Reversible Adders With New Reversible GateabstractReversible logic is emerging as a promising computing paradigm having its applications in low power VLSI design, quantum computing, nanotechnology and optical computing. In this paper, a new 4 times 4 reversible gate termed `OTG' (online testable gate) is proposed suitable for online testability in reversible logic circuits. OTG can also work singly as a reversible full adder with a bare minimum of two garbage outputs. OTG is shown better than the recently proposed R1 gate (introduced for providing online testability in reversible logic circuits), in terms of computation complexity. The proposed reversible gate is combined with the existing 4 times 4 Feynman gate to design online testable reversible adders such as ripple carry adder, carry skip adder and BCD adder. The efficient reversible design of two pair rail checker is also shown in this paper. The testable reversible circuits proposed in this work are shown to be better than the recently proposed testable designs in terms of number of reversible gates, garbage outputs and unit delay Himanshu Thapliyal, A. Prasad Vinod 0001 |
ISCAS | 1 |
| 2006 | Low Power Hierarchical Multiplier and Carry Look-Ahead ArchitectureabstractThis paper proposes a novel 8x8 multiplier architecture based on Wallace Tree, efficient in terms of power and regularity without significant increase in delay and area. The idea involves the generation of partial products in parallel using AND gates. The addition of these partial products is done using Wallace Tree which is hierarchically divided into levels. There will be a significant reduction in the power consumption, since power is provided only to the level that is involved in computation and thereby rendering the remaining two levels switched off (by employing a control circuitry). Furthermore, to improve the speed of addition at the 3rd level of computation, a novel carry look-ahead adder (CLA) is also proposed which is better than the recently proposed CLA architecture when compared its efficiency in terms of area/speed. The efficiency of the proposed multiplier is also tested by embedding it in higher width partition multipliers. Himanshu Thapliyal, Gopi Neela, K. K. Pavan Kumar, M. B. Srinivas |
AICCSA | 1 |
| 2006 | Novel Reversible Multiplier Architecture Using Reversible TSG GateabstractIn the recent years, reversible logic has emerged as a promising technology having its applications in low power CMOS, quantum computing, nanotechnology, and optical computing. The classical set of gates such as AND, OR, and EXOR are not reversible. Recently a 4 * 4 reversible gate called “TSG” is proposed. The most significant aspect of the proposed gate is that it can work singly as a reversible full adder, that is reversible full adder can now be implemented with a single gate only. This paper proposes a NXN reversible multiplier using TSG gate. It is based on two concepts. The partial products can be generated in parallel with a delay of d using Fredkin gates and thereafter the addition can be reduced to log2N steps by using reversible parallel adder designed from TSG gates. A 4x4 architecture of the proposed reversible multiplier is also designed. It is demonstrated that the proposed multiplier architecture using the TSG gate is much better and optimized, compared to its existing counterparts in literature; in terms of number of reversible gates and garbage outputs. Thus, this paper provides the initial threshold to building of more complex system which can execute more complicated operations using reversible logic. Himanshu Thapliyal, M. B. Srinivas |
AICCSA | 1 |