EDBT 2026 Demo / reviewers in the wild / expert
Swaroop Ghosh
dblp:91/2072
· DBLP profile ↗
119ranked-venue papers
21as first author
32since 2021 · last 2026
0000-0001-8753-490XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 109 · 18 first-author · 31 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 4 first-author · 1 since 2021Software engineering, systems software and programming languages · 16 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 3Security and privacy · 1Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Survival of the Optimized: An Evolutionary Approach to T-depth ReductionabstractQuantum Error Correction (QEC) is the corner-stone of practical Fault-Tolerant Quantum Computing (FTQC), but incurs enormous resource overheads. Circuits must decompose into Clifford +T gates, and the non-transversal T gates demand costly magic-state distillation. As circuit complexity grows, sequential T-gate layers (” T-depth”) increase, amplifying the spatiotemporal overhead of QEC. Optimizing T-depth is NP-hard, and existing greedy or brute-force strategies are either inefficient or computationally prohibitive. We frame T-depth reduction as a search optimization problem and present a Genetic Algorithm (GA) framework that approximates optimal layer-merge patterns across the non-convex search space. We introduce a mathematical formulation of the circuit expansion for systematic layer reordering and a greedy initial merge-pair selection, accelerating the convergence and enhancing the solution quality. In our benchmark with $\mathbf{\sim 9 0 - 1 0 0}$ qubits, our method reduces T-depth by 79.23% and overall T-count by 41.86%. Compared to the standard reversible circuit benchmarks, we achieve a $\sim 2.58 \times$ average improvement in T-depth over the state-of-the-art methods, demonstrating its viability for near-term FTQC. Archisman Ghosh 0001, Avimita Chatterjee, Swaroop Ghosh |
ASP-DAC | 3 |
| 2026 | Forensics of Quantum Processors using Program Queue Depth Analysis
Rupshali Roy, Swaroop Ghosh |
VTS | 2 |
| 2025 | Impact of Error Rate Misreporting on Resource Allocation in Multi-tenant Quantum Computing and Defense
Swaroop Ghosh |
ACM Great Lakes Symposium on VLSI | 2 |
| 2025 | Inverse-Transpilation: Reverse-Engineering Quantum Compiler Optimization Passes from Circuit Snapshots
Satwik Kundu, Swaroop Ghosh |
ACM Great Lakes Symposium on VLSI | 2 |
| 2025 | Adversarial Data Poisoning Attack on Quantum Machine Learning in the NISQ Era
Satwik Kundu, Swaroop Ghosh |
ACM Great Lakes Symposium on VLSI | 2 |
| 2025 | Forensics of Transpiled Quantum Circuits
Rupshali Roy, Archisman Ghosh 0001, Swaroop Ghosh |
ACM Great Lakes Symposium on VLSI | 3 |
| 2025 | Invited Paper: Toward Secure In-Sensor Intelligence: Threats and Defenses in SNNsabstractSpiking Neural Networks (SNNs) are inspired by the event-driven and temporally sparse nature of biological neurons, enabling deployment in in-sensor computing systems. The sensing and computation being tightly coupled in the in-sensor devices help in low-latency and energy-efficient data processing. This paradigm introduces a novel security front, exposing vulnerabilities in both neuromorphic hardware and the temporally sparse spike encodings to a variety of emerging attack modalities. This survey offers a comprehensive examination of the security and robustness landscape for SNNs deployed in in-sensor computing environments. It begins by outlining the architectural and algorithmic characteristics that define in-sensor SNN pipelines, with particular focus on temporal coding, asynchronous processing, and hardware constraints. We then review pertinent threat models, including spike-level adversarial perturbations, sensor spoofing, electromagnetic interference, fault injection, and timing-based privacy leakage, considering both white-box and black-box attack scenarios that exploit spatiotemporal vulnerabilities. Existing defense mechanisms, spanning noise shaping, homeostatic control, adversarial training, secure spike encoding, and hardware-level protections, are systematically categorized and assessed in the context of resource-constrained, event-driven platforms. Finally, we highlight emerging research directions in secure neuromorphic learning, such as continual and federated SNN training under adversarial settings, establishing a foundation for advancing the research in secure neuromorphic systems. Archisman Ghosh 0001, Swaroop Ghosh |
ICCAD | 2 |
| 2025 | Forensics of Error Rates of Quantum HardwareabstractThe qubit technologies, basis gate set, noise behavior, speed and coupling architecture are among the various factors that vary among various backends. Although third-party cloud providers offering quantum hardware as a service offer lower cost and flexibility to the users to choose from several qubit technologies, quantum hardware, and coupling maps; the actual execution of the program is not clearly visible to the customer. The success of the user program, in addition to various other metadata such as cost, performance, & number of iterations to converge, depends on the error rate of the backend used. Moreover, the third-party provider and/or tools (e.g., hardware allocator and mapper) may hold insider/outsider adversarial agents to conserve resources and maximize profit by running the quantum circuits on error-prone hardware. Thus it is important to gain visibility of the backend from various perspectives of the computing process e.g., execution, transpilation and outcomes. In this paper, we estimate the error rate of the backend from the original and transpiled circuit. Although many quantum services providers publish the error rates of their backends, we assume that such information may not be accurate and/or correspond to the actual hardware allocated to the program. For the forensics we propose two complementary approaches. First, we exploit the fact that qubit mapping and routing steps of the transpilation process select qubits and qubit pairs with low gate errors to minimize overall error accumulation. We leverage this to rank qubit links into bins and compare with publicly available data we are able to assign a bin rank within a difference of 2 with respect to the actual bin for upto$\mathbf{8 3. 5 \%}$of the qubit links in IBM Sherbrooke and$\mathbf{8 0 \%}$in IBM Brisbane, 127 qubit IBM backends. Second, we derive the error rates of the backends from a pool of programs by solving fidelity equations using numerical nonlinear optimizer. We achieve upto$92.7 \%(97.3 \%)$accuracy for single qubit (2 qubit) gate error rates. Rupshali Roy, Swaroop Ghosh |
ICCD | 2 |
| 2025 | A Primer on Security of Quantum Computing HardwareabstractQuantum computing (QC) is an emerging paradigm with the potential to transform numerous application domains by addressing classically intractable problems. However, its growing presence in cyberspace has introduced new security and privacy challenges. Similar to classical computing systems, the QC stack including software and hardware relies extensively on third parties, many of which are emerging and trust-seeking or less-trusted. This stack often contains sensitive intellectual property (IP) that demands protection. Unique features of quantum systems can enable classical-style attacks: for instance, crosstalk in multitenant settings can facilitate fault-injection attacks, while malicious calibration services can misreport error rates or miscalibrate qubits to induce denial-of-service (DoS) conditions. Given the high cost and limited availability of likely trustworthy quantum hardware, users may be enticed to explore emerging and trust-seeking but cheaper and readily available quantum hardware, which can enable the stealth of IP and tampering of quantum programs and/or computation outcomes. Similarly, emerging compilation services may compromise circuit confidentiality or insert Trojans. Despite the strategic significance of QC and its potential to process sensitive information, its security and privacy concerns remain underexplored. This article presents a comprehensive overview of QC fundamentals, key vulnerabilities, recent attack vectors, and corresponding defenses, and concludes with directions for future research to strengthen the quantum security community. Swaroop Ghosh, Suryansh Upadhyay, Abdullah Ash-Saki |
Proc. IEEE | 1 |
| 2024 | AltGraph: Redesigning Quantum Circuits Using Generative Graph Models for Efficient OptimizationabstractQuantum circuit transformation aims to optimize circuits for depth, gate count, and compatibility with Noisy Intermediate Scale Quantum (NISQ) devices that suffer from various error sources. Prior methods use combinations of expert-defined rules and Reinforcement Learning (RL). We introduce AltGraph, a novel approach employing generative graph models to generate functionally equivalent quantum circuits using—specifically, Direct Acyclic Graph (DAG) Variational Autoencoder (D-VAE) variants (GRU and GCN) and Deep Generative Model for Graphs (DeepGMG). AltGraph perturbs the latent space to generate quantum circuits optimized for hardware coupling maps, reducing gate count by 37.55% and circuit depth by 37.75% post-transpiling, with 0.0074 Mean Squared Error (MSE) in the density matrix—outperforming state-of-the-art methods by 2.56%. Collin Beaudoin, Koustubh Phalak, Swaroop Ghosh |
ACM Great Lakes Symposium on VLSI | 3 |
| 2024 | Q-Embroidery: A Study on Weaving Quantum Error Correction into the Fabric of Quantum ClassifiersabstractQuantum computing holds transformative potential for various fields, yet its practical application is hindered by the susceptibility to errors. This study makes a pioneering contribution by applying quantum error correction codes (QECCs) for complex, multi-qubit classification tasks. We implement 1-qubit and 2-qubit quantum classifiers with QECCs, specifically the Steane code, and the distance 3 & 5 surface codes to analyze 2-dimensional and 4-dimensional datasets. This research uniquely evaluates the performance of these QECCs in enhancing the robustness and accuracy of quantum classifiers against various physical errors, including bit-flip, phase-flip, and depolarizing errors. The results emphasize that the effectiveness of a QECC in practical scenarios depends on various factors, including qubit availability, desired accuracy, and the specific types and levels of physical errors, rather than solely on theoretical superiority. Avimita Chatterjee, Debarshi Kundu, Swaroop Ghosh |
ACM Great Lakes Symposium on VLSI | 3 |
| 2024 | Application of Quantum Tensor Networks for Protein ClassificationabstractComputational methods in drug discovery significantly reduce both time and experimental costs. Nonetheless, certain computational tasks in drug discovery can be daunting with classical computing techniques which can be potentially overcome using quantum computing. A crucial task within this domain involves the functional classification of proteins. However, a challenge lies in adequately representing lengthy protein sequences given the limited number of qubits available in existing noisy quantum computers. We show that protein sequences can be thought of as sentences in natural language processing and can be parsed using the existing Quantum Natural Language framework into parameterized quantum circuits of reasonable qubits, which can be trained to solve various protein-related machine-learning problems. We classify proteins based on their sub-cellular locations—a pivotal task in bioinformatics that is key to understanding biological processes and disease mechanisms. Leveraging the quantum-enhanced processing capabilities, we demonstrate that Quantum Tensor Networks (QTN) can effectively handle the complexity and diversity of protein sequences. We present a detailed methodology that adapts QTN architectures to the nuanced requirements of protein data, supported by comprehensive experimental results. We demonstrate two distinct QTNs, inspired by classical recurrent neural networks (RNN) and convolutional neural networks (CNN), to solve the binary classification task mentioned above. Our top-performing quantum model has achieved a 94% accuracy rate, which is comparable to the performance of a classical model that uses the ESM2 protein language model embeddings. It’s noteworthy that the ESM2 model is extremely large, containing 8 million parameters in its smallest configuration, whereas our best quantum model requires only around 800 parameters. We demonstrate that these hybrid models exhibit promising performance, showcasing their potential to compete with classical models of similar complexity. Debarshi Kundu, Archisman Ghosh 0001, Srinivasan Ekambaram, Jian Wang 0094, Nikolay V. Dokholyan, Swaroop Ghosh |
ACM Great Lakes Symposium on VLSI | 6 |
| 2024 | Evaluating Efficacy of Model Stealing Attacks and Defenses on Quantum Neural NetworksabstractCloud hosting of quantum machine learning (QML) models exposes them to a range of vulnerabilities, the most significant of which is the model stealing attack. In this study, we assess the efficacy of such attacks in the realm of quantum computing. Our findings revealed that model stealing attacks can produce clone models achieving up to 0.9 × and 0.99 × clone test accuracy when trained using Top-1 and Top-k labels, respectively (k: num_classes). To defend against these attacks, we propose: 1) hardware variation-induced perturbation (HVIP) and 2) hardware and architecture variation-induced perturbation (HAVIP). Despite limited success with our defense techniques, it has led to an important discovery: QML models trained on noisy hardwares are naturally resistant to perturbation or obfuscation-based defenses or attacks. Satwik Kundu, Debarshi Kundu, Swaroop Ghosh |
ACM Great Lakes Symposium on VLSI | 3 |
| 2023 | Knowledge Distillation in Quantum Neural Network Using Approximate SynthesisabstractRecent assertions of a potential advantage of Quantum Neural Network (QNN) for specific Machine Learning (ML) tasks have sparked the curiosity of a sizable number of application researchers. The parameterized quantum circuit (PQC), a major building block of a QNN, consists of several layers of single-qubit rotations and multi-qubit entanglement operations. The optimum number of PQC layers for a particular ML task is generally unknown. A larger network often provides better performance in noiseless simulations. However, it may perform poorly on hardware compared to a shallower network. Because the amount of noise varies amongst quantum devices, the optimal depth of PQC can vary significantly. Additionally, the gates chosen for the PQC may be suitable for one type of hardware but not for another due to compilation overhead. This makes it difficult to generalize a QNN design to wide range of hardware and noise levels. An alternate approach is to build and train multiple QNN models targeted for each hardware which can be expensive. To circumvent these issues, we introduce the concept of knowledge distillation in QNN using approximate synthesis. The proposed approach will create a new QNN network with (i) a reduced number of layers or (ii) a different gate set without having to train it from scratch. Training the new network for a few epochs can compensate for the loss caused by approximation error. Through empirical analysis, we demonstrate ≈71.4% reduction in circuit layers, and still achieve ≈16.2% better accuracy under noise. Mahabubul Alam, Satwik Kundu, Swaroop Ghosh |
ASP-DAC | 3 |
| 2023 | Large-Scale Quantum Approximate Optimization via Divide-and-ConquerabstractQuantum approximate optimization algorithm (QAOA) is a promising hybrid quantum-classical algorithm for solving combinatorial optimization problems. However, it cannot overcome qubit limitation for large-scale problems. Furthermore, the simulation time of QAOA scales poorly with the problem size. We propose a divide-and-conquer QAOA (DC-QAOA) to address the above challenges for graph maximum cut (MaxCut) problem. The algorithm works by recursively partitioning a larger graph into smaller ones whose MaxCut solutions are obtained with small-size noisy intermediate-scale quantum computers. The overall solution is retrieved from the subsolutions by applying the combination policy of measurement distribution reconstruction (MDR). The solution quality depends on the graph partitioning algorithm and MDR policy. Multiple partitioning and reconstruction methods are proposed and compared. Results are evaluated by metrics, such as quantum program runtime, measurement expectation value (EV), and approximation ratio (AR). The results show that DC-QAOA achieves 97.14% AR (20.32% higher than classical counterpart), and 94.79% EV (15.80% higher than quantum annealing). DC-QAOA solves large-scale graph instances with a polynomial rate or returns unsuccessful partition if graph connectivity requirement is not fulfilled otherwise. Junde Li, Mahabubul Alam, Swaroop Ghosh |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | Muzzle the Shuttle: Efficient Compilation for Multi-Trap Trapped-Ion Quantum ComputersabstractTrapped-ion systems can have a limited number of ions (qubits) in a single trap. Increasing the qubit count to run meaningful quantum algorithms would require multiple traps where ions need to shuttle between traps to communicate. The existing compiler has several limitations, which result in a high number of shuttle operations and degraded fidelity. In this paper, we target this gap and propose compiler optimizations to reduce the number of shuttles. Our technique achieves a maximum reduction of 51.17% in shuttles (average ~ 33%) tested over 125 circuits. Furthermore, the improved compilation enhances the program fidelity up to 22.68X with a modest increase in the compilation time. Abdullah Ash-Saki, Rasit Onur Topaloglu, Swaroop Ghosh |
DATE | 3 |
| 2022 | Scalable Variational Quantum Circuits for Autoencoder-based Drug DiscoveryabstractThe de novo design of drug molecules is recognized as a time-consuming and costly process, and computational approaches have been applied in each stage of the drug discovery pipeline. Variational autoencoder is one of the computer-aided design methods which explores the chemical space based on an existing molecular dataset. Quantum machine learning has emerged as an atypical learning method that may speed up some classical learning tasks because of its strong expressive power. However, near-term quantum computers suffer from limited num-ber of qubits which hinders the representation learning in high dimensional spaces. We present a scalable quantum generative autoencoder (SQ-VAE) for simultaneously reconstructing and sampling drug molecules, and a corresponding vanilla variant (SQ-AE) for better reconstruction. The architectural strategies in hybrid quantum classical networks such as, adjustable quantum layer depth, heterogeneous learning rates, and patched quantum circuits are proposed to learn high dimensional dataset such as, ligand-targeted drugs. Extensive experimental results are reported for different dimensions including 8x8 and 32x32 after choosing suitable architectural strategies. The performance of quantum generative autoencoder is compared with the corre-sponding classical counterpart throughout all experiments. The results show that quantum computing advantages can be achieved for normalized low-dimension molecules, and that high-dimension molecules generated from quantum generative autoencoders have better drug properties within the same learning period. Junde Li, Swaroop Ghosh |
DATE | 2 |
| 2022 | Analysis of Power-Oriented Fault Injection Attacks on Spiking Neural NetworksabstractSpiking Neural Networks (SNN) are quickly gaining traction as a viable alternative to Deep Neural Networks (DNN). In comparison to DNNs, SNNs are more computationally powerful and provide superior en-ergy efficiency. SNNs, while exciting at first appearance, contain security-sensitive assets (e.g., neuron threshold voltage) and vulnerabilities (e.g., sensitivity of classification accuracy to neuron threshold voltage change) that adversaries can exploit. We investigate global fault injection attacks by employing external power supplies and laser-induced local power glitches to corrupt crucial training parameters such as spike amplitude and neuron's membrane threshold potential on SNNs developed using common analog neurons. We also evaluate the impact of power-based attacks on individual SNN layers for 0% (i.e., no attack) to 100% (i.e., whole layer under attack). We investigate the impact of the attacks on digit classification tasks and find that in the worst-case scenario, classification accuracy is reduced by 85.65%. We also propose defenses e.g., a robust current driver design that is immune to power-oriented attacks, improved circuit sizing of neuron components to reduce/recover the adversarial accuracy degradation at the cost of negligible area and 25% power overhead. We also present a dummy neuron-based voltage fault injection detection system with ~ 1% power and area overhead. Karthikeyan Nagarajan, Junde Li, Sina Sayyah Ensan, Mohammad Nasim Imtiaz Khan, Sachhidh Kannan, Swaroop Ghosh |
DATE | 6 |
| 2022 | Session details: Session 5B: VLSI Design + VLSI Circuits and Power Aware Design 2abstractNo abstract available. Swaroop Ghosh |
ACM Great Lakes Symposium on VLSI | 1 |
| 2022 | Security Aspects of Quantum Machine Learning: Opportunities, Threats and DefensesabstractIn the last few years, quantum computing has experienced a growth spurt. One exciting avenue of quantum computing is quantum machine learning (QML) which can exploit the high dimensional Hilbert space to learn richer representations from limited data and thus can efficiently solve complex learning tasks. Despite the increased interest in QML, there have not been many studies that discuss the security aspects of QML. In this work, we explored the possible future applications of QML in the hardware security domain. We also expose the security vulnerabilities of QML and emerging attack models, and corresponding countermeasures. Satwik Kundu, Swaroop Ghosh |
ACM Great Lakes Symposium on VLSI | 2 |
| 2022 | A Shuttle-Efficient Qubit Mapper for Trapped-Ion Quantum ComputersabstractTrapped-ion (TI) quantum computer is one of the forerunner quantum technologies. Execution of a quantum gate in multiple trap TI system may frequently involve ions from two different traps, hence one of the ions needs to be shuttled (moved) between traps to be co-located, degrading fidelity, and increasing the program execution time. The choice of initial mapping influences the number of shuttles. The existing Greedy policy neglects the depth of the program at which a gate is present. Intuitively, the contribution of the late-stage gates to the initial mapping is less since the ions might have already shuttled to a different trap to satisfy other gate operations. In this paper, we target this gap and propose a new program adaptive policy especially for programs with considerable depth and high number of qubits (valid for practical-scale quantum programs). Our technique achieves an average reduction of 9% shuttles/program (with 21.3% at best) for 120 random circuits and enhances the program fidelity up to 3.3X (1.41X on average). Suryansh Upadhyay, Abdullah Ash-Saki, Rasit Onur Topaloglu, Swaroop Ghosh |
ACM Great Lakes Symposium on VLSI | 4 |
| 2022 | Quantum Machine Learning for Material Synthesis and Hardware Security (Invited Paper)abstractUsing quantum computing, this paper addresses two scientifically-pressing and day to day-relevant problems, namely, chemical retrosynthesis which is an important step in drug/material discovery and security of semiconductor supply chain. We show that Quantum Long Short-Term Memory (QLSTM) is a viable tool for retrosynthesis. We achieve 65% training accuracy with QLSTM whereas classical LSTM can achieve 100%. However, in testing we achieve 80% accuracy with the QLSTM while classical LSTM peaks at only 70% accuracy! We also demonstrate an application of Quantum Neural Network (QNN) in the hardware security domain, specifically in Hardware Trojan (HT) detection using a set of power and area Trojan features. The QNN model achieves detection accuracy as high as 97.27%. Collin Beaudoin, Satwik Kundu, Rasit Onur Topaloglu, Swaroop Ghosh |
ICCAD | 4 |
| 2022 | Optimization of Quantum Read-Only Memory CircuitsabstractQuantum computing is a rapidly expanding field with applications ranging from optimization all the way to complex machine learning tasks. Quantum memories, while lacking in practical quantum computers, have the potential to bring quantum advantage. In quantum machine learning applications for example, a quantum memory can simplify the data loading process and potentially accelerate the learning task. Quantum memory can also store intermediate quantum state of qubits that can be reused for computation. However, the depth, gate count and compilation time of quantum memories such as, Quantum Read Only Memory (QROM) scale exponentially with the number of address lines making them impractical in state-of-the-art Noisy Intermediate-Scale Quantum (NISQ) computers beyond 4-bit addresses. In this paper, we propose techniques such as, pre-decoding logic and qubit reset to reduce the depth and gate count of QROM circuits to target wider address ranges such as, 8-bits. The proposed approach reduces the number of gates and depth count by at least 2X compared to the naive implementation at only 36% qubit overhead. A reduction in circuit depth and gate count as high as 75X and compilation time by 85X at the cost of a maximum of 2.28X qubit overhead is observed. Experimentally, the fidelity with the proposed pre-decoding circuit compared to existing optimization approach is also higher (as much as 73% compared to 40.8%) under reduced error rates. Koustubh Phalak, Mahabubul Alam, Abdullah Ash-Saki, Rasit Onur Topaloglu, Swaroop Ghosh |
ICCD | 5 |
| 2022 | Special Session: On the Reliability of Conventional and Quantum Neural Network HardwareabstractNeural Networks (NNs) are being extensively used in critical applications such as aerospace, healthcare, autonomous driving, and military, to name a few. Limited precision of the underlying hardware platforms, permanent and transient faults injected unintentionally as well as maliciously, and voltage/temperature fluctuations can potentially result in malfunctions in NNs with consequences ranging from substantial reduction in the network accuracy to jeopardizing the correct prediction of the network in worst cases. To alleviate such reliability concerns, this paper discusses the state-of-the-art reliability enhancement schemes that can be tailored for deep learning accelerators. We will discuss the errors associated with the hardware implementation of Deep-Learning (DL) algorithms along with their corresponding countermeasures. An in-field self-test methodology with a high test coverage is introduced, and an accurate high-level framework, so-called FIdelity, is proposed that enables the designers to evaluate DL accelerators in presence of such errors. Then, a state-of-the-art robustness-preserving training algorithm based on the Hessian Regularization is introduced. This algorithm alleviates the perturbations during inference time with negligible degradation in the accuracy of the network. Finally, Quantum Neural Networks (QNNs) and the methods to make them resilient against a variety of vulnerabilities such as fault injection, spatial and temporal variations in Qubits, and noise in QNNs are discussed. Mehdi Sadi, Yi He 0010, Yanjing Li, Mahabubul Alam, Satwik Kundu, Swaroop Ghosh, Javad Bahrami, Naghmeh Karimi |
VTS | 6 |
| 2022 | Addressing Resiliency of In-Memory Floating Point ComputationabstractIn-memory computing (IMC) can eliminate data movement between processor and memory, which is a barrier to the energy efficiency and performance in von Neumann computing. Due to low power consumption, fast operation, and tiny footprint in crossbar architecture, resistive RAM (RRAM) is one of the most promising devices for IMC applications. We present FPCAS, a pipelined floating point (FP) arithmetic (addition/ subtraction) solver based on RRAM crossbars. Although promis- ing, RRAM-based computing may experience random failures, such as the stuck-at fault where RRAM cells are stuck at either a high-resistance state (HRS), i.e., stuck-at-0 (SA0), or a low-resistance state (LRS), i.e., stuck-at-1 (SA1). We propose techniques to prevent SA1 failures, namely, shifting-at-the-output (SATO), force to$V_{\mathrm{ DD}}$(FTV), and force to ground (FTG) since 96% of the RRAMs employed in our architecture are in HRS. Using an extra clock cycle, both strategies employ the memory array’s fault-free RRAMs to conduct the computation. When the failure rate is less than 2%, SATO can manage more than 70% of faults, whereas FTV can handle more than 90% of faults at low power and low area overhead. Simulation results reveal that, for$\mathrm{\scriptstyle NAND}$–$\mathrm{\scriptstyle NAND}$- and$\mathrm{\scriptstyle NOR}$–$\mathrm{\scriptstyle NOR}$-based implementations, FPCAS consumes 335 and 322 pJ, respectively. Both implementations incur a performance overhead of 50% at the array level and 4% for pipelined FP implementation. Sina Sayyah Ensan, Swaroop Ghosh, Seyedhamidreza Motaman, Derek Weast |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2021 | Invited: Drug Discovery Approaches using Quantum Machine LearningabstractTraditional drug discovery pipelines can require multiple years and billions of dollars of investment. Deep generative and discriminative models are widely adopted to assist in drug development. Classical machines cannot efficiently reproduce the atypical patterns of quantum computers, which may improve the quality of learned tasks. We propose a suite of quantum machine learning techniques: incorporating generative adversarial networks (GAN), convolutional neural networks (CNN) and variational auto-encoders (VAE) to generate small drug molecules, classify binding pockets in proteins, and generate large drug molecules, respectively. Junde Li, Mahabubul Alam, Congzhou M. Sha, Jian Wang 0094, Nikolay V. Dokholyan, Swaroop Ghosh |
DAC | 6 |
| 2021 | A Survey and Tutorial on Security and Resilience of Quantum ComputingabstractPresent-day quantum computers suffer from various noises or errors such as, gate error, relaxation, dephasing, readout error, and crosstalk. Besides, they offer a limited number of qubits with restrictive connectivity. Therefore, quantum programs running these computers face resilience issues and low output fidelities. The noise in the cloud-based access of quantum computers also introduce new modes of security and privacy issues. Furthermore, quantum computers face several threat models from insider and outsider adversaries including input tampering, program misallocation, fault injection, Reverse Engineering (RE) and Cloning. This paper provides an overview of various assets embedded in quantum computers and programs, vulnerabilities and attack models and the relation between resilience and security. We also cover countermeasures against the reliability and security issues and present future outlook for security of quantum computing. Abdullah Ash-Saki, Mahabubul Alam, Koustubh Phalak, Aakarshitha Suresh, Rasit Onur Topaloglu, Swaroop Ghosh |
ETS | 6 |
| 2021 | Quantum-Classical Hybrid Machine Learning for Image Classification (ICCAD Special Session Paper)abstractImage classification is a major application domain for conventional deep learning (DL). Quantum machine learning (QML) has the potential to revolutionize image classification. In any typical DL-based image classification, we use convolutional neural network (CNN) to extract features from the image and multi-layer perceptron network (MLP) to create the actual decision boundaries. QML models can be useful in both of these tasks. On one hand, convolution with parameterized quantum circuits (Quanvolution) can extract rich features from the images. On the other hand, quantum neural network (QNN) models can create complex decision boundaries. Therefore, Quanvolution and QNN can be used to create an end-to-end QML model for image classification. Alternatively, we can extract image features separately using classical dimension reduction techniques such as, Principal Components Analysis (PCA) or Convolutional Autoen-coder (CAE) and use the extracted features to train a QNN. We review two proposals on quantum-classical hybrid ML models for image classification namely, Quanvolutional Neural Network and dimension reduction using a classical algorithm followed by QNN. Particularly, we make a case for trainable filters in Quanvolution and CAE-based feature extraction for image datasets (instead of dimension reduction using linear transformations such as, PCA). We discuss various design choices, potential opportunities, and drawbacks of these models. We also release a Python-based framework to create and explore these hybrid models with a variety of design choices. Mahabubul Alam, Satwik Kundu, Rasit Onur Topaloglu, Swaroop Ghosh |
ICCAD | 4 |
| 2021 | Split Compilation for Security of Quantum CircuitsabstractAn efficient quantum circuit (program) compiler aims to minimize the gate-count - through efficient instruction translation, routing, gate, and cancellation - to improve run-time and noise. Therefore, a high-efficiency compiler is paramount to enable the game-changing promises of quantum computers. To date, the quantum computing hardware providers are offering a software stack supporting their hardware. However, several third-party software toolchains, including compilers, are emerging. They support hardware from different vendors and potentially offer better efficiency. As the quantum computing ecosystem becomes more popular and practical, it is only prudent to assume that more companies will start offering software-as-a-service for quantum computers, including high-performance compilers. With the emergence of third-party compilers, the security and privacy issues of quantum intellectual properties (IPs) will follow. A quantum circuit can include sensitive information such as critical financial analysis and proprietary algorithms. Therefore, submitting quantum circuits to untrusted compilers creates opportunities for adversaries to steal IPs. In this paper, we present a split compilation methodology to secure IPs from untrusted compilers while taking advantage of their optimizations. In this methodology, a quantum circuit is split into multiple parts that are sent to a single compiler at different times or to multiple compilers. In this way, the adversary has access to partial information. With analysis of over 152 circuits on three IBM hardware architectures, we demonstrate the split compilation methodology can completely secure IPs (when multiple compilers are used) or can introduce factorial time reconstruction complexity while incurring a modest overhead (~ 3% to ~ 6% on average). Abdullah Ash-Saki, Aakarshitha Suresh, Rasit Onur Topaloglu, Swaroop Ghosh |
ICCAD | 4 |
| 2021 | ReLOPE: Resistive RAM-Based Linear First-Order Partial Differential Equation SolverabstractData movement between memory and processing units poses an energy barrier to Von-Neumann-based architectures. In-memory computing (IMC) eliminates this barrier. RRAM-based IMC has been explored for data-intensive applications, such as artilicial neural networks and matrix-vector multiplications that are considered as “soft” tasks where performance is a more important factor than accuracy. In “hard” tasks such as partial differential equations (PDEs), accuracy is a determining factor. In this brief, we propose ReLOPE, a fully RRAM crossbar-based IMC to solve PDEs using the Runge-Kutta numerical method with 97% accuracy. ReLOPE expands the operating range of solution by exploiting shifters to shift input data and output data. ReLOPE range of operation and accuracy can be expanded by using line-grained step sizes by programming other RRAMs on the BL. Compared to software-based PDE solvers, ReLOPE gains 31.4× energy reduction at only 3% accuracy loss. Sina Sayyah Ensan, Swaroop Ghosh |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2021 | SCARE: Side Channel Attack on In-Memory Computing for Reverse EngineeringabstractIn-memory computing (IMC) architectures provide a much needed solution to energy-efficiency barriers posed by Von-Neumann computing. The functions implemented in such in-memory architectures are often proprietary and constitute confidential intellectual property (IP). Our studies indicate that IMC architectures implemented using resistive RAM (RRAM) are susceptible to side channel attack (SCA). Unlike the conventional SCAs that are aimed to leak private keys from cryptographic implementations, SCA on IMC for reverse engineering (SCARE) can reveal the sensitive IP implemented within the memory through power/timing side channels. Therefore, the adversary does not need to perform invasive reverse engineering (RE) to unlock the functionality. We demonstrate SCARE by taking recent IMC architectures, such as dynamic computing in memory (DCIM) and memristor-aided logic (MAGIC) as test cases. Simulation results indicate that AND, OR, and NOR gates (which are the building blocks of complex functions) yield distinct power and timing signatures based on the number of inputs, making them vulnerable to SCA. We show that adversary can use templates (using foundry-calibrated simulations or fabricating known functions in test chips) and analysis to identify the structure of the implemented function by testing a limited number of patterns. We also propose countermeasures, such as redundant inputs and expansion of literals. Redundant inputs can mask the IP with 25% area and 20% power overhead. However, functions can be found at higher RE effort. Expansion of literals incurs 36% power overhead. However, it imposes a brute force search increasing the adversarial RE effort by$3.04\times $. Sina Sayyah Ensan, Karthikeyan Nagarajan, Mohammad Nasim Imtiaz Khan, Swaroop Ghosh |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2021 | SecNVM: Power Side-Channel Elimination Using On-Chip Capacitors for Highly Secure Emerging NVMabstractEmerging nonvolatile memories (NVMs), such as resistive RAM (RRAM) and spin-transfer-torque RAM (STTRAM), present exciting opportunities for data storage applications and offer improved access speeds, retention times, power consumption, and scalability. However, these technologies leak the Hamming weight of data through power side-channel during read and write operations. We propose a technique leveraging on-chip capacitor and voltage regulator (VR) that powers the NVM read/write operations. The side-channel leakage is eliminated due to the isolation of memory array from the external power supply during read/write operations. The residual charge on capacitor bank is discarded safely to prevent information leakage during capacitor recharging. The VR ensures a steady voltage during the entire read/write operations even though the capacitor discharges. The design presents a performance (instructions per cycle) degradation of 0.53%-1.2% under parsec and splash-2 benchmarks and incurs an area overhead of ~ 3.54×10-5% and an energy overhead of ~ 3.05 ×10-5% for a 4-Mb RRAM memory array. For a 64-bit word, the design improves security by 2.7 × 1019× to 264×. SecNVM should be used in small security-critical memory macros to limit the overhead. SecNVM is generic and could protect any security module such as encryption engines, against power side-channel attacks. Karthikeyan Nagarajan, Farid Uddin Ahmed, Mohammad Nasim Imtiaz Khan, Asmit De, Masud H. Chowdhury, Swaroop Ghosh |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2020 | An Efficient Circuit Compilation Flow for Quantum Approximate Optimization AlgorithmabstractQuantum approximate optimization algorithm (QAOA) is a promising quantum-classical hybrid algorithm to solve hard combinatorial optimization problems. The two-qubits gates used in quantum circuit for QAOA are commutative i.e., the order of gates can be altered without changing the logical output. This re-ordering leads to execution of more gates in parallel and a smaller number of additional gates to compile the QAOA circuit resulting in lower circuit depth and gate-count which is beneficial for circuit run-time and noise. A lower number of gates means a lower accumulation of gate errors, and a lower circuit depth means the quantum bits will have a lower time to decohere (lose state). However, finding the best re-ordered circuit is a difficult problem and does not scale well with circuit size. This paper presents a compilation flow with 3 approaches to find an optimal re-ordered circuit with reduced depth and gate count. Our approaches can reduce gate count up to 23.21% and circuit depth up to 53.65%. Our approaches are compiler agnostic, can be integrated with existing compilers, and scalable. Mahabubul Alam, Abdullah Ash-Saki, Swaroop Ghosh |
DAC | 3 |
| 2020 | Accelerating Quantum Approximate Optimization Algorithm using Machine LearningabstractWe propose a machine learning based approach to accelerate quantum approximate optimization algorithm (QAOA) implementation which is a promising quantum-classical hybrid algorithm to prove the so-called quantum supremacy. In QAOA, a parametric quantum circuit and a classical optimizer iterates in a closed loop to solve hard combinatorial optimization problems. The performance of QAOA improves with increasing number of stages (depth) in the quantum circuit. However, two new parameters are introduced with each added stage for the classical optimizer increasing the number of optimization loop iterations. We note a correlation among parameters of the lower-depth and the higher-depth QAOA implementations and, exploit it by developing a machine learning model to predict the gate parameters close to the optimal values. As a result, the optimization loop converges in a fewer number of iterations. We choose graph MaxCut problem as a prototype to solve using QAOA. We perform a feature extraction routine using 100 different QAOA instances and develop a training data-set with 13, 860 optimal parameters. We present our analysis for 4 flavors of regression models and 4 flavors of classical optimizers. Finally, we show that the proposed approach can curtail the number of optimization iterations by on average 44.9% (up to 65.7%) from an analysis performed with 264 flavors of graphs. Mahabubul Alam, Abdullah Ash-Saki, Swaroop Ghosh |
DATE | 3 |
| 2020 | Quantum-Soft QUBO Suppression for Accurate Object Detection
Junde Li, Swaroop Ghosh |
ECCV (29) | 2 |
| 2020 | Noise Resilient Compilation Policies for Quantum Approximate Optimization AlgorithmabstractQuantum approximate optimization algorithm (QAOA) is a promising quantum-classical hybrid algorithm to solve hard combinatorial optimization problems using noisy quantum devices. The multiqubit CPHASE gates used in the quantum circuit for QAOA are commutative i.e., the order of the gates can be altered without changing the output state. This re-ordering leads to the execution of more gates in parallel and a smaller number of additional SWAP gates to compile the QAOA circuit resulting in lower circuit-depth and gate-count. A less number of gates generally indicates a lower accumulation of gate-errors, and a reduced circuit-depth means less decoherence time for the qubits. However, near-term quantum devices exhibit significant variations in the gate success probabilities. Variation-aware compilation policies (i.e. putting most gate operations on qubits with higher gate success probabilities) can enhance the probability of successful program execution on the hardware. The greater flexibility of QAOA-circuits offer better scope of optimization with QAOA-tailored compilation policies. This paper presents an argument for compilation policies to exploit the unique characteristics of QAOA-circuits alongside the variation-awareness of the noisy devices. We present two procedures - variation-aware qubit placement (VQP) and variation-aware iterative mapping (VIM) that can improve the circuit success probability quite significantly (≈8.408X on average) for a set of QAOA-MaxCut problems on ibmq_16_melbourne. Mahabubul Alam, Abdullah Ash-Saki, Junde Li, Anupam Chattopadhyay, Swaroop Ghosh |
ICCAD | 5 |
| 2020 | Power Side Channel Attack Analysis and DetectionabstractSide Channel Attack (SCA) is a serious threat to the hardware implementation of cryptographic protocols. Various side channels such as, power, timing, electromagnetic emission and acoustic noise have been explored to extract the secret keys. Machine Learning (ML)-based detection of SCA have been proposed in past which incur high design overheads and, require digitization that reduce their accuracy under process variations. We propose a real-time power SCA detection technique using on-chip sensors based on a thorough analysis. The dependency of phase/frequency of Ring Oscillator (RO) on supply voltage is exploited to detect the insertion of a SCA resistance in the power rail. The proposed approach is validated using simulation with a detailed model of Power Delivery Network (PDN) and power grid. The technique can detect a minimum resistance of 1 Ω within 2 μs of attack initiation and incurs a tiny fraction of area/power (0.044%/0.1065%, respectively) compared to ML-based techniques. Navyata Gattu, Mohammad Nasim Imtiaz Khan, Asmit De, Swaroop Ghosh |
ICCAD | 4 |
| 2020 | FAuto: An Efficient GMM-HMM FPGA Implementation for Behavior Estimation in Autonomous SystemsabstractDriving behavior estimation in car-following scenario based on contextual traffic information is an essential capability for autonomous driving systems. Real-time motion planning based on incomplete environment perception requires complicated probabilistic model for interactions with surrounding objects and road conditions. Hidden Markov Model (HMM) with Gaussian emissions has been used to model driving behaviors for its ability of inferring unobserved states. While the high-dimensional contextual data is continuously processed, the system should be high-performance and power-efficient to make real-time decisions for safe operations. Field Programmable Gate Array (FPGA) is being increasingly used on embedded System-on-Chip (SoC) for mobile applications mainly because of its parallel computation and low-power consumption. This paper implements FAuto: the framework of HMM coupled with GMM algorithm on a Xilinx PYNQ-Z2 board for autonomous systems. We design the hybrid GMM-HMM model in python, and train the model using Next Generation SIMulation (NGSIM) trajectory data on a CPU platform. The hardware accelerator is designed through Vivado HLS 2018.2, and verified with Jupiter notebook. FAuto achieves 2.59 TOPS/W power efficiency, and 10.39× speedup compared to Python software implementation running on quad-core i7-7500U CPU. Junde Li, Navyata Gattu, Swaroop Ghosh |
IJCNN | 3 |
| 2020 | Analysis of crosstalk in NISQ devices and security implications in multi-programming regimeabstractThe noisy intermediate-scale quantum (NISQ) computers suffer from unwanted coupling across qubits referred to as crosstalk. Existing literature largely ignores the crosstalk effects which can introduce significant error in circuit optimization. In this work, we present a crosstalk modeling analysis framework for near-term quantum computers after extracting the error-rates experimentally. Our analysis reveals that crosstalk can be of the same order of gate error which is considered a dominant error in NISQ devices. We also propose adversarial fault injection using crosstalk in a multiprogramming environment where the victim and the adversary share the same quantum hardware. Our simulation and experimental results from IBM quantum computers demonstrated that the adversary can inject fault and launch a Denial-of-Service attack. Finally, we propose system- and device-level countermeasures. Abdullah Ash-Saki, Mahabubul Alam, Swaroop Ghosh |
ISLPED | 3 |
| 2020 | Resiliency analysis and improvement of variational quantum factoring in superconducting qubitabstractVariational algorithm using Quantum Approximate Optimization Algorithm (QAOA) can solve the prime factorization problem in near-term noisy quantum computers. Conventional Variational Quantum Factoring (VQF) requires a large number of 2-qubit gates (especially for factoring a large number) resulting in deep circuits. The output quality of the deep quantum circuit is degraded due to errors limiting the computational power of quantum computing. In this paper, we explore various transformations to optimize the QAOA circuit for integer factorization. We propose two criteria to select the optimal quantum circuit that can improve the noise resiliency of VQF. Mahabubul Alam, Abdullah Ash-Saki, Swaroop Ghosh |
ISLPED | 4 |
| 2020 | Assuring Security and Reliability of Emerging Non-Volatile MemoriesabstractAt the end of Silicon roadmap, keeping the leakage power in tolerable limit has become one of the biggest challenges. Several promising Non-Volatile Memories (NVMs) offering high-density, high speed, and competitive reliability/endurance while eliminating leakage issues are being investigated. On one hand, the above-desired properties make emerging NVM suitable candidates to assist or replace conventional memories in memory hierarchy as well as to infuse compute capability to eliminate Von-Neumann bottleneck. On the other hand, their unique features such as high and asymmetric read/write current and persistence bring new threats to data security while compute-capability imposes new fundamentally different security challenges. Some of these memories are already deployed in full systems and as discrete chips. Therefore, it is utmost important to investigate the security issues of NVMs spanning the application space. This work makes pioneering contributions to this challenge through a holistic approach- from devices to circuits and systems using a combination of design and test methodologies to develop secure and resilient NVMs. The proposed attacks and countermeasures are validated on test boards using commercial NVM chips. Finally, this research has been tied to education by converting the test boards to design a modular and reproducible self-learning cybersecurity kit which has been piloted to train graduate and undergraduate students and K-12 teachers. Mohammad Nasim Imtiaz Khan, Swaroop Ghosh |
ITC | 2 |
| 2020 | Circuit Compilation Methodologies for Quantum Approximate Optimization AlgorithmabstractThe quantum approximate optimization algorithm (QAOA) is a promising quantum-classical hybrid algorithm to solve hard combinatorial optimization problems. The multi-qubit CPHASE gates used in the quantum circuit for QAOA are commutative i.e., the order of the gates can be altered without changing the output state. This re-ordering leads to the execution of more gates in parallel and a smaller number of additional SWAP gates to compile the QAOA-circuit. Consequently, the circuit-depth and cumulative gate-count become lower which is beneficial for circuit execution time and noise resilience. A less number of gates indicates a lower accumulation of gate-errors, and a reduced circuit-depth means less decoherence time for the qubits. However, finding the best-ordered circuit is a difficult problem and does not scale well with circuit size. This paper presents four generic methodologies to optimize QAOA-circuits by exploiting gate re-ordering. We demonstrate a reduction in gate-count by ≈23.0% and circuit-depth by ≈53.0% on average over a conventional approach without incurring any compilation-time penalty. We also present a variation-aware compilation which enhances the compiled circuit success probability by ≈62.7% for the target hardware over the variation unaware approach. A new metric, Approximation Ratio Gap (ARG), is proposed to validate the quality of the compiled QAOA-circuit instances on actual devices. Hardware implementation of a number of QAOA instances shows ≈25.8% improvement in the proposed metric on average over the conventional approach on ibmq 16 melbourne. Mahabubul Alam, Abdullah Ash-Saki, Swaroop Ghosh |
MICRO | 3 |
| 2020 | Hardware Assisted Buffer Protection Mechanisms for Embedded RISC-VabstractRISC-V is a promising open-source architecture that targets low-power embedded devices and system-on-chips (SoCs). However, there is a dearth of practical and low-overhead security solutions in the RISC-V architecture. Programs compiled using RISC-V toolchains are still vulnerable to code injection and code reuse attacks, such as buffer overflow and return-oriented programming (ROP). In this article, we propose two hardware-implemented security extensions to RISC-V that provides a defense mechanism against such attacks. We first employ a physically unclonable function (PUF)-based randomized canary generation technique that removes the need to store the sensitive canary words in memory or CPU registers, thereby being more secure, while incurring low overheads. We implement the proposed Canary Engine in RISC-V RocketChip with rocket custom coprocessor (RoCC). The simulation results show 2.2% average execution overhead with a single buffer protection, while a 10× increase in buffer count only increases the overhead by 1.5× when protection is extended to all buffers. We further improve upon this with a dedicated security coprocessor flow integrity extensions for embedded RISC-V (FIXER), implemented on the RoCC. FIXER enforces fine-grained control-flow integrity (CFI) of running programs on backward edges (returns) and forward edges (calls) without requiring any architectural modifications to the processor core. Compared to software-based solutions, FIXER reduces energy overhead by 60% at minimal execution time (1.5%) and area (2.9%) overheads. Asmit De, Aditya Basu, Swaroop Ghosh, Trent Jaeger |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2020 | Threshold-Defined Logic and Interconnect for Protection Against Reverse EngineeringabstractSecuring the intellectual property (IP) from counterfeiting is an important goal toward trustworthy computing. Camouflaging of logic gates is a well-known technique to prevent an adversary from de-layering the chip and stealing IP. In this paper, we propose threshold voltage modulation to realize 2-input static camouflaged logic that can hide six functionalities. We extend the concept of threshold-voltage defined logic to propose multi-input camouflaged gates capable of hiding six 3-input Boolean functions (NAND, NOR, AOI, OAI, XOR, and XNOR). We also propose interconnect camouflaging technique which hides the original connectivity of nets using a novel threshold-voltage defined pass transistor mux. Since threshold voltages are asserted during fabrication and are difficult to identify during optical reverse engineering (RE)-based techniques, the adversary will be forced to launch a brute-force search. We present a thorough analysis of RE effort and overheads associated with the proposed camouflaging techniques. The proposed methodology is demonstrated using fabricated test-chip in 65 nm technology. Jae-Won Jang, Asmit De, Deepak Vontela, Ithihasa Reddy Nirmala, Swaroop Ghosh, Anirudh Iyengar |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2020 | Design, Analysis and Application of Embedded Resistive RAM Based Strong Arbiter PUFabstractResistive Random Access Memory (RRAM) based Physical Unclonable Function (PUF) designs exploit either the probabilistic switching or the resistance variability during forming, SET and RESET processes of RRAM. Memory PUFs using RRAM are typically weak PUFs due to fewer number of challenge response pairs. We propose a strong arbiter PUF based on 1T-1R bit cell which is designed from conventional RRAM memory array with minimally invasive changes. Conventional voltage sense amplifier is repurposed to act like an arbiter and generate the response. Similarly, address and data lines are repurposed to act as challenge and response bits respectively. The PUF is simulated using 65 nm predictive technology models for CMOS and Verilog-A model for a hafnium oxide based RRAM. The proposed PUF architecture is evaluated for uniqueness, uniformity and reliability for various number of stages. It demonstrates mean intra-die Hamming Distance (HD) of 0.135 percent and inter-die HD of 51.4 percent, and passes the NIST tests. We study the vulnerability of proposed PUF to machine learning attacks. We also present an application of proposed PUF for data attestation in the internet of things. Proposed PUF-based data attestation consumes 9.88pJ of total energy per data block of 64-bits and offers a speed of 120.7 kbps. Rekha Govindaraj, Swaroop Ghosh, Srinivas Katkoori |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2020 | Test Methodologies and Test-Time Compression for Emerging Non-Volatile MemoryabstractEmerging nonvolatile memories (NVMs) are considered as suitable candidates to replace conventional memories such as static RAM (SRAM) and dynamic RAM (DRAM) due to high density, high performance, and low (static) power operation. However, NVMs bring new fault issues and call for new tests. For example, NVMs exhibit wide read and write latency distribution, incur high write current (leads to high supply noise), are susceptible to external magnetic/thermal field, show high and stochastic retention time, and are prone to endurance and reliability failures. The conventional tests either cannot capture the faults specific to emerging NVMs or they incur significant test time if implemented on emerging NVMs. In this article, we summarize fault models specific to NVMs and explain the related test issues and challenges. We also propose new tests along with necessary design-for-test techniques to characterize the failures. We further summarize NVM tests proposed in prior works and analyze their test time requirements. Mohammad Nasim Imtiaz Khan, Swaroop Ghosh |
IEEE Trans. Reliab. | 2 |
| 2020 | HarTBleed: Using Hardware Trojans for Data Leakage ExploitsabstractData and information leakage is an important security concern in current systems. Several data leakage prevention (DLP) techniques have been proposed in the literature to prevent external as well as internal data leakage. Most of these solutions try to trace data flow and perform privilege checks to ensure the security of the data at the software and system level. Architecture level leakage vulnerabilities such as Spectre and Meltdown can be mitigated by performance-expensive software patches or by modifying the architecture itself. However, these solutions assume that the underlying hardware platform is secure and free from tampering. In this article, we present HarTBleed, a class of system attacks involving hardware compromised with a Trojan embedded in the CPU. We show that attacks crafted specifically to make use of the Trojan can be used to obtain sensitive information from the address space of a process. We propose the use of a capacitor-based Trojan trigger that exploits the virtual addressing of L1 cache to activate a Trojan payload that resets a target translation lookaside buffer (TLB) entry to maliciously map to sensitive data in memory. Extensive circuit simulation indicates that the proposed Trojan trigger is not activated during test or normal operation even under a wide range of process/temperature conditions. Therefore, it remains undetected. A successful HarTBleed-based exploit is demonstrated using an attack code by modeling the Trojan effects in the GEM5 simulator. Asmit De, Mohammad Nasim Imtiaz Khan, Karthikeyan Nagarajan, Swaroop Ghosh |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2020 | Cache-Out: Leaking Cache Memory Using Hardware TrojanabstractData leakage is an important security concern in current systems. Existing data leakage prevention techniques assume that the underlying hardware platform is secure and free from tampering. In this work, we present Cache-Out, a class of system attacks involving hardware compromised with a Trojan embedded in the CPU. We assume that a memory Trojan trigger is present in L1 d-cache and gets activated if one particular address of L1 d-cache is hammered with a particular data pattern for a certain number of times. Once the Trojan is triggered, accessing another address delivers payloads, such as, read disturb, write disturb, retention failure, and information leakage. We mainly exploit the advanced circuit features employed in the peripherals of nanometer cache memories, such as wordline underdrive (WLUD) (prevents read disturb) and negative bitline (NBL) (assists write) for static RAM (SRAM) to deliver the payloads. Simulation indicates that WLUD and NBL manipulation can inject read and write failures, respectively. We also show that WLUD activation during write operation can inject write failure. Furthermore, NBL along with column multiplexing can also be leveraged to steal data. We validated Cache-Out using GEM5 architectural simulator. We propose L1 address obfuscation, read/write verification, scrambling error correcting code (ECC) bits, and trusted ECC as countermeasures. Results indicate that read/write verification incurs 7.56 μm2of area and 0.1 μW/91.3 μW of static/dynamic power in 22-nm technology for a 64-bit word size. Mohammad Nasim Imtiaz Khan, Asmit De, Swaroop Ghosh |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2019 | QURE: Qubit Re-allocation in Noisy Intermediate-Scale Quantum ComputersabstractConcerted efforts by the academia and the industries e.g., IBM, Google and Intel have brought us to the era of Noisy Intermediate-Scale Quantum (NISQ) computers. Qubits, the basic elements of quantum computer, have been proven extremely susceptible to different noises. Recent experiments have exhibited spatial variations among the qubits in NISQ hardware. Therefore, conventional mapping of qubit done without quality awareness results in significant loss of fidelity for a given workload. In this paper, we have analyzed the effects of various noise sources on the overall fidelity of the given workload for a real NISQ hardware. We have also presented novel optimization technique namely, Qubit Re-allocation (QURE) to maximize the sequence fidelity of a given workload. QURE is scalable and can be applied to future large scale quantum computers. QURE can improve the fidelity of a quantum workload up to 1.54X (1.39X on average) in simulation and up to 1.7X in real device compared to variation oblivious qubit allocation without incurring any physical overhead. Abdullah Ash-Saki, Mahabubul Alam, Swaroop Ghosh |
DAC | 3 |
| 2019 | Sensitivity based Error Resilient Techniques for Energy Efficient Deep Neural Network AcceleratorsabstractWith inherent algorithmic error resilience of deep neural networks (DNNs), supply voltage scaling could be a promising technique for energy efficient DNN accelerator design. In this paper, we propose novel error resilient techniques to enable aggressive voltage scaling by exploiting different amount of error resilience (sensitivity) with respect to DNN layers, filters, and channels. First, to rapidly evaluate filter/channel-level weight sensitivities of large scale DNNs, first-order Taylor expansion is used, which accurately approximates weight sensitivity from actual error injection simulation. With measured timing error probability of each multiply-accumulate (MAC) units considering process variations, the sensitivity variation among filter weights can be leveraged to design DNN accelerator, such that the computations with more sensitive weights are assigned to more robust MAC units, while those with less sensitive weights are assigned to less robust MAC units. Based on post-synthesis timing simulations, 51% energy savings has been achieved with CIFAR-10 dataset using VGG-9 compared to state-of-the-art timing error recovery technique with the same constraint of 3% accuracy loss. Wonseok Choi 0004, Dongyeob Shin, Jongsun Park 0001, Swaroop Ghosh |
DAC | 4 |
| 2019 | FIXER: Flow Integrity Extensions for Embedded RISC-VabstractWith the recent proliferation of Internet of Things (IoT) and embedded devices, there is a growing need to develop a security framework to protect such devices. RISC-V is a promising open source architecture that targets low-power embedded devices and SoCs. However, there is a dearth of practical and low-overhead security solutions in the RISC-V architecture. Programs compiled using RISC-V toolchains are still vulnerable to code injection and code reuse attacks such as buffer overflow and return-oriented programming (ROP). In this paper, we propose FIXER, a hardware implemented security extension to RISC-V that provides a defense mechanism against such attacks. FIXER enforces fine-grained control-flow integrity (CFI) of running programs on backward edges (returns) and forward edges (calls) without requiring any architectural modifications to the RISC-V processor core. We implement FIXER on RocketChip, a RISC-V SoC platform, by leveraging the integrated Rocket Custom Coprocessor (RoCC) to detect and prevent attacks. Compared to existing software based solutions, FIXER reduces energy overhead by 60% at minimal execution time (1.5%) and area (2.9%) overheads. Asmit De, Aditya Basu, Swaroop Ghosh, Trent Jaeger |
DATE | 3 |
| 2019 | Hardware Trojans in Emerging Non-Volatile MemoriesabstractEmerging Non-Volatile Memories (NVMs) possess unique characteristics that make them a top target for deploying Hardware Trojan. In this paper, we investigate such knobs that can be targeted by the Trojans to cause read/write failure. For example, NVM read operation depends on clamp voltage which the adversary can manipulate. Adversary can also use ground bounce generated in NVM write operation to hamper another parallel read/write operation. We have designed a Trojan that can be activated and deactivated by writing a specific data pattern to a particular address. Once activated, the Trojan can couple two predetermined addresses and data written to one address (victim's address space) will get copied to another address (adversary's address space). This will leak sensitive information e.g., encryption keys. Adversary can also create read/write failure to predetermined locations (fault injection). Simulation results indicate that the Trojan can be activated by writing a specific data pattern to a specific address for 1956 times. Once activated, the attack duration can be as low as 52.4μs and as high as 1.1ms (with reset-enable trigger). We also show that the proposed Trojan can scale down the clamp voltage by 400mV from optimum value which is sufficient to inject specific data-polarity read error. We also propose techniques to inject noise in the ground/power rail to cause read/write failure. Mohammad Nasim Imtiaz Khan, Karthikeyan Nagarajan, Swaroop Ghosh |
DATE | 3 |
| 2019 | TOIC: Timing Obfuscated Integrated CircuitsabstractTo counter the threats of reverse engineering (RE) and Trojan in-sertion, researchers have considered gate-level obfuscation in inte-grated circuits (IC) as a viable solution. However, several techniques are present in the literature to crack the obfuscation with varying degree of success raising the concern about their secrecy. In this article, we have presented TOIC (Timing Obfuscated Integrated Circuits), a novel technique where sequential elements are obfuscated to hide the true timing paths in the design. TOIC can act as a standalone countermeasure against IC reverse engineering or can be incorporated with existing gate camouflaging techniques to maximize adversarial RE effort. Previous research has shown that limiting access to internal nodes can improve the adversarial RE effort at the cost of poor testability. TOIC can impose prohibitively large decamouflaging time complexity by limiting the controllability and observability over the internal nodes in an IC while preserving complete testability. Mahabubul Alam, Swaroop Ghosh, Sujay Hosur |
ACM Great Lakes Symposium on VLSI | 2 |
| 2019 | MUQUT: Multi-Constraint Quantum Circuit Mapping on NISQ Computers: Invited PaperabstractRapid advancement in the domain of quantum technologies have opened up researchers to the real possibility of experimenting with quantum circuits, and simulating small-scale quantum programs. Nevertheless, the quality of currently available qubits and environmental noise pose a challenge in smooth execution of the quantum circuits. Therefore, efficient design automation flows for mapping a given algorithm to the Noisy Intermediate Scale Quantum (NISQ) computer becomes of utmost importance. State-of-the-art quantum design automation tools are primarily focused on reducing logical depth, gate count and qubit counts with recent emphasis on topology-aware (nearest-neighbour compliance) mapping. In this work, we extend the technology mapping flows to simultaneously consider the topology and gate fidelity constraints while keeping logical depth and gate count as optimization objectives. We provide a comprehensive problem formulation and multi-tier approach towards solving it. The proposed automation flow is compatible with commercial quantum computers, such as IBM QX and Rigetti. Our simulation results over 10 quantum circuit benchmarks, show that the fidelity of the circuit can be improved up to 3.37 × with an average improvement of 1.87 ×. Debjyoti Bhattacharjee, Abdullah Ash-Saki, Mahabubul Alam, Anupam Chattopadhyay, Swaroop Ghosh |
ICCAD | 5 |
| 2019 | FPCAS: In-Memory Floating Point Computations for Autonomous SystemsabstractAutonomous systems e.g., cars and drones generate vast amount of data from sensors that need to be processed in timely fashion to make accurate and safe decisions. Majority of these computations deal with Floating Point (FP) numbers. Conventional Von-Neumann computing paradigm suffers from overheads associated with data transfer. In-memory computing (IMC) can solve this challenge by processing the data locally. However, in-memory FP computing has not been investigated before. We propose F P arithmetic (adder/subtractor and multiplier) using Resistive RAM (ReRAM) crossbar based IMC. A novel shift circuitry is proposed to lower the shift overhead inherently present in the FP arithmetic. The proposed single precision FP adder consumes 335 pJ and 322 pJ for NAND-NAND and NOR-NOR based implementation for addition/subtraction, respectively. The proposed adder/subtractor improves latency, power and energy by 828X, 3.2X, and 3.7X, respectively, compared to MAGIC [1]. Furthermore, the proposed multiplier reduces energy per operation by 1.13X and improves performance by 4.4X compared to ReVAMP [2]. Sina Sayyah Ensan, Swaroop Ghosh |
IJCNN | 2 |
| 2019 | Meeting the Conflicting Goals of Low-Power and Resiliency Using Emerging Memories : (Invited Paper)abstractEmerging non-volatile memory (NVM) technologies are being aggressively explored to replace and/or assist conventional CMOS technology. Although NVMs can cut down leakage power, achieve low footprint and allow compute capability along with storage, they suffer from new sources of variability. We review the noise sources associated with NVMs and describe resilience enhancement techniques for both memory and computing. We also present security applications where noise and variability is desirable. Karthikeyan Nagarajan, Mohammad Nasim Imtiaz Khan, Sina Sayyah Ensan, Abdullah Ash-Saki, Swaroop Ghosh |
IOLTS | 5 |
| 2019 | Addressing Temporal Variations in Qubit Quality Metrics for Parameterized Quantum CircuitsabstractThe public access to noisy intermediate-scale quantum (NISQ) computers facilitated by IBM, Rigetti, D - Wave, etc., has propelled the development of quantum applications that may offer quantum supremacy in the future large-scale quantum computers. Parameterized quantum circuits (P QC) have emerged as a major driver for the development of quantum routines that potentially improve the circuit's resilience to the noise. PQC's have been applied in both generative (e.g. generative adversarial network) and discriminative (e.g. quantum classifier) tasks in the field of quantum machine learning. PQC's have been also considered to realize high fidelity quantum gates with the available imperfect native gates of a target quantum hardware. Parameters of a P QC are determined through an iterative training process for a target noisy quantum hardware. However, temporal variations in qubit quality metrics affect the performance of a P QC. Therefore, the circuit that is trained without considering temporal variations exhibits poor fidelity over time. In this paper, we present training methodologies for P QC in a completely classical environment that can improve the fidelity of the trained P QC on a target NISQ hardware by as much as 21.91%. Mahabubul Alam, Abdullah Ash-Saki, Swaroop Ghosh |
ISLPED | 3 |
| 2019 | SHINE: A Novel SHA-3 Implementation Using ReRAM-based In-Memory ComputingabstractIn memory-computing (IMC) architectures provide a much needed solution to energy-efficiency barriers posed by Von-Neumann computing due to movement of data between the processor and the memory. Emerging non-volatile memories (NVM) such as Resistive RAM (ReRAM) implemented in a crossbar array are promising substrates to realize IMC due to excellent High Resistance State (HRS) to Low Resistance State (LRS) ratios and high-densities. Hardware security primitives such as SHA-3 require heavy data traffic between processing elements and memory. Therefore, they can be benefited substantially by in-memory acceleration. We propose SHINE, a high performance and area efficient hardware implementation of the Keccak function that forms the core of SHA-3 by exploiting ReRAM-based IMC. SHINE implements various functions in a Sum of Product (SOP) form in the crossbar array architecture. Simulation results show that it cuts down energy by ~90.5% and increases throughput by 1.5X to 2.8X as compared to conventional CMOS based implementations such as [1] and [2]. Karthikeyan Nagarajan, Sina Sayyah Ensan, Mohammad Nasim Imtiaz Khan, Swaroop Ghosh, Anupam Chattopadhyay |
ISLPED | 4 |
| 2019 | TCAD EIC Message: February 2019abstractAs we close out the year 2018, it is time to reflect back a number of milestones achieved throughout the year. The transition to the new EIC and team included a 50-member editorial board with 19 new members selected after an extensive round of open call for editorial board nominations. While, this was a reduction in the editorial board from 66 members previously, the response time remained steady at about two months from submission to first decision. As of this writing in December, we received 469 new manuscripts as well as 371 revised manuscripts in 2018. The top two departments with substantial lead over the rest were “Modeling and Simulation” and “Emerging Technologies and Applications.” Philip Brisk, Claudionor José Nunes Coelho Jr., Abdoulaye Gamatié, Swaroop Ghosh |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2019 | Guest Editorial Special Section on Security Challenges and Solutions With Emerging Computing TechnologiesabstractMultiple emerging computing technologies based on, e.g., graphene, spintronics, resistive RAM, quantum computing, and others are being developed to enhance the capabilities of logic devices and circuits. The rapid growth in these technologies is synchronized with the decline of Moore’s law, thus promises to herald the era of Beyond CMOS technologies with a significant improvement in energy efficiency, reliability, performance, and manufacturability. These devices enable very different computing paradigms, e.g., neuromorphic computing, non-Boolean computing, and in-memory computing, thus making these platforms an interesting playground for circuit and application-developers alike. Anupam Chattopadhyay, Swaroop Ghosh, Wayne P. Burleson, Debdeep Mukhopadhyay |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2018 | Novel application of spintronics in computing, sensing, storage and cybersecurityabstractWith conventional Von Neumann computing struggling to match the energy-efficiency of biological systems, there is pressing need to explore alternative computing models. CMOS switches, although universal, fails to offer additional features to meet this end goal Recent experimental studies have revealed that spintronics possess many promising features that can not only enable non Von Neumann compute models but also high-density storage, sensing of environmental parameters and protection from cybersecurity threats. This paper provides an in-depth study of spintronics and its relation to these novel aspects from device, circuit and system standpoint. Seyedhamidreza Motaman, Mohammad Nasim Imtiaz Khan, Swaroop Ghosh |
DATE | 3 |
| 2018 | How Multi-Threshold Designs Can Protect Analog IPsabstractAnalog Integrated Circuits (ICs) are one of the top targets for counterfeiting. However, the security of analog Intellectual Property (IP) is not well investigated as its digital counterpart. In this paper, we explore the possibility of multi-threshold voltage (VTH) design to protect the analog IP from Reverse Engineering (RE)-based attacks. Analog circuits are sensitive to VTH as the operating region of a transistor can vary with VTH. Furthermore, the VTH of individual transistors cannot be identified during the RE process. The trial-and-error based technique to guess the VTH and validate with a golden IC will ramp up RE effort exponentially. Thus, by carefully including multi-VTH transistors, the designer can ensure that the properties of analog IP e.g., gain, bandwidth, and linearity are protected even though the physical dimensions of the transistors are revealed. We demonstrate this technique by using a case study on a wide-swing cascode amplifier. Simulations show that incorrect VTH inference can lead to substantially degraded performance like 98 dB drop in open-loop gain and up to 19% increase in total harmonic distortion. Based on VTH choice, the proposed technique can save ~ 3% area over conventional design. We show that the reverse engineering effort can be ~1013 years. We propose a technique like transistor splitting to increase the effort even more. Mismatch analysis shows that the proposed technique results in only 1% loss in mean robustness. Abdullah Ash-Saki, Swaroop Ghosh |
ICCD | 2 |
| 2018 | Analysis of Row Hammer Attack on STTRAMabstractIn this paper, we model and investigate the impact of Row Hammering (RH) on Spin-Transfer Torque RAM (STTRAM) by exploiting its write operation. STTRAM suffers from high write current and long write latency which can result in ground bounce. The magnitude of the bounce depends on the old data and the new data that is being written. The bounce can propagate to the nearest word-line drivers and partially turn ON the access transistors making weak current flow through the memory bitcells and reducing their thermal energy barrier. Therefore, continuous write at a particular location can force the massive number of unselected bits to suffer from degraded thermal barrier due to weak RH current. Reduced thermal barrier may lead to retention failures and make the bits sensitive to stray magnetic field/thermal noise. Those bits can also suffer from read disturb if they are read. These issues could be even worse for Short Retention NVM (SRNVM) which is suitable for Last Level Cache (LLC) and has a base retention of only few seconds. The ground bounce can also propagate to bitline/ source-line drivers and the selected cells will experience lower headroom voltage. This will lead to read failure (due to degraded sense margin) and write failure (due to increased write latency). Simulation result indicates that RH attack can flip the bits in just 30.84secs for STTRAM with base retention of 1 min. In presence of elevated temperature, the retention time can be further reduced to 2.46secs and 0.19secs for T=50C and T=75C respectively. RH attack can increase read disturb by 2.09X for bitcell with 1min base retention at T=25C. Simulation result also indicates that RH attack can cause read/write failure if the bitcell being read/written experience 306mV (for data 0)/110mV (for writing 0 –> 1) of bounce. To the best of our knowledge, this is the first RH attack study for STTRAM-based cache. Mohammad Nasim Imtiaz Khan, Swaroop Ghosh |
ICCD | 2 |
| 2018 | Dynamic Computing in Memory (DCIM) in Resistive Crossbar ArraysabstractWith Von-Neumann computing struggling to match the energy-efficiency of biological systems, there is pressing need to explore alternative computing models. Recent experimental studies have revealed that Resistive Random Access Memory (RRAM) is promising alternative for DRAM. Resistive crossbar arrays possess many promising features that can not only enable high-density and low-power storage but also non Von-Neumann compute models. Most recent works focus on dot product operation with RRAM crossbar arrays, and therefore are not flexible to implement various logical functions. We propose a low-power dynamic computing in memory system which can implement various functions in Sum of Product (SOP) form in RRAM crossbar array architecture. We evaluate the proposed technique by performing simulation over wide range of MCNC benchmarks. Simulation results show 1.42X and 20X latency improvement as well as 2.6X and 12.6X power saving compared to static and MAGIC computing in memory methods. Seyedhamidreza Motaman, Swaroop Ghosh |
ICCD | 2 |
| 2018 | Threshold Defined Camouflaged Gates in 65nm Technology for Reverse Engineering ProtectionabstractDue to the ever-increasing threat of Reverse Engineering (RE) of Intellectual Property (IP) for malicious gains, camouflaging of logic gates is becoming very important. In this paper, we present experimental demonstration of transistor threshold voltage-defined switch [2] based camouflaged logic gates that can hide six logic functionalities i.e. NAND, AND, NOR, OR, XOR and XNOR. The proposed gates can be used to design the IP, forcing an adversary to perform brute-force guess-and-verify of the underlying functionality---increasing the RE effort. We propose two flavors of camouflaging, one employing only a pass transistor (NMOS-switch) and the other utilizing a full pass transistor (CMOS-switch). The camouflaged gates are used to design Ring-Oscillators (RO) in ST 65nm technology, one for each functionality, on which we have performed temperature, voltage, and process-variation analysis. We observe that CMOS-switch based camouflaged gate offers a higher performance (~1.5-8X better) than NMOS-switch based gate at an added area cost of only 5%. The proposed gates show functionality till 0.65V. We are also able to reclaim lost performance by dynamically changing the switch gate voltage and show that robust operation can be achieved at lower voltage and under temperature fluctuation. Anirudh Iyengar, Deepak Vontela, Ithihasa Reddy Nirmala, Swaroop Ghosh, Seyedhamidreza Motaman, Jae-Won Jang |
ISLPED | 4 |
| 2018 | Information Leakage Attacks on Emerging Non-Volatile Memory and CountermeasuresabstractEmerging Non-Volatile Memories (NVMs) suffer from high and asymmetric read/write current and long write latency which can result in supply noise, such as supply voltage droop and ground bounce. The magnitude of supply noise depends on the old data and the new data that is being written (for a write operation) or on the stored data (for a read operation). Therefore, victim's write operation creates a supply noise which propagates to adversary's memory space. The adversary can detect victim's write initiation and can leverage faster read latency (compared to write) to further sense the Hamming Weight (HW) of the victim's write data by detecting read failures in his memory space. These attacks are specifically possible if exhaustive testing of the memory for all patterns, all possible location combinations, all possible parallel read/write conditions are not performed under bit-to-bit process variations and specified (-10°C to 90°C) and unspecified temperature ranges (i.e., less than -10°C and greater than 90°C). Simulation result indicates that adversary can sense HW of victim's (near-by) write data = 66.77%, and further narrow the range based on read/write failure characteristics. Side Channel Attacks can utilize this information to strengthen the attacks. Mohammad Nasim Imtiaz Khan, Swaroop Ghosh |
ISLPED | 2 |
| 2018 | A Monolithic-3D SRAM Design with Enhanced Robustness and In-Memory Computation SupportabstractWe present a novel 3D-SRAM cell using a Monolithic 3D integration (M3D-IC) technology for realizing both robustness and In-memory Boolean logic compute support. The proposed two-layer design makes use of additional transistors over the SRAM layer to enable assist techniques as well as provide logic functions (such as AND/NAND, OR/NOR, XNOR/XOR) without degrading cell density. Through analysis, we provide insights into the benefits provided by three memory assist and two logic modes and evaluate the energy efficiency of our proposed design. Assist techniques improve SRAM read stability by 2.2x and increase the write margin by 17.6%, while staying within the SRAM footprint. By virtue of increased robustness, the cell enables seamless operation at lower supply voltages and thereby ensures energy efficiency. Energy Delay Product (EDP) reduces by 1.6x over standard 6T SRAM with a faster data access. Transistor placement and their biasing technique in layer-2 enables In-memory bitwise Boolean computation. When computing bulk In-memory operations, 6.5x energy savings is achieved as compared to computing outside the memory system. Srivatsa Rangachar Srinivasa, Akshay Krishna Ramanathan, Xueqing Li 0002, Wei-Hao Chen, Fu-Kuo Hsueh, Chih-Chao Yang, Chang-Hong Shen, Jia-Min Shieh, Sumeet Kumar Gupta, Meng-Fan Chang, Swaroop Ghosh, Jack Sampson, Narayanan Vijaykrishnan |
ISLPED | 11 |
| 2018 | Test of Supply Noise for Emerging Non-Volatile MemoryabstractEmerging Non-Volatile Memories (NVMs) suffer from high read/write current which can result in supply noise such as voltage droop and ground bounce. The magnitude of supply noise depends on the old data and the new data that is being written (for a write operation) or the stored data (for a read operation). In prior work, it has been shown that the noise generated by one access can affect another parallel access. Therefore, parallel read/write operation should be tested considering the supply noise. However, testing for read/write failure with supply noise considerations can take significant test time. In this work, we show that test time can be reduced by 410.82X for RRAM-based NVM Last Level Cache (LLC) by using Design for Test (DFT) circuits such as wordline overdrive and ending write operation early. We also show that the proposed test can save 79.875J of energy compared to the baseline test method. Mohammad Nasim Imtiaz Khan, Swaroop Ghosh |
ITC | 2 |
| 2018 | Test challenges and solutions for emerging non-volatile memoriesabstractAt the end of Silicon roadmap, keeping the leakage power in tolerable limit has become one of the biggest challenges. Several promising non-volatile memories (NVMs) are being investigated by the scientific community to address the issue. Some of the NVMs such as Spin-Transfer Torque RAM, Magnetic RAM, Resistive RAM, Phase Change Memory and Ferroelectric RAM have already entered the mainstream computing. However, the unique characteristics of these NVMs bring new fault models such as statistical and stochastic retention failures, magnetic and thermal tolerance failures, voltage droop and ground bounce induced read and write failures and long latency failures. In this work, we summarize new test failure mechanisms in NVMs and associated test challenges. We also propose new test methodologies, test patterns and Design-for-Test (DFT) techniques to characterize new failure models and compress test time. Mohammad Nasim Imtiaz Khan, Swaroop Ghosh |
VTS | 2 |
| 2018 | Impact of Process Variation on Self-Reference Sensing Scheme and Adaptive Current Modulation for Robust STTRAM SensingabstractSpin-Transfer-Torque RAM (STTRAM) is a promising technology for high-density on-chip cache due to low standby power and high speed. However, the process variation of the Magnetic Tunnel Junction (MTJ) and access transistor poses a serious challenge to sensing. Nondestructive sensing suffers from reference resistance variation, whereas destructive sensing suffers from failures due to unoptimized selection of data and reference currents. Furthermore, the sense speed is tightly coupled with the reference/data current requirement. In this work, we study the process variation effect on a self-reference sensing scheme to eliminate bit-to-bit process variation in MTJ resistance. Read current modulation is proposed to overcome the failures due to process variation. Simulation results reveal <0.01% failures at the cost of 9ns sense time and 190uW power consumption. Seyedhamidreza Motaman, Swaroop Ghosh, Jaydeep P. Kulkarni |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2018 | CSRO-Based Reconfigurable True Random Number Generator Using RRAMabstractIn this paper, we propose a high-speed (kilohertz-megahertz), reconfigurable current starved ring oscillator (CSRO)-based true random number generator (TRNG) design. The proposed TRNG exploits the intradevice stochastic variations in resistive RAM switching parameters and random telegraph noise (RTN). We demonstrate the effect of RTN on the jitter of CSRO oscillations. We also propose a methodology to reconfigure the TRNG to generate new random numbers. The proposed 10-bit TRNG is validated by NIST test suite for randomness in the data stream. Energy/bit is 22.8 fJ for generation, and the speed of random data generation is 6 MHz. Security vulnerabilities and countermeasures of the proposed TRNG are also investigated. Rekha Govindaraj, Swaroop Ghosh, Srinivas Katkoori |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2018 | Novel Magnetic Burn-In for Retention and Magnetic Tolerance Testing of STTRAM
Mohammad Nasim Imtiaz Khan, Anirudh Iyengar, Swaroop Ghosh |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2017 | Droop mitigating last level cache architecture for STTRAMabstractSpin-Transfer Torque Random Access Memory (STTRAM) is one of the emerging Non-Volatile Memory (NVM) technologies especially preferred for the Last Level Cache (LLC). The amount of current needed to switch the magnetization is high (~100ßA per bit). For a full cache line (512-bit) write, this extremely high current results in a voltage droop in the conventional cache architecture. Due to this droop, the write operation fails especially, when the farthest bank of the cache is accessed. In this paper, we propose a new cache architecture to mitigate this problem of droop and make the write operation successful. Instead of continuously writing the entire cache line (512-bit) in a single bank, the proposed architecture writes 64-bits in multiple physically separated locations across the cache. The simulation results obtained (both circuit and micro-architectural) comparing our proposed architecture against the conventional are found to be 1.96% (IPC) and 5.21% (energy). Radha Krishna Aluru, Swaroop Ghosh |
DATE | 2 |
| 2017 | Novel magnetic burn-in for retention testing of STTRAMabstractSpin-Transfer Torque RAM (STTRAM) is an emerging Non-Volatile Memory (NVM) technology that has drawn significant attention due to complete elimination of bitcell leakage. However, it brings new challenges in characterizing the retention time of the array during test. Significant shift of retention time under static (process variation (PV)) and dynamic (voltage, temperature fluctuation) variability furthers this issue. In this paper, we propose a novel magnetic burn-in (MBI) test which can be implemented with minimal changes in the existing test flow to enable STTRAM retention testing at short test time. The magnetic burn-in is also combined with thermal burn-in (MBI−BI) for further compression of retention and test time. Simulation results indicate MBI with 220Oe (at 25C) can improve the test time by 3.71×1013X while MBI−BI with 220Oe at 125C can improve the test time by 1.97×1014X. Mohammad Nasim Imtiaz Khan, Anirudh Iyengar, Swaroop Ghosh |
DATE | 3 |
| 2017 | Side-Channel Attack on STTRAM Based Cache for Cryptographic ApplicationabstractIn this paper, we propose a Side Channel Attack (SCA) model on Spin-Torque Transfer RAM (STTRAM) where an adversary can monitor the supply current of the memory array consumed during read/write operations and recover the secret key of Advanced Encryption Standard (AES) execution. Simulation results indicate that by monitoring write current, 50% of keys could be extracted using 2000 traces. Further improvement of attacks on write operation is also proposed. The read current is found to be more susceptible to leak the key. It reveals first byte in only 40 traces and leaks the entire key in as low as 400 traces. The results are then compared with Static RAM (SRAM) based cache. The attack model has been experimentally validated on read operation of commercial MRAM chip (STTRAM variant). Experimental results indicate that the attack can reveal correct key in 15 traces compared to 40 in simulation due to less algorithmic noise. To the best of our knowledge, this is the first comprehensive SCA study for STTRAM based cache for cryptographic application. Mohammad Nasim Imtiaz Khan, Shivam Bhasin, Alex Yuan, Anupam Chattopadhyay, Swaroop Ghosh |
ICCD | 5 |
| 2017 | Design and Analysis of STTRAM-Based Ternary Content Addressable Memory CellabstractContent Addressable Memory (CAM) is widely used in applications where searching a specific pattern of data is a major operation. Conventional CAMs suffer from area, power, and speed limitations. We propose Spin-Torque Transfer RAM--based Ternary CAM (TCAM) cells. The proposed NOR-type TCAM cell has a 62.5% (33%) reduction in number of transistor compared to conventional CMOS TCAMs (spintronic TCAMs). We analyzed the sense margin of the proposed TCAM with respect to 16-, 32-, 64-, 128-, and 256-bit word sizes in 22nm predictive technology. Simulations indicated a reliable sense margin of 50mV even at 0.7V supply voltage for 256-bits word. We also explored a selective threshold voltage modulation of transistors to improve the sense margin and tolerate process and voltage variations. The worst-case search latency and sense margin of 256-bit TCAM is found to be 263ps and 220mV, respectively, at 1V supply voltage. The average search power consumed is 13mW, and the search energy is 4.7fJ/bit search. The write time is 4ns, and the write energy is 0.69pJ/bit. We leverage the NOR-type TCAM design to realize a 9T-2 Magnetic Tunnel Junctions NAND-type TCAM cell that has 43.75% less number of transistors than the conventional CMOS TCAM cell. A NAND-type cell can support up to 64-bit words with a maximum sense margin of up to 33mV. We compare the performance metrics of NOR- and NAND-type TCAM cells with other TCAMs in the literature. Rekha Govindaraj, Swaroop Ghosh |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2016 | Data privacy in non-volatile cache: Challenges, attack models and solutionsabstractSpin-Transfer-Torque RAM (STTRAM) is considered to be a strong candidate for last level cache (LLC). Although promising STTRAM LLC brings new security challenges that were absent in conventional volatile memories such as Static RAM (SRAM). The root cause is persistent data and the fundamental dependency of the memory technology on ambient parameters such as magnetic field and temperature that can be exploited to compromise the data. We provide a qualitative analysis of the data privacy issues in the emerging nonvolatile cache. We also propose new attack models to compromise the sensitive data in LLC. The encryption technique used to secure data in main memory and hard disk may not be useful for LLC due to latency overhead. We propose two low-overhead techniques to ensure data privacy in LLC- (a) implementing semi nonvolatile memory (SNVM); and, (b) data erasure at power OFF. Erasing could be energy intensive and may require dedicated battery to work under power failure attacks. To address this concern we reuse the energy stored in power rail after power OFF to erase the bits using a canary circuit to track MTJ write time. The simulation results show 0.6% IPC loss and 1.2% energy overhead during normal operation due to added circuitry. Nitin Rathi, Swaroop Ghosh, Anirudh Iyengar, Helia Naeimi |
ASP-DAC | 2 |
| 2016 | A novel threshold voltage defined switch for circuit camouflagingabstractSemiconductor supply chain is increasingly getting exposed to variety of security attacks such as Trojan insertion, cloning, counterfeiting, reverse engineering (RE), piracy of Intellectual Property (IP) or Integrated Circuit (IC) and side-channel analysis due to involvement of untrusted parties. In this paper, we propose threshold voltage-defined switches that will camouflage the logic gate both logically and physically to resist RE and IP piracy. The proposed gate can function as NAND, AND, NOR, OR, XOR, and XNOR robustly using threshold defined switches. We also propose a flavor of camouflaged gate that represents reduced functionality (NAND, NOR and NOT) at much lower overhead. The camouflaged design operates at nominal voltage and obeys conventional reliability limits. A small fraction of gates can be camouflaged to increase the RE effort extremely high. Simulation results indicate 46-53% area, 59-68% delay and 52-76% power overhead when 5-15% gates are identified and camouflaged using the proposed gate. A significant higher RE effort is achieved when the proposed gate is employed in the netlist using controllability, observability and hamming distance sensitivity based gate selection metrics. Ithihasa Reddy Nirmala, Deepak Vontela, Swaroop Ghosh, Anirudh Iyengar |
ETS | 3 |
| 2016 | Security and privacy threats to on-chip non-volatile memories and countermeasuresabstractNon-volatile memories (NVMs) such as Spin-Transfer Torque RAM (STTRAM) have drawn significant attention due to complete elimination of bitcell leakage. In addition to the plethora of benefits such as density, non-volatility, low-power and high speed, majority of Non-Volatile Memories (NVMs) are also compatible with CMOS technology enabling easy integration. Although promising, NVM brings new security challenges that were absent in their conventional volatile memory counterparts such as Static RAM (SRAM) and embedded Dynamic RAM (eDRAM). The root cause is persistent data that may allow the adversary to retrieve sensitive information like password or cryptographic keys. This is primarily due to the fundamental dependency of these memory technologies on environmental parameters such as magnetic fields and temperature which can be exploited by the adversary to tamper with the stored data. This paper investigates the data security and privacy challenges in NVMs by exploring the security specific properties and novel security primitives realized using spintronic building blocks. A thorough analysis is done on the vulnerabilities, data security and privacy issues, threats and possible countermeasures to enable safe computing environment using spintronics. Swaroop Ghosh, Mohammad Nasim Imtiaz Khan, Asmit De, Jae-Won Jang |
ICCAD | 1 |
| 2016 | A strong arbiter PUF using resistive RAM within 1T-1R memory architectureabstractPhysically Unclonable Function (PUF) is cost effective and reliable security primitives widely used in authentication and in-place secret key generation. With growing research in the area of non-CMOS technologies for memories and circuits, it is important to understand their implications on the design of security primitives. Resistive Random Accessible Memory (RRAM) offers easy integration with CMOS due to minimal changes in the process technology. RRAM also demonstrates resistance variability characteristics due to inherent defects in the conducting filament formed inside the metal oxide layer. RRAM based PUF designs exploit either the probabilistic switching of RRAM or the resistance variability during forming, SET and RESET processes. Memory PUFs using RRAM are typically weak PUFs due to fewer number of Challenge Response Pairs (CRPs). We propose strong arbiter PUF based on 1T-1R bit cell which is obtained from conventional RRAM memory array with minimally invasive changes. Conventional voltage sense amplifier is employed to generate the response. The PUF is simulated using 65nm predictive technology models for CMOS and Verilog-A model for a hafnium oxide based RRAM. The proposed PUF architecture is evaluated for uniqueness, uniformity and reliability and by running NIST benchmarks. It demonstrates mean intra-die Hamming Distance (HD) of 0.13% and inter-die HD of 51.3%, and, passes the NIST tests. Rekha Govindaraj, Swaroop Ghosh |
ICCD | 2 |
| 2016 | Domain Wall Memory based Convolutional Neural Networks for Bit-width Extendability and Energy-EfficiencyabstractIn the hardware implementation of deep learning algorithms such as Convolutional Neural Networks (CNNs), vector-vector multiplications and memories for storing parameters take a significant portion of area and power consumption. In this paper, we propose a Domain Wall Memory (DWM) based design of CNN convolutional layer. In the proposed design, the resistive cell sensing mechanism is efficiently exploited to design a low-cost DWM-based cell arrays for storing parameters. The unique serial access mechanism and small footprint of DWM are also used to reduce the area and power cost of the input registers for aligning inputs. Contrary to the conventional implementation using Memristor-Based Crossbar (MBC), the bit-width of the proposed CNN convolutional layer is extendable for high resolution classifications and training. Simulation results using 65 nm CMOS process show that the proposed design archives 34% of energy savings compared to the conventional MBC based design approach. Jinil Chung, Jongsun Park 0001, Swaroop Ghosh |
ISLPED | 3 |
| 2016 | Performance Impact of Magnetic and Thermal Attack on STTRAM and Low-Overhead Mitigation TechniquesabstractIn this paper, we analyze the fundamental vulnerabilities of Spin-Torque-Transfer RAM on magnetic field and temperature that can be exploited by adversaries with an intent to trigger soft performance failures. We present novel attack vectors and their impact on memory performance (i.e., read, write and retention). We propose a novel low-overhead clock frequency-adaptation technique to mitigate the attack. Our analysis indicate slowing the clock frequency by 85% restores 170 mV of sense margin under 300 Oe DC magnetic field. In addition, 66% operating clock slowdown allows STTRAM to tolerate over 300 Oe AC magnetic field. Jae-Won Jang, Swaroop Ghosh |
ISLPED | 2 |
| 2016 | Spintronic PUFs for Security, Trust, and AuthenticationabstractWe propose spintronic physically unclonable functions (PUFs) to exploit security-specific properties of domain wall memory (DWM) for security, trust, and authentication. We note that the nonlinear dynamics of domain walls (DWs) in the physical magnetic system is an untapped source of entropy that can be leveraged for hardware security. The spatial and temporal randomness in the physical system is employed in conjunction with microscopic and macroscopic properties such as stochastic DW motion, stochastic pinning/depinning, and serial access to realize novel relay-PUF and memory-PUF designs. The proposed PUFs show promising results (∼50% interdie Hamming distance (HD) and 10% to 20% intradie HD) in terms of randomness, stability, and resistance to attacks. We have investigated noninvasive attacks, such as machine learning and magnetic field attack, and have assessed the PUFs resilience. Anirudh Iyengar, Swaroop Ghosh, Kenneth Ramclam, Jae-Won Jang, Cheng-Wei Lin |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2016 | Spintronics and Security: Prospects, Vulnerabilities, Attack Models, and PreventionsabstractThe experimental demonstration of current-driven spin-transfer torque (STT) for switching magnets and push domain walls (DWs) in magnetic nanowires have opened up new avenues for spintronic computations. These devices have shown great promise for logic and memory applications due to superior energy efficiency and nonvolatility. It has been noted that the nonlinear dynamics of DWs in the physical magnetic system is an untapped source of entropy that can be leveraged for hardware security. The inherent noise, spatial, and temporal randomness in the magnetic system can be employed in conjunction with microscopic and macroscopic properties to realize novel hardware security primitives. Due to simplicity of integration, the spintronic circuits can be an add-on to the silicon substrate to complement the existing CMOS-based security and trust infrastructures. This paper investigates the prospects of spintronics in hardware security by exploring the security-specific properties and novel security primitives realized using spintronic building blocks. As spintronic elements enter the mainstream computing platforms, they are exposed to emerging attacks that were infeasible before. This paper covers the security vulnerabilities, security and privacy attack models, and possible countermeasures to enable safe computing environment using spintronics. Swaroop Ghosh |
Proc. IEEE | 1 |
| 2016 | Adaptive Write and Shift Current Modulation for Process Variation Tolerance in Domain Wall CachesabstractDomain wall memory (DWM), also known as racetrack memory, is gaining significant attention for embedded cache application due to low standby power, excellent retention, and the ability to store multiple bits per cell. In addition, it offers fast access time, good endurance, and retention. However, it suffers from poor write latency, shift latency, shift power, and write power. In addition, we observe that process variation can result in a large spread in write and read latency variations. The performance of conventionally designed DWM cache can degrade as much as 13% due to process variations. We propose a novel and adaptive write current and shift current boosting to address this issue. The bits experiencing worst case write latency are fixed through a combination of write and shift boosting, whereas worst case read bits are fixed by shift boosting. Simulations show a 30% dynamic energy improvement compared with boosting all bit-cells and a 18% performance improvement compared with worst case latency due to process variation over a wide range of PARSEC benchmarks. Seyedhamidreza Motaman, Swaroop Ghosh |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2015 | Domain wall memory based digital signal processors for area and energy-efficiencyabstractIn many Digital Signal Processing (DSP) applications such as Viterbi decoder and Fast Fourier Transform (FFT), Static Random Access Memory (SRAM) based embedded memory consumes significant portion of area and power. These DSP units are dominated by sequential memory access where SRAM-based memory is inefficient in terms of area and power. We propose spintronic Domain Wall Memory (DWM) based embedded memories for DSP building blocks e.g., survivor-path memories in Viterbi decoder and First-In-First-Out (FIFO) register files in FFT processor that exploit the unique serial access mechanism, non-volatility and small footprint of the memory for area and power saving. Simulations using 65nm technology show that the DWM based Viterbi decoder achieves 66.4 % area and 59.6 % power savings over the conventional SRAM-based implementation. For 8K point FFT processor, the DWM based design shows 60.6 % area and 60.3 % power savings. Jinil Chung, Kenneth Ramclam, Jongsun Park 0001, Swaroop Ghosh |
DAC | 4 |
| 2015 | Self-correcting STTRAM under magnetic field attacksabstractSpin-Transfer Torque Random Access Memory (STTRAM) is a possible candidate for universal memory due to its high-speed, low-power, non-volatility, and low cost. Although attractive, STTRAM is susceptible to contactless tampering through malicious exposure to magnetic field with the intention to steal or modify the bitcell content. In this paper, for the first time to our knowledge, we analyze the impact of magnetic attacks on STTRAM using micro-magnetic simulations. Next, we propose a novel array-based sensor to detect the polarity and magnitude of such attacks and then propose two design techniques to mitigate the attack, namely, array sleep with encoding and variable strength Error Correction Code (ECC). Simulation results indicate that the proposed sensor can reliably detect an attack and provide sufficient compensation window (few ns to ~100us) to enable proactive protection measures. Finally, we shows that variable-strength ECC can adapt correction capability to tolerate failures with various strength of an attack. Jae-Won Jang, Jongsun Park 0001, Swaroop Ghosh, Swarup Bhunia |
DAC | 3 |
| 2015 | Impact of process-variations in STTRAM and adaptive boosting for robustness
Seyedhamidreza Motaman, Swaroop Ghosh, Nitin Rathi |
DATE | 2 |
| 2015 | Design and analysis of 6-T 2-MTJ ternary Content Addressable MemoryabstractContent Addressable Memory (CAM) is widely used in pattern matching, internet data processing and many other fields where searching a specific pattern of data is a major operation. Conventional CAMs suffer from area, power, and speed limitations. We propose a magnetic tunnel junction (MTJ) based Ternary CAM (TCAM). The proposed TCAM cell is 127 percent (33 percent) area efficient compared to conventional CMOS TCAM (spintronic TCAMs). We analyzed sense margin of the proposed TCAM with respect to 16, 32, 64, 128 and 256-bit words sizes in 22nm predictive technology. Simulations indicated reliable sense margin of 50mV even at 0.7V supply voltage. The worst case sense delay and sense margin of 256-bit TCAM is found to be 263ps and 220mV respectively at 1V supply voltage. The average search power consumed is 13mW and the search energy is 4.7fJ per bit search. The write time is 4ns and the write energy is 0.69pJ per bit. Rekha Govindaraj, Swaroop Ghosh |
ISLPED | 2 |
| 2015 | A novel slope detection technique for robust STTRAM sensingabstractSpin-Torque-Transfer RAM (STTRAM) is a promising technology for high density on-chip cache due to low standby power and high speed. However, the process variation of magnetic tunnel junction (MTJ) and access transistor poses serious challenge to sensing. Nondestructive sensing suffers from reference resistance variation whereas destructive sensing suffers from failures due to unoptimized selection of data and reference currents. We propose a novel slope detection technique to exploit MTJ resistance switching from high to low state using low-overhead sample-and-hold circuit. The proposed sensing technique is destructive in nature and can be combined with double sampling for improved robustness. Simulation results reveal <;0.12% failure under process variation using single sampling (at 0.2% area overhead) and <;0.08% failures with double sampling (at 0.6% area overhead). The overall sense time is found to be 6.8ns. Seyedhamidreza Motaman, Swaroop Ghosh, Jaydeep P. Kulkarni |
ISLPED | 2 |
| 2015 | Emerging Trends in Design and Applications of Memory-Based Computing and Content-Addressable MemoriesabstractContent-addressable memory (CAM) and associative memory (AM) are types of storage structures that allow searching by content as opposed to searching by address. Such memory structures are used in diverse applications ranging from branch prediction in a processor to complex pattern recognition. In this paper, we review the emerging challenges and opportunities in implementing different varieties of CAM/AM structures. Beyond-CMOS silicon and nonsilicon memory technologies hold significant promise in implementing dense, fast, and energy-efficient CAM/AM structures. We describe circuit/architecture level implementations of CAM/AM using these technologies, as well as novel applications in different domains, including informatics, text analytics, data mining, and reconfigurable computing platforms. Robert Karam, Ruchir Puri, Swaroop Ghosh, Swarup Bhunia |
Proc. IEEE | 3 |
| 2014 | Modeling and Analysis of Domain Wall Dynamics for Robust and Low-Power Embedded MemoryabstractNon-volatile memories are gaining significant attention for embedded cache application due to low standby power and excellent retention. Domain wall memory (DWM) is one possible candidate due to its ability to store multiple bits/cell in order to break the density barrier. Additionally, it provides low standby power, fast access time, good endurance and good retention. In this paper, we provide a physics-based model of domain wall that comprehends process variations (PV) and Joule heating. The proposed model has been used for circuit simulation. We also propose techniques to mitigate the impact of variability and Joule heating while enabling low-power and high frequency operation. Anirudh Iyengar, Swaroop Ghosh |
DAC | 2 |
| 2014 | Simultaneous Sizing, Reference Voltage and Clamp Voltage Biasing for Robustness, Self-Calibration and Testability of STTRAM ArraysabstractSpin-Torque Transfer Random Access Memory (STTRAM) is a promising technology for high density on-chip cache due to low standby power and high speed. However, the limited sense-margin poses challenge towards applicability of STTRAM. Reference voltage (Vref) biasing and clamp voltage (Vclamp) biasing are possible techniques to balance '0' and '1' sense margins for improved robustness. In this paper, we show that Vref and Vclamp biasing are more effective when employed on appropriately sized sense circuit. Our investigation also reveals that these two techniques can be used for meeting two different objectives namely, self-calibration and improved testability. We show that the proposed sizing and biasing technique can improve both robustness and testability while sacrificing minimum sense margin compared to conventional sense circuit that is designed to provide best sense margin. Seyedhamidreza Motaman, Swaroop Ghosh |
DAC | 2 |
| 2014 | Design and analysis of robust and wide operating low-power level-shifter for embedded dynamic random access memoryabstractLevel shifters (LS) are crucial components in low power design where the die is segregated in multiple voltage domains. LS are used at the voltage domain interfaces to mitigate sneak path current. Another important application of LS is in high voltage drivers for designs where voltage boosting is needed for performance and functionality. We explore one such application in embedded Dynamic Random Access Memories (eDRAM) where LS is employed in the wordline path. Our investigation reveals that leakage power of LS can pose a serious threat by lowering the wordline voltage and subsequently affecting the speed and retention time of eDRAM. Furthermore the delay of LS under worse case process corners can cause functional discrepancies. We propose low-power pulsed-LS with supply gating to circumvent these issues. Our analysis indicate that pulsed-LS can improve the worst case speed from 2.7%-43%. We also propose power-gating for LSs to improve the retention time and bandwidth with minimal power and area overhead. Kenneth Ramclam, Swaroop Ghosh |
ACM Great Lakes Symposium on VLSI | 2 |
| 2014 | Synergistic circuit and system design for energy-efficient and robust domain wall cachesabstractNon-volatile memories are gaining significant attention for embedded cache application due to their low standby power and excellent retention. Domain wall memory (DWM) is one possible candidate due to its ability to store multiple bits per cell in order to break the density barrier. Additionally, it provides low standby power, fast access time, good endurance and retention. However, it suffers from poor write latency, shift latency, shift power and write power. DWM is sequential in nature and latency of read/write operations depends on the offset of the bit from the read/write head. This paper investigates the circuit design challenges such as bitcell layout, head positioning, utilization factor of the nanowire, shift power, shift latency and provides solutions to deal with these issues. A synergistic system is proposed by combining circuit techniques such as merged read/write heads (for compact layout), flipped-bitcell and shift gating (for shift power optimization), wordline (WL) strapping (for access latency), shift circuit design with micro-architectural techniques such as segmented cache to realize energy-efficient and robust DWM cache. Simulations show 3-33% better performance and 1.25X-14.4X better power over a wide range of PARSEC benchmarks. Seyedhamidreza Motaman, Anirudh Iyengar, Swaroop Ghosh |
ISLPED | 3 |
| 2013 | Path to a TeraByte of on-chip memory for petabit per second bandwidth with < 5watts of powerabstractWe propose a path to achieve an ambitious target that has never been tried before: a terabyte of on-chip memory for petabit/second of bandwidth with < 5W of power. Conventional methodology of on-chip memory design is bottom up where the choice of bitcell topology and associated peripherals are predetermined. The resulting memory is sub-optimal and often suffers from high power and poor bandwidth. We approach this problem from top down where the capacity, bandwidth and power specifications guide the choice of bitcell. Our evaluation shows that domain wall memory (DWM) can be a potential technology that can meet TB capacity and Pb/s bandwidth with shoestring power budget. Swaroop Ghosh |
DAC | 1 |
| 2011 | Integrated Design & Test: Conquering the Conflicting Requirements of Low-Power, Variation-Tolerance and Test CostabstractDesign objectives of robustness and low-power usually do not go hand in hand with the test objectives of maximum test coverage and minimum test cost. Low power robust design techniques such as dual-Vth, dual-VDD, or adaptive body biasing have negative impact on the associated test cost. Similarly, test techniques like enhanced scan have large overhead in terms of area and power. In this paper, we try to mitigate the conflicting design and test requirements using an integrated approach to design and test that utilizes the existing low power and error resilient design techniques and augments them to improve test coverage and cost. Simulation results on an example 8×8 Wallace tree multiplier in 90nm technology node show 20% reduction in operating power, 60% reduction in test power and 99% reduction in critical paths while at the same time improving the yield from 96% to 100%, compared to existing design and test methodologies. All this comes at the cost of a marginal increase in area (7.8%). Ashish Goel, Swaroop Ghosh, Mesut Meterelliyoz, Jeff Parkhurst, Kaushik Roy 0001 |
Asian Test Symposium | 2 |
| 2011 | Novel Low Overhead Post-Silicon Self-Correction Technique for Parallel Prefix Adders Using Selective Redundancy and Adaptive ClockingabstractIn this paper, we present a post-silicon self-correction technique to leverage the redundancy present in parallel prefix adders (PPA). Our technique is based on the fact that a set of carries in PPAs can be made mutually exclusive. Therefore, defects in a set of bits can only corrupt the corresponding set of Sum outputs whereas the remaining Sums are computed correctly. To efficiently utilize the above property of PPAs in presence of defects, we perform addition in multiple clock cycles. In cycle-1, one of the correct set of bits are computed and stored at the output registers. In the subsequent cycles, the operands are shifted by one bit at a time and the remaining sets of bits are recovered. This allows us to compute the correct output at the cost of throughput degradation and minor area and delay overhead while maintaining high frequency and yield. Finally, the proposed technique is used in a superscalar processor, whereby the self-correcting adder is assigned lower priority than fault-free adders to reduce the overall throughput degradation. Swaroop Ghosh, Kaushik Roy 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2010 | Parameter Variation Tolerance and Error Resiliency: New Design Paradigm for the Nanoscale EraabstractVariations in process parameters affect the operation of integrated circuits (ICs) and pose a significant threat to the continued scaling of transistor dimensions. Such parameter variations, however, tend to affect logic and memory circuits in different ways. In logic, this fluctuation in device geometries might prevent them from meeting timing and power constraints and degrade the parametric yield. Memories, on the other hand, experience stability failures on account of such variations. Process limitations are not exhibited as physical disparities only; transistors experience temporal device degradation as well. Such issues are expected to further worsen with technology scaling. Resolving the problems of traditional Si-based technologies by employing non-Si alternatives may not present a viable solution; the non-Si miniature devices are expected to suffer the ill-effects of process/temporal variations as well. To circumvent these nonidealities, there is a need to design ICs that can adapt themselves to operate correctly under the presence of such inconsistencies. In this paper, we first provide an overview of the process variations and time-dependent degradation mechanisms. Next, we discuss the emerging paradigm of variation-tolerant adaptive design for both logic and memories. Interestingly, these resiliency techniques transcend several design abstraction levels-we present circuit and microarchitectural techniques to perform reliable computations in an unreliable environment. Swaroop Ghosh, Kaushik Roy 0001 |
Proc. IEEE | 1 |
| 2010 | Voltage Scalable High-Speed Robust Hybrid Arithmetic Units Using Adaptive ClockingabstractIn this paper, we explore various arithmetic units for possible use in high-speed, high-yield ALUs operated at scaled supply voltage with adaptive clock stretching. We demonstrate that careful logic optimization of the existing arithmetic units (to create hybrid units) indeed make them further amenable to supply voltage scaling. Such hybrid units result from mixing right amount of fast arithmetic into the slower ones. Simulations on differenthybridadder and multipliers in BPTM 70 nm technology show 18%-50% improvements in power compared to standard adders with only 2%-8% increase in die-area at iso-yield. These optimized datapath units can be used to construct voltage scalable robust ALUs that can operate at high clock frequency with minimal performance degradation due to occasional clock stretching. Swaroop Ghosh, Debabrata Mohapatra, Georgios Karakonstantis, Kaushik Roy 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2010 | Trifecta: A Nonspeculative Scheme to Exploit Common, Data-Dependent Subcritical PathsabstractPipelined processor cores are conventionally designed to accommodate the critical paths in the critical pipeline stage(s) in a single clock cycle, to ensure correctness. Such conservative design is wasteful in many cases since critical paths are rarely exercised. Thus, configuring the pipeline to operate correctly for rarely used critical paths targets the uncommon case instead of optimizing for the common case. In this study, we describe Trifecta-an architectural technique that completes common-case, subcritical path operations in a single cycle but uses two cycles when the critical path is exercised. This increases slack for both single-and two-cycle operations and offers a unique advantage under process variation. In contrast with existing mechanisms that trade power or performance for yield, Trifecta improves the yield while preserving performance and power. We applied this technique to the critical pipeline stages of a superscalar out-of-order (OoO) and a single issue in-order processor, namely instruction issue and execute, respectively. Our experiments show that the rare two-cycle operations result in a small decrease (5% for integer and 2% for floating-point benchmarks of SPEC2000) in instructions per cycle. However, the increased delay slack causes an improvement in yield-adjusted-throughput by 20% (12.7%) for an in-order (InO) processor configuration. Patrick Ndai, Nauman Rafique, Mithuna Thottethodi, Swaroop Ghosh, Swarup Bhunia, Kaushik Roy 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2008 | Exploring high-speed low-power hybrid arithmetic units at scaled supply and adaptive clock-stretchingabstractMeeting power and performance requirement is a challenging task in high speed ALUs. Supply voltage scaling is promising because it reduces both switching and active power but it also degrades robustness. Recently, researchers have proposed novel design technique for linear time complexity adders that maintain high yield and high clock frequency even at scaled supply voltage. The idea is based on the fact that the critical paths of arithmetic units are exercised rarely. The technique (a) predicts the set of critical paths, (b) reduces the supply voltage to operate non-critical paths at rated frequency, and; (c) avoids possible delay failures in the critical paths by dynamically stretching the clock period (to say, two-cycles assuming all standard operations are single-cycle), when they are activated. This allows circuits to operate at scaled supply with minimal performance degradation. The off-critical paths operate in single clock cycle while critical paths are operated in stretched clock period. Different classes of adders may benefit differently using such technique. For example, ripple carry adders can reap the benefits more effectively than say, tree adders (balanced paths). However, logic modification may ease the application of supply voltage scaling. In this paper, we explore various arithmetic units for possible use in high speed, high yield ALU design at scaled supply voltage with variable latency operation. We demonstrate that careful logic optimization of the existing arithmetic units indeed make them further suitable for supply voltage scaling with tolerable area overhead Simulation results on different adder and multiplier topologies in BPTM 70nm technology show 18-60% extra improvement in power with only 2-8% increase in die-area at iso-yield We also extend our studies to design low power and high yield multipliers. These optimized low power datapath units can be used to construct low power and robust ALU that can operate at high clock frequency with minimal performance degradation due to occasional clock stretching. Swaroop Ghosh, Kaushik Roy 0001 |
ASP-DAC | 1 |
| 2008 | A Novel Low Overhead Fault Tolerant Kogge-Stone Adder Using Adaptive ClockingabstractAs the feature size of transistors gets smaller, fabricating them becomes challenging. Manufacturing process follows various corrective design-for-manufacturing (DFM) steps to avoid shorts/opens/bridges. However, it is not possible to completely eliminate the possibility of such defects. If spare units are not present to replace the defective parts, then such failures cause yield loss. In this paper, we present a fault tolerant technique to leverage the redundancy present in high speed regular circuits such as Kogge-Stone adder (KSA). Due to its regularity and speed, KSA is widely used in ALU design. In KSA, the carries are computed fast by computing them in parallel. Our technique is based on the fact that even and odd carries are mutually exclusive. Therefore, defect in even bit can only corrupt the even Sum outputs whereas the odd Sums are computed correctly (and vice versa). To efficiently utilize the above property of KSA in presence of defects, we perform addition in two-clock cycles. In cycle-1, one of the correct set of bits (even or odd) are computed and stored at output registers. In cycle-2, the operands are shifted by one bit and the remaining sets of bits (odd or even) are computed and stored. This allows us to tolerate the defect at the cost of throughput degradation while maintaining high frequency and yield. The proposed technique can tolerate any number of faults as long as they are confined to either even or odd bits (but not in both). Further, this technique is applicable for any type of fault model (stuck-at, bridging, complete opens/shorts). We performed simulations on 64-bit KSA using 180 nm devices. The results indicate that the proposed technique incur less that 1 % area overhead. Note that there is very little throughput degradation (<0.3%) for the fault-free adders. The proposed technique utilizes the existing scan flip-flops for storage and shifting operation to minimize the area/performance overhead. Finally, the proposed technique is used in a superscalar processor, whereby the faulty adder is assigned lower priority than fault-free adders to reduce the overall throughput degradation. Experiments performed using Simplescalar for a superscalar pipeline (with four integer adders) show throughput degradation of 0.5% in the presence of a single defective adder. Swaroop Ghosh, Patrick Ndai, Kaushik Roy 0001 |
DATE | 1 |
| 2008 | O2C: occasional two-cycle operations for dynamic thermal management in high performance in-order microprocessorsabstractIn this paper, we propose O2C, a novel non-speculative adaptive thermal management technique that reduces the temperature during die-overheating using supply voltage scaling, while maintaining the rated clock frequency. This is accomplished by (a) scaling down the supply voltage, (b) isolating and predicting the set of critical paths, (c) ensuring (by design) that they are activated rarely, and (d) getting around occasional delay failures (at reduced voltage during die-overheating) in these paths by two-cycle operations (assuming all standard operations are single-cycle). Two-cycle operation is achieved by stalling the pipeline for extra clock cycles whenever the set of critical paths are activated. The rare two-cycle operation results in a small decrease in IPC (instructions per cycle). Since called maintains the rated clock frequency and does not require pipeline stalling during supply voltage ramp-up/ramp-down, it achieves high throughput in a thermally constrained environment. We applied called to the integer execution units of an in-order superscalar pipeline. Standard full-chip Dynamic Voltage-Frequency Scaling (DVFS) is very effective in bringing down the temperature, however; it is associated with large throughput loss due to pipeline stalling and slow operating frequency during thermal management. We integrated O2C with standard (called O2Cα) to demonstrate that it can act as a first step before full-scale thermal management is required. Our simulations indeed reveal that called O2Cα policy can avoid the requirement of full-scale DVFS during execution of programs. Swaroop Ghosh, Jung Hwan Choi, Patrick Ndai, Kaushik Roy 0001 |
ISLPED | 1 |
| 2008 | An alternate design paradigm for low-power, low-cost, testable hybrid systems using scaled LTPS TFTsabstractThis article presents a holistic hybrid design methodology for low-power, low-cost, testable digital designs using low-temperature polycrystalline-silicon thin-film transistors (LTPS TFTs). An alternate scaling rule under low thermal budget (due to flexible substrate) is developed to improve the performance of TFTs in the presence of process variation. We demonstrate that LTPS TFTs can be further optimized for ultralow-power subthreshold operation with performances comparable to contemporary single-crystal silicon-on-insulator (c-Si SOI) devices after process optimization. The optimized LTPS TFTs with high current drivability and less variability can comprise a promising low-cost option to augment Si CMOS technology, opening up a plethora of new hybrid 3D applications. We illustrate one such application: IC testing. Testing of complex VLSI systems is a prime concern due to design cost of DFT circuits, area/delay overheads, and poor test confidence. To harness the benefits of TFT technology, a novel low-power, process-tolerant, generic, and reconfigurable test structure designed using LTPS TFTs is proposed to reduce the test cost, as well as to improve diagnosability and verifiability, of complex VLSI systems. Due to proper optimization of TFT devices, the proposed test structure consumes low power but operates with reasonable performance. Furthermore, the test circuits do not consume any silicon area because they can be integrated on-chip using 3D technology. Since the test architecture is reconfigurable, this eliminates the need to redesign built-in-self-test (BIST) components that may vary from one processor generation to another. We have developed test structures using 200nm TFT devices and evaluated them on designs implemented in 130nm bulk CMOS. For circuit simulations, we have developed a SPICE-compatible model for TFT devices. The BIST components designed using the test structures operate at 0.8--4.3 GHz (compared to 8.2 GHz in bulk CMOS) with low power consumption. The enhanced scan cells partially implemented in TFT (3D hybrid design) consume ∼24% less power and ∼15--20% less area of Si die compared to conventional bulk-Si design (2D planar design), with minimal delay overhead. Jing Jane Li, Aditya Bansal, Swaroop Ghosh, Kaushik Roy 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2007 | Low-overhead circuit synthesis for temperature adaptation using dynamic voltage scheduling
Swaroop Ghosh, Swarup Bhunia, Kaushik Roy 0001 |
DATE | 1 |
| 2007 | Tolerance to Small Delay Defects by Adaptive Clock StretchingabstractBridging defects typically manifest themselves as increased path delays instead of stuck-at failures. On the other hand, parametric variations (both inter- and intra-die) increase the spread of the circuit delay. Low power design techniques such as voltage scaling, dual-Vth etc. deteriorate the delay spread further. These mechanisms for delay variations in nanoscaled technologies significantly affect the parametric yield. We propose a new design methodology to tolerate subtle delay failures that arise both due to manufacturing defects and parameter fluctuations. We synthesize the circuit to (a) isolate and predict the critical paths of a circuit; (b) create timing slack between critical and off-critical paths and ensure that they are activated rarely; and, (c) avoid the delay failures in these paths by adaptively stretching the clock period. Since critical paths are the most sensitive section of the circuit in terms of delay defects, we ensure fault-free operation by isolating them and providing extra computation time by predicting their activation. This allows us to achieve the required yield with small performance penalty (due to occasional clock stretching under critical path activation). We present application of the proposed methodology for both linear and non-linear pipeline designs. We also suggest two possible circuit-level implementations of clock stretching using clock gating and handshaking, respectively. Simulations on MCNC benchmark circuits with BPTM 70 nm devices show that the proposed technique can achieve good yield by tolerating increased path delays (under variations and bridging defects of various sizes) with small overhead in performance and ~14% die-area compared to the conventional design. For performance analysis, we have implemented the proposed methodology in simplein-orderpipeline in Simplescalar. Simulation results on SPEC2000 benchmarks show less that 2% of IPC (instructions-per-cycle) degradation. Swaroop Ghosh, Patrick Ndai, Swarup Bhunia, Kaushik Roy 0001 |
IOLTS | 1 |
| 2007 | A generic and reconfigurable test paradigm using Low-cost integrated Poly-Si TFTsabstractIn this work, we propose a novel low power, process tolerant, generic and reconfigurable test structure to reduce the test cost, improve diagnosability and verifiability of complex VLSI systems. The test structure contains a variety of configurable design-for-test units designed with low cost Low Temperature Polycrystalline Silicon Thin Film Transistors (LTPS TFTs) that are fabricated on a separate substrate (e.g., polymer, glass etc). The proposed test circuits do not consume any silicon area because they can be integrated on the chip using 3-D technology. This reconfigurable test paradigm eliminates the need to re-design the BIST components that may vary from one processor generation to another. Jing Jane Li, Swaroop Ghosh, Kaushik Roy 0001 |
ITC | 2 |
| 2007 | CRISTA: A New Paradigm for Low-Power, Variation-Tolerant, and Adaptive Circuit Synthesis Using Critical Path IsolationabstractDesign considerations for robustness with respect to variations and low-power operations typically impose contradictory design requirements. Low-power design techniques such as voltage scaling, dual- , etc., can have a large negative impact on parametric yield. In this paper, we propose a novel paradigm for low-power variation-tolerant circuit design called critical path isolation for timing adaptiveness (CRISTA), which allows aggressive voltage scaling. The principal idea includes the following: 1) isolate and predict the set of possible paths that may become critical under process variations; 2) ensure that they are activated rarely; and 3) avoid possible delay failures in the critical paths by dynamically switching to two-cycle operation (assuming all standard operations are single cycle), when they are activated. This allows us to operate the circuit at reduced supply voltage while achieving the required yield. Simulation results on a set of benchmark circuits with Berkeley-predictive-technology-model [BPTM 70 nm: Berkeley predictive technology model] 70-nm devices that show an average of 60% improvement in power with small overhead in performance and 18% overhead in die area compared to conventional design. We also present two applications of the proposed methodology that include the following: 1) pipeline design for low power and 2) temperature-adaptive circuit design. Swaroop Ghosh, Swarup Bhunia, Kaushik Roy 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2007 | Low-Power and testable circuit synthesis using Shannon decompositionabstractStructural transformation of a design to enhance its testability while satisfying design constraints on power and performance can result in improved test cost and test confidence. In this article, we analyze the testability in a new style of logic design based on Shannon's decomposition and supply gating . We observe that the tree structure of a logic circuit due to Shannon's decomposition makes it intrinsically more testable than a conventionally synthesized circuit, while at the same time providing an improvement in active power. We have analyzed four different aspects of the testability of a circuit: a) IDDQ test sensitivity, b) test power during scan-based testing, c) test length (for both ATPG-generated deterministic and random patterns), and d) noise immunity. Simulation results on a set of MCNC benchmarks show promising results on all these aspects (an average improvement of 94% in IDDQ sensitivity, 50% in test power, 19% (21%) in test length for deterministic (random) patterns, and 50% in coupling noise immunity). We have also demonstrated that the new logic structure can improve parametric yield (6% on average) of a circuit under process variations when considering a bound on circuit leakage. Swaroop Ghosh, Swarup Bhunia, Kaushik Roy 0001 |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2006 | Self-calibration technique for reduction of hold failures in low-power nano-scaled SRAMabstractIncreasing source voltage (Source-Biasing) is an efficient technique for reducing gate and sub-threshold leakage of SRAM arrays. However, due to process variation, a higher source voltage can significantly increase data flipping in standby mode (Hold Failures) resulting in faulty memories. This imposes serious concerns in reducing standby power with source-bias. In this paper, we analyze the effect of source bias on hold failures under both inter-die and intra-die variations. We propose a self-calibrating SRAM for aggressively reducing leakage while maintaining the hold failures under control. Swaroop Ghosh, Saibal Mukhopadhyay, Keejong Kim, Kaushik Roy 0001 |
DAC | 1 |
| 2006 | A new paradigm for low-power, variation-tolerant circuit synthesis using critical path isolationabstractDesign considerations for robustness with respect to variations and low power operations typically impose contradictory design requirements. Low power design techniques such as voltage scaling, dual-Vth etc. can have a large negative impact on parametric yield. In this paper, we propose a novel paradigm for low-power variationtolerant circuit design, which allows aggressive voltage scaling. The principal idea is to (a) isolate and predict the set of possible paths that may become critical under process variations, (b) ensure that they are activated rarely, and (c) avoid possible delay failures in the critical paths by dynamically switching to two-cycle operation (assuming all standard operations are single cycle), when they are activated. This allows us to operate the circuit at reduced supply voltage while achieving the required yield. Simulation results on a set of benchmark circuits at 70nm process technology show average power reduction of 60% with less than 10% performance overhead and 18% overhead in die-area compared to conventional synthesis. Application of the proposed methodology to pipelined design is also investigated. Swaroop Ghosh, Swarup Bhunia, Kaushik Roy 0001 |
ICCAD | 1 |
| 2006 | Delay Fault Localization in Test-Per-Scan BIST Using Built-In Delay SensorabstractDelay failures are becoming a dominant failure mechanism in nanometer technologies. Diagnosis of such failures is important to ensure yield and robustness of the design. However, the increasing circuit size limits the granularity of diagnosis, resulting in large suspect fault list. In this paper, we present a methodology for improving delay fault localization in test-per-scan BIST using on-die delay sensing at selective test points. It is demonstrated that the proposed technique can improve the resolution of fault localization for both transition and segment delay fault models. Experimental results for a set of ISCAS89 benchmarks show up to 49% (82%) average improvement in fault localization for transition (segment) delay fault models. The area overhead due to delay sensing hardware have been limited to 4% Swaroop Ghosh, Swarup Bhunia, Arijit Raychowdhury, Kaushik Roy 0001 |
IOLTS | 1 |
| 2006 | A Novel Delay Fault Testing Methodology Using Low-Overhead Built-In Delay SensorabstractA novel integrated approach for delay-fault testing in external (automatic-test-equipment-based) and test-per-scan built-in self-test (BIST) using on-die delay sensing and test point insertion is proposed. A robust, low-overhead, and process-tolerant on-chip delay-sensing circuit is designed for this purpose. An algorithm is also developed to judiciously insert delay-sensor circuits at the internal nodes of logic blocks for improving delay-fault coverage with little or no impact on the critical-path delay. The proposed delay-fault testing approach is verified for transition- and segment-delay-fault models. Experimental results for external testing (BIST) show up to 31% (30%) improvement in fault coverage and up to 67.5% (85.5%) reduction in test length for transition faults. An increase in the number of robustly detectable critical-path segments of up to 54% and a reduction in test length for the segment-delay-fault model of up to 76% were also observed. The delay and area overhead due to insertion of the delay-sensing hardware have been limited to 2% and 4%, respectively Swaroop Ghosh, Swarup Bhunia, Arijit Raychowdhury, Kaushik Roy 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2005 | Shannon Expansion Based Supply-Gated Logic for Improved Power and TestabilityabstractStructural transformation of a design to enhance its testability while satisfying design constraints on power and performance, can result in improved test cost and test confidence. In this paper, we analyze the testability in a new style of logic design based on Shannon’s decomposition and supply gating. We observe that tree structure of a logic circuit due to Shannon’s decomposition makes it intrinsically more testable than conventionally synthesized circuit, while at the same time entailing an improvement in active power. We have analyzed three different aspects of testability of a circuit: a) IDDQ test sensitivity b) test power during scan-based testing, and c) test length (for both ATPG-generated deterministic and random patterns). Simulation results on a set of MCNC benchmarks show promising results on all the above aspects. We have also demonstrated that the new logic structure can improve parametric yield of a circuit under process variations when considering a bound on circuit leakage. Swaroop Ghosh, Swarup Bhunia, Kaushik Roy 0001 |
Asian Test Symposium | 1 |
| 2005 | A novel delay fault testing methodology using on-chip low-overhead delay measurement hardware at strategic probe pointsabstractWe propose a delay fault testing methodology using on-chip delay measurement hardware. We have designed a process-tolerant, low-overhead delay measurement hardware and developed an algorithm to judiciously insert the hardware at internal nodes of logic blocks. Experimental results for a set of ISCAS89 benchmarks show up to 16.9% improvement in transition fault coverage and up to 10.5% increase in the number of detected faults for segment delay fault model, with fixed test length. The reduction in test length is up to 59% for transition fault, with fixed target coverage. The delay and area overhead due to additional DFT logic is limited to 2% and 4% respectively. Arijit Raychowdhury, Swaroop Ghosh, Swarup Bhunia, Debjyoti Ghosh, Kaushik Roy 0001 |
ETS | 2 |
| 2005 | A Novel On-Chip Delay Measurement Hardware for Efficient Speed-BinningabstractWith the aggressive scaling of the CMOS technology parametric variation of the transistor threshold voltage causes significant spread in the circuit delay as well as leakage spectrum. Consequently, speed binning of the high performance VLSI chips is essential and it costs significant amount of test application time. Further, the knowledge of the actual delay in the critical path of the circuit enables efficient use of typical low power methodologies e.g., voltage scaling, adaptive body biasing etc. In this paper, the authors have proposed a novel on-chip, low overhead and process tolerant delay measurement circuit which can estimate the critical path delay in a single clock period. This has the advantage of efficient on-chip speed binning. Arijit Raychowdhury, Swaroop Ghosh, Kaushik Roy 0001 |
IOLTS | 2 |
| 2004 | Scan Chain Fault Identification Using Weight-Based Codes for SoC CircuitsabstractRecently, it has been observed that embedded cores in a high-speed SoC circuit have the problem of broken scan chains that cannot shift properly. Also, scan chain intermittent faults caused by hold-time violations and crosstalk noises are pervasive. In this research, an efficient method is proposed to identify the faulty scan chain(s) at the core level. That is, the core where the scan chain is defective can be identified, even if the scan chain is broken. The result can be used to tune up the fabrication process or to guide the fine-grained scan cell identification process. Here, weight-based m-out-of-n codes, which can generate a large number of codewords, with small hardware overhead and high fault detection capability are used to generate the scan chain diagnostic patterns for permanent (and possibly intermittent) faults. An efficient codeword generation method is proposed to maximize the number of codewords, minimize the aliasing probabilities and test application cost. The idea of multiple m-out-of-n codes is also proposed to guarantee that sufficient number of codewords are generated to perturb the scan chains and the associated combinational circuits. Simulation results demonstrate the feasibility of the proposed method. Swaroop Ghosh, K. W. Lai, Wen-Ben Jone, Shih-Chieh Chang 0001 |
Asian Test Symposium | 1 |
| 2003 | Embedded core test generation using broadcast test architecture and netlist scramblingabstractIn this work, based on the concept of test pattern broadcasting, we propose a new core-based testing method which gives core users the maximum level of test freedom. Instead of only using the test patterns delivered by core providers, core users are allowed to broadcast their own test patterns to the cores of a SoC (system on chip) design for parallel scan testing. The fault coverage of each core test, using test patterns developed by any core user, can be evaluated by an enhanced version of a traditional fault simulator. The netlist of each core is scrambled before it is delivered to core users, thus the netlist will not be revealed. The enhanced fault simulator of a core has the capabilities of decoding the scrambled netlist, and performing fault simulation for the test patterns provided by each of the core users. For each core, both random test patterns (applied by a core user), and golden test patterns (delivered by the core provider) jointly achieve high and flexible fault coverage requirements. The enhanced logic simulator of each core can also decrypt the scrambled netlist, and perform logic simulation with the objective of generating fault-free test responses for signature analysis (for example). The proposed method has the advantages of minimizing the number of scan pins, reducing the test application time, and achieving the maximum level of test quality control by core users. Simulation results demonstrate the feasibility of this method. J. H. Jiang, Wen-Ben Jone, Shih-Chieh Chang 0001, Swaroop Ghosh |
IEEE Trans. Reliab. | 4 |