Anupam Chattopadhyay

dblp:99/4535 · DBLP profile ↗
← Back
145ranked-venue papers
15as first author
51since 2021 · last 2026
0000-0002-8818-6983ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 102 · 11 first-author · 31 since 2021Security and privacy · 20 · 1 first-author · 11 since 2021Software engineering, systems software and programming languages · 19 · 4 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 1 since 2021Theory of computation · 8 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 7 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Are Language Models Any Good at Density Modeling?
abstract
Large Language Models (LLMs) surprised the world with their ability to mimic humans in writing and are starting to be used as simulations of human writers for various kinds of linguistic analyses. However, these analyses rest on the belief that LLMs are good density models that accurately capture the underlying probability distribution of the language. In this paper, we question this basic assumption and try to evaluate language models on their density modelling capabilities. Since a ground truth does not exist for the probability distribution of any natural language, we come up with a synthetic language made up of decimal numbers written in words in English. We train language models from scratch on various probability distributions over this synthetic language and compare the distributions learned by the models with the original distributions. Experiments show that language models can learn underlying probability distributions across a wide range of cases, but they fail when those distributions depend on deep semantic properties of numbers that cannot be inferred from syntactic patterns. Additionally, we observed a strong bias in the models towards numbers that frequently occur as substrings within other numbers. This suggests that such a bias possibly exists in real-world natural language models as well, and negatively impacts downstream tasks and analyses that rely on model-generated probabilities.
Sriram Ranga, Sai Shashank Bedampeta, Rui Mao 0010, Anupam Chattopadhyay
AAAI4
2026 Mimi: Dynamically Secure Multi-Keyword Retrieval Scheme With Two-Factor Verification
abstract
Existing privacy-preserving multi-keyword retrieval schemes often suffer from reduced retrieval efficiency, lack robust verification mechanisms in dynamic environments, and are prone to symmetric key leakage issues. To address these shortcomings, we propose a dynamic and secure multi-keyword search scheme with a two-factor verification mechanism, named Mimi. Specifically, Mimi first constructs a dynamic verification tree structure to accelerate the verification of the correctness of returned results. Second, it builds an encrypted searchable index that supports sub-linear search time complexity. Third, Mimi incorporates a secure symmetric key exchange protocol to protect the confidentiality of the symmetric key. Furthermore, Mimi supports multi-user search operations without increasing the index construction costs and accommodates dynamic updates to both user roles and data. Through comprehensive security analysis, we demonstrate that Mimi ensures the security of the encrypted searchable inverted index and maintains query indistinguishability for users. Empirical evaluations show that the Mimi scheme is efficient and effective.
Dong Li 0054, Anupam Chattopadhyay, Qianyu Li 0001, Jiahui Wu 0001, Qingguo Lü, Tao Xiang 0001, Xiaofeng Liao 0001
IEEE Trans. Dependable Secur. Comput.2
2026 Intelligent Penetration Testing Through Integrated Knowledge Graph and Historical Decision Enhancement
abstract
Penetration Testing (PT), a key network security assessment technique that simulates real cyber attacks to identify vulnerabilities, is traditionally manual and expert-dependent, leading to low efficiency and high costs. Automating and intelligentizing PT has thus become a critical research focus, yet current technologies face two core challenges: lack of standardized, reusable simulated network scenarios (hindering unified experiments and result comparison) and intelligent models' failure to integrate historical decision experience or utilize attack chain temporal correlations (restricting adaptability). To address these, this study proposes an intelligent PT method integrating knowledge graph-driven automated scenario construction and historical decision enhancement. Two innovations are introduced: a network knowledge graph-based mechanism to generate standardized, real-characteristic testing environments; and a historical decision enhancement scheme with a collaborative state temporal processing and action filtering architecture. Experimental results show the method reduces average iterations by 69%, eliminates redundant executions, and enhances decision rationality, offering a new path for automated PT advancement.
Qianyu Li 0001, Anupam Chattopadhyay, Cheng Tu, Fan Shi 0003, Min Zhang 0054, Zulie Pan
IEEE Trans. Dependable Secur. Comput.3
2026 FSAT: A Faster Secure Convolutional Neural Network Inference Framework With Adversarial Training in Resource-Constrained Scenarios
abstract
Existing CNN inference frameworks based on FHE often suffer from reduced efficiency and accuracy due to the polynomial approximation of activation functions, and they lack effective mechanisms to prevent sensitive information leakage during the final classification stage. To address these limitations, we propose FSAT, a fast and secure inference framework enhanced with adversarial training. Specifically, FSAT employs a private CNN model architecture, where linear layers are computed through an optimized homomorphic ciphertext convolution operation, while non-linear layer operations are efficiently realized using a secure searchable index and an encrypted look-up table, which replace polynomial activation approximations and significantly improve inference accuracy and latency performance. To further mitigate information leakage, we introduce a dual-constraint adversarial training scheme that makes it substantially more difficult for an adversary to infer sensitive attributes of the input data. Experimental results demonstrate that FSAT achieves high inference accuracy and efficiency while substantially reducing the risk of sensitive data leakage.
Dong Li 0054, Anupam Chattopadhyay, Qingguo Lü, Jiahui Wu 0001, Tao Xiang 0001, Xiaofeng Liao 0001
IEEE Trans. Inf. Forensics Secur.2
2026 Tensor Manipulation Unit (TMU): Reconfigurable, Near-Memory Tensor Manipulation for High-Throughput AI SoC
abstract
While recent advances in AI SoC design have focused heavily on accelerating tensor computation, the equally critical task of tensor manipulation (TM)—centered on high-volume data movement with minimal computation—remains underexplored. This work addresses that gap by introducing the TM unit (TMU): a reconfigurable, near-memory hardware block designed to execute data-movement-intensive (DMI) operators efficiently. The TMU manipulates long datastreams in a memory-to-memory fashion using a RISC-inspired execution model and a unified addressing abstraction, enabling support for both a wide range of coarse- and fine-grained tensor transformations. The proposed architecture integrates the TMU alongside a TPU within a high-throughput AI system-on-chip (SoC), leveraging double buffering and output forwarding to improve pipeline utilization. The TMU, synthesized under the SMIC 40-nm standard cell library, occupies only$0.019~\mathrm {\text {mm}^{2}}$while supporting over 10 representative DMI operators. Benchmarking shows that the TMU alone achieves up to$82.42\times $and$11.06\times $operator-level latency reduction over ARM A72 and NVIDIA Jetson TX2, respectively. When integrated with the in-house TPU, the complete system achieves a 22.89% reduction in end-to-end inference latency, demonstrating the effectiveness in reducing inference latency and the scalability of the TMU architecture across diverse tensor operators.
Weiyu Zhou, Zheng Wang 0027, Chao Chen 0022, Yongkui Yang, Zhuoyu Wu, Anupam Chattopadhyay
IEEE Trans. Very Large Scale Integr. Syst.7
2025 POSTER: Stealthy SWAP-Based Side-Channel Attack on Multi-Tenant Quantum Cloud Systems
abstract
The rapid advancement of quantum computing has led to an increased use of cloud-based quantum devices, raising concerns over the privacy and security of computations on multi-tenant quantum systems. Previous studies have explored crosstalk in these environments, but several questions remain regarding its causes, countermeasures, and applicability. In this work, we revisit the crosstalk effect, tracing it to the SWAP path between qubits, demonstrating its effectiveness even over long distances. Specifically, we introduce a SWAP-based side-channel attack in both active and passive modes, validated on real IBM quantum devices. In the active attack, a single CNOT gate disrupts Grover’s Algorithm from afar with an 81.62% reduction in output accuracy. The passive attack, using a circuit just 6.25% of the victim’s size, achieves 100% accuracy in predicting circuit attributes in Simon’s Algorithm. These findings challenge the existing defense strategy of maximizing topological distance between circuits, demonstrating that attackers can still extract sensitive information or manipulate results remotely.
Wei Jie Bryan Lee, Suman Dutta 0001, Walid El Maouaki, Anupam Chattopadhyay
AsiaCCS5
2025 An Efficient Circuit Synthesis Framework for TFHE via Convex Sub-graph Optimization
Ayantika Chatterjee, Anupam Chattopadhyay, Debdeep Mukhopadhyay
AsiaCCS3
2025 Reducing T-Depth and T-Count in Quantum Multiplication Using Compressor Primitives
abstract
Optimization of quantum multiplication is a critical area of study due to its pivotal role in quantum algorithms such as Shor’s factorization. Every time a quantum multiplier is used, it repeatedly executes several key components to perform the multiplication. Most current works have focused on using components such as the basic half and full-adder designs, which have limited efficiency. In this paper, we demonstrate that by using a generalized (m:k) compressor-based Wallace Tree one can significantly improve efficiency; this method achieves reductions of up to 92.8% in T-Depth and 55.6% in T-Count while maintaining a competitive Qubit-Count through brute-force and dynamic programming optimization.
Suman Dutta 0001, Wei Jie Bryan Lee, Jerrie Feng, Anupam Chattopadhyay
ACM Great Lakes Symposium on VLSI6
2025 AttenPU: An Area Efficient Attention Processor with Reconfigurable FP8 Precision and Dataflow
abstract
Efficient numerical representation is crucial for deep learning accelerators, especially for large language models (LLMs). The 8-bit-floating-point (FP8) data representation achieves higher precision and fewer quantization efforts than integer, which has been proven inevitable in attention-based accelerators for LLMs. Therefore, area-efficient design techniques for FP8 play a central role in lowering LLMs chip’s budget. This paper presents AttenPU, which is built upon reconfigurable FP8 units and supports E4M3 for inference and E5M2 for training. Bidirectional dataflow is exploited to enable AttenPU to interact with FP32 coprocessor to reduce latency. The design achieves a low FP8-to-INT8 area ratio of 1.63×, an area efficiency of 193.5 GFLOPS/mm2, with an 87.05% reduction in the latency of RTX3090 GPU.
Qiawei Zheng, Zheng Wang 0027, Zhuoyu Wu, Zhihao Du, Chao Chen 0022, Yongkui Yang, Wenqi Fang, Anupam Chattopadhyay
ACM Great Lakes Symposium on VLSI10
2025 The Plagiarism Singularity Conjecture
abstract
Sriram Ranga, Rui Mao, Erik Cambria, Anupam Chattopadhyay. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Sriram Ranga, Rui Mao 0010, Erik Cambria, Anupam Chattopadhyay
NAACL (Long Papers)4
2025 Et tu, Brute? Side-Channel Assisted Chosen Ciphertext Attacks Using Valid Ciphertexts on HQC KEM
Thales B. Paiva, Prasanna Ravi, Dirmanto Jap, Shivam Bhasin, Anupam Chattopadhyay
PQCrypto (2)6
2025 AI Attacks AI: Recovering Neural Network Architecture from NVDLA Using AI-Assisted Side Channel Attack
abstract
During the last decade, there has been a stunning progress in the domain of Artificial Intelligence (AI) aided by highly trained Machine Learning (ML) models. Such models are valuable Intellectual Property (IP) and, therefore, have been subjected to various model recovery attacks. In this work, we study the vulnerabilities of commercial, open-source accelerator NVDLA and present the first successful model recovery attack. For this purpose, we used power and timing information from the side-channel leakage of convolutional neural networks (CNN) models to train CNN-based attack models. Utilizing these attack models, we demonstrate that even with a highly pipelined architecture, multiple parallel execution in the accelerator along with Linux OS running tasks in the background, recovery of number of layers, kernel sizes, output neurons and distinguishing different layers, is possible with very high accuracy. This is also the first work to show the impact of differences in hyperparameters on the power traces. Our solution is fully automated, AI-based, and portable to other hardware neural networks, thus presenting a greater threat toward IP protection. Using LeNet as the target victim model, we demonstrate an accuracy of more than 95% in recovering various parameters. This study presents a serious practical threat, in the form of side-channel attack, toward complex commercial architectures. Furthermore, we show that AI-guided attack significantly boosts the attacker capability.
Naina Gupta 0001, Arpan Jati, Anupam Chattopadhyay
ACM Trans. Embed. Comput. Syst.3
2025 Persistence of Backdoor-Based Watermarks for Neural Networks: A Comprehensive Evaluation
abstract
Deep neural networks (DNNs) have gained considerable traction in recent years due to the unparalleled results they gathered. However, the cost behind training such sophisticated models is resource-intensive, resulting in many to consider DNNs to be intellectual property (IP) to model owners. In this era of cloud computing, high-performance DNNs are often deployed all over the Internet so that people can access them publicly. As such, DNN watermarking schemes, especially backdoor-based watermarks, have been actively developed in recent years to preserve proprietary rights. Nonetheless, there lies much uncertainty on the robustness of existing backdoor watermark schemes, toward both adversarial attacks and unintended means such as fine-tuning neural network models. One reason for this is that no complete guarantee of robustness can be assured in the context of backdoor-based watermark. In this article, we extensively evaluate the persistence of recent backdoor-based watermarks within neural networks in the scenario of fine-tuning, and we propose/develop a novel data-driven idea to restore watermark after fine-tuning without exposing the trigger set. Our empirical results show that by solely introducing training data after fine-tuning, the watermark can be restored if model parameters do not shift dramatically during fine-tuning. Depending on the types of trigger samples used, trigger accuracy can be reinstated to up to 100%. This study further explores how the restoration process works using loss landscape visualization, as well as the idea of introducing training data in the fine-tuning stage to alleviate watermark vanishing.
Tu Anh Ngo, Chuan Song Heng, Nandish Chattopadhyay, Anupam Chattopadhyay
IEEE Trans. Neural Networks Learn. Syst.4
2025 Optimal Toffoli-Depth Quantum Adder
abstract
Efficient quantum arithmetic circuits are commonly found in numerous quantum algorithms of practical significance. To date, the logarithmic-depth quantum adders include a constant coefficient k ≥ 2 while achieving the Toffoli-Depth of k log n + 𝒪(1). In this work, 160 alternative compositions of the carry-propagation structure are comprehensively explored to determine the optimal depth structure for a quantum adder. By extensively studying these structures, it is shown that an exact Toffoli-Depth of log n + 𝒪(1) is achievable. This presents a reduction of Toffoli-Depth by almost 50% compared to the best known quantum adder circuits presented to date. We demonstrate a further possible design by incorporating a different expansion of propagate and generate forms, as well as an extension of the modular framework. Our article elaborates on these designs, supported by detailed theoretical analyses and simulation-based studies, firmly substantiating our claims of optimality within all possible configurations outlined in this work. The results also mirror similar improvements, recently reported in classical adder circuit complexity.
Ankit Mondal, Anupam Chattopadhyay
ACM Trans. Quantum Comput.3
2024 POSTER: MalaQ - A Malware Against Quantum Computer
abstract
Quantum computers are set to revolutionize multiple application domains, including financial portfolio optimization, drug discovery, supply chain optimization, and cryptography, by offering algorithmic speed-up over the best-known classical algorithms. This large-scale adoption will also make quantum computers a lucrative target for cyber-criminals. However, to date, there has been minimal experimental study on the possible attack surfaces and the severity of attacks on a real quantum computer. In this work, we introduce MalaQ, a malware specifically developed for quantum computers. MalaQ exploits the classical-computer frontend of a quantum system and causes extensive damages like performance degradation and even complete failure of the quantum circuits. In this paper, we discuss the design, implementation, and experiments using MalaQ in great detail, including drawing parallels with prior works on both classical and quantum cyber-attacks.
Alex Jin, Tarun Dutta, Manas Mukherjee, Anupam Chattopadhyay
AsiaCCS6
2024 RTL Agent: An Agent-Based Approach for Functionally Correct HDL Generation via LLMs
abstract
LLMs as code generators have undergone rapid progress over the past couple of years. However, the models on their own provide no guarantees for the functional correctness of the generated code. Functional tests can not only be used by designers to assess the functional correctness of code, but also to guide them towards the solution of the problem. The same can be applied to LLMs performing automatic code generation through the use of the Reflexion technique. Reflexion is an agent-based workflow where the model generating code iterates over a loop of code generation, getting feedback from the test bench, reflecting on the feedback, and making appropriate changes to the code. The technique is known to drastically improve the performance of LLMs on software code generation. In this work, we adopt the technique for hardware description language (HDL) code generation as RTL Agent - an implementation of the workflow for Verilog generation with Reflexion. We compare multiple LLMs with standard inference vs with Reflexion on the VerilogEval benchmark. Within 5 iterations of the feedback loop, we observe a relative improvement of 33.2% for the low-performance model Llama3 and an average of 17.7% for the high-performance models GPT-4o and GPT-4o-mini in their Pass@1 performance. We present a cost analysis of the technique to enable cost-performance trade-offs.
Sriram Ranga, Rui Mao 0010, Debjyoti Bhattacharjee, Erik Cambria, Anupam Chattopadhyay
ATS5
2024 Formal Verification of Secure Boot Process
abstract
Formal verification is widely used for checking the correctness of a system with respect to the succinct properties at early design phase. In recent times, these methods are also adopted to validate the security of a system by rightly abstracting the security-oriented properties. In this work, we focus on a recently reported attack on the secure boot process of Zynq-7000 platform. We study the vulnerability with formal analysis of its boot loader code. The challenge is to model the system and identify the right set of properties so that, the vulnerability is exposed quickly. It is shown that the UPPAAL model checking analysis can be used, which helps to find the vulnerability instantaneously with the help of four properties and a virtual memory usage peak of about 50MB to test each property. To the best of our knowledge, we perform the first analysis of UPPAAL on a concrete software implementation and perform a case study on a real security vulnerability reported in a CVE [1].
Sriram Vasudevan, Prasanna Ravi, Arpan Jati, Shivam Bhasin, Anupam Chattopadhyay
DATE5
2024 Backdooring Post-Quantum Cryptography: Kleptographic Attacks on Lattice-based KEMs
abstract
Post-quantum Cryptography (PQC) has reached the verge of standardization competition, with Kyber as a winning candidate. In this work, we demonstrate practical backdoor insertion in Kyber through kleptrography. The backdoor can be inserted using classical techniques like ECDH or post-quantum Classic Mceliece. The inserted backdoor targets the key generation procedure where generated output public keys subliminally leak information about the secret key to the owner of the backdoor. We demonstrate first practical instantiations of such attack at the protocol level by validating it on TLS 1.3.
Prasanna Ravi, Shivam Bhasin, Anupam Chattopadhyay, Aikata, Sujoy Sinha Roy
ACM Great Lakes Symposium on VLSI3
2024 PUF-based Lightweight Mutual Authentication Protocol for Internet of Things (IoT) Devices
abstract
The Internet of Things(IoT) can transfer data between the sensor node and the cloud server with the Internet’s help for various automated tasks like remote monitoring and controlling. IoT has various applications in different sections, including healthcare, also called the Internet of Medical Things(IoMT). IoT used in healthcare applications can collect patient’s biomedical data through medical sensors, which is sent to the cloud server with the help of the Internet. IoT/IoMT systems are having serious challenging concerns such as data security and IoT device’s security. Hence, this paper proposes a Physically Unclonable Function(PUF) based lightweight mutual authentication and key agreement protocol for IoT/IoMT devices. A lightweight AEAD cipher, ASCON, and a PUF are used for mutual authentication and key agreement between the IoT/IoMT device and the server. The agreed key is used for encrypting the sensor node data using ASCON cipher. The protocol is implemented in Artix-7 FPGA, and the formal verification is performed using the automated tool Proverif. The proposed protocol takes 912 bits of communication cost, which is 13% less compared to the best existing protocol. Further, the protocol requires a node storage cost of 128 bits, which is only 66% of the best existing protocol.
Kamal Raj, Srinivasu Bodapati 0001, Anupam Chattopadhyay
ISCAS3
2024 Boosting the Efficiency of Quantum Divider through Effective Design Space Exploration
abstract
Rapid progress in the design of scalable, robust quantum computing necessitates efficient quantum circuit implementation for algorithms with practical relevance. For several algorithms, arithmetic kernels, in particular, division plays an important role. In this manuscript, we focus on enhancing the performance of quantum slow dividers by exploring the design choices of its sub-blocks, such as, adders. Through comprehensive design space exploration of state-of-the-art quantum addition building blocks, our work have resulted in an impressive achievement: a reduction in Toffoli Depth of up to 93.90%, accompanied by substantial reductions in both Toffoli and Qubit Count of up to 92.12% and 99.38%, respectively. This paper offers crucial perspectives on efficient design of quantum dividers, and emphasizes the importance of adopting a systematic design space exploration approach.
Eugene Lim, Anupam Chattopadhyay
ISCAS3
2024 A Novel Current Comparator Enabling Large RRAM Crossbars for BNNs and PUFs
abstract
Emerging non-volatile memory (NVM) device technologies are advancing in-memory computing (IMC) applications by providing faster computation speeds and reducing resource overhead. Crossbar structures are commonly used in IMC with NVM devices such as memristors to perform matrix-vector multiplication for deep learning and security primitive applications. Large crossbars are required to implement today's deep learning models, especially for implementing specialized Binarized Neural Networks (BNNs) architectures and to construct security primitives such as physical unclonable functions (PUFs). To digitize crossbar currents, current-sense comparators based on current mirrors are widely used in BNN crossbars. However, conventional comparators have a limited current range due to the decrease of the bit-line voltage as the number of active crossbar elements connected to the bit-line increases. This paper presents a current comparator that employs a regulated cascode sensing stage which boosts the input current range by stabilizing the bit-line voltages. By increasing the current range, a larger crossbar size can be supported. Extensive simulations were carried out to verify the proposed technique, demonstrating that a significantly larger crossbar size can be achieved with the proposed comparator compared to a traditional current-sense comparator for the same area and resolution. Designed in a 180 nm technology, the proposed comparator achieves a resolution of 50nA and 100μA for PUF and BNN applications, respectively, and dissipates 198 μW,
Gokulnath Rajendran, Debajit Basak, Anupam Chattopadhyay
VLSI-SoC5
2024 Minimum Depth Quantum Modular Addition Through Carry-Save Architecture
abstract
Shor's factorization algorithm, as one of the most significant achievements in quantum computing, exhibits an exponential speedup compared to the corresponding classical algorithm. In Shor's factorization algorithm, modular exponentiation is one of the most computationally intensive components, which relies on the modular addition building block. This work aims to explore novel designs for enhancing the efficiency of quantum modular addition. In particular, we introduce a novel quantum modular addition framework based on carry-save architecture, which facilitates the conversion of multiple 2-addend quantum operations within modular addition into a single 3-addend operation, thereby reducing the computational depth. Compared to the most efficient existing quantum modular addition, our design has achieved an impressive result - a reduction in Toffoli Depth by up to 33.33%, while maintaining comparable Toffoli Count and Qubit Count. This research underscores the potential of carry-save architecture as a promising technique for accelerating quantum modular arithmetic as well as advancing the development of quantum computing in general.
Eugene Lim, Xiufan Li, Jerrie Feng, Anupam Chattopadhyay
VLSI-SoC5
2024 Priority Arbiter PUF: Analysis
Meenakshi Kansal, Animesh Roy 0004, Dibyendu Roy 0001, Srinivasu Bodapati 0001, Anupam Chattopadhyay
Discret. Appl. Math.5
2024 A Configurable CRYSTALS-Kyber Hardware Implementation with Side-Channel Protection
abstract
In this work, we present a configurable and side channel resistant implementation of the post-quantum key-exchange algorithm CRYSTALS-Kyber . The implemented design can be configured for different performance and area requirements leading to different trade-offs for different applications. A low area implementation can be achieved in 5,269 LUTs and 2,422 FFs, whereas a high performance implementation required 7,151 LUTs and 3,730 FFs. Due to a deeply pipelined architecture, a high operating speed of more than 250 MHz could be achieved on 28nm Xilinx FPGAs. The side channel resistance is implemented using a carefully chosen set of novel and known techniques such as Fault Detection Hashes, Instruction Randomization, FSM Protection and so on. resulting in a low overhead of less than 5% while being highly configurable. To the best of our knowledge, this work presents the first side-channel and fault attack protected configurable accelerator for CRYSTALS-Kyber . Using TVLA (test vector leakage assessment), we validate the implemented protection techniques and demonstrate that the design does not leak information even after 200 K traces. Furthermore, one of the configuration choices results in the smallest hardware implementation of CRYSTALS-Kyber known in the literature.
Arpan Jati, Naina Gupta 0001, Anupam Chattopadhyay, Somitra Kumar Sanadhya
ACM Trans. Embed. Comput. Syst.3
2024 Side-channel and Fault-injection attacks over Lattice-based Post-quantum Schemes (Kyber, Dilithium): Survey and New Results
abstract
In this work, we present a systematic study of Side-Channel Attacks (SCA) and Fault Injection Attacks (FIA) on structured lattice-based schemes, with main focus on Kyber Key Encapsulation Mechanism (KEM) and Dilithium signature scheme, which are leading candidates in the NIST standardization process for Post-Quantum Cryptography (PQC). Through our study, we attempt to understand the underlying similarities and differences between the existing attacks while classifying them into different categories. Given the wide variety of reported attacks, simultaneous protection against all the attacks requires to implement customized protections/countermeasures for both Kyber and Dilithium. We therefore present a range of customized countermeasures, capable of providing defenses/mitigations against existing SCA/FIA, and incorporate several SCA and FIA countermeasures within a single design of Kyber and Dilithium. Among the several countermeasures discussed in this work, we present novel countermeasures that offer simultaneous protection against several SCA- and FIA-based chosen-ciphertext attacks for Kyber KEM. We implement the presented countermeasures within two well-known public software libraries for PQC: (1) pqm4 library for the ARM Cortex-M4-based microcontroller and (2) liboqs library for the Raspberry Pi 3 Model B Plus based on the ARM Cortex-A53 processor. Our performance evaluation reveals that the presented custom countermeasures incur reasonable performance overheads on both the evaluated embedded platforms. We therefore believe our work argues for usage of custom countermeasures within real-world implementations of lattice-based schemes, either in a standalone manner or as reinforcements to generic countermeasures such as masking.
Prasanna Ravi, Anupam Chattopadhyay, Jan-Pieter D'Anvers, Anubhab Baksi
ACM Trans. Embed. Comput. Syst.2
2024 DynPen: Automated Penetration Testing in Dynamic Network Scenarios Using Deep Reinforcement Learning
abstract
Penetration testing, a crucial industrial practice for securing networked systems and infrastructures, has traditionally depended on the extensive expertise of human professionals. Addressing the scarcity of human experts, the development of automated penetration testing tools emerges as a promising avenue. Against the backdrop of rapid advancements in artificial intelligence technologies, reinforcement learning has demonstrated considerable potential for realizing automated penetration testing. However, existing research predominantly concentrates on reinforcement learning-based automated penetration testing tools within static scenarios, with limited exploration in dynamic network environments. This paper addresses a noteworthy challenge in developing autonomous agents for real-world applications, particularly focusing on scenarios marked by environmental changes. Such alterations necessitate autonomous agents to continuously monitor environmental characteristics, and adapt, and adjust learned actions to ensure the system’s effective operation. Consequently, the paper proposes an automated reinforcement learning-based penetration testing scheme tailored for dynamic network scenarios, named DynPen. DynPen captures observed changes in the scenario, aiding the penetration testing agent in decision-making based on historical experiences. Simulation results demonstrate the proposed scheme’s efficacy in significantly expediting the convergence speed of the penetration testing agent using reinforcement learning algorithms. Furthermore, the scheme successfully maintains the learning agility and adaptability of the agent in dynamic network scenarios.
Qianyu Li 0001, Dong Li 0054, Fan Shi 0003, Min Zhang 0054, Anupam Chattopadhyay, Yi Shen 0012, Yang Li 0215
IEEE Trans. Inf. Forensics Secur.6
2024 Synthesis Techniques for Fault-tolerant Quantum Circuit Implementation using Clifford+ZN-group
abstract
Decoherence jeopardizes the entanglement of fragile quantum states, and is among the foremost challenges towards engineering scalable quantum computers. Realizing quantum circuit implementation with small qubit count and shallow circuit depth is necessary due to the linear scaling of decoherence rate with qubit count and circuit depth. Conversely, reasonable correction of small unitary errors can be achieved by using surface codes along with a transversal gate set to protect quantum information from decoherence. In this paper, we analyze and report the upper bound of non-Clifford phase-depth for different mapping schemes and synthesis approaches. We introduce a synthesis methodology based on lookup-table (LUT) networks, wherein the Boolean logic translates into fault-tolerant quantum logic using Clifford+ Z N group with zero ancillary cost. We also present fault-tolerant synthesis techniques for k -LUT network using additional ancillary lines with exponential phase-count and unit phase-depth.
Laxmidhar Biswal, Debjyoti Bhattacharjee, Amlan Chakrabarti, Anupam Chattopadhyay
ACM Trans. Quantum Comput.4
2023 Hardware Security Primitives Using Passive RRAM Crossbar Array: Novel TRNG and PUF Designs
abstract
With rapid advancements in electronic gadgets, the security and privacy aspects of these devices are significant. For the design of secure systems, physical unclonable function (PUF) and true random number generator (TRNG) are critical hardware security primitives for security applications. This paper proposes novel implementations of PUF and TRNGs on the RRAM crossbar structure. Firstly, two techniques to implement the TRNG in the RRAM crossbar are presented based on write-back and 50% switching probability pulse. The randomness of the proposed TRNGs is evaluated using the NIST test suite. Next, an architecture to implement the PUF in the RRAM crossbar is presented. The initial entropy source for the PUF is used from TRNGs, and challenge-response pairs (CRPs) are collected. The proposed PUF exploits the device variations and sneak-path current to produce unique CRPs. We demonstrate, through extensive experiments, reliability of 100%, uniqueness of 47.78%, uniformity of 49.79%, and bit-aliasing of 48.57% without any post-processing techniques. Finally, the design is compared with the literature to evaluate its implementation efficiency, which is clearly found to be superior to the state-of-the-art.
Simranjeet Singh, Furqan Zahoor, Gokulnath Rajendran, Sachin B. Patkar, Anupam Chattopadhyay, Farhad Merchant
ASP-DAC5
2023 CRYSTALS-Dilithium on RISC-V Processor: Lightweight Secure Boot Using Post-Quantum Digital Signature
abstract
With the ongoing efforts for transitioning towards post-quantum security, NIST has recently selected the digital signature algorithm CRYSTALS-Dilithium for standardization. In this work, we demonstrate the first Dilithium based hardware accelerated secure boot architecture developed around Ariane, an open-source RISC- V core. By utilizing a compact design with novel verification engine, a secure boot flow is implemented with only 3.48ms runtime overhead compared to normal boot, while requiring 10.4K LUTs and 5.7K FFs on an FPGA. Compared to the state-of-the-art we achieve a reduction of 3.42× and 7.88 × for LUTs and FFs respectively. Also, the design when realized in 65nm ASIC requires only 125 kGE and 6.3 mW power at 100 MHz. Further, as secure boot is one of the critical processes and the security of the whole system depends on it, we implemented hardware fault countermeasures and evaluated their effectiveness in preventing secure boot bypass.
Naina Gupta 0001, Arpan Jati, Anupam Chattopadhyay
ICCAD3
2023 Invited Paper: Machine Learning Based Blind Side-Channel Attacks on PQC-Based KEMs - A Case Study of Kyber KEM
abstract
Kyber KEM, the NIST selected PQC standard for Public Key Encryption and Key Encapsulation Mechanisms (KEMs) has been subjected to a variety of side-channel attacks, through the course of the NIST PQC standardization process. However, all these attacks targeting the decapsulation procedure of Kyber KEM either require knowledge of the ciphertexts or require to control the value of ciphertexts for key recovery. However, there are no known attacks in a blind setting, where the attacker does not have access to the ciphertexts. While blind side-channel attacks are known for symmetric key cryptographic schemes, we are not aware of such attacks for Kyber KEM. In this paper, we fill this gap by proposing the first blind side-channel attack on Kyber KEM. We target leakage of the pointwise multiplication operation in the decryption procedure to carry out practical blind side-channel attacks resulting in full key recovery. We perform practical validation of our attack using power side-channel from the reference implementation of Kyber KEM taken from the pqm4 library, implemented on the ARM Cortex-M4 microcontroller. Our experiments clearly indicate the feasibility of our proposed attack in recovering the full key in only a few hundred to few thousand traces, in the presence of a suitably accurate Hamming Weight (HW) classifier.
Prasanna Ravi, Dirmanto Jap, Shivam Bhasin, Anupam Chattopadhyay
ICCAD4
2023 A RISC-V SoC with Hardware Trojans: Case Study on Trojan-ing the On-Chip Protocol Conversion
abstract
Hardware Trojans (HTs) are a serious security threat to the highly-decentralized, multi-stage production flow of today’s Integrated Circuit (IC) industry. Considerable research efforts have gone into developing methodologies for detecting HTs. A significant issue in validating HT detection algorithms is the lack of open-source benchmarks with the complexity of modern-day System-on-Chips (SoCs). The currently available open-source benchmarks are more elementary and, therefore, do not reveal the actual robustness of the algorithms against false positives and false negatives. To address this issue, we present the design and integration of three kinds of HTs (publicly available at [38]) in a RISC-V—based SoC. We explain their functionality and taxonomy in detail. To our knowledge, this work is the first to launch trojan attacks targeting the mismatch in the attributes of two widely used on-chip communication protocols in an SoC. We performed extensive behavioral simulations to verify the functionality of these kinds of SoC-level HTs. We estimated the detectability of these HTs by: (i) synthesizing the HT-infested SoC for FPGA and (ii) evaluating them against a Graph Neural Network-based pre-Silicon HT detection tool, automatic test pattern generation, reverse engineering, and formal verification. In a nutshell, this paper demonstrates the risk of HTs in today’s SoCs and an effective environment for strengthening research on HT detection.
Anupam Chattopadhyay, Avi Mendelson
VLSI-SoC2
2023 Optimized Quantum Circuit Implementation of Payoff Function
abstract
Large-scale quantum computers that can execute practical quantum algorithms have the potential to solve complex problems that are currently challenging for classical computers. This involves converting these problems into a form that can be processed by quantum circuits, a crucial process that requires minimizing quantum resources like qubit count, gate count, and circuit depth. Our work focuses on implementing and optimizing the foundational task of quantum finance, known as option pricing, as a quantum circuit. This enables the utilization of quantum computing benefits, within the financial domain. Specifically, we implement and optimize the function fK(S) = max(S−K, 0). Taking into consideration the significant trade-offs between qubit count and circuit depth, we have developed quantum circuits for the optimized implementation of the fK(S). Our work incorporates various optimization techniques for the circuit, such as selecting the optimal adder, optimizing the S−K operation, parallelization, and qubit reuse. Furthermore, we offer various versions of our quantum circuits for the fK(S), each featuring different adders and Toffoli decompositions, thereby providing flexibility for a wide range of use cases.
Sejin Lim, Kyungbae Jang, Anubhab Baksi, Anupam Chattopadhyay, Hwajeong Seo
VLSI-SoC6
2023 PR-PUF: A Reconfigurable Strong RRAM PUF
abstract
Physical Unclonable Functions (PUFs) offer the natural advantage of built-in key generation, thus eliminating the costly process of embedding unique key after manufacturing millions of integrated circuits. When PUFs are deployed for the application, all the Challenge-Response Pairs (CRP) are collected and stored in a trusted server, and the responses are compared with the one from the device during the run-time - forming the crux of various security protocols. Two issues are commonly faced during PUF designs. First, to enhance the applicability of PUF, larger set of CRP is desirable, which is referred to as a strong PUF. Second, due to the emergence of machine learning-based PUF modelling attacks, it is now imperative to have a PUF demonstrating resistance against such attacks. In this paper, we propose a novel Parity Resistive RAM PUF (PR-PUF) implemented using RRAM crossbar architecture. PR-PUF supports low-overhead reconfiguration, where both the original and reconfigured CRP space enhances the CRP size, with average uniqueness between reconfiguration of 49.98%. The construction also demonstrates excellent robustness against various modeling attacks. We present detailed design analysis and circuit-level simulation studies.
Gokulnath Rajendran, Furqan Zahoor, Simranjeet Singh, Farhad Merchant, Vikas Rana, Anupam Chattopadhyay
VLSI-SoC6
2023 Reducing Depth of Quantum Adder using Ling Structure
abstract
Improving the performance of quantum adder is an important technical challenge with major impact on the implementation of efficient, large-scale quantum computing. Continuing along this research direction, we propose a novel parallel-prefix quantum adder based on Ling expansion. We systematically explored classical structures for parallel-prefix adders assessing their suitability to be realized in quantum domain. Furthermore, Ling adder enforces Logical OR and large fan-out, which require innovative solutions. We addressed these challenges to realize the quantum Ling adder, which results in a T-depth of only $O\left( {\log \frac{n}{2}} \right)$. This represents a substantial improvement over the previous quantum adders based on parallel prefix structure, which require O(log n) T-depth. We present extensive theoretical and simulation-based studies to establish our claims.
Anupam Chattopadhyay
VLSI-SoC2
2023 Improved Linear Decomposition of Majority and Threshold Boolean Functions
abstract
To support efficient design automation for emerging computing fabrics, novel data structures for logic synthesis and technology mapping are being intensively studied. It has been shown that for several promising computing technologies intermediate forms, such as Majority-inverter graph (MIG) and XOR-Majority graph (XMG) can be particularly beneficial. This has propelled the Boolean Majority operator at the forefront of research. Though these structures primarily utilize 3-input Majority nodes, the efficacy of$n$-input Majority operators has been demonstrated as well. A long-standing research problem, in that context and also for theoretical circuit complexity, is to determine efficient decomposition of an$n$-input Majority$({\mathrm{ Maj}}_{n})$function in terms of 3-input Majority$({\mathrm{ Maj}}_{3})$operator. In this manuscript, we make two significant advances in this topic. First, a practically realizable linear decomposition is provided, thus improving the previously reported quadratic bounds. Second, the theoretical upper bound of decomposing${\mathrm{ Maj}}_{n}$, in terms of${\mathrm{ Maj}}_{3}$, is reduced from$5.884n$to$3n$. The erstwhile theoretical upper bound of$5.884n$also lacked a practical construction for${\mathrm{ Maj}}_{n}$decomposition, presumably due to the presence of sequential elements in the algorithm. The proof of the linearity, detailed construction procedure along with experimental studies using state-of-the-art synthesis flows to validate the aforementioned claims are presented in this work. The results are applicable to threshold Boolean functions, too.
Anupam Chattopadhyay, Debjyoti Bhattacharjee, Subhamoy Maitra
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2023 Lightweight Hardware Accelerator for Post-Quantum Digital Signature CRYSTALS-Dilithium
abstract
The looming threat of an adversary with quantum computing capability led to a worldwide research effort towards identifying and standardizing novel post-quantum cryptographic primitives. Post-standardization, all existing security protocols will need to support efficient implementation of these primitives. In this work, we contribute to these efforts by reporting the smallest implementation of CRYSTALS-Dilithium, one of the chosen post-quantum digital signature scheme for NIST standardization process. By invoking multiple optimizations to leverage parallelism, pre-computation and memory access sharing, we obtain an implementation that could be fit into one of the smallest Zynq FPGA. On Zynq Ultrascale+, our design achieves an improvement of about 36.7%/35.4%/42.3% in Area$\times $Time (LUTs$\times \text{s}$) trade-off for KeyGen/Sign/Verify respectively over state-of-the-art implementation. We also evaluate our design as a co-processor on three different hardware platforms and compare the results with software implementation, thus presenting a detailed evaluation of CRYSTALS-Dilithium targeted for embedded applications. Further, on ASIC using TSMC 65nm technology, our design requires 0.227mm2area and can operate at a frequency of 1.176 GHz. As a result, it only requires$53.7\mu \text{s}/96.9\mu \text{s}/57.7\mu \text{s}$for KeyGen/Sign/Verify operation for the best-case scenario.
Naina Gupta 0001, Arpan Jati, Anupam Chattopadhyay, Gautam Jha
IEEE Trans. Circuits Syst. I Regul. Pap.3
2023 Hardware Trojan Detection using Transition Probability with Minimal Test Vectors
abstract
Hardware Trojans (HTs) are malicious manipulations of the standard functionality of an integrated circuit (IC). Sophisticated defense against HT attacks has become the utmost current research endeavor. In particular, the HTs whose operations depend on the rare activation condition are the most critical ones. Among other techniques, logic test by rare net excitation is advocated as one of the viable detection methods due to no extra hardware requirement. However, logic test faces a tremendous challenge of the overhead of testing configuration. This work presents a methodology based on the primary input’s impact over rare nets using transition probability to select the useful test vectors. To generate a test vector, each input’s toggle probability is calculated, which drastically minimizes the search space. The capability of rare-signal generation selects the final list of test vectors. Simulations performed in the presence of different HT triggers on different benchmark circuits, like ISCAS ’85, ISCAS ’89, and ITC ’99, show that the proposed methodology is capable of producing test vectors with significantly improved rare net coverage. Furthermore, compared to an existing technique, the proposed methodology produces average higher rare switching (around 72%) inside a netlist.
Anindan Mondal, Shubrojyoti Karmakar, Mahabub Hasan Mahalat, Suchismita Roy, Bibhash Sen, Anupam Chattopadhyay
ACM Trans. Embed. Comput. Syst.6
2022 TextBack: Watermarking Text Classifiers using Backdooring
abstract
Creating high performance neural networks is ex-pensive, incurring costs that can be attributed to data collection and curation, neural architecture search and training on dedi-cated hardware accelerators. Stakeholders invested in any one or more of these aspects of deep neural network training expect as-surances on ownership and guarantees that unauthorised usage is detectable and therefore preventable. Watermarking the trained neural architectures can prove to be a solution to this. While such techniques have been demonstrated in image classification tasks, we posit that a watermarking scheme can be developed for natural language processing applications as well. In this paper, we propose TextBack, which is a watermarking technique developed for text classifiers using backdooring. We have tested for the functionality preserving properties and verifiable proof of ownership of TextBack on multiple neural architectures and datasets for text classification tasks. The watermarked models consistently generate accuracies within a range of 1 - 2% of models without any watermarking, whilst being reliably verifiable during watermarking verification. TextBack has been tested on two different kinds of Trigger Sets, which can be chosen by the owner as preferred. We have studied the efficiencies of the algorithm that embeds the watermarks by fine tuning using a combination of Trigger samples and clean samples. The benefit of using TextBack's fine tuning approach on pre-trained models from a computational cost perspective against embedding watermarks by training models from scratch is also established experimentally. This watermarking scheme is not computation intensive and adds no additional burden to the neural architecture. This makes TextBack suitable for lightweight applications on edge devices as the watermarked model can be deployed on resource-constrained hardware and SoCs when required.
Nandish Chattopadhyay, Rajan Kataria, Anupam Chattopadhyay
DSD3
2022 Robust Perception for Autonomous Vehicles using Dimensionality Reduction
abstract
Adversarial attacks on machine learning models have proven to be a major contributor for the lack of actual deployment and adoption of ML in many practical use-cases. They have been found to be equally lethal in video streams as they are in images and texts. This finds particular relevance in the domain of autonomous driving tasks, which make use of object recognition neural architectures. In this paper, we take a step back from the cat and mouse chase of novel attacks and ad-hoc defenses and try to explain adversarial attacks from the perspective of the geometry of the high-dimensional spaces that the models operate in. Additionally, we make use of our idea of relating adversarial attacks to dimensionality to propose a counter-measure that uses dimension reduction. We have tested our proposition on state-of-the-art object detection and classification models on video streams including Faster-RCNN and YOLO and corresponding adversarial attacks on these models. Having optimally tuned the hyper-parameter associated with variability preservation upon dimension reduction using simple Singular Value Decomposition, we have shown that the performance of the robust version of the object detector models is within 2−3% of that on the clean samples, despite the presence of adversarial perturbation.
Nandish Chattopadhyay, Anupam Chattopadhyay
TrustCom3
2022 PA-PUF: A Novel Priority Arbiter PUF
abstract
This paper proposes a 3-input arbiter-based novel physically unclonable function (PUF) design. Firstly, a 3-input priority arbiter is designed using a simple arbiter, two multiplexers (2:1), and an XOR logic gate. The priority arbiter has an equal probability of 0’s and 1’s at the output, which results in excellent uniformity (49.45%) while retrieving the PUF response. Secondly, a new PUF design based on priority arbiter PUF (PA-PUF) is presented. The PA-PUF design is evaluated for uniqueness, non-linearity, and uniformity against the standard tests. The proposed PA-PUF design is configurable in challenge-response pairs through an arbitrary number of feed-forward priority arbiters introduced to the design. We demonstrate, through extensive experiments, reliability of 100% after performing the error correction techniques and uniqueness of 49.63%. Finally, the design is compared with the literature to evaluate its implementation efficiency, where it is clearly found to be superior compared to the state-of-the-art.
Simranjeet Singh, Srinivasu Bodapati 0001, Sachin B. Patkar, Rainer Leupers, Anupam Chattopadhyay, Farhad Merchant
VLSI-SoC5
2022 The Bitlet Model: A Parameterized Analytical Model to Compare PIM and CPU Systems
abstract
Currently, data-intensive applications are gaining popularity. Together with this trend, processing-in-memory (PIM)–based systems are being given more attention and have become more relevant. This article describes an analytical modeling tool called Bitlet that can be used in a parameterized fashion to estimate the performance and power/energy of a PIM-based system and, thereby, assess the affinity of workloads for PIM as opposed to traditional computing. The tool uncovers interesting trade-offs between, mainly, the PIM computation complexity (cycles required to perform a computation through PIM), the amount of memory used for PIM, the system memory bandwidth, and the data transfer size. Despite its simplicity, the model reveals new insights when applied to real-life examples. The model is demonstrated for several synthetic examples and then applied to explore the influence of different parameters on two systems — IMAGING and FloatPIM. Based on the demonstrations, insights about PIM and its combination with a CPU are provided.
Ronny Ronen, Adi Eliahu, Orian Leitersdorf, Natan Peled, Kunal Korgaonkar, Anupam Chattopadhyay, Ben Perach, Shahar Kvatinsky
ACM J. Emerg. Technol. Comput. Syst.6
2022 On Exploiting Message Leakage in (Few) NIST PQC Candidates for Practical Message Recovery Attacks
abstract
In this work, we propose generic and practical side-channel attacks for message recovery in post-quantum lattice-based public key encryption (PKE) and key encapsulation mechanisms (KEM). The targeted schemes are based on the well known Learning With Errors (LWE) and Learning With Rounding (LWR) problem and include three finalists and six semi-finalist candidates of the ongoing NIST’s standardization process for post-quantum cryptography. Notably, we propose to exploit inherentciphertext malleabilityproperties of LWE/LWR-based PKEs as a powerful tool for side-channel assisted message recovery attacks. The use of ciphertext malleability widens the scope of previous attacks with the ability to target multiple operations for message recovery. Moreover, our attacks are adaptable to different implementation variants and are also applicable to implementations protected with concreteshufflingandmaskingside-channel countermeasures. Our work mainly highlights the presence of inherent algorithmic properties in LWE/LWR-based schemes that can aid side-channel attacks for message recovery, thereby stressing on the need for strong side-channel countermeasures against message recovery for LWE/LWR-based schemes.
Prasanna Ravi, Shivam Bhasin, Sujoy Sinha Roy, Anupam Chattopadhyay
IEEE Trans. Inf. Forensics Secur.4
2021 On the Cost of ASIC Hardware Crackers: A SHA-1 Case Study
Anupam Chattopadhyay, Mustafa Khairallah, Gaëtan Leurent, Zakaria Najm, Thomas Peyrin, Vesselin Velichkov
CT-RSA1
2021 Feeding Three Birds With One Scone: A Generic Duplication Based Countermeasure To Fault Attacks
abstract
In the current world of the Internet-of-things and edge computing, computations are increasingly performed locally on small connected systems. As such, those devices are often vulnerable to adversarial physical access, enabling a plethora of physical attacks which is a challenge even if such devices are built for security. As cryptography is one of the cornerstones of secure communication among devices, the pertinence of fault attacks is becoming increasingly apparent in a setting where a device can be easily accessed in a physical manner. In particular, two recently proposed fault attacks, Statistical Ineffective Fault Attack (SIFA) and the Fault Template Attack (FTA) are shown to be formidable due to their capability to bypass the common duplication based countermeasures. Duplication based countermeasures, deployed to counter the Differential Fault Attack (DFA), work by duplicating the execution of the cipher followed by a comparison to sense the presence of any effective fault, followed by an appropriate recovery procedure. While a handful of countermeasures are proposed against SIFA, no such countermeasure is known to thwart FTA to date. In this work, we propose a novel countermeasure based on duplication, which can protect against both SIFA and FTA. The proposal is also lightweight with only a marginally additional cost over simple duplication based countermeasures. Our countermeasure further protects against all known variants of DFA, including Selmke, Heyszl, Sigl's attack from FDTC 2016. It does not inherently leak side-channel information and is easily adaptable for any symmetric key primitive. The validation of our countermeasure has been done through gate-level fault simulation.
Anubhab Baksi, Shivam Bhasin, Jakub Breier, Anupam Chattopadhyay, Vinay B. Y. Kumar
DATE4
2021 Perspectives on Emerging Computation-in-Memory Paradigms
abstract
The traditional Von-Neumann architecture is reaching its limits and finding it difficult to cope up with the ever-increasing demands of modern workloads like artificial intelligence. This demand has fueled the search of technologies that can mimic human brain to efficiently combine both memory and computation within a single device. In this work, we present the state-of-the-art research in the domain of computation-in-memory. In particular, we take a look at memristors and its widespread application in neuromorphic computation. We introduce ReRAMs in terms of their novel computing paradigms and present ReRAM-specific design flows. We address the various circuit opportunities and challenges related to reliability and fault tolerance associated with them. Another high-potential candidate to leverage memory and computation from a single device is Ferroelectric Field-effect Transistor (FeFET). Here we present a co-integration of such FeFETs with another emerging nanotechnology concept, called Reconfigurable Field Effect Transistor (RFET) and discuss the impact of the higher amount of states provided by this combination.
Shubham Rai, Anteneh Gebregiorgis, Debjyoti Bhattacharjee, Krishnendu Chakrabarty, Said Hamdioui, Anupam Chattopadhyay, Jens Trommer, Akash Kumar 0001
DATE7
2021 ROWBACK: RObust Watermarking for neural networks using BACKdoors
abstract
Claiming ownership of trained neural networks is critical towards stakeholders investing heavily in high performance neural networks. There is an associated cost for the entire pipeline starting from data curation to high performance computing infrastructure for neural architecture search and training the model. Watermarking neural networks is a potential solution to the problem, but standard techniques suffer from vulnerabilities demonstrated by attackers. In this paper, we propose a robust watermarking mechanism for neural architectures. Our proposed method ROWBACK turns two properties of neural networks, the presence of adversarial examples and the ability to trap backdoors in the network while training, into a scheme that guarantees strong proofs of ownership. We redesign the Trigger Set for watermarking using adversarial examples of the model which needs to be watermarked, and assign specific labels based on adversarial behaviour. We also mark every layer separately, during training, in order to ensure that removing watermarks requires complete retraining. We have tested ROWBACK for satisfying key indicative properties expected of a reliable watermarking scheme (generates accuracies within 1 - 2% of actual model, and a complete 100% match on the Trigger Set for verification), whilst being robust against state-of-the-art watermark removal attacks [1] (requires re-training of all layers with at least 60% samples and for at least more than 45% of epochs of actual training).
Nandish Chattopadhyay, Anupam Chattopadhyay
ICMLA2
2021 In Quest for Fast and Secure SoC
abstract
Over the past few years, edge devices has gained a lot of attention. It is mainly due to significant improvements in technology, available processing power and efficiency. This evolution has resulted in edge devices becoming intelligent, smarter and more responsive. Such a growth has also resulted in many security challenges especially with the rise in side-channel attack possibilities. As these devices collect a lot of data and the decision making process is data driven; security of such devices becomes a necessity for safety-critical and time-critical applications.It is a well known fact that security in any system comes with a cost. As many IoT devices are constrained either due to available resources or the time sensitiveness of the decision they are required to make; therefore, in this work, we focus on individual System on Chip (SoC) components to integrate security measures while maintaining a balance between performance and resource utilization.
Naina Gupta 0001, Anupam Chattopadhyay
VLSI-SoC2
2021 Practical Side-Channel and Fault Attacks on Lattice-Based Cryptography
abstract
The impending threat of large-scale quantum computers to classical RSA and ECC-based public-key cryptographic schemes prompted NIST to initiate a global level standardization process for post-quantum cryptography. This process which started in 2017 with 69 submissions is currently in its third and final round with seven main candidates and eight alternate candidates, out of which seven (7) out of the fifteen (15) candidates are schemes based on hard problems over structured lattices, known as lattice-based cryptographic schemes. Among the various parameters such as theoretical post-quantum (PQ) security guarantees, implementation cost and performance, resistance against physical attacks such as Side-Channel Analysis (SCA) and Fault Injection Analysis (FIA) has also emerged as an important criterion for standardization in the final round [1]. This is especially relevant for adoption of PQC in embedded devices, which are most likely used in environments where an attacker can have unimpeded physical access to the device.
Prasanna Ravi, Anupam Chattopadhyay, Shivam Bhasin
VLSI-SoC2
2021 MemEnc: A Lightweight, Low-Power, and Transparent Memory Encryption Engine for IoT
abstract
Recent advancement in technologies has led to the widespread adoption and deployment of Internet-of-Things devices. Because of the ubiquitous nature of these devices, they process large amounts of personal and sensitive data. These data are typically stored on DRAM chips, and hence becomes an easy target for attackers. Memory encryption is a commonly adopted solution to provide confidentiality. However, realizing a lightweight, low-latency, low-power solution for resource-constrained devices is a challenge. To address this, we designed MemEnc, a purely hardware-based solution that performs encryption on-the-fly and handles memory requests transparently without any OS intervention. MemEnc runs at a maximum frequency of 401 MHz and requires about 23.2 kGE (gate equivalents) on 65-nm ASIC and consumes only 1.9 mW of power at 250 MHz. Using comprehensive benchmarking, we also analyze the applicability of the proposed solution on real-world workloads. Our experiments show that certain real-time applications can run with full memory encryption and still meet system requirements. Moreover, we integrated our memory encryption engine with ARM TrustZone and present comparative results for the different case studies with Intel SGX. We show that static and dynamic efficiency-security tradeoff is necessary for all scenarios.
Naina Gupta 0001, Arpan Jati, Anupam Chattopadhyay
IEEE Internet Things J.3
2021 Autonomous Vehicle: Security by Design
abstract
Security of (semi)-autonomous vehicles is a growing concern, first, due to the increased exposure of the functionality to potential attackers; second, due to the reliance of functionalities on diverse (semi)-autonomous systems; third, due to the interaction of a single-vehicle with myriads of other smart systems in urban traffic infrastructure. Beyond these technical issues, we argue that the security-by-design principle for smart and complex autonomous systems, such as an Autonomous Vehicle (AV) is poorly understood and rarely practiced. Unlike traditional IT systems, where the risk mitigation techniques and adversarial models are well studied and developed with security design principles such as security perimeter and defense-in-depth, the lack of such a framework for connected autonomous systems is plaguing the design and implementation of a secure AV. We attempt to identify the core issues of securing an AV. This is done methodically by developing a security-by-design framework for AV from the first principles. Subsequently, the technical challenges for AV security are identified.
Anupam Chattopadhyay, Kwok-Yan Lam, Yaswanth Tavva
IEEE Trans. Intell. Transp. Syst.1
2021 PQC Acceleration Using GPUs: FrodoKEM, NewHope, and Kyber
abstract
In this article, we present the first GPU implementation for FrodoKEM-976, NewHope-1024, and Kyber-1024. These algorithms belong to three different classes of post-quantum algorithms: Learning with errors (LWE), Ring-LWE, and Module-LWE. We show the practical applicability of the algorithms in different scenarios using two different implementation approaches. Moreover, we achieve highly efficient realization of computationally expensive operations such as NTT (Number Theoretic Transform), matrix multiplication, and Keccak. Since, these are the most common operations in lattice-based cryptographic algorithms, the techniques presented in this article will likely benefit other similar algorithms. Using a NVIDIA QUADRO GV100 graphics card, we undertook a detailed experimental study. For NewHope and Kyber we were able to perform approximately 504K and 473K key exchanges per second, demonstrating a speedup of almost 53.1× and 51.05× compared to the reference C implementation. Compared to the optimized AVX2 versions we obtain speedups of 25.7× and 14.6×, respectively. Further, implementation of FrodoKEM resulted in a speedup of 50.6×, 44.2×, and 36.9× for KeyGen, Encaps and Decaps operations. Compared to its AVX2 counterpart, we achieved a speedup of about 7.3×, 4.7× and 4.9×, respectively. We also show that using multiple streams resulted in further speedup of about 28-38 percent.
Naina Gupta 0001, Arpan Jati, Amit Kumar Chauhan, Anupam Chattopadhyay
IEEE Trans. Parallel Distributed Syst.4
2020 A Novel Duplication Based Countermeasure to Statistical Ineffective Fault Analysis
Anubhab Baksi, Vinay B. Y. Kumar, Banashri Karmakar, Shivam Bhasin, Dhiman Saha, Anupam Chattopadhyay
ACISP6
2020 Towards Secure Composition of Integrated Circuits and Electronic Systems: On the Role of EDA
abstract
Modern electronic systems become evermore complex, yet remain modular, with integrated circuits (ICs) acting as versatile hardware components at their heart. Electronic design automation (EDA) for ICs has focused traditionally on power, performance, and area. However, given the rise of hardware-centric security threats, we believe that EDA must also adopt related notions like secure by design and secure composition of hardware. Despite various promising studies, we argue that some aspects still require more efforts, for example: effective means for compilation of assumptions and constraints for security schemes, all the way from the system level down to the "bare metal"; modeling, evaluation, and consideration of security-relevant metrics; or automated and holistic synthesis of various countermeasures, without inducing negative cross-effects.In this paper, we first introduce hardware security for the EDA community. Next we review prior (academic) art for EDA-driven security evaluation and implementation of countermeasures. We then discuss strategies and challenges for advancing research and development toward secure composition of circuits and systems.
Johann Knechtel, Elif Bilge Kavun, Francesco Regazzoni 0001, Annelie Heuser, Anupam Chattopadhyay, Debdeep Mukhopadhyay, Soumyajit Dey, Yunsi Fei, Yaacov Belenky, Itamar Levi, Tim Güneysu, Patrick Schaumont, Ilia Polian
DATE5
2020 Post-Quantum Secure Boot
abstract
A secure boot protocol is fundamental to ensuring the integrity of the trusted computing base of a secure system. The use of digital signature algorithms (DSAs) based on traditional asymmetric cryptography, particularly for secure boot, leaves such systems vulnerable to the threat of quantum computers. This paper presents the first post-quantum secure boot solution, implemented fully as hardware for reasons of security and performance. In particular, this work uses the eXtended Merkle Signature Scheme (XMSS), a hash-based scheme that has been specified as an IETF RFC. The solution has been integrated into a secure SoC platform around RISC-V cores and evaluated on an FPGA and is shown to be orders of magnitude faster compared to corresponding hardware/software implementations and to compare competitively with a fully hardware elliptic curve DSA based solution.
Vinay B. Y. Kumar, Naina Gupta 0001, Anupam Chattopadhyay, Michael Kasper, Christoph Krauß, Ruben Niederhagen
DATE3
2020 Noise Resilient Compilation Policies for Quantum Approximate Optimization Algorithm
abstract
Quantum approximate optimization algorithm (QAOA) is a promising quantum-classical hybrid algorithm to solve hard combinatorial optimization problems using noisy quantum devices. The multiqubit CPHASE gates used in the quantum circuit for QAOA are commutative i.e., the order of the gates can be altered without changing the output state. This re-ordering leads to the execution of more gates in parallel and a smaller number of additional SWAP gates to compile the QAOA circuit resulting in lower circuit-depth and gate-count. A less number of gates generally indicates a lower accumulation of gate-errors, and a reduced circuit-depth means less decoherence time for the qubits. However, near-term quantum devices exhibit significant variations in the gate success probabilities. Variation-aware compilation policies (i.e. putting most gate operations on qubits with higher gate success probabilities) can enhance the probability of successful program execution on the hardware. The greater flexibility of QAOA-circuits offer better scope of optimization with QAOA-tailored compilation policies. This paper presents an argument for compilation policies to exploit the unique characteristics of QAOA-circuits alongside the variation-awareness of the noisy devices. We present two procedures - variation-aware qubit placement (VQP) and variation-aware iterative mapping (VIM) that can improve the circuit success probability quite significantly (≈8.408X on average) for a set of QAOA-MaxCut problems on ibmq_16_melbourne.
Mahabubul Alam, Abdullah Ash-Saki, Junde Li, Anupam Chattopadhyay, Swaroop Ghosh
ICCAD4
2020 CONTRA: Area-Constrained Technology Mapping Framework For Memristive Memory Processing Unit
abstract
Data-intensive applications are poised to benefit directly from processing-in-memory platforms, such as memristive Memory Processing Units, which allow leveraging data locality and performing stateful logic operations. Developing design automation flows for such platforms is a challenging and highly relevant research problem. In this work, we investigate the problem of minimizing delay under arbitrary area constraint for MAGIC-based in-memory computing platforms. We propose an end-to-end area constrained technology mapping framework, CONTRA. CONTRA uses Look-Up Table (LUT) based mapping of the input function on the crossbar array to maximize parallel operations and uses a novel search technique to move data optimally inside the array. CONTRA supports benchmarks in a variety of formats, along with crossbar dimensions as input to generate MAGIC instructions. CONTRA scales for large benchmarks, as demonstrated by our experiments. CONTRA allows mapping benchmarks to smaller crossbar dimensions than achieved by any other technique before, while allowing a wide variety of area-delay trade-offs. CONTRA improves the composite metric of area-delay product by 2.1× to 13.1× compared to seven existing technology mapping approaches.
Debjyoti Bhattacharjee, Anupam Chattopadhyay, Srijit Dutta, Ronny Ronen, Shahar Kvatinsky
ICCAD2
2020 Deploy-able Privacy Preserving Collaborative ML
abstract
In the data-driven world, emerging technologies like the Internet of Things (IoT) and other crowd-sourced data sources like mobile devices etc. generate a tremendous volume of decentralized data that needs to be analyzed for obtaining useful insights, necessary for reliable decision making. Although the overall data is rich, contributors of such kind of data are reluctant to share their own data due to serious concerns regarding protection of their privacy; while those interested in harvesting the data are constrained by the limited computational resources available with each participant. In this paper, we propose an end-to-end algorithm that puts in coalescence the mechanism of learning collaboratively in a decentralized fashion, using Federated Learning, while preserving differential privacy of each participating client, which are typically conceived as resource-constrained edge devices. We have developed the proposed infrastructure and analyzed its performance from the standpoint of a machine learning task using standard metrics. We observed that the collaborative learning framework actually increases prediction capabilities in comparison to a centrally trained model (by 1-2%), without having to share data amongst the participants, while strong guarantees on privacy (ε, δ) can be provided with some compromise on performance (about 2-4%). Additionally, quantization of the model for deployment on edge devices do not degrade its capability, whilst enhancing the overall system efficiency.
Nandish Chattopadhyay, Ritabrata Maiti, Anupam Chattopadhyay
ICDCS3
2020 Enabling Efficient Mapping of XMG-Synthesized Networks to Spintronic Hardware
abstract
Spintronics presents great promise for efficient processing and storage of information in the post-Moore era, thanks to its attributes of non-volatility, excellent integration-density, near-unlimited endurance and compatibility with CMOS process-technology. Today's state-of-the-art EDA tools primarily use AND-Inverter Graphs (AIGs), Majority-Inverter Graphs (MIGs) and XOR-Majority Graphs (XMGs) for representing any complex Boolean logic. To be able to utilize the existing EDA tools for implementing spin-based logic circuits, it is important that the logic primitives in these data structures can be natively realized by spin devices. This paper demonstrates how the XMGs synthesized by EDA flows can be more-efficiently mapped to spintronic fabric using a novel domain wall motion-based XOR device. We develop a device-to-system simulation-framework to precisely evaluate the post-mapping (to domain-wall gates) performances of synthesized networks. Our study over several challenging benchmark-suites shows that the use of this XOR-gate enables the efficient mapping of XMGs while improving the {size, depth, size·depth, energy, EDP} performances of the network by averages of {31.54%, 19.00%, 41.56%, 38.03%, 45.47%} over those of mapped MIGs.
Anupam Chattopadhyay
ISCAS2
2020 Authentication Protocol for Secure Automotive Systems: Benchmarking Post-Quantum Cryptography
abstract
Ensuring communication security in real-time automotive networks is of paramount importance given the sensitivity of the exchanged information and highly safety critical nature of its operation. The first step towards ensuring security is to securely authenticate all the computational nodes through use of authentication protocols based on public-key cryptography. But, traditional public-key cryptographic primitives we use today are believed to be breakable by large scale quantum computers of the future. Thus, NIST is currently running a global global level standardization process for quantum-resistant public-key cryptography, better known as post-quantum cryptography. In this work, we perform a first of its kind practical implementation of a secure authentication protocol for automotive systems with post-quantum cryptographic algorithms and perform a detailed comparative evaluation of the speed and communication bandwidth performance against their pre-quantum counterparts.
Prasanna Ravi, Vijaya Kumar Sundar, Anupam Chattopadhyay, Shivam Bhasin, Arvind Easwaran
ISCAS3
2020 FlexWatts: A Power- and Workload-Aware Hybrid Power Delivery Network for Energy-Efficient Microprocessors
abstract
Modern client processors typically use one of three commonly-used power delivery network (PDN) architectures: 1) motherboard voltage regulators (MBVR), 2) integrated voltage regulators (IVR), and 3) low dropout voltage regulators (LDO). We observe that the energy-efficiency of each of these PDNs varies with the processor power (e.g, thermal design power (TDP) and dynamic power-state) and workload characteristics (e.g., work-load type and computational intensity). This leads to energy-inefficiency and performance loss, as modern client processors operate across a wide spectrum of power consumption and execute a wide variety of workloads. To address this inefficiency, we propose FlexWatts, a hybrid adaptive PDN for modern client processors whose goal is to provide high energy-efficiency across the processor's wide range of power consumption and workloads. FlexWatts provides high energy-efficiency by intelligently and dynamically allocating PDNs to processor domains depending on the processor's power consumption and workload. FlexWatts is based on three key ideas. First, FlexWatts combines IVRs and LDOs in a novel way to share multiple on-chip and off-chip resources and thus reduce cost, as well as board and die area overheads. This hybrid PDN is allocated for processor domains with a wide power consumption range (e.g., CPU cores and graphics engines) and it dynamically switches between two modes: IVR-Mode and LDO-Mode, depending on the power consumption. Second, for all other processor domains (that have a low and narrow power range, e.g., the IO domain), FlexWatts statically allocates off-chip VRs, which have high energy-efficiency for low and narrow power ranges. Third, FlexWatts introduces a novel prediction algorithm that automatically switches the hybrid PDN to the mode (IVR-Mode or LDO-Mode) that is the most beneficial based on processor power consumption and workload characteristics. To evaluate the tradeoffs of PDNs, we develop and open-source PDNspot, the first validated architectural PDN model that enables quantitative analysis of PDN metrics. Using PDNspot, we evaluate FlexWatts on a wide variety of SPEC CPU2006, graphics (3DMark06), and battery life (e.g., video playback) workloads against IVR, the state-of-the-art PDN in modern client processors. For a 4 W thermal design power (TDP) processor, FlexWatts improves the average performance of the SPEC CPU2006 and 3DMark06 workloads by 22% and 25%, respectively. For battery life workloads, FlexWatts reduces the average power consumption of video playback by 11% across all tested TDPs (4W-50W). FlexWatts has comparable cost and area overhead to IVR. We conclude that FlexWatts provides high energy-efficiency across a modern client processor's wide range of power consumption and wide variety of workloads, with minimal overhead.
Jawad Haj-Yahya, Mohammed Alser, Jeremie S. Kim, Lois Orosa 0001, Efraim Rotem, Avi Mendelson, Anupam Chattopadhyay, Onur Mutlu
MICRO7
2020 Mind the Portability: A Warriors Guide through Realistic Profiled Side-channel Analysis
Shivam Bhasin, Anupam Chattopadhyay, Annelie Heuser, Dirmanto Jap, Stjepan Picek, Ritu Ranjan Shrivastwa
NDSS2
2020 Hierarchical discovery of large-scale and focal copy number alterations in low-coverage cancer genomes
abstract
BACKGROUND: Detection of DNA copy number alterations (CNAs) is critical to understand genetic diversity, genome evolution and pathological conditions such as cancer. Cancer genomes are plagued with widespread multi-level structural aberrations of chromosomes that pose challenges to discover CNAs of different length scales, and distinct biological origins and functions. Although several computational tools are available to identify CNAs using read depth (RD) signal, they fail to distinguish between large-scale and focal alterations due to inaccurate modeling of the RD signal of cancer genomes. Additionally, RD signal is affected by overdispersion-driven biases at low coverage, which significantly inflate false detection of CNA regions. RESULTS: We have developed CNAtra framework to hierarchically discover and classify 'large-scale' and 'focal' copy number gain/loss from a single whole-genome sequencing (WGS) sample. CNAtra first utilizes a multimodal-based distribution to estimate the copy number (CN) reference from the complex RD profile of the cancer genome. We implemented Savitzky-Golay smoothing filter and Modified Varri segmentation to capture the change points of the RD signal. We then developed a CN state-driven merging algorithm to identify the large segments with distinct copy numbers. Next, we identified focal alterations in each large segment using coverage-based thresholding to mitigate the adverse effects of signal variations. Using cancer cell lines and patient datasets, we confirmed CNAtra's ability to detect and distinguish the segmental aneuploidies and focal alterations. We used realistic simulated data for benchmarking the performance of CNAtra against other single-sample detection tools, where we artificially introduced CNAs in the original cancer profiles. We found that CNAtra is superior in terms of precision, recall and f-measure. CNAtra shows the highest sensitivity of 93 and 97% for detecting large-scale and focal alterations respectively. Visual inspection of CNAs revealed that CNAtra is the most robust detection tool for low-coverage cancer data. CONCLUSIONS: . It is freely available at https://github.com/AISKhalil/CNAtra.
Ahmed Ibrahim S. Khalil, Costerwell Khyriem, Anupam Chattopadhyay, Amartya Sanyal
BMC Bioinform.3
2020 Identification and utilization of copy number information for correcting Hi-C contact map of cancer cell lines
abstract
BACKGROUND: Hi-C and its variant techniques have been developed to capture the spatial organization of chromatin. Normalization of Hi-C contact map is essential for accurate modeling and interpretation of high-throughput chromatin conformation capture (3C) experiments. Hi-C correction tools were originally developed to normalize systematic biases of karyotypically normal cell lines. However, a vast majority of available Hi-C datasets are derived from cancer cell lines that carry multi-level DNA copy number variations (CNVs). CNV regions display over- or under-representation of interaction frequencies compared to CN-neutral regions. Therefore, it is necessary to remove CNV-driven bias from chromatin interaction data of cancer cell lines to generate a euploid-equivalent contact map. RESULTS: We developed the HiCNAtra framework to compute high-resolution CNV profiles from Hi-C or 3C-seq data of cancer cell lines and to correct chromatin contact maps from systematic biases including CNV-associated bias. First, we introduce a novel 'entire-fragment' counting method for better estimation of the read depth (RD) signal from Hi-C reads that recapitulates the whole-genome sequencing (WGS)-derived coverage signal. Second, HiCNAtra employs a multimodal-based hierarchical CNV calling approach, which outperformed OneD and HiNT tools, to accurately identify CNVs of cancer cell lines. Third, incorporating CNV information with other systematic biases, HiCNAtra simultaneously estimates the contribution of each bias and explicitly corrects the interaction matrix using Poisson regression. HiCNAtra normalization abolishes CNV-induced artifacts from the contact map generating a heatmap with homogeneous signal. When benchmarked against OneD, CAIC, and ICE methods using MCF7 cancer cell line, HiCNAtra-corrected heatmap achieves the least 1D signal variation without deforming the inherent chromatin interaction signal. Additionally, HiCNAtra-corrected contact frequencies have minimum correlations with each of the systematic bias sources compared to OneD's explicit method. Visual inspection of CNV profiles and contact maps of cancer cell lines reveals that HiCNAtra is the most robust Hi-C correction tool for ameliorating CNV-induced bias. CONCLUSIONS: HiCNAtra is a Hi-C-based computational tool that provides an analytical and visualization framework for DNA copy number profiling and chromatin contact map correction of karyotypically abnormal cell lines. HiCNAtra is an open-source software implemented in MATLAB and is available at https://github.com/AISKhalil/HiCNAtra .
Ahmed Ibrahim S. Khalil, Siti Rawaidah Binte Mohammad Muzaki, Anupam Chattopadhyay, Amartya Sanyal
BMC Bioinform.3
2020 Crossbar-Constrained Technology Mapping for ReRAM Based In-Memory Computing
abstract
In-memory computing has gained significant attention due to the potential for dramatic improvement in speed and energy. Redox-based resistive RAMs (ReRAMs), capable of non-volatile storage and logic operations simultaneously have been used for logic-in-memory computing approaches. To this effect, we propose ReRAM based VLIW Architecture for in-Memory comPuting (ReVAMP), supported by a detailed device-accurate simulation setup with peripheral circuitry. We present theoretical bounds on the minimum area required for in-memory computation of arbitrary Boolean functions specified using structural representation (And-Inverter Graph and Majority-Inverter Graph) and two-level representation (Exclusive-Sum-of-Product). To support the ReVAMP architecture, we present two technology mapping flows that fully exploit the bit-level parallelism offered by the execution of logic using ReRAM crossbar array. The area-constrained mapping (ArC) generates feasible mapping for a variety of crossbar dimensions while the delay-constrained mapping (DeC) focuses primarily on minimizing the latency of mapping. We evaluate the proposed mappings against two state-of-the-art technology in-memory computing architectures, PLiM and MAGIC along with their automation flows (SIMPLE and COMPACT). ArC and DeC outperform state-of-the-art PLiM architecture by 1.46x and 4.3x on average in latency. ArC offers significantly lower area (on average 25.27x and 6.57x), while improving the area-delay product by 1.37x and 1.12x against two mapping approaches for MAGIC respectively. In contrast, DeC achieves average area (1.45x and 3.06x) and area-delay product (1.12x and 6.36x) improvements over the mapping approaches for MAGIC architecture respectively. The proposed mapping techniques allow a variety of runtime efficiency trade-offs.
Debjyoti Bhattacharjee, Yaswanth Tavva, Arvind Easwaran, Anupam Chattopadhyay
IEEE Trans. Computers4
2020 Threshold Implementations of <tt>GIFT</tt>: A Trade-Off Analysis
abstract
Threshold Implementation (TI) is one of the most widely used countermeasure for side channel attacks. Over the years several TI techniques have been proposed for randomizing cipher execution using different variations of secret-sharing and implementation techniques. For instance, sharing without decomposition (4-shares) is the most straightforward implementation of the threshold countermeasure. However, its usage is limited due to its high area requirements. On the other hand, sharing using decomposition (3-shares) countermeasure for cubic non-linear functions significantly reduces area and complexity in comparison to 4-shares. Nowadays, security of ciphers using a side channel countermeasure is of utmost importance. This is due to the wide range of security critical applications from smart cards, battery operated IoT devices, to accelerated crypto-processors. Such applications have different requirements (higher speed, energy efficiency, low latency, small area etc.) and hence need different implementation techniques. Although, many TI strategies and implementation techniques are known for different ciphers, there is no single study comparing these on a single cipher. Such a study would allow a fair comparison of the various methodologies. In this work, we present an in-depth analysis of the various ways in which TI can be implemented for a lightweight cipher. We chose GIFT for our analysis as it is currently one of the most energy-efficient lightweight ciphers. The experimental results show that different implementation techniques have distinct applications. For example, the 4-shares technique is good for applications demanding high throughput whereas 3-shares is suitable for constrained environments with less area and moderate throughput requirements. The techniques presented in the paper are also applicable to other blockciphers. For security evaluation, we performed TVLA (test vector leakage assessment) on all the design strategies. Experiments using up to 50 million traces show that the designs are protected against first-order attacks.
Arpan Jati, Naina Gupta 0001, Anupam Chattopadhyay, Somitra Kumar Sanadhya, Donghoon Chang
IEEE Trans. Inf. Forensics Secur.3
2019 Recruiting Fault Tolerance Techniques for Microprocessor Security
abstract
The growing threat of various attacks on modern microprocessors and systems calls for major design overhauls ranging from plugging micro-architectural side channels such as due to speculative execution to implementing cryptographic accelerators for side-channel and fault attack resistance. In this paper, we suggest to focus on the similarities and the differences between fault tolerance techniques and countermeasures against attacks on security sensitive systems. Modern digital circuits and systems use a diverse set of techniques to ensure operational correctness in the presence of faults. From a security perspective, the goal is to ensure a set of stated security properties hold in the presence of 'security faults' (extending the notion of conventional faults to include injected faults as well as vulnerabilities such as passive side-channels). A point of note here is that under some security faults, the operational correctness may not be compromised. This paper advocates the re-purposing of some of the known fault tolerance techniques, and show how those can be useful for enhancing security in the presence of active side-channel attacks. As a simple illustration of these ideas, we present an experimental case study in fortifying a cryptographic sub-component of a RISC-V based secure system-on-chip, against a formidable fault attack called SIFA.
Vinay B. Y. Kumar, Mustafa Khairallah, Anupam Chattopadhyay, Avi Mendelson
ATS5
2019 Improving Speed of Dilithium's Signing Procedure
Prasanna Ravi, Sourav Sen Gupta 0001, Anupam Chattopadhyay, Shivam Bhasin
CARDIS3
2019 Exploiting Determinism in Lattice-based Signatures: Practical Fault Attacks on pqm4 Implementations of NIST Candidates
abstract
In this paper, we analyze the implementation level fault vulnerabilities of deterministic lattice-based signature schemes. In particular, we extend the practicality of skip-addition fault attacks through exploitation of determinism in Dilithium and qTESLA signature schemes, which are two leading candidates for the NIST standardization of post-quantum cryptography. We show that single targeted faults injected in the signing procedure allow to recover an important portion of the secret key. Though faults injected in the signing procedure do not recover all the secret key elements, we propose a novel forgery algorithm that allows the attacker to sign any given message with only the extracted portion of the secret key. We perform experimental validation of our attack using Electromagnetic fault injection on reference implementations taken from the pqm4 library, a benchmarking and testing framework for post quantum cryptographic implementations for the ARM Cortex-M4 microcontroller. We also show that our attacks break two well known countermeasures known to protect against skip-addition fault attacks. We further propose an efficient mitigation strategy against our attack that exponentially increases the attacker's complexity at almost zero increase in computational complexity.
Prasanna Ravi, Mahabir Prasad Jhanwar, James Howe, Anupam Chattopadhyay, Shivam Bhasin
AsiaCCS4
2019 La Petite Fee Cosmo: Learning Data Structures Through Game-Based Learning
abstract
This research aims to implement the productive failure teaching concept with interactive learning games as a method to nurture innovative teaching and learning. The research also aims to promote innovative approaches to learning and improving students' learning experience, and their understanding of linked list data structure concepts taught in computer science subjects since students do not widely understand this concept. A 2D bridge building puzzle game, “La Petite Fee Cosmo” was developed to assist students in not only understanding the underlying concepts of the linked list but also foster creative usage of the various functionalities of linked list in diverse situations.
Vinayak Teoh Kannappan, Owen Noel Newton Fernando, Anupam Chattopadhyay, Xavier Tan, Jeffrey Hong, Seah Hock Soon, Hui En Lye
CW3
2019 SAID: A Supergate-Aided Logic Synthesis Flow for Memristive Crossbars
abstract
A Memristor is a two-terminal device that can serve as a non-volatile memory element with built-in logic capabilities. Arranged in a crossbar structure, memristive arrays allow to represent complex Boolean logic functions that adhere to the logic-in-memory paradigm, where data and logic gates are glued together on the same piece of hardware. Needless to say, novel and ad-hoc CAD solutions are required to achieve practical and feasible hardware implementations. Existing techniques aim at optimal mapping strategies that account for Boolean logic functions described by means of 2-input NOR and NOT gates, thus overlooking the optimization capabilities that a smart and dedicated technology-aware logic synthesis can provide. In this paper, we introduce a novel library-free supergate-aided (SAID) logic synthesis approach with a dedicated mapping strategy tailored on MAGIC crossbars. Supergates are obtained with a Look-Up Table (LUT)-based synthesis that splits a complex logic network into smaller Boolean functions. Those functions are then mapped on the crossbar array as to minimize latency. The proposed SAID flow allows to (i) maximize supergate-level parallelism, thus reducing the total number of computing cycles, and (ii) relax mapping constraints, allowing an easy and fast mapping of Boolean functions on memristive crossbars. Experimental results obtained on several benchmarks from ISCAS'85 and IWLS'93 suites demonstrate that our solution is capable to outperform other state-of-the-art techniques in terms of speedup (3.89× in the best case), at the expense of a very low area overhead.
Valerio Tenace, Roberto Giorgio Rizzo, Debjyoti Bhattacharjee, Anupam Chattopadhyay, Andrea Calimera
DATE4
2019 MUQUT: Multi-Constraint Quantum Circuit Mapping on NISQ Computers: Invited Paper
abstract
Rapid advancement in the domain of quantum technologies have opened up researchers to the real possibility of experimenting with quantum circuits, and simulating small-scale quantum programs. Nevertheless, the quality of currently available qubits and environmental noise pose a challenge in smooth execution of the quantum circuits. Therefore, efficient design automation flows for mapping a given algorithm to the Noisy Intermediate Scale Quantum (NISQ) computer becomes of utmost importance. State-of-the-art quantum design automation tools are primarily focused on reducing logical depth, gate count and qubit counts with recent emphasis on topology-aware (nearest-neighbour compliance) mapping. In this work, we extend the technology mapping flows to simultaneously consider the topology and gate fidelity constraints while keeping logical depth and gate count as optimization objectives. We provide a comprehensive problem formulation and multi-tier approach towards solving it. The proposed automation flow is compatible with commercial quantum computers, such as IBM QX and Rigetti. Our simulation results over 10 quantum circuit benchmarks, show that the fidelity of the circuit can be improved up to 3.37 × with an average improvement of 1.87 ×.
Debjyoti Bhattacharjee, Abdullah Ash-Saki, Mahabubul Alam, Anupam Chattopadhyay, Swaroop Ghosh
ICCAD4
2019 Curse of Dimensionality in Adversarial Examples
abstract
While machine learning and deep neural networks in particular, have undergone massive progress in the past years, this ubiquitous paradigm faces a relatively newly discovered challenge, adversarial attacks. An adversary can leverage a plethora of attacking algorithms to severely reduce the performance of existing models, therefore threatening the use of AI in many safety-critical applications. Several attempts have been made to try and understand the root cause behind the generation of adversarial examples. In this paper, we try to relate the geometry of the high-dimensional space in which the model operates and optimizes, and the properties and problems therein, to such adversarial attacks. We present the mathematical background, the intuition behind the existence of adversarial examples and substantiate them with empirical results from our experiments.
Nandish Chattopadhyay, Anupam Chattopadhyay, Sourav Sen Gupta 0001, Michael Kasper
IJCNN2
2019 Spintronic Device-Structure for Low-Energy XOR Logic using Domain Wall Motion
abstract
The recent slowdown in CMOS scaling has witnessed the rise of spintronics as a new direction for efficient information-processing due to its non-volatility, excellent integration-density, near-unlimited endurance and compatibility with CMOS process-technology. Nevertheless, research into exploring the potential of spin devices to perform logic operations is still in its nascent stage. In this paper, we demonstrate how a novel, domain wall motion-based spin-device can be used for performing XOR logic. Simulation studies establish the proposed gate to be 63% more energy-efficient than a baseline XOR-gate. A magnetic full-adder implemented using the proposed gate consumes 31.38% less energy and achieves 8.5% improved energy-delay product in comparison to a baseline full-adder.
Anupam Chattopadhyay
ISCAS2
2019 SHINE: A Novel SHA-3 Implementation Using ReRAM-based In-Memory Computing
abstract
In memory-computing (IMC) architectures provide a much needed solution to energy-efficiency barriers posed by Von-Neumann computing due to movement of data between the processor and the memory. Emerging non-volatile memories (NVM) such as Resistive RAM (ReRAM) implemented in a crossbar array are promising substrates to realize IMC due to excellent High Resistance State (HRS) to Low Resistance State (LRS) ratios and high-densities. Hardware security primitives such as SHA-3 require heavy data traffic between processing elements and memory. Therefore, they can be benefited substantially by in-memory acceleration. We propose SHINE, a high performance and area efficient hardware implementation of the Keccak function that forms the core of SHA-3 by exploiting ReRAM-based IMC. SHINE implements various functions in a Sum of Product (SOP) form in the crossbar array architecture. Simulation results show that it cuts down energy by ~90.5% and increases throughput by 1.5X to 2.8X as compared to conventional CMOS based implementations such as [1] and [2].
Karthikeyan Nagarajan, Sina Sayyah Ensan, Mohammad Nasim Imtiaz Khan, Swaroop Ghosh, Anupam Chattopadhyay
ISLPED5
2019 Guest Editorial Special Section on Security Challenges and Solutions With Emerging Computing Technologies
abstract
Multiple emerging computing technologies based on, e.g., graphene, spintronics, resistive RAM, quantum computing, and others are being developed to enhance the capabilities of logic devices and circuits. The rapid growth in these technologies is synchronized with the decline of Moore’s law, thus promises to herald the era of Beyond CMOS technologies with a significant improvement in energy efficiency, reliability, performance, and manufacturability. These devices enable very different computing paradigms, e.g., neuromorphic computing, non-Boolean computing, and in-memory computing, thus making these platforms an interesting playground for circuit and application-developers alike.
Anupam Chattopadhyay, Swaroop Ghosh, Wayne P. Burleson, Debdeep Mukhopadhyay
IEEE Trans. Very Large Scale Integr. Syst.1
2018 Technology-aware logic synthesis for ReRAM based in-memory computing
abstract
Resistive RAMs (ReRAMs) have gained prominence for design of logic-in-memory circuits and architectures due to fast read/write speeds, high endurance, density and logic operation capabilities. ReRAM crossbar arrays allow constrained bit-level parallel operations. In this paper, for the first time, we propose optimization techniques during logic synthesis, which are specifically targeted for leveraging the parallelism offered by ReRAM crossbar arrays. Our method uses Majority-Inverter Graph (MIG) for the internal representation of the Boolean functions. The novel optimization techniques, when applied to the MIG, exposes the bit-level parallelism, and is further coupled with an efficient technology mapping flow. The entire synthesis process is benchmarked exhaustively over large arithmetic functions using a representative ReRAM crossbar architecture, while varying the crossbar dimensions. For the hard benchmarks, we obtained 10% reduction in the number of nodes with 16% reduction in delay on average.
Debjyoti Bhattacharjee, Luca G. Amarù, Anupam Chattopadhyay
DATE3
2018 DFARPA: Differential fault attack resistant physical design automation
abstract
Differential Fault Analysis (DFA), aided by sophisticated mathematical analysis techniques for ciphers and precise fault injection methodologies, has become a potent threat to cryptographic implementations. In this paper, we propose, to the best of the our knowledge, the first “DFA-aware” physical design automation methodology, that effectively mitigates the threat posed by DFA. We first develop a novel floorplan heuristic, which resists the simultaneous corruption of cipher states necessary for successful fault attack, by exploiting the fact that most fault injections are localized in practice. Our technique results in the computational complexity of the fault attack to shoot up to exhaustive search levels, making them practically infeasible. In the second part of the work, we develop a routing mechanism, which tackles more precise and costly fault injection techniques, like laser and electromagnetic guns. We propose a routing technique by integrating a specially designed ring oscillator based sensor circuit around the potential fault attack targets without incurring any performance overhead. We demonstrate the effectiveness of our technique by applying it on state of the art ciphers.
Mustafa Khairallah, Rajat Sadhukhan, Radhamanjari Samanta, Jakub Breier, Shivam Bhasin, Rajat Subhra Chakraborty, Anupam Chattopadhyay, Debdeep Mukhopadhyay
DATE7
2018 Domain Wall Motion-based XOR-like Activation Unit With A Programmable Threshold
abstract
Spintronic devices promise an excellent opportunity for implementing ultra-low power neuromorphic platforms due to the inherent correspondence between their physical characteristics and the required neuronal, synaptic functionalities. Neuromorphic circuits using domain wall motion-based threshold neurons have been demonstrated in previous studies. However, threshold neurons are unable to realize linearly inseparable functions. Our work addresses this challenge by proposing a new domain wall motion-based neural activation unit with XOR-like activation function. We also develop a new learning algorithm for neurons with this activation unit. Offline training is performed on real-world datasets from the UCI machine learning repository. Neuromorphic circuits corresponding to these datasets are also simulated. The results suggest femto-Joule range energy consumption of a neuron with the proposed activation unit and 1.08×-1.82× lower misclassification rate (MCR) of the proposed algorithm in comparison to the traditional perceptron learning algorithm.
Tarun Vatwani, Anupam Chattopadhyay, Arindam Basu, Xuanyao Fong
IJCNN3
2018 Efficient Hardware Accelerator for NORX Authenticated Encryption
abstract
Authenticated encryption with associated data (AEAD) plays a significant role in cryptography due to its ability to provide integrity, confidentiality and authenticity at the same time. There is an unceasing demand of high-performance and area-efficient AEAD ciphers due to the emergence of security at the edge of computing fabric, such as, sensors and smartphone devices. Currently, a worldwide contest, titled CAESAR, is being held to decide on a set of AEAD ciphers, which are distinguished by their security, runtime performance, energy-efficiency and low area budget. In this paper, we focus on optimizing the hardware architecture of NORX by applying a pipeline technique. Our pre-layout results using commercial ASIC TSMC 65 technology library show that optimized NORX is 40.81% faster, 18.01% smaller, and improved the throughput per area by 76.9% when compared with state-of-the-art NORX implementation.
Jawad Haj-Yahya, Anupam Chattopadhyay
ISCAS3
2018 Efficient and Lightweight Quantized Compressive Sensing using μ-Law
abstract
IoT devices for video sensing need to operate within the constraints of limited bandwidth and low computing capabilities. To that effect, Compressive Sensing (CS) emerged as a prominent technique for balancing the quality of images/video and the computing/communication overheads. For CS of video data, the Block-based CS (BCS) is typically used due to low complexity. However, while CS reduces the number of samples to be transmitted, the bit-width of each sample increases due to the linear algebraic operations involved in CS, thus making CS less attractive in its pure and straightforward form. To further optimize the use of CS in IoT devices for video sensing, we explore the use of μ-law quantization technique due to its low hardware implementation overhead. We designed and implemented a complete CS platform with the integration of μ-law quantization, and studied the image quality at different compression ratios. The results show that the proposed quantization technique requires only up to 40 additional LUTs compared to the baseline algorithm, while achieving an additional compression of up to 280% in the best case.
Vikramkumar Pudi, Anupam Chattopadhyay, Kwok-Yan Lam
ISCAS2
2018 A New High Throughput and Area Efficient SHA-3 Implementation
abstract
High performance and area efficient Secure Hash Algorithm (SHA-3) hardware realization is investigated and proposed in this work. In addition to the new and simplified round constant (RC) generator, the presented SHA-3 hash implementations employed architectural optimization approaches based on the concepts of unrolling, pipelining and subpipelining. This has therefore produced a total of five implementations of SHA-3 which are denoted as Cases I-V in both FPGA and ASIC. Considering the trade-offs between the performance and hardware cost, the best architecture in term of the throughput and area efficiency is identified in Case V. The architecture has the highest throughput of 16.51 Gbps and area efficiency of 11.47 Mbps/slices for the FPGA implementation. While in ASIC, our best implementation (Case V) achieves the highest throughput of 48 Gbps.
Ming Ming Wong, Jawad Haj-Yahya, Suman Sau, Anupam Chattopadhyay
ISCAS4
2018 A Security Model for Intelligent Vehicles and Smart Traffic Infrastructure
abstract
Intelligent vehicles require to communicate with other vehicles as well as with the road-side infrastructure for gathering information such as traffic management, cooperative driving, telematics and road construction. However, vehicle-to vehicle (V2V) and vehicle to infrastructure (V2I) communications can impose some serious security threats against vehicles' safety and other sensitive information which can lead to catastrophic consequences. Therefore, there is a pressing need to develop appropriate security protocols facilitating V2I and V2V communication. This paper presents a model for smart traffic infrastructure consisting of numerous entities like smart sensors, intelligent vehicles, base stations along with a new user authentication and key-exchange protocol which aids in the establishment of a secure session for communication between the entities. In the proposed protocol, any SUF-CMA (Strong Unforgeable- Chosen Message Attack) secure digital signature algorithm can be used for authenticating the users and any NM-CCA2 (Non-Malleability Chosen Ciphertext Attack) secure algorithm can also be used for the challenge response phase between the users authenticating themselves.
Sonu Jha, Sumit Kumar Pandey, Anupam Chattopadhyay
Intelligent Vehicles Symposium4
2018 ReRAM-based In-Memory Computation of Galois Field arithmetic
abstract
Robust data communication is a prime need in the age of Internet-of-things (IoT), where multiple connected devices actively exchange information. To permit robustness of this information exchange, error resilient secure communication is necessary. Security, error detection as well as correction are fundamentally based on Galois Field (GF) arithmetic. In this work, we present a novel method for performing GF arithmetic on a state-of-the art ReRAM-based in-memory computing platform. ReRAM devices offer low leakage power, high endurance and non-volatile storage capabilities, coupled with stateful logic operations. The proposed lightweight library presents the mapping of GF element generation, addition and multiplication. We have experimentally verified the results. For GF(24), 3.8 nJ, 0.1 nJ and 3.1 nJ energy are required for element generation, addition and multiplication operations respectively, which demonstrates the efficacy of the mapping.
Swagata Mandal, Debjyoti Bhattacharjee, Yaswanth Tavva, Anupam Chattopadhyay
VLSI-SoC4
2018 Lightweight and High Performance SHA-256 using Architectural Folding and 4-2 Adder Compressor
abstract
The modern era of Internet-of-Things (IoT) is naturally imposing a tight area/runtime constraint on the computing kernels. Security kernels, as part of the standardized protocols as well as custom defense techniques, are among the most common tasks executed on every digital device. Therefore, low area cost and high performance implementation of security kernels is an important goal of current system designers. In this paper, we revisit the state-of-the-art implementations of SHA-256, a standardized security primitive for authentication and propose novel optimizations. Our optimizations, based on architectural folding and 4-2 adder compressor, are geared toward both lightweight and high performance implementations. Detailed experiments of our optimized architecture on different FPGA fabrics clearly demonstrate their benefits. Our presented design point successfully attained the highest hardware efficiency (throughput/area) figures among the published literature so far.
Ming Ming Wong, Vikramkumar Pudi, Anupam Chattopadhyay
VLSI-SoC3
2018 On Hardware Implementation of Tang-Maitra Boolean Functions
Mustafa Khairallah, Anupam Chattopadhyay, Bimal Mandal, Subhamoy Maitra
WAIFI2
2018 Kogge-Stone Adder Realization using 1S1R Resistive Switching Crossbar Arrays
abstract
Low operating voltage, high storage density, non-volatile storage capabilities, and relative low access latencies have popularized memristive devices as storage devices. Memristors can be ideally used for in-memory computing in the form of hybrid CMOS nano-crossbar arrays. In-memory serial adders have been theoretically and experimentally proven for crossbar arrays. To harness the parallelism of memristive arrays, parallel-prefix adders can be effective. In this work, a novel mapping scheme for in-memory Kogge-Stone adder has been presented. The number of cycles increases logarithmically with the bit width N of the operands, i.e., O ( log 2 N ), and the device count is 5 N . We verify the correctness of the proposed scheme by means of TaO × device model-based memristive simulations. We compare the proposed scheme with other proposed schemes in terms of number of cycle and number of devices.
Debjyoti Bhattacharjee, Anne Siemon, Eike Linn, Stephan Menzel, Anupam Chattopadhyay
ACM J. Emerg. Technol. Comput. Syst.5
2018 Wireless Communication and Security Issues for Cyber-Physical Systems and the Internet-of-Things
abstract
Wireless sensors and actuators connected by the Internet-of-Things (IoT) are central to the design of advanced cyber-physical systems (CPSs). In such complex, heterogeneous systems, communication links must meet stringent requirements on throughput, latency, and range, while adhering to tight energy budget and providing high levels of security. In this paper, we first summarize wireless communication principles from the perspective of the connectivity needs of IoT and CPS. Based on these principles, we then review the most relevant wireless communication standards before focusing on the key security issues and features of such systems. In particular, the gap between the security features in the communication standards used in CPSs and IoT and their actual vulnerabilities are pointed out with practical examples and recent attacks. We emphasize the need for a more in-depth study of the security issues across all the protocol layers, including both logical layer security and physical layer security.
Andreas Peter Burg, Anupam Chattopadhyay, Kwok-Yan Lam
Proc. IEEE2
2018 Toward Threat of Implementation Attacks on Substation Security: Case Study on Fault Detection and Isolation
abstract
Modern and future substations are aimed to be more interconnected, leveraging communication standards like IEC 61850-9-2, and associated abstract data models and communication services like generic object oriented substation event, manufacturing message specification, and sampled measured value. Such interconnection would enable fast and secure data transfer, sharing of the analytics information for various purposes like wide area monitoring, faster outage recovery, blackout prevention, distributed state estimation, etc. This would require strong focus on communication security, both at system level as well as at embedded device level. Although communication level security is dealt in IEC 62351, implementation attack on the embedded system is not considered. Since the embedded system makes the core of the smart grid, in this paper, we take a deeper look into impact of implementation attacks on substation security. An overview of potential exploits is first provided. This is followed by a case study, where implementation attacks like malicious fault injection attacks and hardware Trojan are used to compromise a substation level intelligent electronic device. The studied scenario extends implementation attacks beyond its usual exploit of confidentiality to affect power grid integrity and availability.
Anupam Chattopadhyay, Abhisek Ukil, Dirmanto Jap, Shivam Bhasin
IEEE Trans. Ind. Informatics1
2018 Efficient Realization of Householder Transform Through Algorithm-Architecture Co-Design for Acceleration of QR Factorization
abstract
QR factorization is a ubiquitous operation in many engineering and scientific applications. In this paper, we present efficient realization of Householder Transform (HT) based QR factorization through algorithm-architecture co-design where we achieve performance improvement of 3-90x in-terms of Gflops/watt over state-of-the-art multicore, General Purpose Graphics Processing Units (GPGPUs), Field Programmable Gate Arrays (FPGAs), and ClearSpeed CSX700. Theoretical and experimental analysis of classical HT is performed for opportunities to exhibit higher degree of parallelism where parallelism is quantified as a number of parallel operations per level in the Directed Acyclic Graph (DAG) of the transform. Based on theoretical analysis of classical HT, an opportunity to re-arrange computations in the classical HT is identified that results in Modified HT (MHT) where it is shown that MHT exhibits 1.33x times higher parallelism than classical HT. Experiments in off-the-shelf multicore and General Purpose Graphics Processing Units (GPGPUs) for HT and MHT suggest that MHT is capable of achieving slightly better or equal performance compared to classical HT based QR factorization realizations in the optimized software packages for Dense Linear Algebra (DLA). We implement MHT on a customized platform for Dense Linear Algebra (DLA) and show that MHT achieves 1.3x better performance than native implementation of classical HT on the same accelerator. For custom realization of HT and MHT based QR factorization, we also identify macro operations in the DAGs of HT and MHT that are realized on a Reconfigurable Data-path (RDP). We also observe that due to re-arrangement in the computations in MHT, custom realization of MHT is capable of achieving 12 percent better performance improvement over multicore and GPGPUs than the performance improvement reported by General Matrix Multiplication (GEMM) over highly tuned DLA software packages for multicore and GPGPUs which is counter-intuitive.
Farhad Merchant, Tarun Vatwani, Anupam Chattopadhyay, Soumyendu Raha, S. K. Nandy 0001, Ranjani Narayan
IEEE Trans. Parallel Distributed Syst.3
2017 Area-constrained technology mapping for in-memory computing using ReRAM devices
abstract
In-memory computing platforms, such as Resistive RAM (ReRAM), offer natural advantage to data-intensive applications. The benefits of data locality and capability to perform native Boolean operations is exploited for significant performance advantage in multiple contexts ranging across neuromorphic computing, associative memory-based computing, arithmetic benchmarks and general-purpose programmable logic-in-memory computing. Despite these advances, design automation tools supporting in-memory computing are still in a nascent phase. In this work, we investigate for the first time, the problem of minimizing delay under arbitrary area constraint of ReRAM devices. We formulate the problem of area-constrained delay minimization as an Integer Linear Programming (ILP) formulation and further propose heuristics that offers scalability as well as solution close to optimal performance. Area-constrained mapping technology mappings enables unlocking significantly large design space trade-offs.
Debjyoti Bhattacharjee, Arvind Easwaran, Anupam Chattopadhyay
ASP-DAC3
2017 A systematic security analysis of real-time cyber-physical systems
abstract
Security in Cyber-Physical Systems (CPS) has become a serious concern owing to the rapid adoption of technologies such as plug-and-play connectivity, robotics and remote coordination and control. It is well understood that the performance overhead incurred due to security considerations is rather high, which needs to be captured holistically for a real-time CPS with strict timing budget and hard deadlines. Additionally, attacks in real-time CPS may only alter the timing behaviour of system components without any changes in functionality, resulting in serious consequences due to missed deadlines. To address this challenging issue, it is necessary to understand the role of diverse components in a real-time CPS and how those expose the system to a malicious attacker. In this paper, we propose a systematic security analysis flow, using a novel Attack Sequence Diagram (ASD), which links the sources, intermediate components and final manifestations of an attack, thereby clearly delineating the attack surfaces of a complex real-time CPS. Based on the ASD, it is possible to evaluate the complexity of an attack, performance overhead of a countermeasure and explore different design trade-offs for a realtime CPS. With the help of real-world and synthetic examples, we demonstrate that ASD seamlessly enables one to map the existing vulnerabilities and uncover new attack possibilities.
Arvind Easwaran, Anupam Chattopadhyay, Shivam Bhasin
ASP-DAC2
2017 ReVAMP: ReRAM based VLIW architecture for in-memory computing
abstract
With diverse types of emerging devices offering simultaneous capability of storage and logic operations, researchers have proposed novel platforms that promise gains in energy-efficiency. Such platforms can be classified into two domains - application-specific and general-purpose. The application-specific in-memory computing platforms include machine learning accelerators, arithmetic units, and Content Addressable Memory (CAM)-based structures. On the other hand, the general-purpose computing platforms stem from the idea that several in-memory computing logic devices do support a universal set of Boolean logic operation and therefore, can be used for mapping arbitrary Boolean functions efficiently. In this direction, so far, researchers have concentrated on challenges in logic synthesis (e.g. depth optimization), and technology mapping (e.g. device count reduction). The important problem of efficient technology mapping of arbitrary logic network onto a crossbar array structure has been overlooked so far. In this paper, we propose, ReVAMP, a general-purpose computing platform based on Resistive RAM crossbar array, which exploits the parallelism in computing multiple logic operations in the same word. Further, we study the problem of instruction generation and scheduling for such a platform. We benchmark the performance of ReVAMP with respect to the state of the art architecture.
Debjyoti Bhattacharjee, Rajeswari Devadoss, Anupam Chattopadhyay
DATE3
2017 Secure Cyber-Physical Systems: Current trends, tools and open research problems
abstract
To understand and identify the attack surfaces of a Cyber-Physical System (CPS) is an essential step towards ensuring its security. The growing complexity of the cybernetics and the interaction of independent domains such as avionics, robotics and automotive is a major hindrance against a holistic view CPS. Furthermore, proliferation of communication networks have extended the reach of CPS from a user-centric single platform to a widely distributed network, often connecting to critical infrastructure, e.g., through smart energy initiative. In this manuscript, we reflect on this perspective and provide a review of current security trends and tools for secure CPS. We emphasize on both the design and execution flows and particularly highlight the necessity of efficient attack surface detection. We provide a detailed characterization of attacks reported on different cyber-physical systems, grouped according to their application domains, attack complexity, attack source and impact. Finally, we review the current tools, point out their inadequacies and present a roadmap of future research.
Anupam Chattopadhyay, Alok Prakash, Muhammad Shafique 0001
DATE1
2017 A Practical Fault Attack on ARX-Like Ciphers with a Case Study on ChaCha20
abstract
This paper presents the first practical fault attack on the ChaCha family of addition-rotation-XOR (ARX)-based stream ciphers. ChaCha has recently been deployed for speeding up and strengthening HTTPS connections for Google Chrome on Android devices. In this paper, we propose differential fault analysis attacks on ChaCha without resorting to nonce misuse. We use the instruction skip and instruction replacement fault models, which are popularly mounted on microcontroller-based cryptographic implementations. We corroborate the attack propositions via practical fault injection experiments using a laser-based setup targeting an Atmel AVR 8-bit microcontroller-based implementation of ChaCha. Each of the proposed attacks can be repeated with 100% accuracy in our fault injection setup, and can recover the entire 256 bit secret key using 5-8 fault injections on an average.
S. V. Dilip Kumar, Sikhar Patranabis, Jakub Breier, Debdeep Mukhopadhyay, Shivam Bhasin, Anupam Chattopadhyay, Anubhab Baksi
FDTC6
2017 Side-Channel Attack on STTRAM Based Cache for Cryptographic Application
abstract
In this paper, we propose a Side Channel Attack (SCA) model on Spin-Torque Transfer RAM (STTRAM) where an adversary can monitor the supply current of the memory array consumed during read/write operations and recover the secret key of Advanced Encryption Standard (AES) execution. Simulation results indicate that by monitoring write current, 50% of keys could be extracted using 2000 traces. Further improvement of attacks on write operation is also proposed. The read current is found to be more susceptible to leak the key. It reveals first byte in only 40 traces and leaks the entire key in as low as 400 traces. The results are then compared with Static RAM (SRAM) based cache. The attack model has been experimentally validated on read operation of commercial MRAM chip (STTRAM variant). Experimental results indicate that the attack can reveal correct key in 15 traces compared to 40 in simulation due to less algorithmic noise. To the best of our knowledge, this is the first comprehensive SCA study for STTRAM based cache for cryptographic application.
Mohammad Nasim Imtiaz Khan, Shivam Bhasin, Alex Yuan, Anupam Chattopadhyay, Swaroop Ghosh
ICCD4
2017 Designing Parity Preserving Reversible Circuits
Goutam Paul 0001, Anupam Chattopadhyay, Chander Chandak
RC2
2017 Automatic Test Pattern Generation for Multiple Missing Gate Faults in Reversible Circuits - Work in Progress Report
Anmol Surhonne, Anupam Chattopadhyay, Robert Wille
RC2
2017 An analysis of root functions - A subclass of the Impossible Class of Faulty Functions (ICFF)
Enes Pasalic, Anupam Chattopadhyay, Debabani Chowdhury
Discret. Appl. Math.2
2017 Hardware Architectures for Embedded Speaker Recognition Applications: A Survey
abstract
Authentication technologies based on biometrics, such as speaker recognition, are attracting more and more interest thanks to the elevated level of security offered by these technologies. Despite offering many advantages, such as remote use and low vulnerability, speaker recognition applications are constrained by the heavy computational effort and the hard real-time constraints. When such applications are run on an embedded platform, the problem becomes more challenging, as additional constraints inherent to this specific domain are added. In the literature, different hardware architectures were used/designed for implementing a process with a focus on a given particular metric. In this article, we give a survey of the state-of-the-art works on implementations of embedded speaker recognition applications. Our aim is to provide an overview of the different approaches dealing with acceleration techniques oriented towards speaker and speech recognition applications and attempt to identify the past, current, and future research trends in the area. Indeed, on the one hand, many flexible solutions were implemented, using either General Purpose Processors or Digital Signal Processors. In general, these types of solutions suffer from low area and energy efficiency. On the other hand, high-performance solutions were implemented on Application Specific Integrated Circuits or Field Programmable Gate Arrays but at the expense of flexibility. Based on the available results, we compare the application requirements vis-à-vis the performance achieved by the systems. This leads to the projection of new research trends that can be undertaken in the future.
Hasna Bouraoui, Chadlia Jerad, Anupam Chattopadhyay, Nejib Ben Hadj-Alouane
ACM Trans. Embed. Comput. Syst.3
2017 RC4-AccSuite: A Hardware Acceleration Suite for RC4-Like Stream Ciphers
abstract
We present RC4-AccSuite, a hardware accelerator, which combines the flexibility of an application specific instruction set processor and the performance of an application specific IC for the most widely deployed commercial stream cipher RC4 and its eight prominent variants, including Spritz (CRYPTO-2014 Rump-session). Our carefully designed instruction set architecture reuses combinational and sequential logic at its various pipeline stages and memories, saving up to 41% in terms of area, compared with the individual cores, while the power budget being dictated primarily by the variant used. Moreover, using state replication, noticeable throughput performance enhancement in RC4 variants is achieved. RC4-AccSuite possesses extensibility for future variants of RC4 with little or no tweaking.
Ayesha Khalid, Goutam Paul 0001, Anupam Chattopadhyay
IEEE Trans. Very Large Scale Integr. Syst.3
2017 A Flexible Divide-and-Conquer MPSoC Architecture for MIMO Interference Cancellation
abstract
The fast-evolving standards of the wireless communication systems drive the demand for flexible baseband processing platforms. However, with the proliferation of MIMO technologies, traditional single-core-based solutions are hardly able to fulfill requirements with acceptable power and area cost. The reliance on multi-/many-core system is increasing. Different from the computation-limited single-core-based solutions, multi/many-core systems are often communication-limited. In this paper, aiming at MIMO interference cancellation algorithms, we propose a flexible master-slave-based multiprocessor system-on-chiparchitecture based on a systematically divide-and-conquer approach to optimize the communication problems from the application-, architecture- and programming-levels. First, a comprehensively analysis of several typical applications in terms of parallelism, communication patterns and computation patterns is presented. According to the analysis results, a low-complexity and flexible ad hoc point-to-point interconnected fine-grained programmable-element (f -PE) is proposed to execute the arithmetic calculation. In order to reduce the communication traffic, an f-PE-based slave-node is constructed to exploit the data and instruction localities of applications, and a master node that is used to schedule and serve data for the slave nodes is also integrated. Furthermore, to improve the ease of use of the architecture, a multiple instruction multiple data like programming model is adopted and an optimizing mapping strategy is developed. In order to show its flexibility potential, seven linear and nonlinear IC algorithms with distinct computation natures are implemented on the proposed architecture. Finally, the gate-level synthesis and postlayout results are presented to demonstrate the strength and weaknesses of our design.
Luechao Yuan, Cang Liu, Chuan Tang, Anupam Chattopadhyay, Gerd Ascheid, Zuocheng Xing
IEEE Trans. Very Large Scale Integr. Syst.5
2016 Look-ahead schemes for nearest neighbor optimization of 1D and 2D quantum circuits
abstract
Ensuring nearest neighbor compliance of quantum circuits by inserting SWAP gates has heavily been considered in the past. Here, quantum gates are considered which work on non-adjacent qubits. SWAP gates are applied in order to “move” these qubits onto adjacent positions. However, a decision how exactly the SWAPs are “moved” has mainly been made without considering the effect a “movement” of qubits may have on the remaining circuit. In this work, we propose a methodology for nearest neighbor optimization which addresses this problem by means of a look-ahead scheme. To this end, two representative implementations are presented and discussed in detail. Experimental evaluations show that, in the best case, reductions in the number of SWAP gates of 56% (compared to the state-of-the-art methods) can be achieved following the proposed methodology.
Robert Wille, Oliver Keszöcze, Marcel Walter, Patrick Rohrs, Anupam Chattopadhyay, Rolf Drechsler
ASP-DAC5
2016 Runtime NBTI Mitigation for Processor Lifespan Extension via Selective Node Control
abstract
Negative bias temperature instability (NBTI) has become one of the major reliability concerns for nanoscale CMOS technology. The NBTI effect degrades pMOS transistors by stressing them with negatively biased voltage, while the transistors heal themselves as the negative bias is removed. In this paper, we propose a cross-layer mitigation technique for NBTI-induced timing degradation in processors. The NOP (No Operation) instruction is replaced by a custom NOP instruction for healing purpose. Cells that are likely to be stressed under negative bias are classified and their upstream cell will be replaced by the internal node control (INC) logics. Upon encountering a custom NOP instruction, the INC logics will force the NBTI-stressed cell to be in its healing mode. The optimal INC logic insertion through genetic programming approach achieves much greater delay mitigation of 44.3% than prior works in a 10-year span with less than 4% of power and negligible area overhead.
Song Bian 0001, Michihiro Shintani, Zheng Wang 0020, Masayuki Hiromoto, Anupam Chattopadhyay, Takashi Sato 0001
ATS5
2016 Statistical fault injection for impact-evaluation of timing errors on application performance
abstract
This paper proposes a novel approach to modeling of gate level timing errors during high-level instruction set simulation. In contrast to conventional, purely random fault injection, our physically motivated approach directly relates to the underlying circuit structure, hence allowing for a significantly more detailed characterization of application performance under scaled frequency / voltage (including supply noise). The model uses gate level timing statistics extracted by dynamic timing analysis from the post place & route netlist of a general-purpose processor to perform instruction-aware fault injections. We employ a 28 nm OpenRISC core as a case study, to demonstrate how statistical fault injection provides a more accurate and realistic analysis of power vs. error performance.
Jeremy Constantin, Andreas Peter Burg, Zheng Wang 0020, Anupam Chattopadhyay, Georgios Karakonstantis
DAC4
2016 Unlocking efficiency and scalability of reversible logic synthesis using conventional logic synthesis
abstract
Latest quantum technologies promise realization of extremely large circuits, whereas, reversible logic synthesis, the key automation step for quantum computing, suffers from scalability bottleneck. Scalability can be achieved with Decision Diagram (DD)-based synthesis at the cost of significant an-cilla/garbage lines overhead. In this paper, we present a novel hierarchical reversible logic synthesis, where DD-based synthesis is invoked within an And-Inverter Graph (AIG)-based synthesis wrapper, balancing scalability and performance.
Mathias Soeken, Anupam Chattopadhyay
DAC2
2016 The Programmable Logic-in-Memory (PLiM) computer
Pierre-Emmanuel Gaillardon, Luca G. Amarù, Anne Siemon, Eike Linn, Rainer Waser, Anupam Chattopadhyay, Giovanni De Micheli
DATE6
2016 A low overhead error confinement method based on application statistical characteristics
Zheng Wang 0020, Georgios Karakonstantis, Anupam Chattopadhyay
DATE3
2016 Delay-optimal technology mapping for in-memory computing using ReRAM devices
abstract
Recent propositions of diverse In-Memory Computing platforms have shown a promising alternative to classical Von Neumann computing models. Significant benefits, in terms of energy-efficiency and performance, are reported for in-memory arithmetic circuits, neural networks, CAM, cache hierarchy and even fully programmable processors. In contrast, design automation tools supporting the development of such designs are still in a nascent phase. By leveraging the native stateful logic operation capability of ReRAM devices, several logic synthesis flows have been reported. In this paper, we complement these flows with a detailed study on the technology mapping phase for ReRAM devices. We provide a delay-optimal solution for technology mapping without area constraint and propose further heuristics to achieve device count reduction and to support delay optimization under the constraint of parallel instruction dispatch. We report at least 3× less delay compared to the naïve technology mapping adopted in recent studies. The proposed heuristics achieve 56% on average reduction in device count. Finally, a range of performance trade-offs is identified by applying the constraint of parallel instruction dispatch without noticeable degradation of delay.
Debjyoti Bhattacharjee, Anupam Chattopadhyay
ICCAD2
2016 Low-quantum cost circuit constructions for adder and symmetric Boolean functions
abstract
Quantum computing necessitates the design of circuits via reversible logic gates. Efficient reversible circuit can be constructed by achieving low ancilla count, reducing logical depth and lowering Quantum costs. Generalized Peres gates have recently been realized with very low Quantum Cost (QC) by utilizing Quantum rotation gates. This is utilized in recent literature for efficient reversible circuit constructions for symmetric Boolean functions. In this paper, we extend this line of construction further by demonstrating efficient realization of adder circuits. In particular, we revisit the adder construction of Vedral, Barenco and Eckert to show that improvement of gate count and QC is achievable by exploiting a construction based only on Peres gates. We also report improved constructions of symmetric Boolean functions by following an approach recently proposed in the context of Boolean function complexity analysis.
Anupam Chattopadhyay, Anubhab Baksi
ISCAS1
2016 Hardware Accelerator for Stream Cipher Spritz
abstract
RC4, the dominant stream cipher in e-commerce and communication protocols such as, WEP, TLS, is being considered for replacement due to the series of vulnerabilities that have been pointed out in recent past. After a thorough analysis of the possible weaknesses, Spritz, a new stream cipher is proposed to that effect by the author of RC4. The design of Spritz is based on Cryptographic Sponge construction, which permits Spritz to be used in different modes, and therefore, makes it an attractive design choice for security protocols. Initial software performance analysis of Spritz shows that it fares poorly compared to the state-of-the-art hash functions and stream ciphers. In this paper, we extend the analysis to the hardware performance. We propose a fully customized accelerator design for Spritz and identify the highest achievable runtime performance for ASIC and FPGA technology. Our results show that the Spritz accelerator is significantly faster in encryption compared to the software implementation (32.38x speed-up for the SQUEEZE and 64.07x speed-up for the ABSORB function), though fares weakly against hardware implementation of state-of-the-art hash functions and stream ciphers in terms of area-efficiency.
Debjyoti Bhattacharjee, Anupam Chattopadhyay
SECRYPT2
2016 Enabling in-memory computation of binary BLAS using ReRAM crossbar arrays
abstract
Memristive devices, such as ReRAMs, are fast gaining prominence for their low leakage power, high endurance and non-volatile storage capabilities. ReRAM crossbar arrays also found usage as platform for in-memory computing, particularly for data-intensive computations, due to its inherent capability to perform stateful logic operations. Binary matrix and vector operations arise in several applications that require close interaction with the storage, such as Error Correction Codes (ECC), approximate graph mining, and in general, diverse big data applications. In this paper, we explore for the first time, an efficient mapping of Binary Basic Linear Algebra Subprograms (BiBLAS) onto hybrid CMOS-ReRAM crossbar array. We investigate the impact of crossbar configurations on the delay, and area of BiBLAS operations for various vector sizes.
Debjyoti Bhattacharjee, Farhad Merchant, Anupam Chattopadhyay
VLSI-SoC3
2016 A Sound and Complete Axiomatization of Majority-n Logic
abstract
Manipulating logic functions via majority operators recently drew the attention of researchers in computer science. For example, circuit optimization based on majority operators enables superior results as compared to traditional synthesis tools. Also, the Boolean satisfiability problem finds new solution approaches when described in terms of majority decisions. To support computer logic applications based on majority, a sound and complete set of axioms is required. Most of the recent advances in majority logic deal only with ternary majority (MAJ-3) operators because the axiomatization with solely MAJ-3 and complementation operators is well understood. However, it is of interest extending such axiomatization to$n$-ary majority operators (MAJ-$n$) from both the theoretical and practical perspective. In this work, we address this issue by introducing a sound and complete axiomatization of MAJ-$n$logic. Our axiomatization naturally includes existing MAJ-3 and MAJ-5 axiomatic systems. Based on this general set of axioms, computer applications can now fully exploit the expressive power of majority logic.
Luca G. Amarù, Pierre-Emmanuel Gaillardon, Anupam Chattopadhyay, Giovanni De Micheli
IEEE Trans. Computers3
2016 Three Snakes in One Hole: The First Systematic Hardware Accelerator Design for SOSEMANUK with Optional Serpent and SNOW 2.0 Modes
abstract
With increasing usage of hardware accelerators in modern heterogeneous System-on-Chips (SoCs), the distinction between hardware and software is no longer rigid. The domain of cryptography is no exception and efficient hardware design of so-called software ciphers are becoming increasingly popular. In this paper, for the first time we propose an efficient hardware accelerator design for SOSEMANUK, one of the finalists of the eSTREAM stream cipher competition in the software category. Since SOSEMANUK combines the design principles of the block cipher Serpent and the stream cipher SNOW 2.0, we make our design flexible to accommodate the option for independent execution of Serpent and SNOW 2.0. In the process, we identify interesting design points and explore different levels of optimizations. We perform a detailed experimental evaluation for the performance figures of each design point. The best throughput achieved by the combined design is 67.84 Gbps for SOSEMANUK, 33.92 Gbps for SNOW 2.0 and 2.12 Gbps for Serpent. Our design outperforms all existing hardware (as well as software) designs of Serpent, SNOW 2.0 and SOSEMANUK, along with those of all other eSTREAM candidates.
Goutam Paul 0001, Anupam Chattopadhyay
IEEE Trans. Computers2
2016 RunStream: A High-Level Rapid Prototyping Framework for Stream Ciphers
abstract
We present RunStream, a rapid prototyping framework for realizing stream cipher implementations based on algorithmic specifications and architectural customizations desired by the users. In the dynamic world of cryptography where newer recommendations are frequently proposed, the need of such tools is imperative. It carries out design validation and generates an optimized software implementation and a synthesizable Register Transfer Level Verilog description. Our framework enables speedy benchmarking against critical resources like area, throughput, power, and latency and allows exploration of alternatives. Using RunStream, we successfully implemented various stream ciphers and benchmarked the quality of results to be at par with published hand-optimized implementations.
Ayesha Khalid, Goutam Paul 0001, Anupam Chattopadhyay, Faezeh Abediostad, Syed Imad Ud Din, Baishik Biswas, Prasanna Ravi
ACM Trans. Embed. Comput. Syst.3
2015 TriviA: A Fast and Secure Authenticated Encryption Scheme
Avik Chakraborti, Anupam Chattopadhyay, Mridul Nandi
CHES2
2015 Exploiting dynamic timing margins in microprocessors for frequency-over-scaling with instruction-based clock adjustment
Jeremy Constantin, Lai Wang, Georgios Karakonstantis, Anupam Chattopadhyay, Andreas Peter Burg
DATE4
2015 New ASIC/FPGA Cost Estimates for SHA-1 Collisions
abstract
SHA-1 remains, till date, the most widely used hash function, in spite of several successful cryptanalytic attacks against it. These attacks, however, remain impractical due to high computation complexity and associated cost. We endeavor to do cost-time product estimation for an attack by the aid of application-specific hardware acceleration. This work proposes an Application-Specific Instruction-set Processor (ASIP), named Cracken. Cracken is aimed to efficiently realize near collision attack on SHA-1. The estimations of the physical attack complexity is done using 65nm standard CMOS technology and commercial FPGA devices. It is estimated, with post-layout simulations, that Stevens' differential attack with an estimated complexity of 2^57.5, can be executed in 46 days using 4096 Cracken cores at a cost of Euros 15m. Estimation for real collision with complexity 2^61 is also done. Our cost-time estimates reveal that an FPGA-based attack is more efficient compared to ASIC. Previously reported SHA-1 attacks based on ASIC and cloud computing platforms are also compiled and benchmarked for reference.
Ayesha Khalid, Anupam Chattopadhyay, Christian Rechberger, Tim Güneysu, Christof Paar
DSD3
2015 In-memory adder functionality in 1S1R arrays
abstract
Memristive devices enable non-volatile data storage and in-memory computing capabilities. By using stateful logic approaches, hybrid CMOS nano-crossbar arrays offer additional functionalities such as arithmetic operations. To enable storage and computing on large-scale arrays, parasitic current paths within the array must be avoided. Therefore, for example, a complementary resistive switch (1CRS) or a bipolar rectifying element (‘selector’) in series to a resistive switching device (1S1R) is required at each cross-point junction to suppress low-ohmic sneak paths. In this work 1S1R arrays are considered. First, the in-memory adder concept, initially developed for CRS arrays, is adjusted for a 1S1R array. After that an optimized design is presented and verified by means of memristive simulations. Third, the energy consumption of both concepts is evaluated as a function of array size, and the delay of memristive adder designs are compared quantitatively.
Anne Siemon, Stephan Menzel, Anupam Chattopadhyay, Rainer Waser, Eike Linn
ISCAS3
2015 Trace Buffer Attack: Security versus observability study in post-silicon debug
abstract
Since the standardization of AES/Rijndael symmetric-key cipher by NIST in 2001, it gained widespread acceptance in various protocols and withstood intense scrutiny from the theoretical cryptanalysts. From the physical implementation point of view, however, AES remained vulnerable. Practical attacks on AES via fault injection, differential power analysis, scan-chain and cache-access timing have been demonstrated so far. Along this line, in this paper, we propose a novel and effective attack, termed Trace Buffer Attack. Trace buffers are extensively used for post-silicon debug of digital designs. We identify this as a source of information leakage and show that, unless proper countermeasure is taken, Trace Buffer Attack is capable of partially recovering the secret keys of different AES implementations. We report the detailed process of trace-buffer attack with experimental results. We also propose a countermeasure in order to avoid such attack.
Yuanwen Huang, Anupam Chattopadhyay, Prabhat Mishra 0001
VLSI-SoC2
2015 Exploiting scalable CGRA mapping of LU for energy efficiency using the Layers architecture
abstract
A scalable and highly efficient numerical linear algebra kernel mapping for coarse-grained reconfigurable architectures is proposed and applied to a 3D reconfigurable architecture, Layers, which exploits functional parallelism and a functional reconfiguration-based programming model to achieve flexibility, scalability and low energy. Instead of solving the complex problem of mapping an application to fit architectural constraints, in our approach we tailor the mapping scheme for efficiency and scalability and exploit architectural flexibility and reconfigurability to adapt the architecture to match the derived mapping. Thus, kernel execution reaches asymptotically optimal efficiency for various architectural parameters and input matrix sizes, without modification of the derived mapping. Detailed performance and power evaluations were done with input data sets with matrix sizes ranging from 64×64 to 16384×16384. Twelve architectural variants with up to 10×10 processing elements were used to explore scalability of the mapping and the architecture, achieving <10% energy increase for architectures up to 8×8 PEs, coupled with performance speed-ups of more than an order of magnitude.
Zoltán Endre Rákossy, Dominik Stengele, Gerd Ascheid, Rainer Leupers, Anupam Chattopadhyay
VLSI-SoC5
2015 Flexible, Efficient Multimode MIMO Detection by Using Reconfigurable ASIP
abstract
The combination of software flexibility and hardware configurability makes partially reconfigurable application-specific instruction-set processor (rASIP) an attractive architecture, which matches the needs of computation-intensive and fast-evolving wireless receiver algorithms. This paper describes the design of a multimode multiple-input-multiple-output (MIMO) detector by using rASIP, which supports multiple MIMO detection algorithms with different antenna and modulation configurations. The rASIP is mainly constructed using a coarse-grained reconfigurable architecture (CGRA) coupled with a processor. In MIMO detection, some important computation steps (e.g., preprocessing) or even the whole detection algorithm is realized using matrix operations. Therefore, for the rASIP, the CGRA is designed to efficiently support different matrix operations used in MIMO detection, and the processor is integrated with special instructions to implement the control path required by different algorithms. Feasibility of the proposed approach is shown by implementing three noniterative MIMO detection algorithms. To evaluate the flexibility of the proposed approach, a Markov Chain Monte Carlo based MIMO detection is also realized by mapping part of the algorithm by using matrix operations on the CGRA. Postlayout results of the rASIP are generated for the implemented detection algorithms on a 65-nm CMOS technology. Compared with some selected designs based on programmable architectures and dedicated application-specified integrated circuits (ASICs), we show that following the proposed approach, the rASIP-based multimode MIMO detection, is about 1.6-5.4 times more efficient than the programmable architectures, and it approaches the throughput performance to the dedicated ASICs.
Andreas Minwegen, Bilal Syed Hussain, Anupam Chattopadhyay, Gerd Ascheid, Rainer Leupers
IEEE Trans. Very Large Scale Integr. Syst.4
2014 Efficient and scalable CGRA-based implementation of Column-wise Givens Rotation
abstract
Givens Rotation is a key computation-intensive block in embedded wireless applications. In order to achieve an efficient mapping which smoothly scales to the underlying architecture, we propose two new Column-based Givens Rotation algorithms, derived from traditional Fast Givens and Square-root and Division Free Givens algorithms. These algorithms allow annihilation of multiple elements in a column of the input matrix simultaneously, without a dependency bottle-neck allowing increased parallelism, resource sharing and scalability. The ease of mapping and scalability has been tested on a layered coarse-grained reconfigurable architecture reaching close to optimal results for highly parallel architectures.
Zoltán Endre Rákossy, Farhad Merchant, Axel Acosta-Aponte, S. K. Nandy 0001, Anupam Chattopadhyay
ASAP5
2014 Efficient Hardware Accelerator for AEGIS-128 Authenticated Encryption
Debjyoti Bhattacharjee, Anupam Chattopadhyay
Inscrypt2
2014 System-level reliability exploration framework for heterogeneous MPSoC
abstract
Power density of digital circuits increased at alarming rate for deep sub-micron CMOS technology, turning reliability into a serious design concern. On the other hand, ever-growing task complexity with strict performance budget forced designers to adopt complex, heterogeneous MPSoCs as the implementation choice. Several commercial system-level design platforms exist currently for design, exploration and implementation of MPSoC. In this paper, we propose a system-level reliability exploration framework by extending a commercial system-level design flow. Using this framework, a heterogeneous MPSoC is designed which can accept a custom mapping algorithm based on the MPSoC topology before the actual task deployment. The dynamic reliability-aware task management is able to consider the desired reliability constraints of tasks as well as reliability levels of the system components. We report our experimental findings using state-of-the-art benchmark applications.
Zheng Wang 0020, Chao Chen 0022, Piyush Sharma, Anupam Chattopadhyay
ACM Great Lakes Symposium on VLSI4
2014 Constructive Reversible Logic Synthesis for Boolean Functions with Special Properties
Anupam Chattopadhyay, Soumajit Majumder, Chander Chandak, Nahian Chowdhury
RC1
2014 Scalable and energy-efficient reconfigurable accelerator for column-wise givens rotation
abstract
A new layered reconfigurable architecture is proposed which exploits modularity, scalability and flexibility to achieve high energy efficiency and memory bandwidth. Using two flavors of Column-wise Givens rotation, derived from traditional Fast Givens and Square root and Division Free Givens Rotation algorithms the architecture is thoroughly evaluated for scalability, speed, area and energy. Combining an efficient mapping strategy of the highly parallel algorithms capable of annihilation of multiple elements of a column of the input matrix and using the new features of the architecture, 9 architectural variants were explored achieving a clean trade-off of execution speed versus area, while keeping relatively constant energy.
Zoltán Endre Rákossy, Farhad Merchant, Axel Acosta-Aponte, S. K. Nandy 0001, Anupam Chattopadhyay
VLSI-SoC5
2013 CoARX: a coprocessor for ARX-based cryptographic algorithms
abstract
Cryptographic coprocessors are inherent part of modern System-on-Chips. It serves dual purpose - efficient execution of cryptographic kernels and supporting protocols for preventing IP-piracy. Flexibility in such coprocessors is required to provide protection against emerging cryptanalytic schemes and to support different cryptographic functions like encryption and authentication. In this context, a novel crypto-coprocessor, named CoARX, supporting multiple cryptographic algorithms based on Addition (A), Rotation (R) and eXclusive-or (X) operations is proposed. CoARX supports diverse ARX-based cryptographic primitives. We show that compared to dedicated hardware implementations and general-purpose microprocessors, it offers excellent performance-flexibility trade-off including adaptability to resist generic cryptanalysis.
Khawar Shahzad, Ayesha Khalid, Zoltán Endre Rákossy, Goutam Paul 0001, Anupam Chattopadhyay
DAC5
2013 High-level modeling and synthesis for embedded FPGAs
abstract
The fast evolving applications in modern digital signal processing have an increasing demand for components which have high computational power and energy efficiency without compromising the flexibility. Embedded FPGA, which is the customized FPGA with heterogeneous fine-grained application specific operations and routing resources, has shown significantly improved efficiency in terms of throughput, power dissipation and chip area for the target application domain. On the other hand, the complexity of such architecture makes it difficult to perform an efficient architecture exploration and application synthesis without tool support. In this work, we propose a framework for the design of embedded FPGA (eFPGA) architectures, which is extended from an existing framework for Coarse-Grained Reconfigurable Architectures (CGRAs). The framework is composed of a high-level modeling formalism for eFPGAs to explore the mapping space, and a retargetable application synthesis flow. To enable fast design space exploration, a force-directed placement algorithm is proposed. Finally, we demonstrate the efficacy of this framework with demanding application kernels.
Jochen Schleifer, Thomas Coenen, Anupam Chattopadhyay, Gerd Ascheid, Tobias G. Noll
DATE5
2013 Accurate and efficient reliability estimation techniques during ADL-driven embedded processor design
abstract
The downscaling of technology features has brought the system developers an important design criteria, reliability, into prime consideration. Due to external radiation effects and temperature gradients, the CMOS device is not guaranteed anymore to function flawlessly. On the other hand, admission for errors to occur allows extending the power budget. The power-performance-reliability trade-off compounds the system design challenge, for which efficient design exploration framework is needed. In this work, we present a high-level processor design framework extended with two reliability estimation techniques. First, a simulation-based technique, which allows a generic instruction-set simulator to estimate reliability via high-level fault injection capability. Second, a novel analytical technique, which is based on the reliability model for coarse arithmetic logical operator blocks within a processor instruction. The techniques are tested with a RISC processor and several embedded application kernels. Our results show the efficiency and accuracy of these techniques against a HDL-level reliability estimation framework.
Zheng Wang 0020, Kapil Singh, Chao Chen 0022, Anupam Chattopadhyay
DATE4
2013 SI-DFA: Sub-expression integrated Deterministic Finite Automata for Deep Packet Inspection
abstract
Finite automata is widely used for Deep Packet Inspection (DPI) of network traffic. Two types of automata employed for this purpose are Non-deterministic Finite Automata (NFA) and Deterministic Finite Automata (DFA). An NFA suffers from a large memory bandwidth per character due to multiple active states. A DFA, in comparison, ensures a linear processing time of O(1) for memory based architectures. However, the DFA state explosion conditions commonly occurring in today's NIDS rule-sets, render the automata with practically infeasible memory space requirements. To avoid state blowup we propose a semi-deterministic automata, Sub-expression Integrated DFA (SI-DFA), that ensures processing time of a single standard DFA. Rules are broken into sub-expressions at blowup conditions and compiled into a single DFA along with an association table, to correctly encapsulate equivalent automata. We list the rare cases in regular expressions for which sub-expression Integration is incorrect and present methodology to detect their occurrences. We evaluate SI-DFA on real-world rule-sets like Bro, Snort and Linux filters and compare their performance with the state-of-the-art hybrid automata solutions. SI-DFA renders a 66-97% reduction in processing bandwidth, up to 68% lower space requirement and an improvement trend with increasing rule complexity when compared to the traditional solutions.
Ayesha Khalid, Rajat Sen, Anupam Chattopadhyay
HPSR3
2013 High-Performance Hardware Implementation for RC4 Stream Cipher
abstract
RC4 is the most popular stream cipher in the domain of cryptology. In this paper, we present a systematic study of the hardware implementation of RC4, and propose the fastest known architecture for the cipher. We combine the ideas of hardware pipeline and loop unrolling to design an architecture that produces 2 RC4 keystream bytes per clock cycle. We have optimized and implemented our proposed design using VHDL description, synthesized with 130, 90, and 65 nm fabrication technologies at clock frequencies 625 MHz, 1.37 GHz, and 1.92 GHz, respectively, to obtain a final RC4 keystream throughput of 10, 21.92, and 30.72 Gbps in the respective technologies.
Sourav Sen Gupta 0001, Anupam Chattopadhyay, Koushik Sinha, Subhamoy Maitra, Bhabani P. Sinha
IEEE Trans. Computers2
2012 FLEXDET: Flexible, Efficient Multi-Mode MIMO Detection Using Reconfigurable ASIP
abstract
This paper describes the implementation of a multi-mode MIMO detector based on the concept of partially reconfigurable ASIP (rASIP). The multi-mode detector can support three different detection algorithms which are the Maximum Ratio Combining, the linear Minimum Mean Square Error (MMSE) detection, and the MMSE Successive Interference Cancellation. The detection algorithms also support different antenna configurations and modulation schemes. The rASIP is based on a Coarse-Grained Reconfigurable Architecture (CGRA), which is designed for efficient architectural support of matrix operations. A matrix inversion algorithm, which is used for the preprocessing of different detection algorithms, is mapped on the CGRA. By integrating a processor with the CGRA, the variations in the control path of different algorithm configurations can be handled efficiently. To the best of our knowledge, we show, for the first time that, a CGRA-based multi-mode MIMO detection is extremely efficient and matches the performance of dedicated ASIC implementation.
Andreas Minwegen, Yahia Hassan, David Kammler, Torsten Kempf, Anupam Chattopadhyay, Gerd Ascheid
FCCM7
2012 Designing high-throughput hardware accelerator for stream cipher HC-128
abstract
Due to ubiquitous deployment of embedded systems, security and privacy are emerging as major design concerns and new stream ciphers are being proposed by the cryptographic community. HC-128 is one of the recent stream ciphers that received attention after its selection as an eStream candidate. Till date, the cipher is believed to have a good security margin. In this paper we study several implementation issues for HC-128 in a disciplined manner. We first discuss the experience on embedded and customizable processors. Then we consider a dedicated hardware accelerator implementation. Further we explore several parallelization strategies for improving throughput. To the best of our knowledge such a detailed implementation exercise has not been presented in the literature. Our novel implementation strategies mark the fastest HC-128 execution reported till date.
Anupam Chattopadhyay, Ayesha Khalid, Subhamoy Maitra, Shashwat Raizada
ISCAS1
2012 Exploring security-performance trade-offs during hardware accelerator design of stream cipher RC4
Anupam Chattopadhyay, Goutam Paul 0001
VLSI-SoC1
2011 Combinational logic synthesis for material implication
abstract
The smooth scaling of technology over past decades is returning diminished profits as researchers are trying to cope with several challenges posed by CMOS devices. As a result, quest for novel physical media for storage and computing is currently an important research pursuit. Recently a new kind of passive electrical device called memristor is proposed, which can retain its state via the resistance in a non-volatile fashion. It is also experimentally demonstrated to perform material implication, a fundamental logical operation. The capability of a memristive device to do logical operations as well as to retain its state makes it a promising candidate for future technologies. In this paper, we investigate the approximate implementation cost of a multi-level combinational logic while using memristive switches as the target technology. Traditional synthesis algorithms are extended and new heuristics are suggested to reduce the costs significantly.
Anupam Chattopadhyay, Zoltán Endre Rákossy
VLSI-SoC1
2008 High-level Modelling and Exploration of Coarse-grained Re-configurable Architectures
abstract
The increasing complexity of today's multimedia and wireless applications is motivating the system designers to innovate continuously. With the challenge to keep various performance metrics in a tight balance while designing a complex system, an entire range of components are now being offered as choices for system building blocks. Coarse-Grained Re-configurable Architecture (CGRA), a strongly emerging class, is currently receiving due attention for offering excellent performance as well as flexibility post fabrication. Compared to the programmable and flexible microprocessors these architectures are shown to yield stronger performance, especially in case of regular and data-driven applications. A variety of system designs are proposed of late, with CGRA as one of the key building blocks. Most of the research initiatives taken in this area have resorted to a template-based approach, where the structure of the re-configurable architecture is partially fixed with several tunable parameters. In this paper, we present a language-driven modelling and exploration framework for CGRAs. In the domain of CGRAs, this framework attempts to bring modelling ease, genericity, early exploration and path to implementation together. The modelling formalism proposed in this paper as well as the exploration capabilities are demonstrated via experiments with several algorithmic kernels.
Anupam Chattopadhyay, Harold Ishebabi, Rainer Leupers, Gerd Ascheid, Heinrich Meyr
DATE1
2008 Prefabrication and postfabrication architecture exploration for partially reconfigurable VLIW processors
abstract
Modern application-specific instruction-set processors (ASIPs) face the daunting task of delivering high performance for a wide range of applications. For enhancing the performance, architectural features, for example, pipelining, VLIW, are often employed in ASIPs, leading to high design complexity. Integrated ASIP design environments, like template-based approaches and language-driven approaches, provide an answer to this growing design complexity. At the same time, increasing hardware design costs have motivated the processor designers to introduce high flexibility in the processor. Flexibility, in its most effective form, can be introduced to the ASIP by coupling a reconfigurable unit to the base processor. Because of its obvious benefits, several reconfigurable ASIPs (rASIPs) have been designed for years. This design paradigm gained momentum with the advent of coarse-grained FPGAs, where the lack of domain-specific performance common in general-purpose FPGAs are largely overcome by choosing application-dependent basic functional units. These rASIP designs lack a generic flow from high-level specification, resulting in intuitive design decisions and hard-to-retarget processor design tools. Although partial, template-based approaches for rASIP design is existent, a clear design methodology especially for the prefabrication architecture exploration is not present. In order to address this issue, a high-level specification and design methodology for partially reconfigurable VLIW processors is proposed in this article. To show the benefit of this approach, a commercial VLIW processor is used as the base architecture and two domains of applications are studied for potential performance gain.
Anupam Chattopadhyay, Harold Ishebabi, Zoltán Endre Rákossy, Kingshuk Karuri, David Kammler, Rainer Leupers, Gerd Ascheid, Heinrich Meyr
ACM Trans. Embed. Comput. Syst.1
2008 A Design Flow for Architecture Exploration and Implementation of Partially Reconfigurable Processors
abstract
During the last years, the growing application complexity, design, and mask costs have compelled embedded system designers to increasingly consider partially reconfigurable application-specific instruction set processors (rASIPs) which combine a programmable base processor with a reconfigurable fabric. Although such processors promise to deliver excellent balance between performance and flexibility, their design remains a challenging task. The key to the successful design of a rASIP is combined architecture exploration of all the three major components: the programmable core, the reconfigurable fabric, and the interfaces between these two. This work presents a design flow that supports fast architecture exploration for rASIPs. The design flow is centered around a unified description of an entire rASIP in an architecture description language (ADL). This ADL description facilitates consistent modeling and exploration of all three components of a rASIP through automatic generation of the software tools (compiler tool chain and instruction set simulator) and the RTL hardware model. The generated software tools and the RTL model can be used either for final implementation of the rASIP or can serve as a preoptimized starting point for implementation that can be hand optimized afterward. The design flow is further enhanced by a number of automatic application analysis tools, including a fine-grained application profiler, an instruction set extension (ISE) generator, and a data path mapper for coarse grained reconfigurable architectures (CGRAs). We present some case studies on embedded benchmarks to show how the design space exploration process helps to efficiently design an application domain specific rASIP.
Kingshuk Karuri, Anupam Chattopadhyay, David Kammler, Ling Hao, Rainer Leupers, Heinrich Meyr, Gerd Ascheid
IEEE Trans. Very Large Scale Integr. Syst.2
2007 Design space exploration of partially re-configurable embedded processors
abstract
In today's embedded processors, performance and flexibility have become the two key attributes. These attributes are often conflicting. The best performance is obtained from custom designed integrated circuits. In contrast, the maximum flexibility is delivered by a general purpose processor. Among the architecture types emerged over the past years to strike an optimum balance between these two attributes, two are prominent. The first ones are field programmable gate array (FPGA)-based architectures and the second ones are application-specific instruction-set processors (ASIPs). Depending on the type of application (i.e. stream-like or control-dominated) either one of the above mentioned architecture types is able to deliver high performance or flexibility or both. Consequently, a new design approach with partial re-configurability on the application-specific processor is attracting strong research interest. We call this architecture re-configurable ASIP (rASIP). Currently, the lack of a high-level abstraction of the rASIP limits the designer from trying out various design alternatives because of long and tedious exploration cycles. To address this issue, in this paper, a high-level specification for re-configurable processors is proposed. Furthermore, a seamless design space exploration methodology using this specification is proposed
Anupam Chattopadhyay, W. Ahmed, Kingshuk Karuri, David Kammler, Rainer Leupers, Gerd Ascheid, Heinrich Meyr
DATE1
2007 Increasing data-bandwidth to instruction-set extensions through register clustering
abstract
The conflicting requirements of performance and flexibility in today 's embedded system market are forcing system designers to use more and more of the so called configurable or customizable processor cores. Such processors tend to meet the demanding performance constraints by accommodating application specific instruction set extensions (ISEs) which have, naturally, become a vital component of current processor customization flows. One major bottleneck in maximizing ISE performance is the limitation on the data-bandwidth between the general purpose register (GPR) file and the ISEs. For improved performance, it is desirable to have a large data-bandwidth from the GPRs to ISEs. However, the tight area constraints of modern embedded processors often restrict the GPR I/O of ISEs to save port area of the register files. This paper presents a novel approach to increase the GPR I/O of ISEs without significantly increasing the size of the GPR files. This is achieved by applying the concept of register clustering, common in many VLIW architectures, to single-issue processors with high performance ISEs. Such clustering often causes extra register moves in compiled code. This work also presents an algorithm to minimize such register moves. The benchmark results presented in this paper show that our solution can significantly reduce the area overhead of many-port GPR files without sacrificing the performance improvements through ISEs.
Kingshuk Karuri, Anupam Chattopadhyay, Manuel Hohenauer, Rainer Leupers, Gerd Ascheid, Heinrich Meyr
ICCAD2
2006 Automatic ADL-based operand isolation for embedded processors
abstract
Cutting-edge applications of future embedded systems demand highest processor performance with low power consumption to get acceptable battery-life times. Therefore, low power optimization techniques are strongly applied during the development of modern application specific instruction set processors (ASIPs). Electronic system level design tools based on architecture description languages (ADL) offer a significant reduction in design time and effort by automatically generating the software tool-suite as well as the register transfer level (RTL) description of the processor. In this paper, the automation of power optimization in ADL-based RTL generation is addressed. Operand isolation is a well-known power optimization technique applicable at all stages of processor development. With increasing design complexity several efforts have been undertaken to automate operand isolation. In pipelined datapaths, where isolating signals are often implicitly available, the traditional RTL-based approach introduces unnecessary overhead. We propose an approach which extracts high-level structural information from the ADL representation and systematically uses the available control signals. Our experiments with state-of-the-art embedded processors show a significant power reduction (improvement in power efficiency)
Anupam Chattopadhyay, Benedikt Geukes, David Kammler, Ernst Martin Witte, Oliver Schliebusch, Harold Ishebabi, Rainer Leupers, Gerd Ascheid, Heinrich Meyr
DATE1
2005 Instruction Set Customization of Application Specific Processors for Network Processing: A Case Study
abstract
The growth of the Internet in the last decade has made current networking applications immensely complex. Systems running such applications need special architectural support to meet the tight constraints of power and performance. This paper presents a case study of architecture exploration and optimization of an application specific instruction set processor (ASIP) for networking applications. The case study particularly focuses on the effects of instruction set customization for applications from different layers of the protocol stack. Using a state-of-the-art VLIW processor as the starting template, and architecture description language (ADL) based architecture exploration tools, this case study suggests possible instruction set and architectural modifications that can speed-up some networking applications up to 6.8 times. Moreover, this paper also shows that there exist very few similarities between diverse networking applications. Our results suggest that, it is extremely difficult to have a common set of architectural features for efficient network protocol processing and, ASIPs with specialized instruction sets can become viable solutions for such an application domain.
Mohammad Mostafizur Rahman Mozumdar, Kingshuk Karuri, Anupam Chattopadhyay, Stefan Kraemer, Hanno Scharwächter, Heinrich Meyr, Gerd Ascheid, Rainer Leupers
ASAP3
2005 A framework for automated and optimized ASIP implementation supporting multiple hardware description languages
abstract
Architecture Description Languages (ADLs) are widely used to perform design space exploration for Application Specific Instruction Set Processors (ASIPs). While the design space exploration is well supported by numerous tools providing high flexibility and quality, the methodology of automated implementation is limited to simple transformations. Assuming fixed architectural templates, information given in the ADL is directly mapped to a hardware description on Register Transfer Level (RTL). Gate-Level synthesis tools are not able to perform potential optimizations, as the computational complexity grows exponential with the size of the architecture. Information such as exclusiveness, parallelism or boolean relations are spread over multiple modules and therefore hard to determine. In this paper, we present an ASIP synthesis approach from architecture description languages, based on an Intermediate Representation (IR). The IR is the key technology to provide new language-independent high-level optimizations and to realize different hardware description language backends. The feasibility of our approach is proven in a case-study.
Oliver Schliebusch, Anupam Chattopadhyay, David Kammler, Gerd Ascheid, Rainer Leupers, Heinrich Meyr, Tim Kogel
ASP-DAC2
2005 Applying Resource Sharing Algorithms to ADL-driven Automatic ASIP Implementation
abstract
Presently, architecture description languages (ADLs) are widely used to raise the abstraction level of the design space exploration of application specific instruction-set processors (ASIPs), benefiting from automatically generated software tool suite and RTL implementation. The increase of abstraction level and automated implementation traditionally comes at the cost of low area, delay or power efficiency. The standard synthesis flow starting at RTL abstraction fails to compensate for this loss of performance. Thus, high level optimizations during RTL synthesis from ADLs are obligatory. Currently, ADL-based optimization schemes do not perform resource sharing. In this paper, we present an iterative algorithm for performing resource sharing on the basis of global dataflow graph matching criteria. This ADL-based resource sharing optimization is performed over a RISC and a VLIW architecture and two industrial embedded processors. The results indicate a significant improvement in overall performance. A comparative study with manually written RTL code is presented, too.
Ernst Martin Witte, Anupam Chattopadhyay, Oliver Schliebusch, David Kammler
ICCD2
2004 RTL Processor Synthesis for Architecture Exploration and Implementation
abstract
Architecture description languages are widely used to perform architecture exploration for application-driven designs, whereas the RT-level is the commonly accepted level for hardware implementation. For this reason, design parameters such as timing, area or power consumption cannot be taken into consideration accurately during design space exploration. Design automation tools currently used to bridge this gap are either limited in the flexibility provided or only generate fragments of the architecture. This paper presents a synthesis tool which preserves the full flexibility of the architecture description language LISA, while being able to generate the complete architecture on RT-level using systemC. This paper also presents two real world architecture case studies to prove the feasibility of our approach.
Oliver Schliebusch, Anupam Chattopadhyay, Rainer Leupers, Gerd Ascheid, Heinrich Meyr, Mario Steinert, Gunnar Braun, Achim Nohl
DATE2