VLDB 2026 Research / reviewers in the wild / expert
Farhad Merchant
dblp:141/0649
· DBLP profile ↗
35ranked-venue papers
2as first author
25since 2021 · last 2026
0000-0002-3708-5621ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 33 · 2 first-author · 24 since 2021Software engineering, systems software and programming languages · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NEURAL: An Elastic Neuromorphic Architecture with Hybrid Data-Event Execution and On-the-fly Attention DataflowabstractSpiking neural networks (SNNs) have emerged as a promising alternative to artificial neural networks (ANNs), offering improved energy efficiency by leveraging sparse and event-driven computation. However, existing hardware implementations of SNNs still suffer from the inherent spike sparsity and multi-timestep execution, which significantly increase latency and reduce energy efficiency. This study presents NEURAL, a novel neuromorphic architecture based on a hybrid dataevent execution paradigm by decoupling sparsity-aware processing from neuron computation and using elastic first-in-firstout (FIFO). NEURAL supports on-the-fly execution of spiking QKFormer by embedding its operations within the baseline computing flow without requiring dedicated hardware units. It also integrates a novel window-to-time-to-first-spike (W2TTFS) mechanism to replace average pooling and enable full-spike execution. Furthermore, we introduce a knowledge distillation (KD)-based training framework to construct single-timestep SNN models with competitive accuracy. NEURAL is implemented on a Xilinx Virtex-7 FPGA and evaluated using ResNet-11, QKFResNet-11, and VGG-11. Experimental results demonstrate that, at the algorithm level, the VGG-11 model trained with KD improves accuracy by $3.20 \%$ on CIFAR-10 and $5.13 \%$ on CIFAR-100. At the architecture level, compared to existing SNN accelerators, NEURAL achieves a $50 \%$ reduction in resource utilization and a $1.97 \times$ improvement in energy efficiency. Yuehai Chen, Farhad Merchant |
ASP-DAC | 2 |
| 2026 | A Comprehensive Review of Tsetlin Machines: Concepts, Applications, Analysis, and the FutureabstractWith the rise of large language models, two significant challenges are their high power consumption during training and their black-box nature. The Tsetlin Machine (TM) offers a logic-based, interpretable alternative to traditional Machine Learning (ML) models to address these issues. TM is used in various ML-based edge computing and Internet of Things applications, such as batteryless sensors, resource-constrained intrusion detection, and on-device training. It shows 30.5× reduction in computation cost, 36.6× reduction in storage memory footprint, 13.5× latency and energy reductions, and converges faster than neural networks. It can be implemented on different hardware platforms, including field-programmable gate arrays, superconducting circuits, and memristor-transistor arrays. Filling a gap, we provide a holistic reference containing more than 160 implementations of TMs. In this tutorial and survey, we explain the algorithm of the basic TM in detail with a background of Learning Automata. Next, we discuss different TM variants, including the Regression TM, Federated Learning-based TMs, the Coalesced TM, Real and Integer-Weighted TMs, and the Convolutional TM for text and image classification. We classify these software and hardware implementations based on their algorithmic similarities and compare them using key performance metrics such as accuracy, training time, energy consumption per classification, and operating frequency. Additionally, we discuss our observations on general trends, insights from our comparative analysis, limitations of existing implementations, and potential directions for future research to further extend its applicability in energy-efficient Internet of Things and edge computing applications. Souraja Kundu, Shruti Patkar, Saras Mani Mishra, Gaurav Trivedi, Farhad Merchant |
IEEE Internet Things J. | 5 |
| 2026 | veriSiM: Formal Verification of SPICE Netlists for MAGIC-Based Logic-in-MemoryabstractAdvancements in emerging technologies have recently increased the traction of non-von Neumann design styles. One of the most popular design styles in this domain involves using memristors to perform logic operations in memory, known as Logic-in-Memory (LiM). Memristor Aided Logic (MAGIC) is one of such LiM based design style that is widely used given its benefits in latency and energy. Several prior works have focused on the generation of logic operations, also called microoperations, for LiM based on the MAGIC design style. Recently, the generation of SPICE netlists for MAGIC design style has been achieved by the MemSPICE tool. While this represents a significant step forward, verifying the correctness of the generated netlists still depends on SPICE-level simulations. These simulations become particularly impractical for medium-to-large designs presenting a bottleneck in the validation process. To address this limitation, in this paper, we introduce veriSiM, an automated formal verification methodology for MAGIC-based LiM. More concretely, it ensures the correctness of the generated LiM SPICE netlists against the golden reference Verilog design. Our methodology involves generating clauses from the SPICE netlists and verifying them against clauses generated from the golden reference Verilog design, using the high-performance Z3 solver to perform the equivalence checking. The clause generation process from the SPICE netlists needs to be based on several conditions, which have been identified and discussed in detail. We have used several benchmarks from ISCAS’85, ISCAS’89, and ITC’99 to demonstrate the efficacy of the veri Chandan Kumar Jha 0001, Simranjeet Singh, Khushboo Qayyum, Ankit Bende, Muhammad Hassan 0002, Vikas Rana, Farhad Merchant, Rolf Drechsler |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2025 | Dependable Neuromorphic Computing-in-Memory Architectures
Farhad Merchant, Ankit Bende, Markus Fritscher, Shahar Kvatinsky, Simranjeet Singh, Vikas Rana, Regina Dittmann, Keerthi Dorai Swamy Reddy, Christian Wenger, Fouwad Jamil Mir, Mottaqiallah Taouil, Manil Dev Gomony, Said Hamdioui, Henk Corporaal |
ETS | 1 |
| 2025 | Neuromorphic Edge Computing: Challenges, Opportunities, and Current SolutionsabstractNeuromorphic computing is emerging as a paradigm for high-performance, energy-efficient edge intelligence. Yet the transition from laboratory prototypes to deployable edge platforms is slowed by four intertwined obstacles: (1) complex near-sensor integration, where spiking inference must co-locate with analogue sensing to minimise latency and data-movement energy; (2) novel event-based optimisation, requiring weight compression and sparsity techniques tailored to event-driven workloads; (3) heterogeneous integration of emerging devices, such as RRAM and other non-volatile memories, into reliable, manufacturable stacks; and (4) novel security risks, including spike-pattern side channels and model-specific attacks that demand to develop neuromorphic security primitives. This paper surveys the state of the art across these four fronts, drawing on recent advances in spiking microcontrollers, mixed-precision compute-in-memory fabrics, sparsity-aware compilation, and hardware-anchored security primitives (physical unclonable functions, true random number generators, computing-in-memory-based cryptography). By distilling lessons from academic research and industrial prototyping, the paper outlines current solutions and future research directions aimed at accelerating the adoption of neuromorphic platforms in real-world edge AI systems. Federico Corradi, Amir Zjajo, Letícia Maria Veiras Bolzani, Milos Krstic, Orlando Moreira, Zeqi Zhu, Farhad Merchant |
ISLPED | 7 |
| 2024 | MemSPICE: Automated Simulation and Energy Estimation Framework for MAGIC-Based Logic-in-MemoryabstractExisting logic-in-memory (LiM) research is limited to generating mappings and micro-operations. In this paper, we present MemSPICE, a novel framework that addresses this gap by automatically generating both the netlist and testbench needed to evaluate the LiM on a memristive crossbar. MemSPICE goes beyond conventional approaches by providing energy estimation scripts to calculate the precise energy consumption of the testbench at the SPICE level. We propose an automated framework that utilizes the mapping obtained from the SIMPLER tool to perform accurate energy estimation through SPICE simulations. To the best of our knowledge, no existing framework is capable of generating a SPICE netlist from a hardware description language. By offering a comprehensive solution for SPICE-based netlist generation, testbench creation, and accurate energy estimation, MemSPICE empowers researchers and engineers working on memristor-based LiM to enhance their understanding and optimization of energy usage in these systems. Finally, we tested the circuits from the ISCAS’85 benchmark on MemSPICE and conducted a detailed energy analysis. Simranjeet Singh, Chandan Kumar Jha 0001, Ankit Bende, Vikas Rana, Sachin B. Patkar, Rolf Drechsler, Farhad Merchant |
ASPDAC | 7 |
| 2024 | Error Detection and Correction Codes for Safe In-Memory ComputationsabstractIn-Memory Computing (IMC) introduces a new paradigm of computation that offers high efficiency in terms of latency and power consumption for AI accelerators. However, the non-idealities and defects of emerging technologies used in advanced IMC can severely degrade the accuracy of inferred Neural Networks (NN) and lead to malfunctions in safety-critical applications. In this paper, we investigate an architectural-level mitigation technique based on the coordinated action of multiple checksum codes, to detect and correct errors at run-time. This implementation demonstrates higher efficiency in recovering accuracy across different AI algorithms and technologies compared to more traditional methods such as Triple Modular Redundancy (TMR). The results show that several configurations of our implementation recover more than 91% of the original accuracy with less than half of the area required by TMR and less than 40% of latency overhead. Luca Parrini, Taha Soliman, Benjamin Hettwer, Jan Micha Borrmann, Simranjeet Singh, Ankit Bende, Vikas Rana, Farhad Merchant, Norbert Wehn |
ETS | 8 |
| 2024 | In-Memory Mirroring: Cloning Without ReadingabstractIn-memory computing (IMC) has gained signifi- cant attention recently as it attempts to reduce the impact of memory bottlenecks. Numerous schemes for digital IMC are presented in the literature, focusing on logic operations. Often, an application's description has data dependencies that must be resolved. Contemporary IMC architectures perform read followed by write operations for this purpose, which results in performance and energy penalties. To solve this fundamental problem, this paper presents in-memory mirroring (IMM). IMM eliminates the need for read and write-back steps, thus avoiding energy and performance penalties. Instead, we perform data movement within memory, involving row-wise and column-wise data transfers. Additionally, the IMM scheme enables parallel cloning of entire row (word) with a complexity of O(1). Moreover, we analyzed the energy consumption of the proposed technique on an RRAM crossbar with an experimentally validated JART VCM v1b model. The IMM increases energy efficiency and shows 2x performance improvement compared to conventional data movement methods. Simranjeet Singh, Ankit Bende, Chandan Kumar Jha 0001, Vikas Rana, Rolf Drechsler, Sachin B. Patkar, Farhad Merchant |
VLSI-SoC | 7 |
| 2024 | veriSIMPLER: An Automated Formal Verification Methodology for SIMPLER MAGIC Design Style Based In-Memory ComputingabstractIn-Memory Computing (IMC) using memristors has gained significant interest in recent years as it addresses the issue of memory bottleneck in the von Neumann architectures. One of the most popular design styles that have been developed to perform memristor-based IMC is Memristor-Aided loGIC (MAGIC). MAGIC design style based NOR and NOT operations can be used to perform IMC on memristor crossbars. The state-of-the-art SIMPLER MAGIC tool is used to generate a mapping of any arbitrary Boolean function to MAGIC operations that are suitable for high-throughput applications. The correctness of the mapping is examined by tedious manual inspections and functional simulations which may not account for all the edge cases which is not desirable. In this work, we alleviate this issue to the best of our knowledge for the first time by proposing veriSIMPLER. veriSIMPLER is an automated formal verification methodology to ensure the functional correctness of the mapping obtained using the SIMPLER MAGIC tool. The veriSIMPLER methodology generates Boolean Satisfiability Formulas (SAT) of the mapping obtained using the SIMPLER MAGIC tool and the golden reference Verilog designs. These SAT formulas are verified against each other using the Z3 solver. The veriSIMPLER methodology identified a critical bug in the mapping obtained from the SIMPLER MAGIC tool when buffers are connected between input and output. We also propose a methodology to patch this bug to generate the correct mapping, which in turn extends the capability of the SIMPLER MAGIC tool to handle buffers on top of providing a formal verification methodology. We have used a variety of benchmark circuits from the widely used ISCAS’85, ISCAS’89, ITC’99, and IWLS’93 to show the efficacy of the veriSIMPLER methodology. We aim to make the formally verified mapping obtained using the veriSIMPLER methodology open-source to promote further research in this direction. Chandan Kumar Jha 0001, Khushboo Qayyum, Kemal Çaglar Coskun, Simranjeet Singh, Muhammad Hassan 0002, Rainer Leupers, Farhad Merchant, Rolf Drechsler |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2023 | Hardware Security Primitives Using Passive RRAM Crossbar Array: Novel TRNG and PUF DesignsabstractWith rapid advancements in electronic gadgets, the security and privacy aspects of these devices are significant. For the design of secure systems, physical unclonable function (PUF) and true random number generator (TRNG) are critical hardware security primitives for security applications. This paper proposes novel implementations of PUF and TRNGs on the RRAM crossbar structure. Firstly, two techniques to implement the TRNG in the RRAM crossbar are presented based on write-back and 50% switching probability pulse. The randomness of the proposed TRNGs is evaluated using the NIST test suite. Next, an architecture to implement the PUF in the RRAM crossbar is presented. The initial entropy source for the PUF is used from TRNGs, and challenge-response pairs (CRPs) are collected. The proposed PUF exploits the device variations and sneak-path current to produce unique CRPs. We demonstrate, through extensive experiments, reliability of 100%, uniqueness of 47.78%, uniformity of 49.79%, and bit-aliasing of 48.57% without any post-processing techniques. Finally, the design is compared with the literature to evaluate its implementation efficiency, which is clearly found to be superior to the state-of-the-art. Simranjeet Singh, Furqan Zahoor, Gokulnath Rajendran, Sachin B. Patkar, Anupam Chattopadhyay, Farhad Merchant |
ASP-DAC | 6 |
| 2023 | DeepAttack: A Deep Learning Based Oracle-less Attack on Logic LockingabstractLogic locking is one of the most promising design-for-trust technique for protecting intellectual property from reverse engineering, IP piracy, and modification throughout the electronic supply chain. However, oracle-less deobfuscation attacks that do not require an activated chip have been successful in obtaining the secret key of locked designs. This requires a detailed determination of the extent of vulnerability available in obfuscated circuitry. In this paper, we propose the oracle-less DeepAttack: an attack on logic locking that is capable of extracting the activation key of the locked netlist using a deep learning model. Based on the ISCAS-85 and EPFL benchmarks evaluation, DeepAttack achieves an average key prediction accuracy of 93.39%, outperforming the oracle-less state-of-the-art attacks SAIL, SnapShot, and OMLA by 21.28, 10.73, and 3.84 percentage points, respectively. Anand Raj, Nikhitha Avula, Pabitra Das, Dominik Germek, Farhad Merchant, Amit Acharyya |
ISCAS | 5 |
| 2023 | IMBUE: In-Memory Boolean-to-CUrrent Inference ArchitecturE for Tsetlin MachinesabstractIn-memory computing for Machine Learning (ML) applications remedies the von Neumann bottlenecks by organizing computation to exploit parallelism and locality. Non-volatile memory devices such as Resistive RAM (ReRAM) offer integrated switching and storage capabilities showing promising performance for ML applications. However, ReRAM devices have design challenges, such as nonlinear digital-analog conversion and circuit overheads. This paper proposes an In-Memory Boolean-to-Current Inference Architecture (IMBUE) that uses ReRAM-transistor cells to eliminate the need for such conversions. IMBUE processes Boolean feature inputs expressed as digital voltages and generates parallel current paths based on resistive memory states. The proportional column current is then translated back to the Boolean domain for further digital processing. The IMBUE architecture is inspired by the Tsetlin Machine (TM), an emerging ML algorithm based on intrinsically Boolean logic. The IMBUE architecture demonstrates significant performance improvements over binarized convolutional neural networks and digital TM in-memory implementations, achieving up to a 12.99x and 5.28x increase, respectively. Omar Ghazal, Simranjeet Singh, Tousif Rahman, Shengqi Yu, Yujin Zheng, Domenico Balsamo, Sachin B. Patkar, Farhad Merchant, Fei Xia 0001, Alexandre Yakovlev, Rishad A. Shafik |
ISLPED | 8 |
| 2023 | PR-PUF: A Reconfigurable Strong RRAM PUFabstractPhysical Unclonable Functions (PUFs) offer the natural advantage of built-in key generation, thus eliminating the costly process of embedding unique key after manufacturing millions of integrated circuits. When PUFs are deployed for the application, all the Challenge-Response Pairs (CRP) are collected and stored in a trusted server, and the responses are compared with the one from the device during the run-time - forming the crux of various security protocols. Two issues are commonly faced during PUF designs. First, to enhance the applicability of PUF, larger set of CRP is desirable, which is referred to as a strong PUF. Second, due to the emergence of machine learning-based PUF modelling attacks, it is now imperative to have a PUF demonstrating resistance against such attacks. In this paper, we propose a novel Parity Resistive RAM PUF (PR-PUF) implemented using RRAM crossbar architecture. PR-PUF supports low-overhead reconfiguration, where both the original and reconfigured CRP space enhances the CRP size, with average uniqueness between reconfiguration of 49.98%. The construction also demonstrates excellent robustness against various modeling attacks. We present detailed design analysis and circuit-level simulation studies. Gokulnath Rajendran, Furqan Zahoor, Simranjeet Singh, Farhad Merchant, Vikas Rana, Anupam Chattopadhyay |
VLSI-SoC | 4 |
| 2023 | SoftFlow: Automated HW-SW Confidentiality Verification for Embedded ProcessorsabstractDespite its ever-increasing impact, security is not considered as a design objective in commercial electronic design automation (EDA) tools. This results in vulnerabilities being overlooked during the software-hardware design process. Specifically, vulnerabilities that allow leakage of sensitive data might stay unnoticed by standard testing, as the leakage itself might not result in evident functional changes. Therefore, EDA tools are needed to elaborate the confidentiality of sensitive data during the design process. However, state-of-the-art implementations either solely consider the hardware or restrict the expressiveness of the security properties that must be proven. Consequently, more proficient tools are required to assist in the software and hardware design. To address this issue, we propose SoftFlow, an EDA tool that allows determining whether a given software exploits existing leakage paths in hardware. Based on our analysis, the leakage paths can be retained if proven not to be exploited by software. This is desirable if the removal significantly impacts the design’s performance or functionality, or if the path cannot be removed as the chip is already manufactured. We demonstrate the feasibility of SoftFlow by identifying vulnerabilities in OpenSSL cryptographic C programs, and redesigning them to avoid leakage of cryptographic keys in a RISC-V architecture. Lennart M. Reimann, Jonathan Wiesner, Dominik Germek, Farhad Merchant, Rainer Leupers |
VLSI-SoC | 4 |
| 2023 | CLARINET: A quire-enabled RISC-V-based framework for posit arithmetic empiricism
Niraj N. Sharma, Riya Jain, Mohana Madhumita Pokkuluri, Sachin B. Patkar, Rainer Leupers, Rishiyur S. Nikhil, Farhad Merchant |
J. Syst. Archit. | 7 |
| 2023 | FPGA-based Deep Learning Inference Accelerators: Where Are We Standing?abstractRecently, artificial intelligence applications have become part of almost all emerging technologies around us. Neural networks, in particular, have shown significant advantages and have been widely adopted over other approaches in machine learning. In this context, high processing power is deemed a fundamental challenge and a persistent requirement. Recent solutions facing such a challenge deploy hardware platforms to provide high computing performance for neural networks and deep learning algorithms. This direction is also rapidly taking over the market. Here, FPGAs occupy the middle ground regarding flexibility, reconfigurability, and efficiency compared to general-purpose CPUs, GPUs, on one side, and manufactured ASICs on the other. FPGA-based accelerators exploit the features of FPGAs to increase the computing performance for specific algorithms and algorithm features. Filling a gap, we provide holistic benchmarking criteria and optimization techniques that work across several classes of deep learning implementations. This article summarizes the current state of deep learning hardware acceleration: More than 120 FPGA-based neural network accelerator designs are presented and evaluated based on a matrix of performance and acceleration criteria, and corresponding optimization techniques are presented and discussed. In addition, the evaluation criteria and optimization techniques are demonstrated by benchmarking ResNet-2 and LSTM-based accelerators. Anouar Nechi, Lukas Groth, Saleh Mulhem, Farhad Merchant, Rainer Buchty, Mladen Berekovic |
ACM Trans. Reconfigurable Technol. Syst. | 4 |
| 2022 | NEUROTEC I: Neuro-inspired Artificial Intelligence Technologies for the Electronics of the FutureabstractThe field of neuromorphic computing is approaching an era of rapid adoption driven by the urgent need of a substitute for the von Neumann computing architecture. NEUROTEC I: “Neuro-inspired Artificial Intelligence Technologies for the Elec-tronics of the Future” project is an initiative sponsored by the German Federal Ministry of Education and Research (BMBF for its initials in German), that aims to effectively advance the foundations for the utilization and exploitation of neuromorphic computing. NEUROTEC I stands at its successful “final stage” driven by the collaboration from more than 8 institutes from the Jiilich Research Center and the RWTH Aachen University, as well as collaboration from several high-tech industry partners. The NEUROTEC I project considers the field interplay among materials, circuits, design and simulation tools. This paper provides an overview of the project's overall structure and discusses the scientific achievements of its individual activities. Melvin Galicia, Stephan Menzel, Farhad Merchant, Maximilian Müller, Qing-Tai Zhao, Felix Cüppers, Abdur R. Jalil, Qi Shu, Peter Schüffelgen, Gregor Mussler, Carsten Funck, Christian Lanius, Stefan Wiefels, Moritz von Witzleben, Christopher Bengel, Nils Kopperberg, Tobias Ziegler 0005, R. Walied Ahmad, Alexander Krüger, Letícia Maria Veiras Bolzani, Regina Dittmann, Susanne Hoffmann-Eifert, Vikas Rana, Detlev Grützmacher, Matthias Wuttig, Dirk J. Wouters, Andrei Vescan, Tobias Gemmeke, Joachim Knoch, Max Christian Lemme, Rainer Leupers, Rainer Waser |
DATE | 3 |
| 2022 | NeuroHammer: Inducing Bit-Flips in Memristive Crossbar MemoriesabstractEmerging non-volatile memory (NVM) technologies offer unique advantages in energy efficiency, latency, and features such as computing-in-memory. Consequently, emerging NVM technologies are considered an ideal substrate for computation and storage in future-generation neuromorphic platforms. These technologies need to be evaluated for fundamental reliability and security issues. In this paper, we present NeuroHammer, a security threat in ReRAM crossbars caused by thermal crosstalk between memory cells. We demonstrate that bit-flips can be deliberately induced in ReRAM devices in a crossbar by systematically writing adjacent memory cells. A simulation flow is developed to evaluate NeuroHammer and the impact of physical parameters on the effectiveness of the attack. Finally, we discuss the security implications in the context of possible attack scenarios. Felix Staudigl, Hazem Al Indari, Daniel Schön, Dominik Germek, Farhad Merchant, Jan Moritz Joseph, Vikas Rana, Stephan Menzel, Rainer Leupers |
DATE | 5 |
| 2022 | PA-PUF: A Novel Priority Arbiter PUFabstractThis paper proposes a 3-input arbiter-based novel physically unclonable function (PUF) design. Firstly, a 3-input priority arbiter is designed using a simple arbiter, two multiplexers (2:1), and an XOR logic gate. The priority arbiter has an equal probability of 0’s and 1’s at the output, which results in excellent uniformity (49.45%) while retrieving the PUF response. Secondly, a new PUF design based on priority arbiter PUF (PA-PUF) is presented. The PA-PUF design is evaluated for uniqueness, non-linearity, and uniformity against the standard tests. The proposed PA-PUF design is configurable in challenge-response pairs through an arbitrary number of feed-forward priority arbiters introduced to the design. We demonstrate, through extensive experiments, reliability of 100% after performing the error correction techniques and uniqueness of 49.63%. Finally, the design is compared with the literature to evaluate its implementation efficiency, where it is clearly found to be superior compared to the state-of-the-art. Simranjeet Singh, Srinivasu Bodapati 0001, Sachin B. Patkar, Rainer Leupers, Anupam Chattopadhyay, Farhad Merchant |
VLSI-SoC | 6 |
| 2022 | Deceptive Logic Locking for Hardware Integrity Protection Against Machine Learning AttacksabstractLogic locking has emerged as a prominent key-driven technique to protect the integrity of integrated circuits. However, novel machine-learning-based attacks have recently been introduced to challenge the security foundations of locking schemes. These attacks are able to recover a significant percentage of the key without having access to an activated circuit. This article address this issue through two focal points. First, we present a theoretical model to test locking schemes for key-related structural leakage that can be exploited by machine learning. Second, based on the theoretical model, we introduce D-MUX: a deceptive multiplexer-based logic-locking scheme that is resilient against structure-exploiting machine learning attacks. Through the design of D-MUX, we uncover a major fallacy in the existing multiplexer-based locking schemes in the form of a structural-analysis attack. Finally, an extensive cost evaluation of D-MUX is presented. To the best of our knowledge, D-MUX is the first machine-learning-resilient locking scheme capable of protecting against all known learning-based attacks. Hereby, the presented work offers a starting point for the design and evaluation of future-generation logic locking in the era of machine learning. Dominik Germek, Farhad Merchant, Lennart M. Reimann, Rainer Leupers |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2021 | Nano Security: From Nano-Electronics to Secure SystemsabstractThe field of computer hardware stands at the verge of a revolution driven by recent breakthroughs in emerging nanodevices. “Nano Security” is a new Priority Program recently approved by DFG, the German Research Council. This initial-stage project initiative at the crossroads of nano-electronics and hardware-oriented security includes 11 projects with a total of 23 Principal Investigators from 18 German institutions. It considers the interplay between security and nano-electronics, focusing on a dichotomy which emerging nano-devices (and their architectural implications) have on system security. The projects within the Priority Program consider both: potential security threats and vulnerabilities stemming from novel nano-electronics, and innovative approaches to establishing and improving system security based on nano-electronics. This paper provides an overview of the Priority Program's overall philosophy and discusses the scientific objectives of its individual projects. Ilia Polian, Frank Altmann, Tolga Arul, Christian Boit, Ralf Brederlow, Lucas Davi, Rolf Drechsler, Nan Du 0004, Thomas Eisenbarth 0001, Tim Güneysu, Sascha Hermann, Matthias Hiller, Rainer Leupers, Farhad Merchant, Thomas Mussenbrock, Stefan Katzenbeisser 0001, Akash Kumar 0001, Wolfgang Kunz, Thomas Mikolajick, Vivek Pachauri, Jean-Pierre Seifert, Frank Sill, Jens Trommer |
DATE | 14 |
| 2021 | Vertical IP Protection of the Next-Generation Devices: Quo Vadis?abstractWith the advent of 5G and IoT applications, there is a greater thrust in terms of hardware security due to imminent risks caused by high amount of intercommunication between various subsystems. Security gaps in integrated circuits, thus represent high risks for both-the manufacturers and the users of electronic systems. Particularly in the domain of Intellectual Property (IP) protection, there is an urgent need to devise security measures at all levels of abstraction so that we can be one step ahead of any kind of adversarial attacks. This work presents IP protection measures from multiple perspectives-from system-level down to device-level security measures, from discussing various attack methods such as reverse engineering and hardware Trojan insertions to proposing new-age protection measures such as multi-valued logic locking and secure information flow tracking. This special session will give a holistic overview at the current state-of-the-art measures and how well we are prepared for the next generation circuits and systems. Shubham Rai, Siddharth Garg, Christian Pilato, Vladimir Herdt, Elmira Moussavi, Dominik Germek, Ramesh Karri, Rolf Drechsler, Farhad Merchant, Akash Kumar 0001 |
DATE | 9 |
| 2021 | QFlow: Quantitative Information Flow for Security-Aware Hardware Design in VerilogabstractThe enormous amount of code required to design modern hardware implementations often leads to critical vulnerabilities being overlooked. Especially vulnerabilities that compromise the confidentiality of sensitive data, such as cryptographic keys, have a major impact on the trustworthiness of an entire system. Information flow analysis can elaborate whether information from sensitive signals flows towards outputs or untrusted components of the system. But most of these analytical strategies rely on the non-interference property, stating that the untrusted targets must not be influenced by the source’s data, which is shown to be too inflexible for many applications. To address this issue, there are approaches to quantify the information flow between components such that insignificant leakage can be neglected. Due to the high computational complexity of this quantification, approximations are needed, which introduce mispredictions. To tackle those limitations, we reformulate the approximations. Further, we propose a tool QFlow with a higher detection rate than previous tools. It can be used by non-experienced users to identify data leakages in hardware designs, thus facilitating a security-aware design process. Lennart M. Reimann, Luca Hanel, Dominik Germek, Farhad Merchant, Rainer Leupers |
ICCD | 4 |
| 2021 | Logic Locking at the Frontiers of Machine Learning: A Survey on Developments and OpportunitiesabstractIn the past decade, a lot of progress has been made in the design and evaluation of logic locking; a premier technique to safeguard the integrity of integrated circuits throughout the electronics supply chain. However, the widespread proliferation of machine learning has recently introduced a new pathway to evaluating logic locking schemes. This paper summarizes the recent developments in logic locking attacks and countermeasures at the frontiers of contemporary machine learning models. Based on the presented work, the key takeaways, opportunities, and challenges are highlighted to offer recommendations for the design of next-generation logic locking. Dominik Germek, Lennart M. Reimann, Elmira Moussavi, Farhad Merchant, Rainer Leupers |
VLSI-SoC | 4 |
| 2021 | Challenging the Security of Logic Locking Schemes in the Era of Deep Learning: A Neuroevolutionary ApproachabstractLogic locking is a prominent technique to protect the integrity of hardware designs throughout the integrated circuit design and fabrication flow. However, in recent years, the security of locking schemes has been thoroughly challenged by the introduction of various deobfuscation attacks. As in most research branches, deep learning is being introduced in the domain of logic locking as well. Therefore, in this article we present SnapShot, a novel attack on logic locking that is the first of its kind to utilize artificial neural networks to directly predict a key bit value from a locked synthesized gate-level netlist without using a golden reference. Hereby, the attack uses a simpler yet more flexible learning model compared to existing work. Two different approaches are evaluated. The first approach is based on a simple feedforward fully connected neural network. The second approach utilizes genetic algorithms to evolve more complex convolutional neural network architectures specialized for the given task. The attack flow offers a generic and customizable framework for attacking locking schemes using machine learning techniques. We perform an extensive evaluation of SnapShot for two realistic attack scenarios, comprising both reference combinational and sequential benchmark circuits as well as silicon-proven RISC-V core modules. The evaluation results show that SnapShot achieves an average key prediction accuracy of 82.60% for the selected attack scenario, with a significant performance increase of 10.49 percentage points compared to the state of the art. Moreover, SnapShot outperforms the existing technique on all evaluated benchmarks. The results indicate that the security foundation of common logic locking schemes is built on questionable assumptions. Based on the lessons learned, we discuss the vulnerabilities and potentials of logic locking uncovered by SnapShot. The conclusions offer insights into the challenges of designing future logic locking schemes that are resilient to machine learning attacks. Dominik Germek, Farhad Merchant, Lennart M. Reimann, Harshit Srivastava, Ahmed Hallawa, Rainer Leupers |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2020 | Next Generation Arithmetic for Edge ComputingabstractArithmetic is a key component and is ubiquitous in today’s digital world, ranging from embedded to high-performance computing systems. With machine learning at the fore in a wide range of application domains from wearables to automotive to avionics to weather prediction, sufficiently accurate yet low-cost arithmetic is the need for the day. Recently, there have been several advances in the domain of computer arithmetic, which includes high-precision anchored numbers from ARM, posit arithmetic, bfloat16, etc. as an alternative to IEEE 754-2008 compliant arithmetic. Optimizations on fixed-point and integer arithmetic are also being pursued actively for low-power computing architectures. Furthermore, approximate computing and transprecision/mixed-precision computing have been exciting areas of research forever. While academic research in the domain of computer arithmetic has a long history, industrial adoption of some of these new data types and techniques is in its early stages and expected to increase in the future. bfloat16 is an excellent example for this. In this paper, we bring academia and industry together to discuss the latest results and future directions for research in the domain of next-generation computer arithmetic, especially for edge computing. Andre Guntoro, Cecilia De la Parra, Farhad Merchant, Florent de Dinechin, John L. Gustafson, Martin Langhammer, Rainer Leupers, Sangeeth Nambiar |
DATE | 3 |
| 2020 | A secure hardware-software solution based on RISC-V, logic locking and microkernelabstractIn this paper we present the first generation of a secure platform developed by following a security-by-design approach. The security of the platform is built on top of two pillars: a secured hardware design flow and a secure microkernel. The hardware design is protected against the insertion of hardware Trojans during the production phase through netlist obfuscation provided by logic locking. The software stack is based on a trustworthy and verified microkernel. Moreover, the system is expected to work in an environment which does not allow physical access to the device. Therefore, on-the-field attacks are only possible via software. We present a solution whose security has been achieved by relying on simple and open hardware and software solutions, namely a RISC-V processor core, open-source peripherals and an seL4--based operating system. Dominik Germek, Farhad Merchant, Lennart M. Reimann, Rainer Leupers, Massimiliano Giacometti, Sascha Kegreiss |
SCOPES | 2 |
| 2019 | Inter-Lock: Logic Encryption for Processor Cores Beyond Module BoundariesabstractThe lack of technical resources and the high cost of establishing a semiconductor foundry has forced most integrated circuit design houses to rely on outsourcing part of their design and fabrication services to off-site companies. The involvement of external parties has given rise to major security threats, ranging from intellectual property piracy to the insertion of malicious circuits. Logic encryption has emerged as a popular mitigation technique against these threats. In recent years, a vast amount of logic encryption algorithms has been proposed. However, existing approaches strongly focus on isolated circuit components without taking the complexity and modular structure of modern circuit designs into account. In this paper, we propose Inter-Lock, a novel logic encryption framework tailored towards scaling logic encryption to larger designs by leveraging their complexity and exploiting multi-module interdependencies to exponentially enhance the security of sequential circuits. To showcase the applicability of the approach, we present an extensive evaluation on a real-life 32-bit RISC-V core together with the analysis of the security-cost trade-off. Our recommended Inter-Lock configuration of the core features a 1024-bit key and 100% functional corruption for 13.2% area overhead, less than 1% power overhead and 22.2% delay penalty in the worst case. Dominik Germek, Farhad Merchant, Rainer Leupers, Gerd Ascheid, Sascha Kegreiss |
ETS | 2 |
| 2019 | Control-Lock: Securing Processor Cores Against Software-Controlled Hardware TrojansabstractMalicious circuit modifications known as hardware Trojans represent a rising threat to the integrated circuit supply chain. As many Trojans are activated based on a specific sequence of circuit states, we have recognized the ease of utilizing an instruction sequence for Trojan activation inside a processor core as a significant security issue. To protect against this threat, we propose Control-Lock: a novel methodology for securing inter-module control signals against software-controlled hardware Trojans, even if the signals are known to the adversary during fabrication. We demonstrate the approach with a RISC-V processor infected with a denial of service Trojan. We evaluate different Control-Lock encryption schemes with regards to the security-cost trade-off. Our results show that protecting a processor against a software-controlled hardware Trojan exploiting code execution implies an area overhead of only 4.75% as well as a negligible delay and power overhead. Dominik Germek, Farhad Merchant, Rainer Leupers, Gerd Ascheid, Sascha Kegreiss |
ACM Great Lakes Symposium on VLSI | 2 |
| 2018 | Parameterized Posit Arithmetic Hardware GeneratorabstractHardware implementation of Floating Point Units (FPUs) has been a key area of research due to their massive area and energy footprints. Recently, a proposal was made to replace IEEE 754-2008 technical standard compliant FPUs with Posit Arithmetic Units (PAUs) due to the greater accuracy, speed, and simpler hardware design. In this paper, we present the architecture of a parameterized PAU generator that can generate PAU adders and PAU multipliers of any bit-width pre-synthesis. We synthesize generated arithmetic units using the parameterized PAU generator for 8-bit, 16-bit, and 32-bit adders and multipliers and compare them with IEEE 754-2008 compliant adders and multipliers. Both, synthesis for Field Programmable Gate Array (FPGA) and Application Specific Integrated Circuit (ASIC) are performed. In our comparison of m-bit PAU units with n-bit IEEE 754-2008 compliant units, it is observed that the area and energy of a PAU adder and multiplier are comparable to their IEEE 754-2008 compliant counterparts where m=n. We argue that an n-bit IEEE 754-2008 adder and multiplier can be safely replaced with an m-bit PAU adder and multiplier where m Rohit Chaurasiya, John L. Gustafson, Rahul Shrestha, Jonathan Neudorfer, Sangeeth Nambiar, Kaustav Niyogi, Farhad Merchant, Rainer Leupers |
ICCD | 7 |
| 2018 | Efficient Realization of Householder Transform Through Algorithm-Architecture Co-Design for Acceleration of QR FactorizationabstractQR factorization is a ubiquitous operation in many engineering and scientific applications. In this paper, we present efficient realization of Householder Transform (HT) based QR factorization through algorithm-architecture co-design where we achieve performance improvement of 3-90x in-terms of Gflops/watt over state-of-the-art multicore, General Purpose Graphics Processing Units (GPGPUs), Field Programmable Gate Arrays (FPGAs), and ClearSpeed CSX700. Theoretical and experimental analysis of classical HT is performed for opportunities to exhibit higher degree of parallelism where parallelism is quantified as a number of parallel operations per level in the Directed Acyclic Graph (DAG) of the transform. Based on theoretical analysis of classical HT, an opportunity to re-arrange computations in the classical HT is identified that results in Modified HT (MHT) where it is shown that MHT exhibits 1.33x times higher parallelism than classical HT. Experiments in off-the-shelf multicore and General Purpose Graphics Processing Units (GPGPUs) for HT and MHT suggest that MHT is capable of achieving slightly better or equal performance compared to classical HT based QR factorization realizations in the optimized software packages for Dense Linear Algebra (DLA). We implement MHT on a customized platform for Dense Linear Algebra (DLA) and show that MHT achieves 1.3x better performance than native implementation of classical HT on the same accelerator. For custom realization of HT and MHT based QR factorization, we also identify macro operations in the DAGs of HT and MHT that are realized on a Reconfigurable Data-path (RDP). We also observe that due to re-arrangement in the computations in MHT, custom realization of MHT is capable of achieving 12 percent better performance improvement over multicore and GPGPUs than the performance improvement reported by General Matrix Multiplication (GEMM) over highly tuned DLA software packages for multicore and GPGPUs which is counter-intuitive. Farhad Merchant, Tarun Vatwani, Anupam Chattopadhyay, Soumyendu Raha, S. K. Nandy 0001, Ranjani Narayan |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2016 | Enabling in-memory computation of binary BLAS using ReRAM crossbar arraysabstractMemristive devices, such as ReRAMs, are fast gaining prominence for their low leakage power, high endurance and non-volatile storage capabilities. ReRAM crossbar arrays also found usage as platform for in-memory computing, particularly for data-intensive computations, due to its inherent capability to perform stateful logic operations. Binary matrix and vector operations arise in several applications that require close interaction with the storage, such as Error Correction Codes (ECC), approximate graph mining, and in general, diverse big data applications. In this paper, we explore for the first time, an efficient mapping of Binary Basic Linear Algebra Subprograms (BiBLAS) onto hybrid CMOS-ReRAM crossbar array. We investigate the impact of crossbar configurations on the delay, and area of BiBLAS operations for various vector sizes. Debjyoti Bhattacharjee, Farhad Merchant, Anupam Chattopadhyay |
VLSI-SoC | 2 |
| 2014 | Efficient and scalable CGRA-based implementation of Column-wise Givens RotationabstractGivens Rotation is a key computation-intensive block in embedded wireless applications. In order to achieve an efficient mapping which smoothly scales to the underlying architecture, we propose two new Column-based Givens Rotation algorithms, derived from traditional Fast Givens and Square-root and Division Free Givens algorithms. These algorithms allow annihilation of multiple elements in a column of the input matrix simultaneously, without a dependency bottle-neck allowing increased parallelism, resource sharing and scalability. The ease of mapping and scalability has been tested on a layered coarse-grained reconfigurable architecture reaching close to optimal results for highly parallel architectures. Zoltán Endre Rákossy, Farhad Merchant, Axel Acosta-Aponte, S. K. Nandy 0001, Anupam Chattopadhyay |
ASAP | 2 |
| 2014 | Scalable and energy-efficient reconfigurable accelerator for column-wise givens rotationabstractA new layered reconfigurable architecture is proposed which exploits modularity, scalability and flexibility to achieve high energy efficiency and memory bandwidth. Using two flavors of Column-wise Givens rotation, derived from traditional Fast Givens and Square root and Division Free Givens Rotation algorithms the architecture is thoroughly evaluated for scalability, speed, area and energy. Combining an efficient mapping strategy of the highly parallel algorithms capable of annihilation of multiple elements of a column of the input matrix and using the new features of the architecture, 9 architectural variants were explored achieving a clean trade-off of execution speed versus area, while keeping relatively constant energy. Zoltán Endre Rákossy, Farhad Merchant, Axel Acosta-Aponte, S. K. Nandy 0001, Anupam Chattopadhyay |
VLSI-SoC | 2 |
| 2014 | A framework for post-silicon realization of arbitrary instruction extensions on reconfigurable data-paths
Saptarsi Das, Kavitha T. Madhu, Madhav Krishna, Nalesh S 0001, Farhad Merchant, Santhi Natarajan, Ipsita Biswas, Adithya Pulli, S. K. Nandy 0001, Ranjani Narayan |
J. Syst. Archit. | 5 |