Kamalika Datta

dblp:24/11161 · DBLP profile ↗
← Back
45ranked-venue papers
8as first author
26since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 33 · 5 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 4 since 2021Theory of computation · 9 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 8 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2026 Late Breaking Results: Efficient Formal Verification of Highly Optimized MAC Units
abstract
The demand for compute-intensive applications such as AI/ML has led to the development of processors with complex functionalities. The Multiply Accumulate (MAC) unit is a vital component in these processors, but its verification is very challenging due to the highly optimized designs used to implement the MAC operation. In this paper, we show some interesting results for optimized MAC design verification using a formal proof engine, Symbolic Computer Algebra (SCA). For the first time, we exploit the combined benefit of phase and dynamic ordering in verifying MAC circuits, a capability not possible using state-of-the-art SCA proof engines.
Jan Kleinekathöfer, Lennart Weingarten, Kamalika Datta, Rolf Drechsler
DATE3
2026 Fan-In Aware Graph-Based Optimization for MAC-Based in-Memory Computing
abstract
Resistive RAM (RRAM) has emerged as a promising technology for in-memory computing, allowing both storage and computation within the same physical substrate. Although its ability to perform analog computations, especially multiplyaccumulate (MAC) operations, has been effectively utilized in neuromorphic systems, there has been limited research on its applicability to Boolean logic synthesis. Existing approaches typically rely on graph-based representations of Boolean functions that are mapped to column-wise MAC operations on standard RRAM crossbars. However, these representations largely inherit binary fan-in constraints from conventional logic synthesis flows, resulting in limited exploitation of MAC-level parallelism and underutilization of available crossbar resources. In this work, we address this limitation by introducing the concept of multi-input OR-Inverter Graphs (m-OIGs), which allow OR nodes with fanin greater than two to better match the accumulation semantics of MAC operations. Experimental results on standard benchmark suites demonstrate that increasing OR fan-in consistently reduces both crossbar area and total evaluation cycles, leading to improved performance and more efficient use of RRAM crossbar resources, highlighting the importance of fan-in-aware logic representations.
Fatemeh Shirinzadeh, Abhoy Kole, Kamalika Datta, Saeideh Shirinzadeh, Rolf Drechsler
DDECS3
2025 Late Breaking Results: Towards Efficient Formal Verification of Dot Product Architectures
abstract
The popularity of compute intensive applications, like AI/ML, has driven the design of processors with complex functionality. The Dot Product (DP) is one of the most essential operations in modern neural processors, although no complete formal verification technique exists that can ensure its 100% correctness. In this paper we show the first step towards formally verifying DP using Symbolic Computer Algebra (SCA). The verification process is performed without the need of a reference model generation which is a key factor in verification. Experimental results show the efficiency and scalability of SCA-based verification for DP architectures.
Lennart Weingarten, Kamalika Datta, Rolf Drechsler
DATE2
2025 Towards an Automated Debugging Approach for Fault Identification in Quantum Circuits
abstract
In this paper, we propose a novel method for locating and diagnosing bugs in quantum circuits. Debugging in the quantum domain is especially challenging due to the inherent inability of assessing the quantum state of a program. Moreover, explaining the root cause behind unexpected outcomes is hard due to the limited information gain provided by measurements. Our approach aims to address both of these issue: Firstly, the bug site is identified using a standard circuit slicing technique combined with an associated measurement strategy. Secondly, we provide information about the nature of the bug, generated through repeated measurements. To minimize the number of measurements, we introduce a notion of equivalence classes based on unitary operations. This allows us to partition the gate library into classes that produce indistinguishable results under certain measurements. Finally, We assess the effectiveness and measurement complexity of our method by applying it to relevant primitive gate components and well-known quantum algorithms. Our empirical results shows that in 95.79% of all cases, our approach reveals the correct location of the bug along with a valid set of fault candidates. Furthermore, we demonstrate that the required number of circuit executions scales logarithmically with the circuit depth or linearly with the number of qubits.
Anton Maidl, Abhoy Kole, Kamalika Datta, Jannis Stoppe, Rolf Drechsler
DDECS3
2025 A Comprehensive Synthesis and Verification Approach for RRAM-Based Neuromorphic Computing
abstract
Resistive RAM (RRAM) has emerged as a promising technology for in-memory computing by enabling storage and computation within the same physical substrate. While its analog computation capability, particularly the multiply-accumulate (MAC) operation, has been effectively used in neuromorphic systems, its potential for logic synthesis remains underexplored. Logic synthesis using MAC not only unlocks new efficiency gains but also aligns with hardware already present in neuromorphic accelerators. In this work, we present the first automated framework for evaluating arbitrary Boolean functions on standard RRAM crossbars using highly parallel MAC operations. The proposed method introduces a logic computation core for RRAM-based neuromorphic architectures without requiring additional hardware, leveraging existing peripheral circuitry. To ensure functional correctness, we further integrate a formal verification approach based on equivalence checking via SAT solvers. Experimental results on standard benchmarks demonstrate substantial reductions in computation cycles and improved efficiency compared to existing RRAM-based logic synthesis methods, highlighting the practical potential of MAC-based logic in emerging computing systems.
Fatemeh Shirinzadeh, Abhoy Kole, Kamalika Datta, Saeideh Shirinzadeh, Rolf Drechsler
DSD3
2025 ForMAt: Formal Verification of Scalable Multiply and Accumulate Units
abstract
With the increasing popularity of compute intensive applications like AI, processors with complex functionalities are designed. Multiply and Accumulate (MAC) is one of the essential operations in modern Neural Processor Units (NPUs), but no sound formal verification technique exists that can efficiently ensure correctness. In this paper we analyze almost 200 configurations of MAC instances for various bit-widths starting from 8 up to several hundred bits. On top of the classical area-delay trade-off, we study verifiability as an additional parameter. It is shown that surprisingly the fastest and smallest instances are not the ones that are the hardest to verify. Exploiting Symbolic Computer Algebra (SCA) we provide a technique that allows scalable verification for large bit-width and classifies the set of MAC units.
Lennart Weingarten, Kamalika Datta, Rolf Drechsler
FDL2
2025 qSAT: Design of an Efficient Quantum Satisfiability Solver for Hardware Equivalence Checking
abstract
The use of Boolean Satisfiability (SAT) solver for hardware verification incurs exponential runtime in several instances. In this work, we have proposed an efficient quantum SAT (qSAT) solver for equivalence checking of Boolean circuits employing Grover’s algorithm. The Exclusive-Sum-of-Product (ESOP)-based generation of the Conjunctive Normal Form (CNF) equivalent clauses demands less qubits and minimizes the gates and depth of quantum circuit interpretation. The consideration of reference circuits for verification affecting Grover’s iterations and quantum resources are also presented as a case study. Experimental results are presented assessing the benefits of the proposed verification approach using open source Qiskit platform and IBM quantum computer.
Abhoy Kole, Mohammed E. Djeridane, Lennart Weingarten, Kamalika Datta, Rolf Drechsler
ACM J. Emerg. Technol. Comput. Syst.4
2024 Towards Formal Verification for MAC-based In-Memory Computing
abstract
Resistive RAM (RRAM) is a non-volatile memory technology with an abrupt switching property that enables it to perform basic logic operations. RRAM also possesses analog computational features by means of the so-called Multiply and Accumulate (MAC) operation that can be performed in all memory columns simultaneously. The MAC operation is particularly interesting for neuromorphic computing as it enables highly parallelized calculation of complex matrix-vector multiplications on standard RRAM crossbars.So far, several forms of universal logic are executed within RRAM devices, which have been the basis for a variety of logic-in-memory synthesis approaches. Recent research has addressed the mapping of logical functions to RRAM crossbars using the MAC operation, which allows for the facilitation of RRAM-based neuromorphic architectures with a basic logical core. Recently, a few formal verification methods have been introduced, which are tailored for synthesis approaches using certain RRAM logic primitives, such as in-memory styles based on the three-input majority operation and NOR gates. This paper analyzes these methods and, for the first time, proposes a verification method customized for MAC-based in-memory computing. A case study has been conducted to compare the proposed method with the existing methods, which reveals the superior performance of our method.
Fatemeh Shirinzadeh, Kamalika Datta, Saeideh Shirinzadeh, Abhoy Kole, Rolf Drechsler
ATS2
2024 Improving Self-Fault-Tolerance Capability of Memristor Crossbar Using a Weight-Sharing Approach
abstract
The ability of resistive memory (ReRAM) to naturally conduct vector-matrix multiplication (VMM), the primary operation carried out in neural networks, has caught the interest of researchers. The memristor crossbar is a suitable architecture to perform VMM and additionally offers benefits like in-memory computation (IMC), low power, and high density. Memristor-based neural networks are typically trained using a mechanism where weight computations are carried out on a host machine and downloaded into the crossbar. However, due to faulty memristors in the crossbar, a cell may not be able to store the exact weight values, which may lead to inference errors. In this paper, we propose a weight-sharing method to improve the self-fault-tolerance capability of memristor crossbar. In order to reduce the impact of faulty memristors, the weights are shared among different layers of memristors in a 3D crossbar. Simulation analyses show considerable improvements in the fault-tolerance capability of the crossbar.
Dev Narayan Yadav, Phrangboklang Lyngton Thangkhiew, F. Lalchhandama, Kamalika Datta, Rolf Drechsler, Indranil Sengupta 0001
ATS4
2024 Dynamic Realization of Multiple Control Toffoli Gate
abstract
Dynamic Quantum Circuits (DQC) is an inevitable solution for today's Noisy Intermediate Scale Quantum (NISQ) systems. This enables realization of an n-qubit (where,$n > 2$) quantum circuit using only 2-qubits with the aid of additional non-unitary operations which is evident from the recent dynamic realizations of algorithms like Quantum Phase Estimation (QPE) and Bernstein- Vazirani (BV) as well as 3-qubit Toffoli operation. In this work, we introduce two different dynamic realization schemes for Multiple Control Toffoli (MCT) gates, for the first time to the best of our knowledge. We compare the respective realizations in terms of resources (e.g., gate, depth and nearest neighbor overhead) and computational accuracy. For this purpose, we apply the proposed dynamic MCT gates in Deutsch-Jozsa (DJ) algorithm, thereby realizing the traditional DJ algorithm as DQCs. Experimental evaluations show that one dynamic scheme for MCT gates leads to DQCs with better computational accuracy, while the other one results in DQCs with better computational resources.
Abhoy Kole, Arighna Deb, Kamalika Datta, Rolf Drechsler
DATE3
2024 Complete and Efficient Verification for a RISC-V Processor Using Formal Verification
abstract
Formal verification techniques are computationally complex and the exact time and space complexities are in general not known, which makes the performance of the process unpredictable. Some of the recent works have shown that it is possible to carry out formal verification with polynomial time and space complexities for specific designs like arithmetic circuits. However, the methodology used cannot be directly extended to complex designs like processors. A recent work has shown polynomial verification of a single-cycle RISC- V processor with limited functionality, which considers only the combinational parts of the AL U. In this paper we propose for the first time a complete verification approach that covers all the functional units of the processor, and at the same time considers its sequential behavior. Experimental results show that the verification can be carried out in polynomial time, and also demonstrate significant improvement over previous methods.
Lennart Weingarten, Kamalika Datta, Abhoy Kole, Rolf Drechsler
DATE2
2024 Exploring the Potential of Decision Diagrams for Efficient In-Memory Design Verification
abstract
In this paper we present the first Decision Diagrams (DDs) based methodology for verifying the Resistive Random Access Memory (ReRAM) synthesis process. In particular, we propose a methodology which leverages Binary Decision Diagrams (BDDs), Multiplicative Binary Moment Diagrams (*BMDs), and Kronecker Multiplicative BMDs (K*BMDs) for verification. We introduce a synthesis tool for ReRAM-compatible micro-operations and a DD generation process for equivalence checking. Experimental results on a large set of arithmetic adders demonstrate that our DD-based approach significantly outperforms SAT solvers in verification speed, offering a more efficient and scalable solution.
Khushboo Qayyum, Abhoy Kole, Kamalika Datta, Muhammad Hassan 0002, Rolf Drechsler
ACM Great Lakes Symposium on VLSI3
2024 Is Simulation the only Alternative for Effective Verification of Dynamic Quantum Circuits?
Liam Hurwitz, Kamalika Datta, Abhoy Kole, Rolf Drechsler
RC2
2024 A new design of parity-preserving reversible multipliers based on multiple-control toffoli synthesis targeting emerging quantum circuits
Mojtaba Noorallahzadeh, Mohammad Mosleh, Kamalika Datta
Frontiers Comput. Sci.3
2024 Exploiting the Extended Neighborhood of Hexagonal Qubit Architecture for Mapping Quantum Circuits
abstract
In this work mapping of quantum circuits to regular hexagonal grid with coupling degree of six has been investigated. Architectures involving superconducting qubits impose restrictions on 2-qubit gate operations to be carried out only between physically coupled qubits, also referred to as nearest-neighbor (NN) constraint. The noise introduced by the 2-qubit gates and the execution time greatly affect the computational reliability. Existing mapping techniques suffer either from the adopted approach to reduce gate overhead or from their inability to take advantage of such architectural regularity. We outlined three different qubit mapping approaches using Remote-CNOT templates, Swap gates and combination of both. We show the benefits of assigning the Cartesian coordinate system in hexagonal grid for runtime elevation and devised approaches for reduction in gate overheads. While the template-based approach gives a strict upper bound of additional gate overheads for a particular qubit mapping, the combined approach provides better result employing a larger lookahead window. Experiments on benchmark quantum circuits confirm that the proposed Swap-based method provides an average \(25\%\) improvement in gate overheads over a recent work and the combined approach contributes further \(15\%\) average improvement on the result at the expense of a little higher runtime.
Abhoy Kole, Kamalika Datta, Indranil Sengupta 0001, Rolf Drechsler
ACM J. Emerg. Technol. Comput. Syst.2
2024 ReSG: A Data Structure for Verification of Majority-based In-memory Computing on ReRAM Crossbars
abstract
Recent advancements in the fabrication of Resistive Random Access Memory (ReRAM) devices have led to the development of large-scale crossbar structures. In-memory computing architectures relying on ReRAM crossbars aim to mitigate the processor-memory bottleneck that exists with current complementary metal-oxide semiconductor technology. With this motivation, several synthesis and mapping approaches focusing on the realizations of Boolean functions in the ReRAM crossbars have been proposed earlier. Thus far, the verification of the designs realized on ReRAM crossbars is done either through manual inspection or using simulation-based approaches. Since manual inspections and simulation-based approaches are limited to smaller designs, they cannot be applied to the verification of complex designs on large-scale ReRAM crossbars. Motivated by this, we propose, for the first time, an automatic equivalence checking flow that determines the equivalence between the original function specification (e.g., Majority-inverter Graph ) and the crossbar micro-operations file formats. We consider two crossbar structures, zero-transistor, one-memristor (0T1R) and one-transistor, one-memristor (1T1R) to implement the micro-operations. While the micro-operations file format exists for 0T1R crossbar structures, no representations for micro-operations to be executed in 1T1R crossbars exist yet. In this work, we introduce the micro-operation file format for 1T1R crossbar structures to efficiently represent the micro-operations as ReRAM crossbar netlists. Afterwards, we introduce two intermediate data structures, ReRAM Sequence Graph for 0T1R crossbars (ReSG-0T1R) and for 1T1R crossbars (ReSG-1T1R) , that are derived from the 0T1R and 1T1R crossbar micro-operations file formats, respectively. These ReSGs are then translated into Boolean Satisfiability (SAT) formula, and then the verification is done by checking the generated SAT formulae against the golden functional specification (represented in Verilog) using Z3 Satisfiability solver. Experimental evaluations confirm the effectiveness of the proposed verification methodology on MCNC and ISCAS benchmarks.
Kousik Bhunia, Arighna Deb, Kamalika Datta, Muhammad Hassan 0002, Saeideh Shirinzadeh, Rolf Drechsler
ACM Trans. Embed. Comput. Syst.3
2023 Automated Equivalence Checking Method for Majority Based In-Memory Computing on ReRAM Crossbars
abstract
Recent progress in the fabrication of Resistive Random Access Memory (ReRAM) devices has paved the way for large scale crossbar structures. In particular, in-memory computing on ReRAM crossbars helps in bridging the processor-memory speed gap for current CMOS technology. To this end, synthesis and mapping of Boolean functions to such crossbars have been investigated by researchers. However the verification of simple designs on crossbar is still done through manual inspection or sometimes complemented by simulation based techniques. Clearly this is an important problem as real world designs are complex and have higher number of inputs. As a result manual inspection and simulation based methods for these designs are not practical.
Arighna Deb, Kamalika Datta, Muhammad Hassan 0002, Saeideh Shirinzadeh, Rolf Drechsler
ASP-DAC2
2023 Extending the Design Space of Dynamic Quantum Circuits for Toffoli based Network
abstract
Recent advances in fault tolerant quantum systems allow to perform non-unitary operations like mid-circuit measurement, active reset and classically controlled gate operations in addition to the existing unitary gate operations. Real quantum devices that support these non-unitary operations enable us to execute a new class of quantum circuits, known as Dynamic Quantum Circuits (DQC). This helps to enhance the scalability, thereby allowing execution of quantum circuits comprising of many qubits by using at least two qubits. Recently DQC realizations of multi-qubit Quantum Phase Estimation (QPE) and Bernstein-Vazirani (BV) algorithms have been demonstrated in two separate experiments. However the dynamic transformation of complex quantum circuits consisting of Toffoli gate operations have not been explored yet. This motivates us to: (a) explore the dynamic realization of Toffoli gates by extending the design space of DQC for Toffoli networks, and (b) propose a general dynamic transformation algorithm for the first time to the best of our knowledge. More precisely, we introduce two dynamic transformation schemes (dynamic-1 and dynamic-2) for Toffoli gates, that differ with respect to the required number of classically controlled gate operations. For evaluation, we consider the Deutsch-Jozsa (DJ) algorithm composed of one or more Toffoli gates. Experimental results demonstrate that dynamic DJ circuits based on dynamic-2 Toffoli realization scheme provides better computational accuracy over the dynamic-1 scheme. Further, the proposed dynamic transformation scheme is generic and can also be applied to non-Toffoli quantum circuits, e.g. BV algorithm.
Abhoy Kole, Arighna Deb, Kamalika Datta, Rolf Drechsler
DATE3
2023 Improved Cost-Metric for Nearest Neighbor Mapping of Quantum Circuits to 2-Dimensional Hexagonal Architecture
Kamalika Datta, Abhoy Kole, Indranil Sengupta 0001, Rolf Drechsler
RC1
2023 Exploiting the Benefits of Clean Ancilla Based Toffoli Gate Decomposition Across Architectures
Abhoy Kole, Kamalika Datta, Philipp Niemann 0001, Indranil Sengupta 0001, Rolf Drechsler
RC2
2022 Unlocking Sneak Path Analysis in Memristor Based Logic Design Styles
abstract
Memristors or Resistive Random Access Memory (RRAM) are emerging non-volatile memory devices that can be used for both storage and computing. In this type of memory the information is stored in memory cells in the form of resistance. One of the very important challenges in memristive crossbars is the existence of Sneak Paths, which result in erroneous reading of memory cells. Most of the logic in-memory techniques have emphasized on improving the logic design perspective, but have given minor importance to the sneak path issue. In this paper we show the effect of sneak paths on crossbars of various sizes, and then try to analyze the logic design approaches like MAGIC and MAJORITY with respect to their immunity to sneak paths. Experimental result shows that with some extra overhead we can eliminate the sneak path effect in various logic design methods.
Kamalika Datta, Saeideh Shirinzadeh, Phrangboklang Lyngton Thangkhiew, Indranil Sengupta 0001, Rolf Drechsler
DSD1
2022 SAT-based Exact Synthesis of Ternary Reversible Circuits using a Functionally Complete Gate Library
abstract
The problem of synthesis and optimization of reversible and quantum circuits have drawn the attention of researchers for the last two decades due to increasing interest in quantum computing. Although lot of works have been done on the synthesis of binary reversible circuits, very less works have been reported on the synthesis of ternary reversible circuits. Ternary circuits have lower cost of implementation as compared to their binary counterparts. However, the synthesis approaches that exist for ternary reversible circuits either use too many circuit lines (qutrits) or too many gates. Only one prior work has discussed the problem of generating cost-optimal ternary reversible circuits, but for a very restrictive gate library, which limits the approach to a specific subset of ternary reversible functions and often the solution becomes sub-optimal due to the imposed restrictions. The present paper overcomes that restriction, and uses multiple control ternary Toffoli gates with all possible ternary target operations as the gate library. This gate library is functionally complete and can be used to synthesize any arbitrary function. The proposed SAT-based synthesis approach provides low cost solutions in terms of the number of gates for any arbitrary ternary reversible function. Experimental results on various randomly generated permutations as well as standard ternary benchmarks establish this claim. The results can be used as template for other synthesis approaches by observing how far they deviate from the optimal solutions.
Abhoy Kole, Kamalika Datta, Indranil Sengupta 0001, Rolf Drechsler
DSD2
2022 Unlocking High Resolution Arithmetic Operations within Memristive Crossbars for Error Tolerant Applications
abstract
Memristor-based crossbar architectures have been explored by researchers for neuromorphic computing, where analog vector-matrix multiplication can be carried out in a single time step. In this paper we explore such architectures for carrying out various arithmetic operations. Since the computations are carried out in analog domain, they are affected by fabrication and performance variability of the manufactured devices. As a result, there can be inherent errors during the computation. However, the architecture can be suitable for approximate computing applications where some errors can be tolerated. We have proposed a method for carrying out arithmetic operations with any multiple of k-bit resolution on the crossbar, for some limited values of k. The fault tolerant capability of the proposed architecture is evaluated through experimentation on benchmark datasets. We also perform case studies to analyze the performance of the approach with particular emphasis on approximate computing. The results of the case studies show that certain applications indeed exhibit fault tolerance in presence of faulty memristors.
Kamalika Datta, Saman Fröhlich, Saeideh Shirinzadeh, Dev Narayan Yadav, Indranil Sengupta 0001, Rolf Drechsler
VLSI-SoC1
2022 FAMCroNA: Fault Analysis in Memristive Crossbars for Neuromorphic Applications
Dev Narayan Yadav, Phrangboklang Lyngton Thangkhiew, Kamalika Datta, Sandip Chakraborty 0001, Rolf Drechsler, Indranil Sengupta 0001
J. Electron. Test.3
2022 CoMIC: Complementary Memristor based in-memory computing in 3D architecture
F. Lalchhandama, Kamalika Datta, Sandip Chakraborty 0001, Rolf Drechsler, Indranil Sengupta 0001
J. Syst. Archit.2
2022 Feed-Forward learning algorithm for resistive memories
Dev Narayan Yadav, Phrangboklang Lyngton Thangkhiew, Kamalika Datta, Sandip Chakraborty 0001, Rolf Drechsler, Indranil Sengupta 0001
J. Syst. Archit.3
2020 Fledge: Flexible Edge Platforms Enabled by In-memory Computing
abstract
The proliferation of advanced analytics and artificial intelligence has been driven by huge volumes of data that are mostly generated at the edge. Simultaneously, there is a rising demand to perform analytics on edge platforms (i.e., near-sensor data analytics). However, conventional architectures of such platforms may not execute the targeted applications in an energy-efficient manner. Emerging near and in-memory computing paradigms can increase the energy efficiency of edge platforms by relying on emerging logic and memory devices. More importantly, these paradigms enable the possibility of performing computations on unconventional platforms, namely flexible computing systems. In this paper, we explore the benefits of in-memory computing at the edge on a flexible substrate enabled by thin-film transistors (TFTs) and resistive RAM (RRAM). As a case study, we consider bio-signal processing application workloads, i.e., compressive sensing and anomaly detection. We model the device, circuit, and architecture of our targeted platform and evaluate the corresponding system-level performance. Preliminary results indicate that in-memory computing enabled by flexible electronic devices enables a new class of edge platforms with lower power consumption, compared to that of rigid TFT devices.
Kamalika Datta, Arko Dutt, Ahmed Zaky, Umesh Chand, Yida Li 0004, Jackson Chun-Yang Huang, Aaron Thean, Mohamed M. Sabry
DATE1
2020 Quantifying the Benefits of Monolithic 3D Computing Systems Enabled by TFT and RRAM
abstract
Current data-centric workloads, such as deep learning, expose the memory-access inefficiencies of current computing systems. Monolithic 3D integration can overcome this limitation by leveraging fine-grained and dense vertical connectivity to enable massively-concurrent accesses between compute and memory units. Thin-Film Transistors (TFTs) and Resistive RAM (RRAM) naturally enable monolithic 3D integration as they are fabricated in low temperature (a crucial requirement). In this paper, we explore ZnO-based TFTs and HfO2-based RRAM to build a 1TFT-1R memory subsystem in the upper tiers. The TFT-based memory subsystem is stacked on top of a Si-FET bottom tier that can include compute units and SRAM. System-level simulations for various deep learning workloads show that our TFT-based monolithic 3D system achieves up to 11.4× system-level energy-delay product benefits compared to 2D baseline with off-chip DRAM—5.8× benefits over interposer-based 2.5D integration and 1.25× over 3D stacking of RRAM on silicon using through-silicon vias. These gains are achieved despite the low density of TFT-based RRAM and the higher energy consumption versus 3D stacking with RRAM, due to inherent TFT limitations.
Abdallah M. Felfel, Kamalika Datta, Arko Dutt, Hasita Veluri, Ahmed Zaky, Aaron Thean, Mohamed M. Sabry
DATE2
2020 An efficient memristor crossbar architecture for mapping Boolean functions using Binary Decision Diagrams (BDD)
Phrangboklang Lyngton Thangkhiew, Alwin Zulehner, Robert Wille, Kamalika Datta, Indranil Sengupta 0001
Integr.4
2020 Improved Mapping of Quantum Circuits to IBM QX Architectures
abstract
Quantum computers are becoming a reality today due to the rapid progress made by researchers in the last years. In the process of building quantum computers, IBM has developed several versions-starting from 5-qubit architectures like IBM QX2 and IBM QX4 to larger 16- or 20-qubit architectures. These architectures support arbitrary rotations of a single qubit and a controlled negation (CNOT) involving two qubits. The two qubit operations come with added coupling-map restrictions that only allow specific physical qubits to be the control and target qubits of the operation. In order to execute a quantum circuit on the IBM QX architecture, CNOT gates must satisfy the so-called coupling constraints of the architecture. Previous works addressed this issue with the objective of reducing the number of gates and the circuit depth. However, in this article, we show that further improvements are possible. To this end, we present a general approach for further improving the number of gate operations and depth of the mapped circuit. The proposed approach encompasses the selection of physical qubits, determining initial and local permutations efficiently to obtain the final circuit mapped to the given IBM QX architecture. Through experiments, improvements are observed over existing methods in terms of the number of gates and circuit depth.
Abhoy Kole, Stefan Hillmich, Kamalika Datta, Robert Wille, Indranil Sengupta 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2019 A staircase structure for scalable and efficient synthesis of memristor-aided logic
abstract
The identification of the memristor as fourth fundamental circuit element and, eventually, its fabrication in the HP labs provide new capabilities for in-memory computing. While there already exist sophisticated methods for realizing logic gates with memristors, mapping them to crossbar structures (which can easily be fabricated) still constitutes a challenging task. This is particularly the case since several (complementary) design objectives have to be satisfied, e.g. the design method has to be scalable, should yield designs requiring a low number of timesteps and utilized memristors, and a layout should result that is hardly skewed. However, all solutions proposed thus far only focus on one of these objectives and hardly address the other ones. Consequently, rather imperfect solutions are generated by state-of-the-art design methods for memristor-aided logic thus far. In this work, we propose a corresponding automatic design solution which addresses all these design objectives at once. To this end, a staircase structure is utilized which employs an almost square-like layout and remains perfectly scalable while, at the same time, keeps the number of timesteps and utilized memristors close to the minimum. Experimental evaluations confirm that the proposed approach indeed allows to satisfy all design objectives at once.
Alwin Zulehner, Kamalika Datta, Indranil Sengupta 0001, Robert Wille
ASP-DAC2
2019 Look-ahead mapping of Boolean functions in memristive crossbar array
Dev Narayan Yadav, Phrangboklang Lyngton Thangkhiew, Kamalika Datta
Integr.3
2018 Scalable in-memory mapping of Boolean functions in memristive crossbar array using simulated annealing
Phrangboklang Lyngton Thangkhiew, Kamalika Datta
J. Syst. Archit.2
2018 A New Heuristic for N-Dimensional Nearest Neighbor Realization of a Quantum Circuit
abstract
One of the main challenges in quantum computing is to ensure error-free operation of the basic quantum gates. There are various implementation technologies of quantum gates for which the distance between interacting qubits must be kept within a limit for reliable operation. This leads to the so-called requirement of neighborhood arrangements of the interacting qubits, often referred to as nearest neighbor (NN) constraint. This is typically achieved by inserting SWAP gates in the quantum circuits, where a SWAP gate between two qubits exchanges their states. Minimizing the number of SWAP gates to provide NN compliance is an important problem to solve. A number of approaches have been proposed in this regard, based on local and global ordering techniques. In this paper, a generalized approach for combined local and global ordering of qubits have been proposed that is based on an improved heuristic for cost estimation and is also scalable. The approach can be extended to N -dimensional arrangement of qubits, for any arbitrary values of N . Practical constraints, however, restrict the maximum value of N to 3. Extensive experiments on benchmark functions have been carried out to evaluate the performance in terms of SWAP gate requirements. 3-D organization of qubits shows average reductions of 6.7% and 37.4%, respectively, in the number of SWAP gates over 2-D and 1-D organizations. Also compared to the best 2-D and 1-D results reported in the literature, on the average 8.7% and 8.4% reductions, respectively, are observed.
Abhoy Kole, Kamalika Datta, Indranil Sengupta 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2018 A Scalable In-Memory Logic Synthesis Approach Using Memristor Crossbar
abstract
Because of their resistive switching properties and ease of controlling the resistive states, memristors have been proposed in nonvolatile storage as well as logic design applications. Memristors can be fabricated in a crossbar and suitable voltages applied to the row and column nanowires to control their states. This makes it possible to move toward new non-von Neumann-type architectures, usually referred to as in-memory computing, where logic operations can be performed directly on the storage fabric. In this paper, a scalable design flow for in-memory computing has been proposed, where a given multioutput logic function is synthesized as a netlist of NOT/NOR gates and then mapped to the crossbar using the Memristor-Aided loGIC (MAGIC) design style. The memristors corresponding to the primary inputs are initialized a priori. Subsequently, the required gate operations are performed by applying suitable row and column voltages in sequence. Two alternate mapping schemes have been analyzed. The switching characteristics of MAGIC NOR gates have been evaluated using circuit simulation under the Cadence Virtuoso environment. Experimental evaluation on ISCAS'85 benchmarks reports the average improvements of 27.7%, 34.6%, and 26.2%, respectively over a recently published work with respect to the number of memristors, number of cycles, and total energy dissipation, respectively. It may be noted that the energy consumption of the gates used in the proposed approach (NOT and NOR) is significantly higher than that using CMOS technology.
Rahul Gharpinde, Phrangboklang Lyngton Thangkhiew, Kamalika Datta, Indranil Sengupta 0001
IEEE Trans. Very Large Scale Integr. Syst.3
2017 Test Pattern Generation Effort Evaluation of Reversible Circuits
Abhoy Kole, Robert Wille, Kamalika Datta, Indranil Sengupta 0001
RC3
2017 Design of Efficient Quantum Circuits Using Nearest Neighbor Constraint in 2D Architecture
Leniency Marbaniang, Abhoy Kole, Kamalika Datta, Indranil Sengupta 0001
RC3
2017 Improved Decomposition of Multiple-Control Ternary Toffoli Gates Using Muthukrishnan-Stroud Quantum Gates
P. Mercy Nesa Rani, Abhoy Kole, Kamalika Datta, Indranil Sengupta 0001
RC3
2015 Towards a Cost Metric for Nearest Neighbor Constraints in Reversible Circuits
Abhoy Kole, Kamalika Datta, Indranil Sengupta 0001, Robert Wille
RC2
2015 A Post-Synthesis Optimization Technique for Reversible Circuits Exploiting Negative Control Lines
abstract
Recent works in the synthesis of reversible logic circuits have been motivated by ever increasing emphasis on low-power design alternatives, and recent developments in quantum computing. Although most of the synthesis approaches use multiple-control Toffoli (MCT) gates with positive control lines, a few recent works have also considered MCT gates with negative control lines resulting in better circuit realizations. Some of the works have also tried to carry out post-synthesis optimization of given MCT gate netlists with positive control lines, using template matching and similar netlist transformation techniques. However, only one work is reported that attempts to optimize netlists containing negative control MCT gates. This paper proposes an efficient optimization technique for MCT gate netlists with both positive and negative control lines, which is based on repeated applications of a small set of pairwise gate merging and replacement rules. Experiments carried out on reversible circuit benchmarks show that it is possible to achieve significant reductions in number of gates and quantum costs.
Kamalika Datta, Indranil Sengupta 0001, Hafizur Rahaman 0001
IEEE Trans. Computers1
2014 Optimizing DD-based synthesis of reversible circuits using negative control lines
abstract
Synthesis of reversible circuits has attracted the attention of many researchers. In particular, approaches based on Decision Diagrams (DDs) have been shown beneficial since they enable the realization of corresponding circuits for large functions. However, all existing approaches rely on a gate library composed of positive control lines only. Recently, it has been shown that the additional use of negative control lines enables significant reductions of the respective circuit costs. In this paper, we aim for exploiting this potential. To this end, two complementary schemes are investigated. First, a post-synthesis optimization that exploits the power of negative control lines is utilized to optimize the circuits generated by previously proposed DD-based methods. Second, negative control lines are explicitly considered during synthesis. Experimental results demonstrate that the proposed approaches result in a significant reduction with respect to gate count as well as quantum costs.
Eleonora Schönborn, Kamalika Datta, Robert Wille, Indranil Sengupta 0001, Hafizur Rahaman 0001, Rolf Drechsler
DDECS2
2014 RevVis: Visualization of Structures and Properties in Reversible Circuits
Robert Wille, Jannis Stoppe, Eleonora Schönborn, Kamalika Datta, Rolf Drechsler
RC4
2014 An Improved Reversible Circuit Synthesis Approach using Clustering of ESOP Cubes
abstract
The problem of reversible logic synthesis has drawn the attention of many researchers over the last two decades with growing emphasis on low-power design. Among the various synthesis approaches that have been reported, the ones based on compact circuit representations like Binary Decision Diagrams (BDD) and Exclusive-or Sum-Of-Products (ESOP) are interesting in the sense that they can handle large circuits with more than 100 inputs. The drawback of these approaches, however, is that the generated netlists are sub-optimal, and there is lot of scope for optimizing them. One of the best methods in this regard is an approach, where the ESOP cubes are grouped into sublists based on sharing among more than one outputs. In the work reported in this article, in contrast, an approach based on clustering the ESOP cubes based on their similarity with respect to input variables is presented, along with a technique to map each of the clusters into reversible gate netlists. This approach results in a significant reduction in quantum cost of the final netlist, but requires one additional garbage line. Experimental results on a number of reversible circuit benchmarks have been presented in support of the claim and also demonstrate that the method is very fast.
Kamalika Datta, Gaurav Rathi, Indranil Sengupta 0001, Hafizur Rahaman 0001
ACM J. Emerg. Technol. Comput. Syst.1
2013 Exploiting Negative Control Lines in the Optimization of Reversible Circuits
Kamalika Datta, Gaurav Rathi, Robert Wille, Indranil Sengupta 0001, Hafizur Rahaman 0001, Rolf Drechsler
RC1
2013 Partial encryption and watermarking scheme for audio files with controlled degradation of quality
Kamalika Datta, Indranil Sengupta 0001
Multim. Tools Appl.1