EDBT 2026 Demo / reviewers in the wild / expert
Debjyoti Bhattacharjee
dblp:160/7918
· DBLP profile ↗
27ranked-venue papers
13as first author
12since 2021 · last 2026
0000-0001-6561-8934ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 23 · 11 first-author · 10 since 2021Software engineering, systems software and programming languages · 8 · 3 first-author · 4 since 2021Security and privacy · 2 · 2 first-authorTheory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evaluating Cross-Architecture Performance Modeling of Distributed ML Workloads Using StableHLOabstractPredicting the performance of large-scale distributed machine learning (ML) workloads across multiple accelerator architectures remains a central challenge in ML system design. Existing GPU and TPU focused simulators are typically architecture-specific, while distributed training simulators rely on workload-specific analytical models or costly post-execution traces, limiting portability and cross-platform comparison. This work evaluates whether MLIR’s StableHLO dialect can serve as a unified workload representation for cross-architecture and crossfidelity performance modeling of distributed ML workloads. The study establishes a StableHLO-based simulation methodology that maps a single workload representation onto multiple performance models, spanning analytical, profiling-based, and simulator-driven predictors. Using this methodology, workloads are evaluated across GPUs and TPUs without requiring access to scaled-out physical systems, enabling systematic comparison across modeling fidelities. An empirical evaluation covering distributed GEMM kernels, ResNet, and large language model training workloads demonstrates that StableHLO preserves relative performance trends across architectures and fidelities, while exposing accuracy trade-offs and simulator limitations. Across evaluated scenarios, prediction errors remain within practical bounds for early-stage design exploration, and the methodology reveals fidelity-dependent limitations in existing GPU simulators. These results indicate that StableHLO provides a viable foundation for unified, distributed ML performance modeling across accelerator architectures and simulators, supporting reusable evaluation workflows and crossvalidation throughout the ML system design process. Jonas Svedas, Nathan Laubeuf, Ryan Harvey, Changhai Man, Abubakr Nada, Tushar Krishna, James Myers, Debjyoti Bhattacharjee |
ISPASS | 9 |
| 2025 | A System Level Performance Evaluation for Superconducting Digital SystemsabstractSuperconducting Digital (SCD) technology offers significant potential for enhancing the performance of next generation large scale compute workloads. By leveraging advanced lithography and a 300 mm platform, SCD devices can reduce energy consumption and boost computational power. This paper presents a cross-layer modeling approach to evaluate the system-level performance benefits of SCD architectures for Large Language Model (LLM) training and inference. Our findings, based on experimental data and Pulse Conserving Logic (PCL) design principles, demonstrate substantial performance gain in both training and inference. We are, thus, able to convincingly show that the SCD technology can address memory and interconnect limitations of present day solutions for next-generation compute systems. Joyjit Kundu, Debjyoti Bhattacharjee, Nathan Josephsen, Ankit Pokhrel, Udara De Silva, Wenzhe Guo, Steven Van Winckel, Steven Brebels, Quentin Herr, Anna Herr, Manu Perumkunnil Komalan |
DATE | 2 |
| 2024 | Analyzing GPU Energy Consumption in Data Movement and StorageabstractGPUs are the prevailing solution to execute high-performance tasks (e.g., machine learning training). As the peak performance of modern GPUs increases with each generation, so does their thermal design power (TDP). Hence, identifying energy bottlenecks in the GPU architecture is crucial to designing more efficient architectures in the future. However, due to the complex proprietary nature of modern GPU architectures, providing a detailed breakdown of the GPU energy consumption is not trivial. The goal of this work is to estimate a lower bound for the energy consumed by data movement and storage in modern GPU architectures, leveraging internal power sensors. We establish a basic energy model for modern GPUs, focused on data movement to/from the hardware-managed caches and software-managed memories. We propose a methodology to calibrate the energy model using microbenchmarks, performance counters, and the internal power sensor. We experimentally calibrate the model on an A100 NVIDIA GPU. Then, we challenge the consistency of the results by cross-validating with modified microbenchmarks with additional instructions. Finally, we use the calibrated energy model to evaluate breakdowns for workloads of increasing complexity (e.g., a ResNet-50 training iteration with different software optimizations). Our results show that data movement dominates the dynamic energy consumption of the GPU (up to 84%), with DRAM accesses being the main contributor. Paul Delestrac, Jonathan Miquel, Debjyoti Bhattacharjee, Diksha Moolchandani, Francky Catthoor, Lionel Torres, David Novo |
ASAP | 3 |
| 2024 | RTL Agent: An Agent-Based Approach for Functionally Correct HDL Generation via LLMsabstractLLMs as code generators have undergone rapid progress over the past couple of years. However, the models on their own provide no guarantees for the functional correctness of the generated code. Functional tests can not only be used by designers to assess the functional correctness of code, but also to guide them towards the solution of the problem. The same can be applied to LLMs performing automatic code generation through the use of the Reflexion technique. Reflexion is an agent-based workflow where the model generating code iterates over a loop of code generation, getting feedback from the test bench, reflecting on the feedback, and making appropriate changes to the code. The technique is known to drastically improve the performance of LLMs on software code generation. In this work, we adopt the technique for hardware description language (HDL) code generation as RTL Agent - an implementation of the workflow for Verilog generation with Reflexion. We compare multiple LLMs with standard inference vs with Reflexion on the VerilogEval benchmark. Within 5 iterations of the feedback loop, we observe a relative improvement of 33.2% for the low-performance model Llama3 and an average of 17.7% for the high-performance models GPT-4o and GPT-4o-mini in their Pass@1 performance. We present a cost analysis of the technique to enable cost-performance trade-offs. Sriram Ranga, Rui Mao 0010, Debjyoti Bhattacharjee, Erik Cambria, Anupam Chattopadhyay |
ATS | 3 |
| 2024 | Multi-Level Analysis of GPU Utilization in ML Training WorkloadsabstractTraining time has become a critical bottleneck due to the recent proliferation of large-parameter ML models. GPUs continue to be the prevailing architecture for training ML models. However, the complex execution flow of ML frameworks makes it difficult to understand GPU computing resource utilization. Our main goal is to provide a better understanding of how efficiently ML training workloads use the computing resources of modern GPUs. To this end, we first describe an ideal reference execution of a GPU-accelerated ML training loop and identify relevant metrics that can be measured using existing profiling tools. Second, we produce a coherent integration of the traces obtained from each profiling tool. Third, we leverage the metrics within our integrated trace to analyze the impact of different software optimizations (e.g., mixed-precision, various ML frameworks, and execution modes) on the throughput and the associated utilization at multiple levels of hardware abstraction (i.e., whole GPU, SM subpartitions, issue slots, and tensor cores). In our results on two modern GPUs, we present seven takeaways and show that although close to 100% utilization is generally achieved at the GPU level, average utilization of the issue slots and tensor cores always remains below 50% and 5.2%, respectively. Paul Delestrac, Debjyoti Bhattacharjee, Simei Yang, Diksha Moolchandani, Francky Catthoor, Lionel Torres, David Novo |
DATE | 2 |
| 2024 | Adaptive Block-Scaled GeMMs on Vector Processors for DNN Training at the EdgeabstractReduced precision datatypes have become essential to the efficient training and deployment of Deep Neural Networks (DNNs). A recent development in the field has been the emergence of block-scaled datatypes: tensor representation formats derived from floating-point, that share a common exponent across multiple elements. While these formats are being broadly adopted and optimised for by DNN-specific inference accelerators, the potential benefits for training workloads on general-purpose (GP) vector processors has yet to be thoroughly explored. This work proposes a benchmarked implementation of block-scaled general matrix multiplications (GeMM) for DNN training at the edge using commercially available vector instruction sets (ARM SVE). Using this implementation, we highlight an accuracy-speed trade-off involving the shape of shared exponent blocks - vectors or squares. We exploit this result to optimize the training of fully connected networks by dynamically adapting the shared exponent block shapes during training. This strategy yields on average around$1.95 \times$faster training with$2\times$lower memory footprint compared to standard IEEE 32-bit floating point (FP32), while achieving similar accuracy. Nitish Satya Murthy, Nathan Laubeuf, Debjyoti Bhattacharjee, Francky Catthoor, Marian Verhelst |
VLSI-SoC | 3 |
| 2024 | Synthesis Techniques for Fault-tolerant Quantum Circuit Implementation using Clifford+ZN-groupabstractDecoherence jeopardizes the entanglement of fragile quantum states, and is among the foremost challenges towards engineering scalable quantum computers. Realizing quantum circuit implementation with small qubit count and shallow circuit depth is necessary due to the linear scaling of decoherence rate with qubit count and circuit depth. Conversely, reasonable correction of small unitary errors can be achieved by using surface codes along with a transversal gate set to protect quantum information from decoherence. In this paper, we analyze and report the upper bound of non-Clifford phase-depth for different mapping schemes and synthesis approaches. We introduce a synthesis methodology based on lookup-table (LUT) networks, wherein the Boolean logic translates into fault-tolerant quantum logic using Clifford+ Z N group with zero ancillary cost. We also present fault-tolerant synthesis techniques for k -LUT network using additional ancillary lines with exponential phase-count and unit phase-depth. Laxmidhar Biswal, Debjyoti Bhattacharjee, Amlan Chakrabarti, Anupam Chattopadhyay |
ACM Trans. Quantum Comput. | 2 |
| 2023 | Improved Linear Decomposition of Majority and Threshold Boolean FunctionsabstractTo support efficient design automation for emerging computing fabrics, novel data structures for logic synthesis and technology mapping are being intensively studied. It has been shown that for several promising computing technologies intermediate forms, such as Majority-inverter graph (MIG) and XOR-Majority graph (XMG) can be particularly beneficial. This has propelled the Boolean Majority operator at the forefront of research. Though these structures primarily utilize 3-input Majority nodes, the efficacy of$n$-input Majority operators has been demonstrated as well. A long-standing research problem, in that context and also for theoretical circuit complexity, is to determine efficient decomposition of an$n$-input Majority$({\mathrm{ Maj}}_{n})$function in terms of 3-input Majority$({\mathrm{ Maj}}_{3})$operator. In this manuscript, we make two significant advances in this topic. First, a practically realizable linear decomposition is provided, thus improving the previously reported quadratic bounds. Second, the theoretical upper bound of decomposing${\mathrm{ Maj}}_{n}$, in terms of${\mathrm{ Maj}}_{3}$, is reduced from$5.884n$to$3n$. The erstwhile theoretical upper bound of$5.884n$also lacked a practical construction for${\mathrm{ Maj}}_{n}$decomposition, presumably due to the presence of sequential elements in the algorithm. The proof of the linearity, detailed construction procedure along with experimental studies using state-of-the-art synthesis flows to validate the aforementioned claims are presented in this work. The results are applicable to threshold Boolean functions, too. Anupam Chattopadhyay, Debjyoti Bhattacharjee, Subhamoy Maitra |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | Tiny ci-SAR A/D Converter for Deep Neural Networks in Analog in-Memory ComputationabstractThis paper presents a tiny charge injection-Successive Approximation (ci-SAR) A/D converter (ADC) to be integrated at the periphery of analog Matrix Vector Multiplication (MVM) accelerators for Deep Neural Network (DNN) inference. Derived from the ci-SAR ADC, this converter exploits a single charge injecting cell to minimize area and energy consumption. The ADC exhibits a signal-to-noise and distortion ratio of 30.5 dB, at 5 bits of nominal resolution. The energy per conversion is 86 fJ, running at 34 MS/s, with a silicon area of $75 \mu m ^{2}$, in 22 nm technology node. From the results of our analytical framework, an SRAM-based Analog in-Memory Compute (AiMC) array, including the proposed ADC at 5 bits of resolution, can achieve an energy efficiency of 1650 TOPs/W. Michele Caselli, Debjyoti Bhattacharjee, Arindam Mallik, Peter Debacker, Diederik Verkest |
ISCAS | 2 |
| 2022 | Dynamic Quantization Range Control for Analog-in-Memory Neural Networks AccelerationabstractAnalog in Memory Computing (AiMC) based neural network acceleration is a promising solution to increase the energy efficiency of deep neural networks deployment. However, the quantization requirements of these analog systems are not compatible with state-of-the-art neural network quantization techniques. Indeed, while the quantization of the weights and activations is considered by modern deep neural network quantization techniques, AiMC accelerators also impose the quantization of each Matrix Vector Multiplication (MVM) result. In most demonstrated AiMC implementations, the quantization range of MVM results is considered a fixed parameter of the accelerator. This work demonstrates that dynamic control over this quantization range is possible but also desirable for analog neural networks acceleration. An AiMC compatible quantization flow coupled with a hardware aware quantization range driving technique is introduced to fully exploit these dynamic ranges. Using CIFAR-10 and ImageNet as benchmarks, the proposed solution results in networks that are both more accurate and more robust to the inherent vulnerability of analog circuits than fixed quantization range based approaches. Nathan Laubeuf, Jonas Doevenspeck, Ioannis A. Papistas, Michele Caselli, Stefan Cosemans, Peter Vrancx, Debjyoti Bhattacharjee, Arindam Mallik, Peter Debacker, Diederik Verkest, Francky Catthoor, Rudy Lauwereins |
ACM Trans. Design Autom. Electr. Syst. | 7 |
| 2021 | Perspectives on Emerging Computation-in-Memory ParadigmsabstractThe traditional Von-Neumann architecture is reaching its limits and finding it difficult to cope up with the ever-increasing demands of modern workloads like artificial intelligence. This demand has fueled the search of technologies that can mimic human brain to efficiently combine both memory and computation within a single device. In this work, we present the state-of-the-art research in the domain of computation-in-memory. In particular, we take a look at memristors and its widespread application in neuromorphic computation. We introduce ReRAMs in terms of their novel computing paradigms and present ReRAM-specific design flows. We address the various circuit opportunities and challenges related to reliability and fault tolerance associated with them. Another high-potential candidate to leverage memory and computation from a single device is Ferroelectric Field-effect Transistor (FeFET). Here we present a co-integration of such FeFETs with another emerging nanotechnology concept, called Reconfigurable Field Effect Transistor (RFET) and discuss the impact of the higher amount of states provided by this combination. Shubham Rai, Anteneh Gebregiorgis, Debjyoti Bhattacharjee, Krishnendu Chakrabarty, Said Hamdioui, Anupam Chattopadhyay, Jens Trommer, Akash Kumar 0001 |
DATE | 4 |
| 2021 | Design-Technology Space Exploration for Energy Efficient AiMC-Based Inference AccelerationabstractExtremely energy-efficient convolutional neural network inference (CNN) is recently enabled by analog in-memory compute (AiMC). The integration of AiMC in a primarily digital inference system brings new challenges ranging from device specifications to defining novel system architectures. A novel framework to evaluate the impact of AiMC array at the system level is presented. The framework is used to model a SRAM-based 1024 × 512 prototype AiMC array, capable of energy efficiency of upto 675 TMACs/W. The proposed framework allows modelling different compute cells, array dimensions, operating voltage, and activation buffer energy models which can be used to determine the overall energy efficiency for various CNN workloads. Debjyoti Bhattacharjee, Nathan Laubeuf, Stefan Cosemans, Ioannis A. Papistas, Arindam Mallik, Peter Debacker, Myung Hee Na, Diederik Verkest |
ISCAS | 1 |
| 2020 | CONTRA: Area-Constrained Technology Mapping Framework For Memristive Memory Processing UnitabstractData-intensive applications are poised to benefit directly from processing-in-memory platforms, such as memristive Memory Processing Units, which allow leveraging data locality and performing stateful logic operations. Developing design automation flows for such platforms is a challenging and highly relevant research problem. In this work, we investigate the problem of minimizing delay under arbitrary area constraint for MAGIC-based in-memory computing platforms. We propose an end-to-end area constrained technology mapping framework, CONTRA. CONTRA uses Look-Up Table (LUT) based mapping of the input function on the crossbar array to maximize parallel operations and uses a novel search technique to move data optimally inside the array. CONTRA supports benchmarks in a variety of formats, along with crossbar dimensions as input to generate MAGIC instructions. CONTRA scales for large benchmarks, as demonstrated by our experiments. CONTRA allows mapping benchmarks to smaller crossbar dimensions than achieved by any other technique before, while allowing a wide variety of area-delay trade-offs. CONTRA improves the composite metric of area-delay product by 2.1× to 13.1× compared to seven existing technology mapping approaches. Debjyoti Bhattacharjee, Anupam Chattopadhyay, Srijit Dutta, Ronny Ronen, Shahar Kvatinsky |
ICCAD | 1 |
| 2020 | Crossbar-Constrained Technology Mapping for ReRAM Based In-Memory ComputingabstractIn-memory computing has gained significant attention due to the potential for dramatic improvement in speed and energy. Redox-based resistive RAMs (ReRAMs), capable of non-volatile storage and logic operations simultaneously have been used for logic-in-memory computing approaches. To this effect, we propose ReRAM based VLIW Architecture for in-Memory comPuting (ReVAMP), supported by a detailed device-accurate simulation setup with peripheral circuitry. We present theoretical bounds on the minimum area required for in-memory computation of arbitrary Boolean functions specified using structural representation (And-Inverter Graph and Majority-Inverter Graph) and two-level representation (Exclusive-Sum-of-Product). To support the ReVAMP architecture, we present two technology mapping flows that fully exploit the bit-level parallelism offered by the execution of logic using ReRAM crossbar array. The area-constrained mapping (ArC) generates feasible mapping for a variety of crossbar dimensions while the delay-constrained mapping (DeC) focuses primarily on minimizing the latency of mapping. We evaluate the proposed mappings against two state-of-the-art technology in-memory computing architectures, PLiM and MAGIC along with their automation flows (SIMPLE and COMPACT). ArC and DeC outperform state-of-the-art PLiM architecture by 1.46x and 4.3x on average in latency. ArC offers significantly lower area (on average 25.27x and 6.57x), while improving the area-delay product by 1.37x and 1.12x against two mapping approaches for MAGIC respectively. In contrast, DeC achieves average area (1.45x and 3.06x) and area-delay product (1.12x and 6.36x) improvements over the mapping approaches for MAGIC architecture respectively. The proposed mapping techniques allow a variety of runtime efficiency trade-offs. Debjyoti Bhattacharjee, Yaswanth Tavva, Arvind Easwaran, Anupam Chattopadhyay |
IEEE Trans. Computers | 1 |
| 2020 | SIMPLER MAGIC: Synthesis and Mapping of In-Memory Logic Executed in a Single Row to Improve ThroughputabstractIn-memory processing can dramatically improve the latency and energy consumption of computing systems by minimizing the data transfer between the memory and the processor. Efficient execution of processing operations within the memory is therefore, a highly motivated objective in modern computer architecture. This article presents a novel automatic framework for efficient implementation of arbitrary combinational logic functions within a memristive memory. Using tools from logic design, graph theory and compiler register allocation technology, we developed synthesis and in-memory mapping of logic execution in a single row (SIMPLER), a tool that optimizes the execution of in-memory logic operations in terms of throughput and area. Given a logical function, SIMPLER automatically generates a sequence of atomic memristor-aided logic (MAGIC) NOR operations and efficiently locates them within a single sizelimited memory row, reusing cells to save area when needed. This approach fully exploits the parallelism offered by the MAGIC NOR gates. It allows multiple instances of the logic function to be performed concurrently, each compressed into a single row of the memory. This virtue makes SIMPLER an attractive candidate for designing in-memory single instruction, multiple data (SIMD) operations. Compared to the previous work (that optimizes latency rather than throughput for a single function), SIMPLER achieves an average throughput improvement of 435×. When the previous tools are parallelized similarly to SIMPLER, SIMPLER achieves higher throughput of at least 5×, with 23× improvement in area and 20× improvement in area efficiency. These improvements more than fully compensate for the increase (up to 17% on average) in latency. Rotem Ben Hur, Ronny Ronen, Ameer Haj-Ali, Debjyoti Bhattacharjee, Adi Eliahu, Natan Peled, Shahar Kvatinsky |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2019 | SAID: A Supergate-Aided Logic Synthesis Flow for Memristive CrossbarsabstractA Memristor is a two-terminal device that can serve as a non-volatile memory element with built-in logic capabilities. Arranged in a crossbar structure, memristive arrays allow to represent complex Boolean logic functions that adhere to the logic-in-memory paradigm, where data and logic gates are glued together on the same piece of hardware. Needless to say, novel and ad-hoc CAD solutions are required to achieve practical and feasible hardware implementations. Existing techniques aim at optimal mapping strategies that account for Boolean logic functions described by means of 2-input NOR and NOT gates, thus overlooking the optimization capabilities that a smart and dedicated technology-aware logic synthesis can provide. In this paper, we introduce a novel library-free supergate-aided (SAID) logic synthesis approach with a dedicated mapping strategy tailored on MAGIC crossbars. Supergates are obtained with a Look-Up Table (LUT)-based synthesis that splits a complex logic network into smaller Boolean functions. Those functions are then mapped on the crossbar array as to minimize latency. The proposed SAID flow allows to (i) maximize supergate-level parallelism, thus reducing the total number of computing cycles, and (ii) relax mapping constraints, allowing an easy and fast mapping of Boolean functions on memristive crossbars. Experimental results obtained on several benchmarks from ISCAS'85 and IWLS'93 suites demonstrate that our solution is capable to outperform other state-of-the-art techniques in terms of speedup (3.89× in the best case), at the expense of a very low area overhead. Valerio Tenace, Roberto Giorgio Rizzo, Debjyoti Bhattacharjee, Anupam Chattopadhyay, Andrea Calimera |
DATE | 3 |
| 2019 | MUQUT: Multi-Constraint Quantum Circuit Mapping on NISQ Computers: Invited PaperabstractRapid advancement in the domain of quantum technologies have opened up researchers to the real possibility of experimenting with quantum circuits, and simulating small-scale quantum programs. Nevertheless, the quality of currently available qubits and environmental noise pose a challenge in smooth execution of the quantum circuits. Therefore, efficient design automation flows for mapping a given algorithm to the Noisy Intermediate Scale Quantum (NISQ) computer becomes of utmost importance. State-of-the-art quantum design automation tools are primarily focused on reducing logical depth, gate count and qubit counts with recent emphasis on topology-aware (nearest-neighbour compliance) mapping. In this work, we extend the technology mapping flows to simultaneously consider the topology and gate fidelity constraints while keeping logical depth and gate count as optimization objectives. We provide a comprehensive problem formulation and multi-tier approach towards solving it. The proposed automation flow is compatible with commercial quantum computers, such as IBM QX and Rigetti. Our simulation results over 10 quantum circuit benchmarks, show that the fidelity of the circuit can be improved up to 3.37 × with an average improvement of 1.87 ×. Debjyoti Bhattacharjee, Abdullah Ash-Saki, Mahabubul Alam, Anupam Chattopadhyay, Swaroop Ghosh |
ICCAD | 1 |
| 2018 | Technology-aware logic synthesis for ReRAM based in-memory computingabstractResistive RAMs (ReRAMs) have gained prominence for design of logic-in-memory circuits and architectures due to fast read/write speeds, high endurance, density and logic operation capabilities. ReRAM crossbar arrays allow constrained bit-level parallel operations. In this paper, for the first time, we propose optimization techniques during logic synthesis, which are specifically targeted for leveraging the parallelism offered by ReRAM crossbar arrays. Our method uses Majority-Inverter Graph (MIG) for the internal representation of the Boolean functions. The novel optimization techniques, when applied to the MIG, exposes the bit-level parallelism, and is further coupled with an efficient technology mapping flow. The entire synthesis process is benchmarked exhaustively over large arithmetic functions using a representative ReRAM crossbar architecture, while varying the crossbar dimensions. For the hard benchmarks, we obtained 10% reduction in the number of nodes with 16% reduction in delay on average. Debjyoti Bhattacharjee, Luca G. Amarù, Anupam Chattopadhyay |
DATE | 1 |
| 2018 | ReRAM-based In-Memory Computation of Galois Field arithmeticabstractRobust data communication is a prime need in the age of Internet-of-things (IoT), where multiple connected devices actively exchange information. To permit robustness of this information exchange, error resilient secure communication is necessary. Security, error detection as well as correction are fundamentally based on Galois Field (GF) arithmetic. In this work, we present a novel method for performing GF arithmetic on a state-of-the art ReRAM-based in-memory computing platform. ReRAM devices offer low leakage power, high endurance and non-volatile storage capabilities, coupled with stateful logic operations. The proposed lightweight library presents the mapping of GF element generation, addition and multiplication. We have experimentally verified the results. For GF(24), 3.8 nJ, 0.1 nJ and 3.1 nJ energy are required for element generation, addition and multiplication operations respectively, which demonstrates the efficacy of the mapping. Swagata Mandal, Debjyoti Bhattacharjee, Yaswanth Tavva, Anupam Chattopadhyay |
VLSI-SoC | 2 |
| 2018 | Kogge-Stone Adder Realization using 1S1R Resistive Switching Crossbar ArraysabstractLow operating voltage, high storage density, non-volatile storage capabilities, and relative low access latencies have popularized memristive devices as storage devices. Memristors can be ideally used for in-memory computing in the form of hybrid CMOS nano-crossbar arrays. In-memory serial adders have been theoretically and experimentally proven for crossbar arrays. To harness the parallelism of memristive arrays, parallel-prefix adders can be effective. In this work, a novel mapping scheme for in-memory Kogge-Stone adder has been presented. The number of cycles increases logarithmically with the bit width N of the operands, i.e., O ( log 2 N ), and the device count is 5 N . We verify the correctness of the proposed scheme by means of TaO × device model-based memristive simulations. We compare the proposed scheme with other proposed schemes in terms of number of cycle and number of devices. Debjyoti Bhattacharjee, Anne Siemon, Eike Linn, Stephan Menzel, Anupam Chattopadhyay |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2017 | Area-constrained technology mapping for in-memory computing using ReRAM devicesabstractIn-memory computing platforms, such as Resistive RAM (ReRAM), offer natural advantage to data-intensive applications. The benefits of data locality and capability to perform native Boolean operations is exploited for significant performance advantage in multiple contexts ranging across neuromorphic computing, associative memory-based computing, arithmetic benchmarks and general-purpose programmable logic-in-memory computing. Despite these advances, design automation tools supporting in-memory computing are still in a nascent phase. In this work, we investigate for the first time, the problem of minimizing delay under arbitrary area constraint of ReRAM devices. We formulate the problem of area-constrained delay minimization as an Integer Linear Programming (ILP) formulation and further propose heuristics that offers scalability as well as solution close to optimal performance. Area-constrained mapping technology mappings enables unlocking significantly large design space trade-offs. Debjyoti Bhattacharjee, Arvind Easwaran, Anupam Chattopadhyay |
ASP-DAC | 1 |
| 2017 | ReVAMP: ReRAM based VLIW architecture for in-memory computingabstractWith diverse types of emerging devices offering simultaneous capability of storage and logic operations, researchers have proposed novel platforms that promise gains in energy-efficiency. Such platforms can be classified into two domains - application-specific and general-purpose. The application-specific in-memory computing platforms include machine learning accelerators, arithmetic units, and Content Addressable Memory (CAM)-based structures. On the other hand, the general-purpose computing platforms stem from the idea that several in-memory computing logic devices do support a universal set of Boolean logic operation and therefore, can be used for mapping arbitrary Boolean functions efficiently. In this direction, so far, researchers have concentrated on challenges in logic synthesis (e.g. depth optimization), and technology mapping (e.g. device count reduction). The important problem of efficient technology mapping of arbitrary logic network onto a crossbar array structure has been overlooked so far. In this paper, we propose, ReVAMP, a general-purpose computing platform based on Resistive RAM crossbar array, which exploits the parallelism in computing multiple logic operations in the same word. Further, we study the problem of instruction generation and scheduling for such a platform. We benchmark the performance of ReVAMP with respect to the state of the art architecture. Debjyoti Bhattacharjee, Rajeswari Devadoss, Anupam Chattopadhyay |
DATE | 1 |
| 2016 | EAST: Efficient Assertion Simulation techniques
Debjyoti Bhattacharjee, Soumi Chattopadhyay, Ansuman Banerjee |
DATE | 1 |
| 2016 | Delay-optimal technology mapping for in-memory computing using ReRAM devicesabstractRecent propositions of diverse In-Memory Computing platforms have shown a promising alternative to classical Von Neumann computing models. Significant benefits, in terms of energy-efficiency and performance, are reported for in-memory arithmetic circuits, neural networks, CAM, cache hierarchy and even fully programmable processors. In contrast, design automation tools supporting the development of such designs are still in a nascent phase. By leveraging the native stateful logic operation capability of ReRAM devices, several logic synthesis flows have been reported. In this paper, we complement these flows with a detailed study on the technology mapping phase for ReRAM devices. We provide a delay-optimal solution for technology mapping without area constraint and propose further heuristics to achieve device count reduction and to support delay optimization under the constraint of parallel instruction dispatch. We report at least 3× less delay compared to the naïve technology mapping adopted in recent studies. The proposed heuristics achieve 56% on average reduction in device count. Finally, a range of performance trade-offs is identified by applying the constraint of parallel instruction dispatch without noticeable degradation of delay. Debjyoti Bhattacharjee, Anupam Chattopadhyay |
ICCAD | 1 |
| 2016 | Hardware Accelerator for Stream Cipher SpritzabstractRC4, the dominant stream cipher in e-commerce and communication protocols such as, WEP, TLS, is being
considered for replacement due to the series of vulnerabilities that have been pointed out in recent past. After
a thorough analysis of the possible weaknesses, Spritz, a new stream cipher is proposed to that effect by
the author of RC4. The design of Spritz is based on Cryptographic Sponge construction, which permits
Spritz to be used in different modes, and therefore, makes it an attractive design choice for security protocols.
Initial software performance analysis of Spritz shows that it fares poorly compared to the state-of-the-art hash
functions and stream ciphers. In this paper, we extend the analysis to the hardware performance. We propose
a fully customized accelerator design for Spritz and identify the highest achievable runtime performance for
ASIC and FPGA technology. Our results show that the Spritz accelerator is significantly faster in encryption
compared to the software implementation (32.38x speed-up for the SQUEEZE and 64.07x speed-up for the
ABSORB function), though fares weakly against hardware implementation of state-of-the-art hash functions
and stream ciphers in terms of area-efficiency. Debjyoti Bhattacharjee, Anupam Chattopadhyay |
SECRYPT | 1 |
| 2016 | Enabling in-memory computation of binary BLAS using ReRAM crossbar arraysabstractMemristive devices, such as ReRAMs, are fast gaining prominence for their low leakage power, high endurance and non-volatile storage capabilities. ReRAM crossbar arrays also found usage as platform for in-memory computing, particularly for data-intensive computations, due to its inherent capability to perform stateful logic operations. Binary matrix and vector operations arise in several applications that require close interaction with the storage, such as Error Correction Codes (ECC), approximate graph mining, and in general, diverse big data applications. In this paper, we explore for the first time, an efficient mapping of Binary Basic Linear Algebra Subprograms (BiBLAS) onto hybrid CMOS-ReRAM crossbar array. We investigate the impact of crossbar configurations on the delay, and area of BiBLAS operations for various vector sizes. Debjyoti Bhattacharjee, Farhad Merchant, Anupam Chattopadhyay |
VLSI-SoC | 1 |
| 2014 | Efficient Hardware Accelerator for AEGIS-128 Authenticated Encryption
Debjyoti Bhattacharjee, Anupam Chattopadhyay |
Inscrypt | 1 |