EDBT 2026 Demo / reviewers in the wild / expert
Shahar Kvatinsky
dblp:92/10504
· DBLP profile ↗
46ranked-venue papers
3as first author
15since 2021 · last 2025
0000-0001-7277-7271ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 41 · 3 first-author · 14 since 2021Software engineering, systems software and programming languages · 6 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Artificial intelligence and machine learning · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Dependable Neuromorphic Computing-in-Memory Architectures
Farhad Merchant, Ankit Bende, Markus Fritscher, Shahar Kvatinsky, Simranjeet Singh, Vikas Rana, Regina Dittmann, Keerthi Dorai Swamy Reddy, Christian Wenger, Fouwad Jamil Mir, Mottaqiallah Taouil, Manil Dev Gomony, Said Hamdioui, Henk Corporaal |
ETS | 4 |
| 2025 | A Highly Reliable Dual-Mode RRAM PUF With Key Concealment SchemeabstractPhysical unclonable function (PUF) has been widely used in the Internet of Things (IoT) as a promising hardware security primitive. In recent years, PUFs based on resistive random access memory (RRAM) have demonstrated excellent reliability and integration density. Most previous designs store PUF keys directly in RRAMs, increasing vulnerability to attacks. This article proposes a dual-mode RRAM PUF, named differential mode and flexible mode, utilizing the difference in switching capability between RRAMs during parallel SET operations as the entropy source. The proposed PUF can reliably reproduce keys between cycles, so a key concealment scheme is used to protect PUF keys from being continuously exposed, improving the security of the RRAM PUF. The proposed RRAM PUF exhibits high reliability over ±10% VDD and a wide temperature range from −25°C to 125°C through post-processing operations. The flexible mode can generate a significant number of keys for high-security applications. Since the PUF keys can be concealed, the proposed PUF is compatible with in-memory computing. It can be implemented using the same RRAM array as experimentally validated using a MAGIC operation, thus reducing the hardware overhead. Jiang Li 0012, Yijun Cui, Chongyan Gu, Chenghua Wang, Weiqiang Liu 0001, Shahar Kvatinsky |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2024 | A Concealable RRAM Physical Unclonable Function Compatible with In-Memory ComputingabstractResistive random access memory (RRAM) has been widely used in physical unclonable function (PUF) design due to its low power consumption, fast read/write speed, and significant intrinsic randomness. However, existing RRAM PUFs cannot overcome the cycle-to-cycle (C2C) variations of RRAM, leading to poor reproducibility of PUF keys across cycles. Most prior designs directly store PUF keys in RRAMs, increasing vulnerability to attacks. In this paper, we propose a concealable RRAM PUF based on an RRAM crossbar array, utilizing the differential resistive switching characteristics of two RRAMs to generate keys. By enabling the reproducibility of PUF keys across cycles, a concealment scheme is proposed to prevent the exposure of PUF keys, thus enhancing the security of the RRAM PUF. Through post-processing operations, the proposed PUF exhibits high reliability over ±10% VDD and a wide temperature range from 248K to 373K. Furthermore, this RRAM PUF is compatible with in-memory computing (IMC), and they can be implemented using the same RRAM crossbar array. Jiang Li 0012, Yijun Cui, Chenghua Wang, Weiqiang Liu 0001, Shahar Kvatinsky |
DATE | 5 |
| 2024 | Assessing the Performance of Stateful Logic in 1-Selector-1-RRAM Crossbar ArraysabstractResistive Random Access Memory (RRAM) crossbar arrays are an attractive memory structure for emerging nonvolatile memory due to their high density and excellent scalability. Their ability to perform logic operations using RRAM devices makes them a critical component in non-von Neumann processing-in-memory architectures. Passive RRAM crossbar arrays (1-RRAM or 1R), however, suffer from a major issue of sneak path currents, leading to a lower readout margin and increasing write failures. To address this challenge, active RRAM arrays have been proposed, which incorporate a selector device in each memory cell (termed 1-selector-1-RRAM or 1S1R). The selector eliminates currents from unselected cells and therefore effectively mitigates the sneak path phenomenon. Yet, there is a need for a comprehensive analysis of 1S1R arrays, particularly concerning in-memory computation. In this paper, we introduce a 1S1R model tailored to a VO2-based selector and TiN/TiOx/HfOx/Pt RRAM device. We also present simulations of 1S1R arrays, incorporating all parasitic parameters, across a range of array sizes from 4 × 4 to 512 × 512. We evaluate the performance of Memristor-Aided Logic (MAGIC) gates in terms of switching delay, power consumption, and readout margin, and provide a comparative evaluation with passive 1R arrays. Arjun Tyagi, Shahar Kvatinsky |
ISCAS | 2 |
| 2024 | PyPIM: Integrating Digital Processing-in-Memory from Microarchitectural Design to Python TensorsabstractDigital processing-in-memory (PIM) architectures mitigate the memory wall problem by facilitating parallel bitwise operations directly within the memory. Recent works have demonstrated their algorithmic potential for accelerating data-intensive applications; however, there remains a significant gap in the programming model and microarchitectural design. This is further exacerbated by aspects unique to memristive PIM such as partitions and operations across both directions of the memory array. To address this gap, this paper provides an end-to-end architectural integration of digital memristive PIM from a high-level Python library for tensor operations (similar to NumPy and PyTorch) to the low-level microarchitectural design. We begin by proposing an efficient microarchitecture and instruction set architecture (ISA) that bridge the gap between the low-level control periphery and an abstraction of PIM parallelism. We subsequently propose a PIM development library that converts high-level Python to ISA instructions and a PIM driver that translates ISA instructions into PIM micro-operations. We evaluate PyPIM via a cycle-accurate simulator on a wide variety of benchmarks that both demonstrate the versatility of the Python library and the performance compared to theoretical PIM bounds. Overall, PyPIM drastically simplifies the development of PIM applications and enables the conversion of existing tensor-oriented Python programs to PIM with ease. Orian Leitersdorf, Ronny Ronen, Shahar Kvatinsky |
MICRO | 3 |
| 2024 | TDPP: 2-D Permutation-Based Protection of Memristive Deep Neural NetworksabstractThe execution of deep neural network (DNN) algorithms suffers from significant bottlenecks due to the separation of the processing and memory units in traditional computer systems. Emerging memristive computing systems introduce an in situ approach that overcomes this bottleneck. The nonvolatility of memristive devices, however, may expose the DNN weights stored in memristive crossbars to potential theft attacks. Therefore, this article proposes a 2-D permutation-based protection (TDPP) method that thwarts such attacks. We first introduce the underlying concept that motivates the TDPP method: permuting both the rows and columns of the DNN weight matrices. This contrasts with previous methods, which focused solely on permuting a single dimension of the weight matrices, either the rows or columns. While it is possible for an adversary to access the matrix values, the original arrangement of rows and columns in the matrices remains concealed. As a result, the extracted DNN model from the accessed matrix values would fail to operate correctly. We consider two different memristive computing systems (designed for layer-by-layer and layer-parallel processing, respectively), and demonstrate the design of the TDPP method that could be embedded into the two systems. Finally, we present a security analysis. Our experiments demonstrate that TDPP can achieve comparable effectiveness to prior approaches, with a high level of security when appropriately parameterized. In addition, TDPP is more scalable than previous methods and results in reduced area and power overheads. The area and power are reduced by, respectively,$1218\times $and$2815\times $for the layer-by-layer system and by$178\times $and$203\times $for the layer-parallel system compared to prior works. Minhui Zou, Zhenhua Zhu 0002, Tzofnat Greenberg-Toledo, Orian Leitersdorf, Jiang Li 0012, Junlong Zhou, Yu Wang 0002, Nan Du 0004, Shahar Kvatinsky |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 9 |
| 2023 | On Consistency for Bulk-Bitwise Processing-in-MemoryabstractProcessing-in-memory (PIM) architectures allow software to explicitly initiate computation in the memory. This effectively makes PIM operations a new class of memory operations, alongside standard memory operations (e.g., load, store). For software correctness, it is crucial to have ordering rules for a PIM operation with other PIM operations and other memory operations, i.e., a consistency model that takes into account PIM operations is vital. To the best of our knowledge, little attention to PIM operation consistency has been given in existing works. In this paper, we focus on a specific PIM approach, named bulk-bitwise PIM. In bulk-bitwise PIM, large bitwise operations are performed directly and stored in the memory array. We show that previous solutions for the related topic of maintaining coherency of bulk-bitwise PIM have broken the host native consistency model and prevent any guaranteed correctness. As a solution, we propose and evaluate four consistency models for bulk-bitwise PIM, from strict to relaxed. Our designs also preserve coherency between PIM and the host processor. Evaluating the proposed designs’ performance with a gem5 simulation, using the YCSB short-range scan benchmark and TPC-H queries, shows that the run time overhead of guaranteeing correctness is at most 6%, and in many cases the run time is even improved. The hardware overhead of our design is less than 0.22%. Ben Perach, Ronny Ronen, Shahar Kvatinsky |
HPCA | 3 |
| 2023 | ClaPIM: Scalable Sequence Classification Using Processing-in-MemoryabstractDeoxyribonucleic acid (DNA) sequence classification is a fundamental task in computational biology with vast implications for applications such as disease prevention and drug design. Therefore, fast high-quality sequence classifiers are significantly important. This article introduces ClaPIM, a scalable DNA sequence classification architecture based on the emerging concept of hybrid in-crossbar and near-crossbar memristive processing-in-memory (PIM). We enable efficient and high-quality classification by uniting the filter and search stages within a single algorithm. Specifically, we propose a custom filtering technique that drastically narrows the search space and a search approach that facilitates approximate string matching through a distance function. ClaPIM is the first PIM architecture for scalable approximate string matching that benefits from the high density of memristive crossbar arrays and the massive computational parallelism of PIM. Compared with Kraken2, a state-of-the-art software classifier, ClaPIM provides significantly higher classification quality (up to$20 \times $improvement in F1 score) and also demonstrates a$1.8 \times $throughput improvement. Compared with edit distance tolerant approximate matching (EDAM), a recently proposed static random-access memory (SRAM)-based accelerator that is restricted to small datasets, we observe both a$30.4 \times $improvement in normalized throughput per area and a 7% increase in classification precision. Marcel Khalifa, Barak Hoffer, Orian Leitersdorf, Robert Hanhan, Ben Perach, Leonid Yavits, Shahar Kvatinsky |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2022 | MatPIM: Accelerating Matrix Operations with Memristive Stateful LogicabstractThe emerging memristive Memory Processing Unit (mMPU) overcomes the memory wall through memristive devices that unite storage and logic for real processing-in-memory (PIM) systems. At the core of the mMPU is stateful logic, which is accelerated with memristive partitions to enable logic with massive inherent parallelism within crossbar arrays. This paper vastly accelerates the fundamental operations of matrix-vector multiplication and convolution in the mMPU, with either full-precision or binary elements. These proposed algorithms establish an efficient foundation for large-scale mMPU applications such as neural-networks, image processing, and numerical methods. We overcome the inherent asymmetry limitation in the previous in-memory full-precision matrix-vector multiplication solutions by utilizing techniques from block matrix multiplication and reduction. We present the first fast in-memory binary matrix-vector multiplication algorithm by utilizing memristive partitions with a tree-based popcount reduction (39$\times$ faster than previous work). For convolution, we present a novel in-memory input-parallel concept which we utilize for a full-precision algorithm that overcomes the asymmetry limitation in convolution, while also improving latency (2$\times$ faster than previous work), and the first fast binary algorithm (12$\times$ faster than previous work). Orian Leitersdorf, Ronny Ronen, Shahar Kvatinsky |
ISCAS | 3 |
| 2022 | The Bitlet Model: A Parameterized Analytical Model to Compare PIM and CPU SystemsabstractCurrently, data-intensive applications are gaining popularity. Together with this trend, processing-in-memory (PIM)–based systems are being given more attention and have become more relevant. This article describes an analytical modeling tool called Bitlet that can be used in a parameterized fashion to estimate the performance and power/energy of a PIM-based system and, thereby, assess the affinity of workloads for PIM as opposed to traditional computing. The tool uncovers interesting trade-offs between, mainly, the PIM computation complexity (cycles required to perform a computation through PIM), the amount of memory used for PIM, the system memory bandwidth, and the data transfer size. Despite its simplicity, the model reveals new insights when applied to real-life examples. The model is demonstrated for several synthetic examples and then applied to explore the influence of different parameters on two systems — IMAGING and FloatPIM. Based on the demonstrations, insights about PIM and its combination with a CPU are provided. Ronny Ronen, Adi Eliahu, Orian Leitersdorf, Natan Peled, Kunal Korgaonkar, Anupam Chattopadhyay, Ben Perach, Shahar Kvatinsky |
ACM J. Emerg. Technol. Comput. Syst. | 8 |
| 2022 | C-AND: Mixed Writing Scheme for Disturb Reduction in 1T Ferroelectric FET MemoryabstractFerroelectric field effect transistor (FeFET) memory has shown the potential to meet the requirements of the growing need for fast, dense, low-power, and non-volatile memories. In this paper, we propose a memory architecture named crossed-AND (C-AND), in which each storage cell consists of a single ferroelectric transistor. The write operation is performed using different write schemes and different absolute voltages, to account for the asymmetric switching voltages of the FeFET. It enables writing an entire wordline in two consecutive cycles and prevents current and power through the channel of the transistor. During the read operation, the current and power are mostly sensed at a single selected device in each column. The read scheme additionally enables reading an entire word without read errors, even along long bitlines. Our Simulations demonstrate that, in comparison to the previously proposed AND architecture, the C-AND architecture diminishes read errors, reduces write disturbs, enables the usage of longer bitlines, and saves up to 2.92X in memory cell area. Mor M. Dahan, Evelyn T. Breyer, Stefan Slesazeck, Thomas Mikolajick, Shahar Kvatinsky |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2021 | Efficient Error-Correcting-Code Mechanism for High-Throughput Memristive Processing-in-MemoryabstractInefficient data transfer between computation and memory inspired emerging processing-in-memory (PIM) technologies. Many PIM solutions enable storage and processing using memristors in a crossbar-array structure, with techniques such as memristor-aided logic (MAGIC) used for computation. This approach provides highly-paralleled logic computation with minimal data movement. However, memristors are vulnerable to soft errors and standard error-correcting-code (ECC) techniques are difficult to implement without moving data outside the memory. We propose a novel technique for efficient ECC implementation along diagonals to support reliable computation inside the memory without explicitly reading the data. Our evaluation demonstrates an improvement of over eight orders of magnitude in reliability (mean time to failure) for an increase of about 26% in computation latency. Orian Leitersdorf, Ben Perach, Ronny Ronen, Shahar Kvatinsky |
DAC | 4 |
| 2021 | multiPULPly: A Multiplication Engine for Accelerating Neural Networks on Ultra-low-power ArchitecturesabstractComputationally intensive neural network applications often need to run on resource-limited low-power devices. Numerous hardware accelerators have been developed to speed up the performance of neural network applications and reduce power consumption; however, most focus on data centers and full-fledged systems. Acceleration in ultra-low-power systems has been only partially addressed. In this article, we present multiPULPly, an accelerator that integrates memristive technologies within standard low-power CMOS technology, to accelerate multiplication in neural network inference on ultra-low-power systems. This accelerator was designated for PULP, an open-source microcontroller system that uses low-power RISC-V processors. Memristors were integrated into the accelerator to enable power consumption only when the memory is active, to continue the task with no context-restoring overhead, and to enable highly parallel analog multiplication. To reduce the energy consumption, we propose novel dataflows that handle common multiplication scenarios and are tailored for our architecture. The accelerator was tested on FPGA and achieved a peak energy efficiency of 19.5 TOPS/W, outperforming state-of-the-art accelerators by 1.5× to 4.5×. Adi Eliahu, Ronny Ronen, Pierre-Emmanuel Gaillardon, Shahar Kvatinsky |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2021 | Improving Efficiency and Lifetime of Logic-in-Memory by Combining IMPLY and MAGIC Families
Minhui Zou, Junlong Zhou, Jin Sun 0001, Chengliang Wang 0002, Shahar Kvatinsky |
J. Syst. Archit. | 5 |
| 2021 | Radiofrequency Switches Based on Emerging Resistive Memory Technologies - A SurveyabstractHigh-performance radio frequency (RF) switches play a critical role in allowing radio transceivers to provide access to shared resources such as antennas. They are important, especially in reconfigurable radios, where the connections within and between the different blocks (e.g., filters, amplifiers, and mixers) can be changed to customize and compose RF functions for distinct frequency bands. To realize this potential, new concepts for RF switch devices that can be integrated with standard transistor process, while providing better transmission performance, lower power consumption, smaller area, and lower actuation voltage than current RF switch technologies [e.g., field-effect transistors (FETs) and microelectromechanical systems (MEMSs)] are needed. Recently, RF switches based on emerging memory technologies, such as conductive-bridge RAM (CBRAM), resistive RAM (ReRAM), and phase-change memory (PCM), have been proposed. These nanoscale switches present major advantages, for example, high cutoff frequency, nonvolatility, fast and low-energy switching, small footprint, and compatibility with back-end-of-line (BEOL) of the standard CMOS process, thus placing them as possible contenders for high-performance RF switches. This article surveys high-performance RF switches based on resistive memories, comparing them with mature RF switching technologies like RF MEMS, p-i-n diodes, and FETs. We discuss the physical mechanisms, device structure, performance characteristics, and applications of these novel RF switches. Furthermore, we examine the prospects and future research directions of these technologies that could lead to their adoption at the industrial level. Nicolás Wainstein, Gina C. Adam, Eilam Yalon, Shahar Kvatinsky |
Proc. IEEE | 4 |
| 2020 | Modeling a Floating-Gate Memristive Device for Computer Aided Design of Neuromorphic ComputingabstractMemristive technology is still not mature enough for the very large-scale integration necessary to obtain practical value from neuromorphic computing. While nonvolatile floating-gate "synapse transistors" have been implemented in very large-scale integrated neuromorphic systems, their large footprint still constrains an upper bound on the overall performance. A two-terminal floating-gate memristive device can combine the technological maturity of the floating-gate transistor and the conceptual novelty of the memristor using a standard CMOS process. In this paper, we present a top-down computer aided design framework of the floating-gate memristive device and show its potential in neuromorphic computing. Our framework includes a Verilog-A model, small-signal schematics, a stochastic model, Monte-Carlo simulations, layout, DRC, LVS, and RC extraction. Loai Danial, Vasu Gupta, Evgeny Pikhay, Yakov Roizin, Shahar Kvatinsky |
DATE | 5 |
| 2020 | CONTRA: Area-Constrained Technology Mapping Framework For Memristive Memory Processing UnitabstractData-intensive applications are poised to benefit directly from processing-in-memory platforms, such as memristive Memory Processing Units, which allow leveraging data locality and performing stateful logic operations. Developing design automation flows for such platforms is a challenging and highly relevant research problem. In this work, we investigate the problem of minimizing delay under arbitrary area constraint for MAGIC-based in-memory computing platforms. We propose an end-to-end area constrained technology mapping framework, CONTRA. CONTRA uses Look-Up Table (LUT) based mapping of the input function on the crossbar array to maximize parallel operations and uses a novel search technique to move data optimally inside the array. CONTRA supports benchmarks in a variety of formats, along with crossbar dimensions as input to generate MAGIC instructions. CONTRA scales for large benchmarks, as demonstrated by our experiments. CONTRA allows mapping benchmarks to smaller crossbar dimensions than achieved by any other technique before, while allowing a wide variety of area-delay trade-offs. CONTRA improves the composite metric of area-delay product by 2.1× to 13.1× compared to seven existing technology mapping approaches. Debjyoti Bhattacharjee, Anupam Chattopadhyay, Srijit Dutta, Ronny Ronen, Shahar Kvatinsky |
ICCAD | 5 |
| 2020 | A Pipelined Memristive Neural Network Analog-to-Digital ConverterabstractWith the advent of high-speed, high-precision, and low-power mixed-signal systems, there is an ever-growing demand for accurate, fast, and energy-efficient analog-to-digital (ADCs) and digital-to-analog converters (DACs). Unfortunately, with the downscaling of CMOS technology, modern ADCs trade off speed, power and accuracy. Recently, memristive neuromorphic architectures of four-bit ADC/DAC have been proposed. Such converters can be trained in real-time using machine learning algorithms, to break through the speed-power-accuracy trade-off while optimizing the conversion performance for different applications. However, scaling such architectures above four bits is challenging. This paper proposes a scalable and modular neural network ADC architecture based on a pipeline of four-bit converters, preserving their inherent advantages in application reconfiguration, mismatch self-calibration, noise tolerance, and power optimization, while approaching higher resolution and throughput in penalty of latency. SPICE evaluation shows that an 8-bit pipelined ADC achieves 0.18 LSB INL, 0.20 LSB DNL, 7.6 ENOB, and 0.97 fJ/conv FOM. This work presents a significant step towards the realization of large-scale neuromorphic data converters. Loai Danial, Kanishka Sharma, Shahar Kvatinsky |
ISCAS | 3 |
| 2020 | An Asynchronous and Low-Power True Random Number Generator using STT-MTJabstractThe emerging spin-transfer torque magnetic tunnel junction (STT-MTJ) technology exhibits interesting stochastic behavior combined with small area and low operation energy. It is, therefore, a promising technology for security applications, specifically the generation of random numbers. In this paper, STT-MTJ is used to construct an asynchronous true random number generator (TRNG) with low power and a high entropy rate. The asynchronous design enables the decoupling of the random number generation from the system clock, allowing it to be embedded in low-power devices. The proposed TRNG is evaluated by a numerical simulation, using the Landau-Lifshitz-Gilbert (LLG) equation as the model of the STT-MTJ devices. Design considerations, attack analysis, and process variation are discussed and evaluated. We show that our design is robust to process variation, thus achieving an entropy generating rate between 99.7 and 127.8 Mb/s with 6-7.7 pJ per bit for 90% of the instances. Ben Perach, Shahar Kvatinsky |
ISCAS | 2 |
| 2020 | abstractPIM: Bridging the Gap Between Processing-In-Memory Technology and Instruction Set ArchitectureabstractThe von Neumann architecture, in which the memory and the computation units are separated, demands massive data traffic between the memory and the CPU. To reduce data movement, new technologies and computer architectures have been explored. The use of memristors, which are devices with both memory and computation capabilities, has been considered for different processing-in-memory (PIM) solutions, including using memristive stateful logic for a programmable digital PIM system. Nevertheless, all previous work has focused on a specific stateful logic family, and on optimizing the execution for a certain target machine. These solutions require new compiler and compilation when changing the target machine, and provide no backward compatibility with other target machines. In this paper, we present abstractPIM, a new compilation concept and flow which enables executing any function within the memory, using different stateful logic families and different instruction set architectures (ISAs). By separating the code generation into two independent components, intermediate representation of the code using target independent ISA and then microcode generation for a specific target machine, we provide a flexible flow with backward compatibility and lay foundations for a PIM compiler. Using abstractPIM, we explore various logic technologies and ISAs and how they impact each other, and discuss the challenges associated with it, such as the increase in execution time. Adi Eliahu, Rotem Ben Hur, Ronny Ronen, Shahar Kvatinsky |
VLSI-SOC | 4 |
| 2020 | X-MAGIC: Enhancing PIM Using Input Overwriting CapabilitiesabstractProcessing-in-memory (PIM) using memristive technologies is an attractive solution for the memory wall problem. PIM can improve the performance and energy efficiency of computing systems by reducing the data transfer between the memory and the processor. Memristor Aided loGIC (MAGIC) is a popular memristive PIM technique that can perform any combinational logic as a sequence of atomic NOR/NOT operations. These NOR/NOT operations rely on initializing their output cell prior to computation. In this paper, we explore input overwriting: the use of the MAGIC gate output cell as an additional input without initializing it. We extend MAGIC and introduce X-MAGIC (eXtended MAGIC) which uses input overwriting, and demonstrate it by two gates, A·(B+C) and A.B̅, where A is an overwritten input. We show that input overwriting improves functionality, performance, and effective lifetime of the system. Due to algorithmic difficulties, available PIM synthesis tools do not support input overwriting. We address these difficulties by modifying an existing synthesis tool for MAGIC (SIMPLER), and presenting several general principles and methods for supporting input overwriting. We examine two configurations of the modified synthesis tool using X-MAGIC gates, differing in their performance/area trade-off. Both configurations achieve a geomean improvement of over 16.5% in performance, and over 20% in effective lifetime compared to standard MAGIC. Due to algorithmic difficulties, available PIM synthesis tools do not support input overwriting. We address these difficulties by modifying an existing synthesis tool for MAGIC (SIMPLER), and presenting several general principles and methods for supporting input overwriting. We examine two configurations of the modified synthesis tool using X-MAGIC gates, differing in their performance/area trade-off. Both configurations achieve a geomean improvement of over 16.5% in performance, and over 20% in effective lifetime compared to standard MAGIC. Natan Peled, Rotem Ben Hur, Ronny Ronen, Shahar Kvatinsky |
VLSI-SOC | 4 |
| 2020 | SIMPLER MAGIC: Synthesis and Mapping of In-Memory Logic Executed in a Single Row to Improve ThroughputabstractIn-memory processing can dramatically improve the latency and energy consumption of computing systems by minimizing the data transfer between the memory and the processor. Efficient execution of processing operations within the memory is therefore, a highly motivated objective in modern computer architecture. This article presents a novel automatic framework for efficient implementation of arbitrary combinational logic functions within a memristive memory. Using tools from logic design, graph theory and compiler register allocation technology, we developed synthesis and in-memory mapping of logic execution in a single row (SIMPLER), a tool that optimizes the execution of in-memory logic operations in terms of throughput and area. Given a logical function, SIMPLER automatically generates a sequence of atomic memristor-aided logic (MAGIC) NOR operations and efficiently locates them within a single sizelimited memory row, reusing cells to save area when needed. This approach fully exploits the parallelism offered by the MAGIC NOR gates. It allows multiple instances of the logic function to be performed concurrently, each compressed into a single row of the memory. This virtue makes SIMPLER an attractive candidate for designing in-memory single instruction, multiple data (SIMD) operations. Compared to the previous work (that optimizes latency rather than throughput for a single function), SIMPLER achieves an average throughput improvement of 435×. When the previous tools are parallelized similarly to SIMPLER, SIMPLER achieves higher throughput of at least 5×, with 23× improvement in area and 20× improvement in area efficiency. These improvements more than fully compensate for the increase (up to 17% on average) in latency. Rotem Ben Hur, Ronny Ronen, Ameer Haj-Ali, Debjyoti Bhattacharjee, Adi Eliahu, Natan Peled, Shahar Kvatinsky |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2019 | Real Processing-in-Memory with Memristive Memory Processing Unit (mMPU)abstractMemristive technologies are attractive candidates to replace conventional memory technologies and can also be used to perform logic and arithmetic operations. In this paper, we show how memristors are used to combine data storage and computation in the memory, thus enabling a novel non-von Neumann architecture called the 'memristive memory processing unit' (mMPU). The mMPU relies on a memristive logic technique called 'memristor aided logic' (MAGIC) that requires no modification to the memory array structure. By greatly reducing the data transfer between the CPU and the memory, the mMPU alleviates the primary restriction on performance and energy efficiency in modern computing systems. This paper describes basic principles and design considerations of MAGIC and the mMPU and presents a case study of digital image processing to demonstrate the benefits. Shahar Kvatinsky |
ASAP | 1 |
| 2019 | STT-ANGIE: Asynchronous True Random Number GEnerator Using STT-MTJabstractThe Spin Transfer Torque Magnetic Tunnel Junction (STT-MTJ) is an emerging memory technology whose interesting stochastic behavior might benefit security applications. In this paper, we leverage this stochastic behavior to construct a true random number generator (TRNG), the basic module in the process of encryption key generation. Our proposed TRNG operates asynchronously and thus can use small and fast STT MTJ devices. As such, it can be embedded in low-power and low-frequency devices without loss of entropy. We evaluate the proposed TRNG using a numerical simulation, solving the Landau-Lifshitz-Gilbert (LLG) equation system of the STT-MTJ devices. Design considerations, attack analysis, and process variation are discussed and evaluated. The evaluation shows that our solution is robust to process variation, achieving a Shannon-entropy generating rate between 99.7Mbps and 127.8Mbps for 90% of the instances. Ben Perach, Shahar Kvatinsky |
DATE | 2 |
| 2019 | The Missing Applications Found: Robust Design Techniques and Novel Uses of MemristorsabstractResistive memory, also known as memristor, is an emerging potential successor to traditional CMOS charge based memories. Memristors have also recently been proposed as a promising candidate for several additional applications such as logic design, sensing, non-volatile storage, neuromorphic computing, Physically Unclonable Functions (PUFs), Content-addressable memory (CAM) and reconfigurable computing. In this paper, we explore three unique applications of memristor technology based implementations, specifically from the perspective of sensing, logic, in-memory computing and their solutions. We review solar cell health monitoring and diagnosis, describe the proposed solutions, and provide directions in memristive gas sensing and in-memory computing. For the gas sensor application, in order to determine the number of memristors to ensure a certain level of accuracy in sensitivity, a technique to optimize the sensor array based on an acceptable sensitivity variation and minimum sensitivity margin is presented. These “out-of-the-box” emerging ideas for applications of memristive devices in enhancing robustness and, at the same time, how the requirements of robust design are enabling unconventional use of the devices. To this end, the papers considers some examples of this mutual interaction. Marco Ottavi, Vishal Gupta 0002, Saurabh Khandelwal, Shahar Kvatinsky, Jimson Mathew, Eugenio Martinelli, Abusaleh M. Jabir |
IOLTS | 4 |
| 2019 | Delta-Sigma Modulation Neurons for High-Precision Training of Memristive Synapses in Deep Neural NetworksabstractThe spike generation mechanism and information coding process of biological neurons can be emulated by the amplitude-to-frequency modulation property of delta-sigma modulators (ΔΣ). Oversampling, averaging, and noise-shaping features of the ΔΣ allow high neural coding accuracy and mitigate the intrinsic noise level in neural networks. In this paper, a ΔΣ is proposed as a neuron activation function for inference and training of artificial analog neural networks. The inherent dithering of the ΔΣ prevents the weights from being stuck in a spurious local minimum, and its nonlinear transfer function makes it attractive for multi-layer architectures. Memristive synapses are used as weights, which are trained by supervised/unsupervised machine learning (ML) algorithms, using stochastic gradient descent (SGD) or biologically plausible spike-time-dependent plasticity (STDP). Our ΔΣ networks outperform the prevalent power-hungry pulse width modulator counterparts, with 97.37% training accuracy and 3.2X speedup in MNIST using SGD. These findings constitute a milestone in closing the cultural gap between brain-inspired models and ML using analog neuromorphic hardware. Loai Danial, Sidharth Thomas, Shahar Kvatinsky |
ISCAS | 3 |
| 2019 | A Product Engine for Energy-Efficient Execution of Binary Neural Networks Using Resistive MemoriesabstractThe need for running complex Machine Learning (ML) algorithms, such as Convolutional Neural Networks (CNNs), in edge devices, which are highly constrained in terms of computing power and energy, makes it important to execute such applications efficiently. The situation has led to the popularization of Binary Neural Networks (BNNs), which significantly reduce execution time and memory requirements by representing the weights (and possibly the data being operated) using only one bit. Because approximately 90% of the operations executed by CNNs and BNNs are convolutions, a significant part of the memory transfers consists of fetching the convolutional kernels. Such kernels are usually small (e.g., 3×3 operands), and particularly in BNNs redundancy is expected. Therefore, equal kernels can be mapped to the same memory addresses, requiring significantly less memory to store them. In this context, this paper presents a custom Binary Dot Product Engine (BDPE) for BNNs that exploits the features of Resistive Random-Access Memories (RRAMs). This new engine allows accelerating the execution of the inference phase of BNNs. The novel BDPE locally stores the most used binary weights and performs binary convolution using computing capabilities enabled by the RRAMs. The system-level gem5 architectural simulator was used together with a C-based ML framework to evaluate the system's performance and obtain power results. Results show that this novel BDPE improves performance by 11.3%, energy efficiency by 7.4% and reduces the number of memory accesses by 10.7% at a cost of less than 0.3% additional die area, when integrated with a 28 nm Fully Depleted Silicon On Insulator ARMv8 in-order core, in comparison to a fully-optimized baseline of YoloV3 XNOR-Net running in a unmodified Central Processing Unit. João Vieira, Edouard Giacomin, Yasir Mahmood Qureshi, Marina Zapater, Xifan Tang, Shahar Kvatinsky, David Atienza 0001, Pierre-Emmanuel Gaillardon |
VLSI-SoC | 6 |
| 2019 | An Asynchronous and Low-Power True Random Number Generator Using STT-MTJabstractThe emerging spin-transfer torque magnetic tunnel junction (STT-MTJ) technology exhibits interesting stochastic behavior combined with small area and low operation energy. It is, therefore, a promising technology for security applications, specifically the generation of random numbers. In this paper, STT-MTJ is used to construct an asynchronous true random number generator (TRNG) with low power and a high entropy rate. The asynchronous design enables the decoupling of the random number generation from the system clock, allowing it to be embedded in low-power devices. The proposed TRNG is evaluated by a numerical simulation, using the Landau-Lifshitz-Gilbert (LLG) equation as the model of the STT-MTJ devices. Design considerations, attack analysis, and process variation are discussed and evaluated. We show that our design is robust to process variation, thus achieving an entropy generating rate between 99.7 and 127.8 Mb/s with 6-7.7 pJ per bit for 90% of the instances. Ben Perach, Shahar Kvatinsky |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2018 | Practical challenges in delivering the promises of real processing-in-memory machinesabstractProcessing-in-Memory (PiM) machines promise to overcome the von Neumann bottleneck in order to further scale performance and energy efficiency of computing systems by reducing the extent of data transfer and offering ample parallelism. In this paper, we take the memristive Memory Processing Unit (mMPU) as a case study of a PiM machine and scrutinize it in practical scenarios. Specifically, we explore the limitations of parallelism and data transfer elimination. We argue that lack of operand locality and arrangement might make data transfer inevitable in the mMPU. We then devise techniques to move data within the mMPU, without transferring it off-chip, and quantify their costs. Additionally, we present electrical parameters that might limit the parallelism offered by the mMPU and evaluate their impact. Using benchmarks from the LGsynth91 suite, their vector extensions, and a few synthetic data-parallel workloads, we show that the internal data transfer results in an increase of up to 1.5× in the execution time, while the parallelism can be limited in some cases to 256 gates, resulting in an increase in execution time by 1.1× to 2×. Nishil Talati, Ameer Haj-Ali, Rotem Ben Hur, Nimrod Wald, Ronny Ronen, Pierre-Emmanuel Gaillardon, Shahar Kvatinsky |
DATE | 7 |
| 2018 | Efficient Algorithms for In-Memory Fixed Point Multiplication Using MAGICabstractThe growing disparity between processor and memory performance poses significant limits on system performance and energy efficiency. To address this widely investigated problem, modern computing systems attempt to minimize data transfer by means of a memory hierarchy. Yet the benefit from such a solution for data-intensive applications is limited. Emerging non-volatile resistive memory technologies (memristors) offer the ability to both store and process data within the memristive memory cells, with almost no data transfer. In this paper, we propose algorithms for performing fixed point multiplication within the memristive memory using Memristor Aided Logic (MAGIC) gates and execute them in a cycle-accurate simulator to verify and evaluate them. Previously proposed implementations were not feasible for execution within memory because the required number of memory cells for the computation was too large to fit the size-limited memristive memory arrays. The algorithms proposed in this paper not only improve the latency as compared to previously proposed algorithms by 1.8× on average, but their significantly better area efficiency now makes it possible to perform numerous fixed point multiplications simultaneously within memristive memory arrays. Ameer Haj-Ali, Rotem Ben Hur, Nimrod Wald, Shahar Kvatinsky |
ISCAS | 4 |
| 2017 | Memristor for computing: Myth or reality?abstractCMOS technology and its sustainable scaling have been the enablers for the design and manufacturing of computer architectures that have been fuelling a wider range of applications. Today, however, both the technology and the computer architectures are suffering from serious challenges/ walls making them incapable to deliver the right computing power at pre-defined constraints. This motivates the need of exploring new architectures and new technologies; not only to maintain the economic benefit of scaling, but also to enable the solutions of emerging computer power and data storage hungry applications such as big-data and data-intensive applications. This paper discusses the emerging memristor device as complementary (or alternative) to CMOS device and shows how this device can enable new ways of computing that will at least solve the challenges of today's architectures for some applications. The paper shows not only the potential of memristor devices in enabling new memory technologies and new logic design styles, but also their potential in enabling memory intensive architectures as well as neuromorphic computing due to their unique properties such as the tight integration with CMOS and the ability to learn and adapt. Said Hamdioui, Shahar Kvatinsky, Gert Cauwenberghs, Lei Xie 0005, Nimrod Wald, Siddharth Joshi 0001, Hesham Mostafa Elsayed, Henk Corporaal, Koen Bertels |
DATE | 2 |
| 2017 | Simple magic: Synthesis and in-memory Mapping of logic execution for memristor-aided logicabstractThis paper presents a novel approach for designing and implementing in-memory logic operations. The uniqueness of this work is the development of SIMPLE, a framework that optimizes the execution of an arbitrary logic function, while considering all the constraints involved in performing it within a memristive memory. SIMPLE automatically generates a defined sequence of atomic memristor-aided logic NOR operations, whose implementation can be facilitated efficiently within the memory. Motivated to overcome the memory-CPU bottleneck, this approach designs an optimal solution in terms of performance by exploiting the parallelism of the memristor-aided logic NOR gates. SIMPLE achieves performance speedups of 1.94x compared to a previous work and 1.48x compared to a naïve optimization based on standard synthesis tools. Rotem Ben Hur, Nimrod Wald, Nishil Talati, Shahar Kvatinsky |
ICCAD | 4 |
| 2017 | Rate-compatible and high-throughput architecture designs for encoding LDPC codesabstractLow-density parity-check (LDPC) codes are known for superior performance over a wide range of codes for communication and memory systems. In many practical scenarios, adaptive ECC system is preferred that can adapt to various codes with varying channel conditions since the behavior of errors changes with time and space. This paper presents two architectural designs for efficient encoding of LDPC codes to support different code rates and lengths, which can be used for several applications. The proposed designs allow switching among different codes without any hardware modification. The first proposed design achieves extremely high throughput by removing the memory from the encoder, while still being able to adapt to a few predefined codes. The other architecture can adapt to any arbitrary code by using the memory for configuration, and yet, it achieves up to 12.9x throughput and 17.5x area improvement as compared to fully-reconfigurable encoders proposed in literature. Nishil Talati, Zhiying Wang 0001, Shahar Kvatinsky |
ISCAS | 3 |
| 2017 | An RF memristor model and memristive single-pole double-throw switchesabstractIn this paper, we present a scalable physics-based model that accurately predicts the steady-state high-frequency behavior of memristive RF switches. This model is, to the best of our knowledge, the first lumped RF memristor model that includes device parasitics obtained from physical measurements reported in the literature. Furthermore, we propose two topologies (series and shunt) for non-volatile single-voltage-controlled Single-Pole Double-Throw switches using the proposed lumped RF memristor model. These topologies exhibit low insertion loss and high isolation. Adding non-volatility to RF switches will result in reduced power consumption. Nicolás Wainstein, Shahar Kvatinsky |
ISCAS | 2 |
| 2016 | A fully analog memristor-based neural network with online gradient trainingabstractIn recent years, Neural Networks (NNs) have become widely popular for the execution of different machine learning algorithms. Training an NN is computationally intensive since it requires numerous multiplications of matrices that represent synaptic weights. It is therefore appealing to build a hardware-based NN accelerator to gain parallelism and efficient computation. Recently, we have proposed a compact circuit of a non-volatile synaptic weight based on two CMOS transistors and a memristor. In this paper, we present a fully analog NN design based on our previously proposed synapse with a full design of the different layers and their supporting CMOS circuits. We show that the presented NN significantly reduces the area as compared to a CMOS-based NN, while executing online gradient training with similar accuracy and computational speed improvement as a software implementation. Eyal Rosenthal, Sergey Greshnikov, Daniel Soudry, Shahar Kvatinsky |
ISCAS | 4 |
| 2016 | Write sneak-path constraints avoiding disturbs in memristor crossbar arraysabstractWe study the problem of write disturbs due to write sneak paths in memristor crossbar arrays. A write sneak path is a bit configuration in the array that causes a write of one cell to undesirably flip the value of another cell. We study the configurations that cause such write sneak paths, and characterize them in terms of tight constraints to prevent them. We show that thanks to the flexibility to choose the write order, the resulting constraints are milder compared to known similar ones for read sneak paths. In addition, we derive the array constraints when parallel write is allowed in the rows or column only, and in both the rows and columns. Yuval Cassuto, Shahar Kvatinsky, Eitan Yaakobi |
ISIT | 2 |
| 2016 | Improving energy efficiency of DRAM by exploiting half page row accessabstractDRAM energy is an important component to optimize in modern computing systems. One outstanding source of DRAM energy is the energy to fetch data stored on cells to the row buffer, which occurs during two DRAM operations, row activate and refresh. This work exploits previously proposed half page row access, modifying the wordline connections within a bank to halve the number of cells fetched to the row buffer, to save energy in both cases. To accomplish this, we first change the data wire connections in the sub-array to reduce the cost of row buffer overfetch in multi-core systems which yields a 12% energy savings on average and a slight performance improvement in quad-core systems. We also propose charge recycling refresh, which reuses charge left over from a prior half page refresh to refresh another half page. Our charge recycling scheme is capable of reducing both auto- and self-refresh energy, saving more than 15% of refresh energy at 85°C, and provides even shorter refresh cycle time. Finally, we propose a refresh scheduling scheme that can dynamically adjust the number of charge recycled half pages, which can save up to 30% of refresh energy at 85°C. Heonjae Ha, Ardavan Pedram, Stephen Richardson, Shahar Kvatinsky, Mark Horowitz |
MICRO | 4 |
| 2016 | Evaluating programmable architectures for imaging and vision applicationsabstractAlgorithms for computational imaging and computer vision are rapidly evolving, and hardware must follow suit: the next generation of image signal processors (ISPs) must be “programmable” to support new algorithms created with high-level frameworks. In this work, we compare flexible ISP architectures, using applications written in the Darkroom image processing language. We target two fundamental architecture classes: programmable in time, as represented by SIMD, and programmable in space, as typified by coarse grain reconfigurable array architectures (CGRA). We consider several optimizations on these two base architectures, such as register file partitioning for SIMD, bus based routing and pipelined wires for CGRA, and line buffer variations. After these optimizations on average, CGRA provides 1.6x better energy efficiency and 1.4x better compute density versus a SIMD solution, and 1.4x the energy efficiency and 3.1x the compute density of an FPGA. However the cost of providing general programmability is still high: compared to an ASIC, CGRA has 6x worse energy and area efficiency, and this ratio would be roughly 10x if memory dominated applications were excluded. Artem Vasilyev, Nikhil Bhagdikar, Ardavan Pedram, Stephen Richardson, Shahar Kvatinsky, Mark Horowitz |
MICRO | 5 |
| 2016 | Logic design with unipolar memristorsabstractMemristors are novel devices that could naturally be designed as memory elements. Recently, several methods of designing memristors for logic operations have been proposed. These methods, mostly, make use of bipolar memristors. In this paper, we propose a method for performing logic with unipolar memristors based on OR and NOT logic gates. An integration of the basic building blocks into more complicated logic functions is described and demonstrated. Our results indicate that any logic function could be performed using an external controller. Thus, adding the capability of performing logic computations to numerous types of unipolar memristive materials in addition to their memory capability. Elad Amrani, Avishay Drori, Shahar Kvatinsky |
VLSI-SoC | 3 |
| 2016 | Resistive GP-SIMD Processing-In-MemoryabstractGP-SIMD, a novel hybrid general-purpose SIMD architecture, addresses the challenge of data synchronization by in-memory computing, through combining data storage and massive parallel processing. In this article, we explore a resistive implementation of the GP-SIMD architecture. In resistive GP-SIMD, a novel resistive row and column addressable 4F 2 crossbar is utilized, replacing the modified CMOS 190F 2 SRAM storage previously proposed for GP-SIMD architecture. The use of the resistive crossbar allows scaling the GP-SIMD from few millions to few hundred millions of processing units on a single silicon die. The performance, power consumption and power efficiency of a resistive GP-SIMD are compared with the CMOS version. We find that PiM architectures and, specifically, GP-SIMD benefit more than other many-core architectures from using resistive memory. A framework for in-place arithmetic operation on a single multivalued resistive cell is explored, demonstrating a potential to become a building block for next-generation PiM architectures. Amir Morad, Leonid Yavits, Shahar Kvatinsky, Ran Ginosar |
ACM Trans. Archit. Code Optim. | 3 |
| 2016 | Information-Theoretic Sneak-Path Mitigation in Memristor Crossbar ArraysabstractIn a memristor crossbar array, functioning as a memory array, a memristor is positioned on each row-column intersection, and its resistance, low or high, represents two logical states. The state of every memristor can be sensed by the current flowing through the memristor. In this paper, we study the sneak path problem in crossbar arrays, in which current can sneak through other cells, resulting in reading a wrong state of the memristor. Our main contributions are modeling the error channel induced by sneak paths, a new characterization of arrays free of sneak paths, and efficient methods to read the array cells while avoiding sneak paths. To each read method, we match a constraint on the array content that guarantees sneak-path free readout, determine the resulting capacity, and provide an efficient encoder that achieves the capacity. Yuval Cassuto, Shahar Kvatinsky, Eitan Yaakobi |
IEEE Trans. Inf. Theory | 2 |
| 2015 | Memristor-Based Multilayer Neural Networks With Online Gradient Descent TrainingabstractLearning in multilayer neural networks (MNNs) relies on continuous updating of large matrices of synaptic weights by local rules. Such locality can be exploited for massive parallelism when implementing MNNs in hardware. However, these update rules require a multiply and accumulate operation for each synaptic weight, which is challenging to implement compactly using CMOS. In this paper, a method for performing these update operations simultaneously (incremental outer products) using memristor-based arrays is proposed. The method is based on the fact that, approximately, given a voltage pulse, the conductivity of a memristor will increment proportionally to the pulse duration multiplied by the pulse magnitude if the increment is sufficiently small. The proposed method uses a synaptic circuit composed of a small number of components per synapse: one memristor and two CMOS transistors. This circuit is expected to consume between 2% and 8% of the area and static power of previous CMOS-only hardware alternatives. Such a circuit can compactly implement hardware MNNs trainable by scalable algorithms based on online gradient descent (e.g., backpropagation). The utility and robustness of the proposed memristor-based circuit are demonstrated on standard supervised learning tasks. Daniel Soudry, Dotan Di Castro, Asaf Gal, Avinoam Kolodny, Shahar Kvatinsky |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2015 | Multistate Register Based on Resistive RAMabstractIn recent years, memristive technologies, such as resistive random access memory (RRAM), have emerged. These technologies are usually considered as alternates for static RAM, dynamic RAM, and Flash. In this paper, a novel digital circuit, the multistate register, is proposed. The multistate register is different from conventional types of memory, and is used to store multiple data bits, where only a single bit is active and the remaining data bits are idle. The active bit is stored within a CMOS flip flop, while the idle bits are stored in an RRAM crossbar co-located with the flip flop. It is demonstrated that additional states require an area overhead of 1.4% per state for a 64-state register. The use of multistate registers as pipeline registers is demonstrated for a novel multithreading architecture-continuous flow multithreading (CFMT), where the total area overhead in the CPU pipeline is only 2.5% for 16 threads compared with a single thread CMOS pipeline. The use of multistate registers in the CFMT microarchitecture enables higher performance processors (40% average performance improvement) with relatively low energy (6.5% average energy reduction) and area overhead. Ravi Patel 0001, Shahar Kvatinsky, Eby G. Friedman, Avinoam Kolodny |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2014 | Memristor-Based Material Implication (IMPLY) Logic: Design Principles and MethodologiesabstractMemristors are novel devices, useful as memory at all hierarchies. These devices can also behave as logic circuits. In this paper, the IMPLY logic gate, a memristor-based logic circuit, is described. In this memristive logic family, each memristor is used as an input, output, computational logic element, and latch in different stages of the computing process. The logical state is determined by the resistance of the memristor. This logic family can be integrated within a memristor-based crossbar, commonly used for memory. In this paper, a methodology for designing this logic family is proposed. The design methodology is based on a general design flow, suitable for all deterministic memristive logic families, and includes some additional design constraints to support the IMPLY logic family. An IMPLY 8-bit full adder based on this design methodology is presented as a case study. Shahar Kvatinsky, Guy Satat, Nimrod Wald, Eby G. Friedman, Avinoam Kolodny, Uri C. Weiser |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2013 | Sneak-path constraints in memristor crossbar arraysabstractIn a memristor crossbar array, a memristor is positioned on each row-column intersection, and its resistance, low or high, represents two logical states. The state of every memristor can be sensed by the current flowing through the memristor. In this work, we study the sneak path problem in crossbars arrays, in which current can sneak through other cells, resulting in reading a wrong state of the memristor. Our main contributions are a new characterization of arrays free of sneak paths, and efficient methods to read the array cells while avoiding sneak paths. To each read method we match a constraint on the array content that guarantees sneak-path free readout, and calculate the resulting capacity. Yuval Cassuto, Shahar Kvatinsky, Eitan Yaakobi |
ISIT | 2 |
| 2011 | Memristor-based IMPLY logic design procedureabstractMemristors can be used as logic gates. No design methodology exists, however, for memristor-based combinatorial logic. In this paper, the design and behavior of a memristive-based logic gate - an IMPLY gate - are presented and design issues such as the tradeoff between speed (fast write times) and correct logic behavior are described, as part of an overall design methodology. A memristor model is described for determining the write time and state drift. It is shown that the widely used memristor model - a linear ion drift memristor - is impractical for characterizing an IMPLY logic gate, and a different memristor model is necessary such as a memristor with a current threshold. Shahar Kvatinsky, Avinoam Kolodny, Uri C. Weiser, Eby G. Friedman |
ICCD | 1 |