Dmitri B. Strukov

dblp:39/6944 · DBLP profile ↗
← Back
33ranked-venue papers
3as first author
4since 2021 · last 2025
0000-0002-4526-4347ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 29 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 6 · 1 first-authorArtificial intelligence and machine learning · 4Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2025 KLIMA: Low-latency mixed-signal In-Memory Computing accelerator for solving arbitrary-order Boolean Satisfiability
Tinish Bhattacharya, Dongseok Kwon, George Higgins Hutchinson, Xiangyi Zhang, Giacomo Pedretti, Fabian Böhm, John Paul Strachan, Thomas Van Vaerenbergh, Raymond G. Beausoleil, Ignacio Rozada, Dmitri B. Strukov
HCS11
2024 Memristor-based hardware and algorithms for higher-order Hopfield optimization solver outperforming quadratic Ising machines
abstract
Ising solvers offer a promising physics-based approach to tackle the challenging class of combinatorial optimization problems. However, typical solvers operate in a quadratic energy space, having only pair-wise coupling elements which already dominate area and energy. We show that such quadratization can cause severe problems: increased dimensionality, a rugged search landscape, and misalignment with the original objective function. Here, we design and quantify a higher-order Hopfield optimization solver, with 28nm CMOS technology and memristive couplings for lower area and energy computations. We combine algorithmic and circuit analysis to show quantitative advantages over quadratic Ising Machines (IM)s, yielding 48x and 72x reduction in time-to-solution (TTS) and energy-to-solution (ETS) respectively for Boolean satisfiability problems of 150 variables, with favorable scaling.
Mohammad Hizzani, Arne Heittmann, George Higgins Hutchinson, Dmitrii Dobrynin, Thomas Van Vaerenbergh, Tinish Bhattacharya, Adrien Renaudineau, Dmitri B. Strukov, John Paul Strachan
ISCAS8
2024 FPIA: Field-Programmable Ising Arrays with In-Memory Computing
abstract
Ising Machines, a promising approach for solving combinatorial optimization problems, are naturally suited for energy-saving and compact in-memory computing implementations with emerging memories. A naïve in-memory computing implementation of a quadratic Ising Machine requires an array of coupling weights that grows quadratically with problem size. This approach, however, uses resources inefficiently due to the inherent sparsity of practical optimization problems. We first show that this issue can be addressed by partitioning a coupling array into smaller sub-arrays. This technique, however, requires interconnecting sub-arrays, which incurs overhead. In response, we present FPIA, an in-memory computing architecture for quadratic Ising Machines inspired by island-type field programmable gate arrays. We adapt open-source tools to optimize problem embedding and model overhead. Modeling results of benchmark problems for the developed architecture show up to 10x increase in density and speed compared to the baseline approach. Finally, we discuss algorithm/circuit co-design techniques for further improvements.
George Higgins Hutchinson, Ethan Sifferman, Tinish Bhattacharya, Dongseok Kwon, Dmitri B. Strukov
ISLPED5
2021 The Impact of Device Uniformity on Functionality of Analog Passively-Integrated Memristive Circuits
abstract
Passively-integrated memristors are the most prospective candidates for designing high-speed, energy-efficient, and compact neuromorphic circuits. Despite all the promising properties, experimental demonstrations of passive memristive crossbars have been limited to circuits with few thousands of devices until now, which stems from the strict uniformity requirements on theIVcharacteristics of memristors. This paper expands upon this vital challenge and investigates how uniformity impacts the computing accuracy of analog memristive circuits, focusing on neuromorphic applications. Specifically, the paper explores the tradeoffs between computing accuracy, crossbar size, switching threshold variations, and target precision. All-embracing simulations of matrix multipliers and deep neural networks on CIFAR-10 and ImageNet datasets have been carried out to evaluate the role of uniformity on the accuracy of computing systems. Further, we study three post-fabrication methods that increase the accuracy of nonuniform 0T1R neuromorphic circuits: hardware-aware training, improved tuning algorithm, and switching threshold modification. The application of these techniques allows us to implement advanced deep neural networks with almost no accuracy drop, using current state-of-the-art analog 0T1R technology.
Z. Fahimi, Mohammad Reza Mahmoodi, Michael Klachko, Hussein Nili, Dmitri B. Strukov
IEEE Trans. Circuits Syst. I Regul. Pap.5
2020 Mixed-Signal Vector-by-Matrix Multiplier Circuits Based on 3D-NAND Memories for Neurocomputing
abstract
We propose an extremely dense, energy-efficient mixed-signal vector-by-matrix-multiplication (VMM) circuits based on the existing 3D-NAND flash memory blocks, without any need for their modification. Such compatibility is achieved using time-domain-encoded VMM design. We have performed rigorous simulations of such a circuit, taking into account non-idealities such as drain-induced barrier lowering, capacitive coupling, charge injection, parasitics, process variations, and noise. Our results, for example, show that the 4-bit VMM of 200-element vectors, using the commercially available 64-layer gate-all-around macaroni-type 3D-NAND memory blocks designed in the 55-nm technology node, may provide an unprecedented area efficiency of 0.14 pm2/byte and energy efficiency of ~11 fJ/Op, including the input/output and other peripheral circuitry overheads.
Mohammad Bavandpour, Shubham Sahay, Mohammad Reza Mahmoodi, Dmitri B. Strukov
DATE4
2020 Efficient Mixed-Signal Neurocomputing Via Successive Integration and Rescaling
abstract
The widespread and ever-increasing demand for performing in situ inference, signal processing, and other computationally intensive applications in mobile Internet-of-Things (IoT) devices requires fast, compact, and energy-efficient vector-by-matrix multipliers (VMMs). The time-domain VMMs based on emerging nonvolatile memory devices exhibit significantly higher circuit density and energy efficiency than their current-mode counterparts. However, the load capacitors used to accumulate the weighted summation of the inputs in the time-domain-based circuits dominate their energy dissipation and footprint area. The true potential of the time-domain-based VMMs may be realized only when this overhead is minimized. To this end, in this brief, we propose a novel successive integration and rescaling (SIR) approach for implementing a highly efficient mixed-signal time-domain VMM for low-to-medium-precision computing. For a proof of concept, we quantitatively evaluated the performance of the proposed SIR VMM and compared it with the results of the conventional time-domain VMM, using a similar 1T-1R array. Preliminary simulation results for the 4-bit $200\, \times \, 200$ VMM, implemented using a 55-nm technology node, show area and energy efficiencies of 1.33 bits/m2and ~1.3 POp/J-the numbers, respectively, $\sim 2.5\times $ and $\sim 2.65\times $ higher than those for the prior-work time-domain VMM. Furthermore, we analyze the system-level performance of the proposed SIR VMM engine in the neuromorphic accelerator architectures and provide the preliminary estimates for various deep/recurrent neural network (DNN/RNN) applications.
Mohammad Bavandpour, Shubham Sahay, Mohammad Reza Mahmoodi, Dmitri B. Strukov
IEEE Trans. Very Large Scale Integr. Syst.4
2019 Boosted Race Trees for Low Energy Classification
abstract
When extremely low-energy processing is required, the choice of data representation makes a tremendous difference. Each representation (e.g. frequency domain, residue coded, log-scale) comes with a unique set of trade-offs --- some operations are easier in that domain while others are harder. We demonstrate that race logic, in which temporally coded signals are getting processed in a dataflow fashion, provides interesting new capabilities for in-sensor processing applications. Specifically, with an extended set of race logic operations, we show that tree-based classifiers can be naturally encoded, and that common classification tasks can be implemented efficiently as a programmable accelerator in this class of logic. To verify this hypothesis, we design several race logic implementations of ensemble learners, compare them against state-of-the-art classifiers, and conduct an architectural design space exploration. Our proof-of-concept architecture, consisting of 1,000 reconfigurable Race Trees of depth 6, will process 15.2M frames/s, dissipating 613mW in 14nm CMOS.
Georgios Tzimpragos, Advait Madhavan, Dilip P. Vasudevan, Dmitri B. Strukov, Timothy Sherwood
ASPLOS4
2019 ChipSecure: A Reconfigurable Analog eFlash-Based PUF with Machine Learning Attack Resiliency in 55nm CMOS
abstract
We exploit randomness in static I-V characteristics and reconfigurability of embedded flash memories to design very efficient physically unclonable function. Leakage current and subthreshold slope variations, nonlinearity, nondeterministic tuning error, and sneak path current in the redesigned commercial flash memory arrays are exploited to create a unique digital fingerprint. A time-multiplexed architecture is designed to enhance the security and expand the challenge-response pair space to 10211. Experimental results demonstrate 50.3% average uniformity, 49.99% average diffuseness, and native <5% bit error rate. The analysis of the measured data also shows strong resilience against machine learning attacks and possibility for extremely energy efficient, 0.56 pJ/b operation.
Mohammad Reza Mahmoodi, Hussein Nili, Shabnam Larimian, Xinjie Guo, Dmitri B. Strukov
DAC5
2019 Improving Noise Tolerance of Mixed-Signal Neural Networks
abstract
Mixed-signal hardware accelerators for deep learning achieve orders of magnitude better power efficiency than their digital counterparts. In the ultra-low power consumption regime, limited signal precision inherent to analog computation becomes a challenge. We perform a case study of a 6-layer convolutional neural network running on a mixed-signal accelerator and evaluate its sensitivity to hardware specific noise. We apply various methods to improve noise robustness of the network and demonstrate an effective way to optimize useful signal ranges through adaptive signal clipping. The resulting model is robust enough to achieve 80.2% classification accuracy on CIFAR-10 dataset with just 1.4 mW power budget, while 6 mW budget allows us to achieve 87.1% accuracy, which is within 1% of the software baseline. For comparison, the unoptimized version of the same model achieves only 67.7% accuracy at 1.4 mW and 78.6% at 6 mW.
Michael Klachko, Mohammad Reza Mahmoodi, Dmitri B. Strukov
IJCNN3
2019 Towards the Development of Analog Neuromorphic Chip Prototype with 2.4M Integrated Memristors
abstract
We have designed and fabricated a neuromorphic accelerator chip which features 180-nm CMOS circuitry and 2.4-million 0.04-μm2-footprint Al2O3/TiO2-xmemristors. Memristors were passively integrated on top of the CMOS wafer into 48×48 crossbar circuits, with a total of 1032 crossbars on a chip. Each memristive crossbar is accessed via on-chip CMOS interface circuits which are controlled by a custom FPGA board. The whole system is designed to implement a variety of deep neural networks by performing vector-by-matrix multiplication, a core operation of any neural network, with a memristive crossbar circuit in analog domain. In this this paper we provide details of the on-chip CMOS circuits and the FPGA board, both of which have already been successfully tested. We also briefly discuss integration results and the on-going work.
Irina Kataeva, Shigeki Ohtsuka, Hussein Nili, Hyungjin Kim 0001, Yoshihiko Isobe, Koichi Yako, Dmitri B. Strukov
ISCAS7
2019 Preliminary Results Towards Reinforcement Learning with Mixed-Signal Memristive Neuromorphic Circuits
abstract
As the end of Moore's law seems to be imminent, emerging technologies that enable high performance neuromorphic hardware systems are attracting increasing attention. A very promising approach is to utilize memristors, programmable nonvolatile memory devices, as synaptic weights in neuromorphic circuits. One of the challenges for memristive hardware with integrated learning capabilities is prohibitively larger number of write cycles that might be required during learning process. In this work we propose a memristive neuromorphic hardware implementation for reinforcement learning based on temporal difference actor-critic algorithm. As a case study, we consider a task of balancing an inverted pendulum, a classical problem in both reinforcement learning and control theory. We introduce training techniques that significantly reduce the number of weight updates and are suitable for efficient in-situ learning hardware implementations. We believe that this study shows the promise of using memristor-based hardware neural networks for handling complex tasks through in-situ reinforcement learning.
Nan Wu 0009, Adrien F. Vincent, Dmitri B. Strukov
ISCAS3
2018 An ultra-low energy internally analog, externally digital vector-matrix multiplier based on NOR flash memory technology
abstract
Vector-matrix multiplication (VMM) is a core operation in many signal and data processing algorithms. Previous work showed that analog multipliers based on nonvolatile memories have superior energy efficiency as compared to digital counterparts at low-to-medium computing precision. In this paper, we propose extremely energy efficient analog mode VMM circuit with digital input/output interface and configurable precision. Similar to some previous work, the computation is performed by gate-coupled circuit utilizing embedded floating gate (FG) memories. The main novelty of our approach is an ultra-low power sensing circuitry, which is designed based on translinear Gilbert cell in topological combination with a floating resistor and a low-gain amplifier. Additionally, the digital-to-analog input conversion is merged with VMM, while current-mode algorithmic analog-to-digital circuit is employed at the circuit backend. Such implementations of conversion and sensing allow for circuit operation entirely in a current domain, resulting in high performance and energy efficiency. For example, post-layout simulation results for 400×400 5-bit VMM circuit designed in 55 nm process with embedded NOR flash memory, show up to 400 MHz operation, 1.68 POps/J energy efficiency, and 39.45 TOps/mm2 computing throughput. Moreover, the circuit is robust against process-voltage-temperature variations, in part due to inclusion of additional FG cells that are utilized for offset compensation.
Mohammad Reza Mahmoodi, Dmitri B. Strukov
DAC2
2018 Mixed-Signal POp/J Computing with Nonvolatile Memories
abstract
The present-day revolution in deep learning was triggered not by any significant algorithm breakthrough, but by the use of more powerful GPU hardware [1]. Though this revolution has stimulated the development of even more powerful dedicated digital systems [2, 3], their speed and energy efficiency are still insufficient for ultrafast pattern classification and more ambitious cognitive tasks. The main reason is that the use of digital operations for the implementation of neuromorphic networks, with their high redundancy and noise/variability tolerance, is inherently unnatural. On the other hand, the network performance may be dramatically improved using mixed-signal integrated circuits, where the key inference-stage operation, the vector-by-matrix multiplication, is implemented on the physical level by utilization of the fundamental Ohm and Kirchhoff laws [4-6].
Mohammad Reza Mahmoodi, Dmitri B. Strukov
ACM Great Lakes Symposium on VLSI2
2018 Breaking POps/J Barrier with Analog Multiplier Circuits Based on Nonvolatile Memories
abstract
Low-to-medium resolution analog vector-by-matrix multipliers (VMMs) offer a remarkable energy/area efficiency as compared to their digital counterparts. Still, the maximum attainable performance in analog VMMs is often bounded by the overhead of the peripheral circuits. The main contribution of this paper is the design of novel sensing circuitry which improves energy-efficiency and density of analog multipliers. The proposed circuit is based on translinear Gilbert cell, which is topologically combined with a floating nonlinear resistor and a low-gain amplifier. Several compensation techniques are employed to ensure reliability with respect to process, temperature, and supply voltage variations. As a case study, we consider implementation of couple-gate current-mode VMM with embedded split-gate NOR flash memory. Our simulation results show that a 4-bit 100x100 VMM circuit designed in 55 nm CMOS technology achieves the record-breaking performance of 3.63 POps/J.
Mohammad Reza Mahmoodi, Dmitri B. Strukov
ISLPED2
2018 High-Performance Mixed-Signal Neurocomputing With Nanoscale Floating-Gate Memory Cell Arrays
abstract
Potential advantages of analog- and mixed-signal nanoelectronic circuits, based on floating-gate devices with adjustable conductance, for neuromorphic computing had been realized long time ago. However, practical realizations of this approach suffered from using rudimentary floating-gate cells of relatively large area. Here, we report a prototype $28\times28$ binary-input, ten-output, three-layer neuromorphic network based on arrays of highly optimized embedded nonvolatile floating-gate cells, redesigned from a commercial 180-nm nor flash memory. All active blocks of the circuit, including 101 780 floating-gate cells, have a total area below 1 mm2. The network has shown a 94.7% classification fidelity on the common Modified National Institute of Standards and Technology benchmark, close to the 96.2% obtained in simulation. The classification of one pattern takes a sub-1- $\mu \text{s}$ time and a sub-20-nJ energy-both numbers much better than in the best reported digital implementations of the same task. Estimates show that a straightforward optimization of the hardware and its transfer to the already available 55-nm technology may increase this advantage to more than $10^{2}\times $ in speed and $10^{4}\times $ in energy efficiency.
Farnood Merrikh-Bayat, Xinjie Guo, Michael Klachko, Mirko Prezioso, Konstantin Likharev, Dmitri B. Strukov
IEEE Trans. Neural Networks Learn. Syst.6
2018 High-Throughput Pattern Matching With CMOL FPGA Circuits: Case for Logic-in-Memory Computing
abstract
In this paper, we propose a novel CMOS+ MOLecular (CMOL) field-programmable gate array (FPGA) circuit architecture to perform massively parallel, high-throughput computations, which is especially useful for pattern matching tasks and multidimensional associative searches. In the new architecture, patterns are stored as resistive states of emerging nonvolatile memory nanodevices, while the analyzed data are streamed via CMOS subsystem. The main improvements over prior work offered by the proposed circuits are increased nanodevice utilization and, as a result, substantially higher throughput, which is demonstrated by a detailed analysis of the implementation of pattern matching task on the new architecture. For example, our estimates show that the proposed CMOL FPGA circuits based on the 22-nm CMOS technology and one crossbar layer with 22-nm nanowire half-pitch allows up to 12.5% average nanodevice utilization, i.e., the fraction of the devices turned to the high conductive state, as compared to a typical ~0.1% of the original CMOL FPGA circuits. This in turn enables throughput close to 7.1 × 1016bits/s/cm2at ~ 1 fJ/bit energy efficiency, for matching of ~ 107250-bit patterns stored locally on a 1 cm2chip. These numbers are at least 2 orders of magnitude better throughput as compared to that of other state-of-the-art FPGA methods, and begin to approach ternary content-addressable memory -like performance at similar CMOS technology nodes. More generally, we argue that the proposed concept combines the versatility of reconfigurable architectures and density of the associative memories. It can be viewed as a very tight symbiotic integration of memory and logic functions for high-performance logic-in-memory computing.
Advait Madhavan, Timothy Sherwood, Dmitri B. Strukov
IEEE Trans. Very Large Scale Integr. Syst.3
2017 3D-DPE: A 3D high-bandwidth dot-product engine for high-performance neuromorphic computing
abstract
We present and experimentally validate 3D-DPE, a general-purpose dot-product engine, which is ideal for accelerating artificial neural networks (ANNs). 3D-DPE is based on a monolithically integrated 3D CMOS-memristor hybrid circuit and performs a high-dimensional dot-product operation (a recurrent and computationally expensive operation in ANNs) within a single step, using analog current-based computing. 3D-DPE is made up of two subsystems, namely a CMOS subsystem serving as the memory controller and an analog memory subsystem consisting of multiple layers of high-density memristive crossbar arrays fabricated on top of the CMOS subsystem. Their integration is based on a high-density area-distributed interface, resulting in much higher connectivity between the two subsystems, compared to the traditional interface of a 2D system or a 3D system integrated using through silicon vias. As a result, 3D-DPE's single-step dot-product operation is not limited by the memory bandwidth, and the input dimension of the operations scales well with the capacity of the 3D memristive arrays. To demonstrate the feasibility of 3D-DPE, we designed and fabricated a CMOS memory controller and monolitically integrated 2 layers of titanium-oxide memristive crossbars. Then we performed the analog dot-product operation under different input conditions in two scenarios: (1) with devices within the same crossbar layer and (2) with devices from different layers. In both cases, the devices exhibited low voltage operation and analog switching behavior with high tuning accuracy.
Miguel Angel Lastras-Montaño, Bhaswar Chakrabarti, Dmitri B. Strukov, Kwang-Ting Cheng
DATE3
2017 Memristor-based perceptron classifier: Increasing complexity and coping with imperfect hardware
abstract
We experimentally demonstrate classification of 4×4 binary images into 4 classes, using a 3-layer mixed-signal neuromorphic network (“MLP perceptron”), based on two passive 20×20 memristive crossbar arrays, board-integrated with discrete CMOS components. The network features 10 hidden-layer and 4 output-layer analog CMOS neurons and 428 metal-oxide memristors, i.e. is almost an order of magnitude more complex than any previously reported functional passive (0T1R) memristor classifier. Moreover, the inference operation of this classifier is performed entirely in the integrated hardware. To deal with larger crossbar arrays, we have developed a semiautomatic approach to their forming and testing, and compared several memristor training schemes for coping with imperfect behavior of these devices, as well as with variability of analog CMOS neurons. The effectiveness of the proposed schemes for defect and variation tolerance was verified experimentally using the implemented network and, additionally, by modeling the operation of a larger network, with 300 hidden-layer neurons, on the MNIST benchmark. Finally, we propose a simple modification of the implemented memristor-based vector-by-matrix multiplier to allow its operation in a wider temperature range.
Farnood Merrikh-Bayat, Mirko Prezioso, Bhaswar Chakrabarti, Irina Kataeva, Dmitri B. Strukov
ICCAD5
2017 Exponential-weight multilayer perceptron
abstract
Analog integrated circuits may increase the neuromorphic network performance dramatically, leaving far behind their digital and biological counterparts, while approaching the energy efficiency of the brain. The key component of the most advanced analog circuit implementations is a nanodevice with adjustable conductance - essentially an analog nonvolatile memory cell, which could mimic synaptic transmission function by multiplying signal from the input neuron (e.g. encoded as voltage applied to the memory device) by its analog weight (device conductance) and passing the product (the resulting current) to the output neuron. Such functionality enables very dense, fast, and low power implementation of dot-product computation, the most common operation in many artificial neural networks. The most promising analog memory devices, however, have nonlinear, typically exponential, I-V characteristics, which result in nonlinear synaptic transmission, thus limiting their application in analog dot-product circuits. Here we investigate multilayer perceptron with exponential transmission function synapses which maps naturally to the most advanced analog neuromorphic circuits. Our simulation results show that the proposed exponential-weight multilayer perceptron with 300 hidden neurons achieves classification performance comparable to the similar-size linear-weight network when benchmarked on MNIST dataset. Moreover, we verify the proposed idea experimentally by implementing small-scale single-layer exponential-weight perceptron classifier with an NOR-flash memory integrated circuit.
Farnood Merrikh-Bayat, Xinjie Guo, Dmitri B. Strukov
IJCNN3
2016 Energy efficient computation with asynchronous races
abstract
By encoding information as digital signal propagation delay, rather than conventional logic levels, some basic processing operations become exceedingly energy efficient to implement. The result of such a computation can then be observed by relative timing differences between injected signals. We demonstrate the embodiment of such an approach utilizing current starved inverters as delay elements and characterize application-level artifacts of circuit-level variance. Specifically we chose the well-studied DNA sequence alignment problem for comparison and we show that, for the synthesized design, asynchronous races are 10× more energy efficient and 4× denser at comparable speeds as compared to prior approaches.
Advait Madhavan, Timothy Sherwood, Dmitri B. Strukov
DAC3
2016 Mellow Writes: Extending Lifetime in Resistive Memories through Selective Slow Write Backs
abstract
Emerging resistive memory technologies, such as PCRAM and ReRAM, have been proposed as promising replacements for DRAM-based main memory, due to their better scalability, low standby power, and non-volatility. However, limited write endurance is a major drawback for such resistive memory technologies. Wear leveling (balancing the distribution of writes) and wear limiting (reducing the number of writes) have been proposed to mitigate this disadvantage, but both techniques only manage a fixed budget of writes to a memory system rather than increase the number available. In this paper, we propose a new type of wear limiting technique, Mellow Writes, which reduces the wearout of individual writes rather than reducing the number of writes. Mellow Writes is based on the fact that slow writes performed with lower dissipated power can lead to longer endurance (and therefore longer lifetimes). For non-volatile memories, an N1to N3times endurance can be achieved if the write operation is slowed down by N times. We present three microarchitectural mechanisms (BankAware Mellow Writes, Eager Mellow Writes, and Wear Quota) that selectively perform slow writes to increase memory lifetime while minimizing performance impact. Assuming a factor N2advantage in cell endurance for a factor N slower write, our best Mellow Writes mechanism can achieve 2.58× lifetime and 1.06× performance of the baseline system. In addition, its performance is almost the same as a system aggressively optimized for performance (at the expense of endurance). Finally, Wear Quota guarantees a minimal lifetime (e.g., 8 years) by forcing more slow writes in presence of heavy workloads. We also perform sensitivity analysis on the endurance advantage factor for slow writes, from N1to N3, and find that our technique is still useful for factors as low as N1.
Lunkai Zhang, Brian Neely, Diana Franklin, Dmitri B. Strukov, Yuan Xie 0001, Fred Chong
ISCA4
2016 Spiking neuromorphic networks with metal-oxide memristors
abstract
This is a brief review of our recent work on memristor-based spiking neuromorphic networks. We first describe the recent experimental demonstration of several most biology-plausible spike-time-dependent plasticity (STDP) windows in integrated metal-oxide memristors and, for the first time, the observed self-adaptive STDP, which may be crucial for spiking neural network applications. We then discuss recent theoretical work in which an analytical, data-verified STDP model was used to simulate operation of a spiking classifier of spatial-temporal patterns, and the capacity-to-fidelity tradeoff and noise immunity o f spiking spatial-temporal associative memories with local and global recording was evaluated.
Mirko Prezioso, Y. Zhong, D. Gavrilov, Farnood Merrikh-Bayat, Brian Hoskins, Gina C. Adam, Konstantin Likharev, Dmitri B. Strukov
ISCAS8
2015 Efficient training algorithms for neural networks based on memristive crossbar circuits
abstract
We have adapted backpropagation algorithm for training multilayer perceptron classifier implemented with memristive crossbar circuits. The proposed training approach takes into account switching dynamics of a particular, though very typical, type of memristive devices and weight update restrictions imposed by crossbar topology. The simulation results show that for crossbar-based multilayer perceptron with one hidden layer of 300 neurons misclassification rate on MNIST benchmark could be as low as 1.47% and 4.06% for batch and stochastic algorithms, respectively, which is comparable to the best reported results for similar neural networks.
Irina Kataeva, Farnood Merrikh-Bayat, Elham Zamanidoost, Dmitri B. Strukov
IJCNN4
2015 Redesigning commercial floating-gate memory for analog computing applications
abstract
We have modified a commercial NOR flash memory array to enable high-precision tuning of individual floating-gate cells for analog computing applications. The modified array area per cell in a 180 nm process is about 1.5 μm2. While this area is approximately twice the original cell size, it is still at least an order of magnitude smaller than in state-of-the-art analog circuit implementations. The new memory cell arrays have been successfully tested, in particular confirming that each cell may be automatically tuned, with ~1% precision, to any desired subthreshold readout current value within an almost three-orders-of-magnitude dynamic range, even using an unoptimized tuning algorithm. Preliminary results for a four-quadrant vector-by-matrix multiplier, implemented with the modified memory array, gate-coupled with additional peripheral floating-gate transistors, show highly linear transfer characteristics over a broad range of input currents.
Farnood Merrikh-Bayat, Xinjie Guo, H. A. Ommani, N. Do, Konstantin Likharev, Dmitri B. Strukov
ISCAS6
2015 A configurable CMOS memory platform for 3D-integrated memristors
abstract
Memristors are emerging as powerful nanoscale devices for diverse applications, such as high-density memories and neuromorphic applications. However, this nascent technology requires considerable advancement before this vision is realized. We present a highly configurable CMOS interface chip which enables the characterization of on-chip memristors, especially for memory applications. The chip was fabricated in On-Semi 3M2P 0.5 μm occupying 2×2 mm2. The chip design allows for post-CMOS fabrication of memristors. The interface between the memristor and the CMOS circuitry was provided via a top metal contact. The chip was designed to support an area-distributed interface decoupling CMOS pitch and memristor pitch, enabling high-density memristor integration. Measurement results on post-CMOS fabricated Ag/SiO2/Pt memristive devices are reported. Though we have shown the results from one memristive material stack, thorough chip characterization demonstrates the versatility of the chip enabling its use with a wide variety of materials stacks.
Melika Payvand, Advait Madhavan, Miguel Angel Lastras-Montaño, Amirali Ghofrani, Justin Rofeh, Kwang-Ting Cheng, Dmitri B. Strukov, Luke Theogarajan
ISCAS7
2014 SpongeDirectory: flexible sparse directories utilizing multi-level memristors
abstract
Cache-coherent shared memory is critical for programmability in many-core systems. Several directory-based schemes have been proposed, but dynamic, non-uniform sharing make efficient directory storage challenging, with each giving up storage space, performance or energy.
Lunkai Zhang, Dmitri B. Strukov, Heba Saadeldeen, Dongrui Fan, Mingzhe Zhang 0005, Diana Franklin
PACT2
2014 Race Logic: A hardware acceleration for dynamic programming algorithms
abstract
We propose a novel computing approach, dubbed “Race Logic”, in which information, instead of being represented as logic levels, as is done in conventional logic, is represented as a timing delay. Under this new information representation, computations can be performed by observing the relative propagation times of signals injected into the circuit (i.e. the outcome of races). Race Logic is especially suited for solving problems related to the traversal of directed acyclic graphs commonly used in dynamic programming algorithms. The main advantage of this novel approach is that information processing (min-max and addition operations) can be very efficiently expressed through the manipulation of the natural delay chaining inherent to digital designs, which then results in superior latency, throughput, and energy efficiency. To verify this hypothesis, we designed several Race Logic implementations of a DNA global sequence alignment engine and compared it to the state-of-the-art conventional systolic array implementation. Our synthesized design shows that synchronous Race Logic is up to 4× faster when both approaches are mapped to a 0.5μm CMOS standard cell technology. At the same time the throughput for sequence matching per circuit area is about 3× higher at 5× lower power density for 20-long-symbol DNA sequences.
Advait Madhavan, Timothy Sherwood, Dmitri B. Strukov
ISCA3
2012 3D CMOS-memristor hybrid circuits: devices, integration, architecture, and applications
abstract
In this paper, we give an overview of our recent research efforts on monolithic 3D integration of CMOS and memristive nanodevices. These hybrid circuits combine a CMOS subsystem with several layers of nanowire crossbars, consisting of arrays of two-terminal memristors, all connected by an area-distributed interface between the CMOS subsystem and the crossbars. This approach combines the advantages of CMOS technology, including its high flexibility, functionality and yield, with the extremely high density of nanowires, nanodevices and interface vias. As a result, the 3D hybrids can overcome limitations pertinent to other 3D integration techniques (such as through-silicon vias) and enable 3D circuits with unprecedented memory density (up to 1014 bits on a single 1-cm2 chip) and aggregate interlayer communication bandwidth (up to 1018 bits per second per cm2) at manageable power dissipation. Such performance represents a significant step towards addressing the most pressing needs of modern compact electronic systems.
Kwang-Ting Cheng, Dmitri B. Strukov
ISPD2
2012 Analog-input analog-weight dot-product operation with Ag/a-Si/Pt memristive devices
Ligang Gao, Fabien Alibart, Dmitri B. Strukov
VLSI-SoC3
2012 Mapping of image and network processing tasks on high-throughput CMOL FPGA circuits
Advait Madhavan, Dmitri B. Strukov
VLSI-SoC2
2010 Monolithically stackable hybrid FPGA
abstract
The paper introduces novel field programmable gate array (FPGA) circuits based on hybrid CMOS/resistive switching device (memristor) technology and explores several logic architectures. The novel FPGA structure is based on the combination of CMOL (Cmos + MOLecular scale devices) FPGA circuits and recent improvements and generalization of the CMOL concept to allow multilayer crossbar integration, compatibility with state-of-the-art foundries, and a wide range of available memristive crosspoint devices. Preliminary results indicate that with no optimization and only conventional CMOS technology, the proposed circuits can be at least ten times denser (and potentially faster) than CMOS FPGAs with the same design rules and similar power density. The second part of this paper shows that this performance can be further improved using optimal MUX-based logic architecture.
Dmitri B. Strukov, Alan Mishchenko
DATE1
2010 Hybrid CMOS/memristor circuits
abstract
This is a brief review of recent work on the prospective hybrid CMOS/memristor circuits. Such hybrids combine the flexibility, reliability and high functionality of the CMOS subsystem with very high density of nanoscale thin film resistance switching devices operating on different physical principles. Simulation and initial experimental results demonstrate that performance of CMOS/memristor circuits for several important applications is well beyond scaling limits of conventional VLSI paradigm.
Dmitri B. Strukov, Duncan R. Stewart, Julien Borghetti, Xuema Li, Matthew D. Pickett, Gilberto Medeiros-Ribeiro, Warren Robinett, Gregory S. Snider, John Paul Strachan, Qiangfei Xia, J. Joshua Yang, R. Stanley Williams
ISCAS1
2006 A reconfigurable architecture for hybrid CMOS/Nanodevice circuits
abstract
This report describes a preliminary evaluation of performance of a cell-FPGA-like architecture for future hybrid "CMOL" circuits. Such circuits will combine a semiconduc-tor-transistor (CMOS) stack and a two-level nanowire crossbar with molecular-scale two-terminal nanodevices (program-mable diodes) formed at each crosspoint. Our cell-based architecture is based on a uniform CMOL fabric of "tiles". Each tile consists of 12 four-transistor basic cells and one (four times larger) latch cell. Due to high density of nanodevices, which may be used for both logic and routing functions, CMOL FPGA may be reconfigured around defective nanodevices to provide high defect tolerance. Using a semi-custom set of design automation tools we have evaluated CMOL FPGA performance for the Toronto 20 benchmark set, so far without optimization of several parameters including the power supply voltage and nanowire pitch. The results show that even without such optimization, CMOL FPGA circuits may provide a density advantage of more than two orders of magnitude over the traditional CMOS FPGA with the same CMOS design rules, at comparable time delay, acceptable power consumption and potentially high defect tolerance.
Dmitri B. Strukov, Konstantin Likharev
FPGA1