VLDB 2026 Research / reviewers in the wild / expert
John Paul Strachan
dblp:34/9429
· DBLP profile ↗
20ranked-venue papers
0as first author
8since 2021 · last 2025
0000-0002-1382-3677ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 20 · 8 since 2021Software engineering, systems software and programming languages · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | KLIMA: Low-latency mixed-signal In-Memory Computing accelerator for solving arbitrary-order Boolean Satisfiability
Tinish Bhattacharya, Dongseok Kwon, George Higgins Hutchinson, Xiangyi Zhang, Giacomo Pedretti, Fabian Böhm, John Paul Strachan, Thomas Van Vaerenbergh, Raymond G. Beausoleil, Ignacio Rozada, Dmitri B. Strukov |
HCS | 7 |
| 2025 | Analog Softmax with Wide Input Current Range for In-Memory ComputingabstractThe Softmax activation function plays a pivotal role in both the attention mechanism of Transformers and in the final layer of neural networks performing classification. The Softmax function outputs probabilities by normalizing the input values, emphasizing differences among them to highlight the largest values. In digital implementations, the complexity of softmax grows linearly with the number of inputs. In contrast, analog implementations enable parallel computations with lower latency. In this work, we demonstrate that this approach achieves a more efficient linear scaling of latency as vector size increases logarithmically. This analog softmax circuits are implemented in TSMC 28 nm PDK technology, capable of driving up to 128 inputs and producing an analog current output spanning three orders of magnitude. The study examines the circuit’s power consumption, latency, and error, emphasizing its efficiency compared to the alternative approach of converting outputs to digital signals via ADCs and performing the softmax calculation digitally. By reducing reliance on these power-intensive operations, this work aims to significantly enhance energy efficiency in in-memory computing systems. Aradhana Dube, Paul Manea, Paolo Gibertini, Erika Covi, John Paul Strachan |
ISCAS | 5 |
| 2025 | Impact of Variability Compensation on the Performance of an RRAM-based 3-SAT SolverabstractThe impact of variability on the performance of an RRAM-Computing-in-Memory-based architecture solving 3-SAT unconstrained optimization problems is evaluated. Realistic CMOS variability parameters were extracted from a 28nm CMOS technology, while RRAM variability parameters were taken from measured devices reported in the literature. It is shown by extensive Monte Carlo simulations that a naïve implementation shows two different error characteristics with respect to (I) produce faulty solutions and (II) tendency for premature termination of iteration loops. A new variability compensation scheme is proposed, elaborated and evaluated for 3-SAT problems up to 200 variables and 860 clauses. By application of the proposed compensation scheme errors of type (I) and (II) are completely eliminated. Arne Heittmann, Mohammad Hizzani, John Paul Strachan |
ISCAS | 3 |
| 2025 | Gain Cell-based Analog Content Addressable Memory for Dynamic Associative Tasks in AIabstractanalog Content Addressable Memories (aCAMs) have proven useful for associative Compute-in-Memory (CIM) applications like Decision Trees, Finite State Machines, and Hyper-dimensional Computing. While non-volatile implementations using FeFETs and ReRAM devices offer speed, power, and area advantages, they suffer from slow write speeds and limited write cycles, making them less suitable for computations involving fully dynamic data patterns. To address these limitations, in this work, we propose a capacitor gain cell-based aCAM designed for dynamic processing, where frequent memory updates are required. Our system compares analog input voltages to boundaries stored in capacitors, enabling efficient dynamic tasks. We demonstrate the application of aCAM within transformer attention mechanisms by replacing the softmax-scaled dot-product similarity with aCAM similarity, achieving competitive results. Circuit simulations on a TSMC 28 nm node show promising performance in terms of energy efficiency, precision, and latency, making it well-suited for fast, dynamic AI applications. Paul-Philipp Manea, Nathan Leroux, Emre Neftci, John Paul Strachan |
ISCAS | 4 |
| 2025 | IMSSA: Deploying modern state-space models on memristive in-memory compute hardwareabstractProcessing long temporal sequences is a key challenge in deep learning. In recent years, Transformers have become state-of-the-art for this task, but suffer from excessive memory requirements due to the need to explicitly store the sequences. To address this issue, structured state-space sequential (S4) models recently emerged, offering a fixed memory state while still enabling the processing of very long sequence contexts. The recurrent linear update of the state in these models makes them highly efficient on modern graphics processing units (GPU) by unrolling the recurrence into a convolution. However, this approach demands significant memory and massively parallel computation, which is only available on the latest GPUs.In this work, we aim to bring the power of S4 models to edge hardware by significantly reducing the size and computational demand of an S4D model through quantization-aware training, even achieving ternary weights for a simple real-world task. To this end, we extend conventional quantization-aware training to tailor it for analog in-memory compute hardware. We then demonstrate the deployment of recurrent S4D kernels on memrisitve crossbar arrays, enabling their computation in an in-memory compute fashion. To our knowledge, this is the first implementation of S4 kernels on in-memory compute hardware. Sebastian Siegel, Ming-Jay Yang, John Paul Strachan |
ISCAS | 3 |
| 2024 | Memristor-based hardware and algorithms for higher-order Hopfield optimization solver outperforming quadratic Ising machinesabstractIsing solvers offer a promising physics-based approach to tackle the challenging class of combinatorial optimization problems. However, typical solvers operate in a quadratic energy space, having only pair-wise coupling elements which already dominate area and energy. We show that such quadratization can cause severe problems: increased dimensionality, a rugged search landscape, and misalignment with the original objective function. Here, we design and quantify a higher-order Hopfield optimization solver, with 28nm CMOS technology and memristive couplings for lower area and energy computations. We combine algorithmic and circuit analysis to show quantitative advantages over quadratic Ising Machines (IM)s, yielding 48x and 72x reduction in time-to-solution (TTS) and energy-to-solution (ETS) respectively for Boolean satisfiability problems of 150 variables, with favorable scaling. Mohammad Hizzani, Arne Heittmann, George Higgins Hutchinson, Dmitrii Dobrynin, Thomas Van Vaerenbergh, Tinish Bhattacharya, Adrien Renaudineau, Dmitri B. Strukov, John Paul Strachan |
ISCAS | 9 |
| 2022 | A general tree-based machine learning accelerator with memristive analog CAMabstractDeep learning models have reached high accuracy in multiple classification tasks. However these models lack explainability, namely the capability of understanding why a certain class is chosen along with the class predicted. On the other hand, tree-based models are top performers in several applications, particularly when the training set is limited, while also being more explainable. However, tree-based models are difficult to accelerate with conventional digital hardware due to irregular memory access patterns. Here we show a tree-based ML accelerator based on a novel analog content addressable memory with memristor devices, capable of handling multiple types of bagging and boosting techniques common in tree-based algorithms. Our results show a large improvement of $\sim 60 \times $ lower latency and $160 \times $ reduced energy consumption compared to the state of the art, demonstrating the promise of our accelerator approach. Giacomo Pedretti, Sergey Serebryakov, John Paul Strachan, Catherine Graves |
ISCAS | 3 |
| 2021 | Mixed Precision Quantization for ReRAM-based DNN Inference AcceleratorsabstractReRAM-based accelerators have shown great potential for accelerating DNN inference because ReRAM crossbars can perform analog matrix-vector multiplication operations with low latency and energy consumption. However, these crossbars require the use of ADCs which constitute a significant fraction of the cost of MVM operations. The overhead of ADCs can be mitigated via partial sum quantization. However, prior quantization flows for DNN inference accelerators do not consider partial sum quantization which is not highly relevant to traditional digital architectures. To address this issue, we propose a mixed precision quantization scheme for ReRAM-based DNN inference accelerators where weight quantization, input quantization, and partial sum quantization are jointly applied for each DNN layer. We also propose an automated quantization flow powered by deep reinforcement learning to search for the best quantization configuration in the large design space. Our evaluation shows that the proposed mixed precision quantization scheme and quantization flow reduce inference latency and energy consumption by up to 3.89x and 4.84x, respectively, while only losing 1.18% in DNN inference accuracy. Sitao Huang, Aayush Ankit, Plínio Silveira, Rodrigo Antunes, Sai Rahul Chalamalasetti, Izzat El Hajj, Dong Eun Kim, Glaucimar Aguiar, Pedro Bruel, Sergey Serebryakov, Can Li 0024, Paolo Faraboschi, John Paul Strachan, Deming Chen, Kaushik Roy 0001, Wen-Mei W. Hwu, Dejan S. Milojicic |
ASP-DAC | 14 |
| 2020 | PANTHER: A Programmable Architecture for Neural Network Training Harnessing Energy-Efficient ReRAMabstractThe wide adoption of deep neural networks has been accompanied by ever-increasing energy and performance demands due to the expensive nature of training them. Numerous special-purpose architectures have been proposed to accelerate training: both digital and hybrid digital-analog using resistive RAM (ReRAM) crossbars. ReRAM-based accelerators have demonstrated the effectiveness of ReRAM crossbars at performing matrix-vector multiplication operations that are prevalent in training. However, they still suffer from inefficiency due to the use of serial reads and writes for performing the weight gradient and update step. A few works have demonstrated the possibility of performing outer products in crossbars, which can be used to realize the weight gradient and update step without the use of serial reads and writes. However, these works have been limited to low precision operations which are not sufficient for typical training workloads. Moreover, they have been confined to a limited set of training algorithms for fully-connected layers only. To address these limitations, we propose a bit-slicing technique for enhancing the precision of ReRAM-based outer products, which is substantially different from bit-slicing for matrix-vector multiplication only. We incorporate this technique into a crossbar architecture with three variants catered to different training algorithms. To evaluate our design on different types of layers in neural networks (fully-connected, convolutional, etc.) and training algorithms, we develop PANTHER, an ISA-programmable training accelerator with compiler support. Our design can also be integrated into other accelerators in the literature to enhance their efficiency. Our evaluation shows that PANTHER achieves up to 8.02×, 54.21×, and 103× energy reductions as well as 7.16×, 4.02×, and 16× execution time reductions compared to digital accelerators, ReRAM-based accelerators, and GPUs, respectively. Aayush Ankit, Izzat El Hajj, Sai Rahul Chalamalasetti, Sapan Agarwal, Matthew J. Marinella, Martin Foltin, John Paul Strachan, Dejan S. Milojicic, Wen-Mei W. Hwu, Kaushik Roy 0001 |
IEEE Trans. Computers | 7 |
| 2019 | PUMA: A Programmable Ultra-efficient Memristor-based Accelerator for Machine Learning InferenceabstractMemristor crossbars are circuits capable of performing analog matrix-vector multiplications, overcoming the fundamental energy efficiency limitations of digital logic. They have been shown to be effective in special-purpose accelerators for a limited set of neural network applications. We present the Programmable Ultra-efficient Memristor-based Accelerator (PUMA) which enhances memristor crossbars with general purpose execution units to enable the acceleration of a wide variety of Machine Learning (ML) inference workloads. PUMA's microarchitecture techniques exposed through a specialized Instruction Set Architecture (ISA) retain the efficiency of in-memory computing and analog circuitry, without compromising programmability. We also present the PUMA compiler which translates high-level code to PUMA ISA. The compiler partitions the computational graph and optimizes instruction scheduling and register allocation to generate code for large and complex workloads to run on thousands of spatial cores. We have developed a detailed architecture simulator that incorporates the functionality, timing, and power models of PUMA's components to evaluate performance and energy consumption. A PUMA accelerator running at 1 GHz can reach area and power efficiency of 577 GOPS/s/mm 2 and 837~GOPS/s/W, respectively. Our evaluation of diverse ML applications from image recognition, machine translation, and language modelling (5M-800M synapses) shows that PUMA achieves up to 2,446× energy and 66× latency improvement for inference compared to state-of-the-art GPUs. Compared to an application-specific memristor-based accelerator, PUMA incurs small energy overheads at similar inference latency and added programmability. Aayush Ankit, Izzat El Hajj, Sai Rahul Chalamalasetti, Geoffrey Ndu, Martin Foltin, R. Stanley Williams, Paolo Faraboschi, Wen-Mei W. Hwu, John Paul Strachan, Kaushik Roy 0001, Dejan S. Milojicic |
ASPLOS | 9 |
| 2018 | Computing In-Memory, RevisitedabstractThe Von Neumann's architecture has been the dominant computing paradigm ever since its inception in the mid-forties. It revolves around the concept of a "stored program" in memory, and a central processing unit that executes the program. As an alternative, Processing-In-Memory (PIM) ideas have been around for at least two decades, however with very limited adoption. Today, three trends are creating a compelling motivation to take a second look. Novel devices such as memristor blur the boundary between memory and compute, effectively providing both in the same element. Power efficiency has become very important, both in the datacenter and at the edge. Machine learning applications driven by a data-flow model have become ubiquitous. In this paper, we sketch our Computing-In-Memory (CIM) vision, and its substantial performance and power improvement potential. Compared to PIM models, CIM more clearly separates computing from memory. We then discuss the programming model, which we consider the biggest challenge. We close by describing how CIM impacts non-functional characteristics, such as reliability, scale, and configurability. Dejan S. Milojicic, Kirk Bresniker, Gary Campbell, Paolo Faraboschi, John Paul Strachan, Stan Williams |
ICDCS | 5 |
| 2018 | Large Memristor Crossbars for Analog ComputingabstractMemristor with tunable non-volatile resistance offers in-memory computing capability that avoids the von-Neumann bottleneck. However, large-scale experimental demonstration to this end is yet to be implemented due to the immaturity of the device and integration technologies. Here in this paper we report our recent process in analog computing using analog-voltage-amplitude-vector input and analog-memristor-conductance matrix, with applications in signal and image processing. The vector matrix multiplication is processed in the memristor crossbars in one step, with 5-8 bit precision depending on the array size. The demonstration is made possible by high memristor yield (99.8%), stable multilevel memresistance states, linear current-voltage (IV) relation in the operation range, and low wire resistance between the cells. Can Li 0024, Yunning Li, Hao Jiang 0017, J. Joshua Yang, Qiangfei Xia, Miao Hu 0002, Eric Montgomery, Noraica Dávila, Catherine Graves, John Paul Strachan, R. Stanley Williams, Ning Ge 0001, Mark Barnell, Qing Wu 0002 |
ISCAS | 15 |
| 2017 | Rescuing Memristor-based Neuromorphic Design with High DefectsabstractMemristor-based synaptic network has been widely investigated and applied to neuromorphic computing systems for the fast computation and low design cost. As memristors continue to mature and achieve higher density, bit failures within crossbar arrays can become a critical issue. These can degrade the computation accuracy significantly. In this work, we propose a defect rescuing design to restore the computation accuracy. In our proposed design, significant weights in a specified network are first identified and retraining and remapping algorithms are described. For a two layer neural network with 92.64% classification accuracy on MNIST digit recognition, our evaluation based on real device testing shows that our design can recover almost its full performance when 20% random defects are present. Miao Hu 0002, John Paul Strachan, Hai Li 0001 |
DAC | 3 |
| 2016 | Dot-product engine for neuromorphic computing: programming 1T1M crossbar to accelerate matrix-vector multiplicationabstractVector-matrix multiplication dominates the computation time and energy for many workloads, particularly neural network algorithms and linear transforms (e.g, the Discrete Fourier Transform). Utilizing the natural current accumulation feature of memristor crossbar, we developed the Dot-Product Engine (DPE) as a high density, high power efficiency accelerator for approximate matrix-vector multiplication. We firstly invented a conversion algorithm to map arbitrary matrix values appropriately to memristor conductances in a realistic crossbar array, accounting for device physics and circuit issues to reduce computational errors. The accurate device resistance programming in large arrays is enabled by close-loop pulse tuning and access transistors. To validate our approach, we simulated and benchmarked one of the state-of-the-art neural networks for pattern recognition on the DPEs. The result shows no accuracy degradation compared to software approach (99 % pattern recognition accuracy for MNIST data set) with only 4 Bit DAC/ADC requirement, while the DPE can achieve a speed-efficiency product of 1,000× to 10,000× compared to a custom digital ASIC. Miao Hu 0002, John Paul Strachan, Emmanuelle M. Grafals, Noraica Dávila, Catherine Graves, Sity Lam, Ning Ge 0001, J. Joshua Yang, R. Stanley Williams |
DAC | 2 |
| 2016 | Fading memory effects in a memristor for Cellular Nanoscale Network applications
Alon Ascoli, Ronald Tetzlaff, Leon O. Chua, John Paul Strachan, R. Stanley Williams |
DATE | 4 |
| 2016 | ISAAC: A Convolutional Neural Network Accelerator with In-Situ Analog Arithmetic in CrossbarsabstractA number of recent efforts have attempted to design accelerators for popular machine learning algorithms, such as those involving convolutional and deep neural networks (CNNs and DNNs). These algorithms typically involve a large number of multiply-accumulate (dot-product) operations. A recent project, DaDianNao, adopts a near data processing approach, where a specialized neural functional unit performs all the digital arithmetic operations and receives input weights from adjacent eDRAM banks. This work explores an in-situ processing approach, where memristor crossbar arrays not only store input weights, but are also used to perform dot-product operations in an analog manner. While the use of crossbar memory as an analog dot-product engine is well known, no prior work has designed or characterized a full-fledged accelerator based on crossbars. In particular, our work makes the following contributions: (i) We design a pipelined architecture, with some crossbars dedicated for each neural network layer, and eDRAM buffers that aggregate data between pipeline stages. (ii) We define new data encoding techniques that are amenable to analog computations and that can reduce the high overheads of analog-to-digital conversion (ADC). (iii) We define the many supporting digital components required in an analog CNN accelerator and carry out a design space exploration to identify the best balance of memristor storage/compute, ADCs, and eDRAM storage on a chip. On a suite of CNN and DNN workloads, the proposed ISAAC architecture yields improvements of 14.8×, 5.5×, and 7.5× in throughput, energy, and computational density (respectively), relative to the state-of-the-art DaDianNao architecture. Ali Shafiee, Anirban Nag, Naveen Muralimanohar, Rajeev Balasubramonian, John Paul Strachan, Miao Hu 0002, R. Stanley Williams, Vivek Srikumar |
ISCA | 5 |
| 2013 | Physics-based memristor modelsabstractIn order to utilize memristors in circuits, one needs high-quality predictive models that can be used for simulations to act as a design aid. Whenever possible, we base our models on the known physics of the memristors we are using. To that end, we perform a wide range of materials characterizations and electronic measurements on which to base the model. However, given the complexity of the physical processes that occur in the devices, such as drift-diffusion-thermophoresis in ion-migration based memristors and Mott transitions in locally active memristors, the corresponding detailed mathematical descriptions are far too complex to solve analytically and numerical solutions are too time consuming to include in a simulation. We thus need to find simpler, analytical approximations that can match the measured behavior of the memristors over many orders of magnitude in time and a wide range of applied voltage. We present some new models that we have developed and describe how they are derived. R. Stanley Williams, Matthew D. Pickett, John Paul Strachan |
ISCAS | 3 |
| 2012 | Designing memristors: Physics, materials science and engineeringabstractRecently, memory and storage have taken a front seat in computer hardware as it experiences an explosive growth at a rate faster than Moore's law for the past 10 years. With the upcoming challenges for further FLASH scaling into the next generations, emerging technologies have appeared portraying perspectives with the potential to shift computer architecture concepts. Here we present a brief overview and progress in our quest to create a device with competitive attributes, with the ultimate goal of achieving a universal, non-volatile data storage solution. Gilberto Medeiros-Ribeiro, J. Joshua Yang, Janice H. Nickel, Antonio Torrezan, John Paul Strachan, R. Stanley Williams |
ISCAS | 5 |
| 2011 | CMOS interface circuits for reading and writing memristor crossbar arrayabstractThis paper describes CMOS interface circuits in 350nm 3.3V/5.0V TSMC process for memristor crossbar array. These circuits are applicable for non-volatile resistive memories. The architecture is targeted for low power and high speed applications. We have demonstrated sense amplifiers for reading the state of a memristor bit. Voltage divider and transimpedence amplifier is used for DC sensing while sigma delta approach is used for averaging. Current limiting write amplifier has also been designed for increasing the device endurance and reliability. Half select array architecture is used to minimize dc leakage current in the crossbar array. Muhammad Shakeel Qureshi, Matthew D. Pickett, Feng Miao, John Paul Strachan |
ISCAS | 4 |
| 2010 | Hybrid CMOS/memristor circuitsabstractThis is a brief review of recent work on the prospective hybrid CMOS/memristor circuits. Such hybrids combine the flexibility, reliability and high functionality of the CMOS subsystem with very high density of nanoscale thin film resistance switching devices operating on different physical principles. Simulation and initial experimental results demonstrate that performance of CMOS/memristor circuits for several important applications is well beyond scaling limits of conventional VLSI paradigm. Dmitri B. Strukov, Duncan R. Stewart, Julien Borghetti, Xuema Li, Matthew D. Pickett, Gilberto Medeiros-Ribeiro, Warren Robinett, Gregory S. Snider, John Paul Strachan, Qiangfei Xia, J. Joshua Yang, R. Stanley Williams |
ISCAS | 9 |