VLDB 2026 Research / reviewers in the wild / expert
R. Stanley Williams
dblp:67/802 · also Richard Stanley Williams
· DBLP profile ↗
25ranked-venue papers
3as first author
10since 2021 · last 2023
0000-0003-0213-4259ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 23 · 3 first-author · 10 since 2021Software engineering, systems software and programming languages · 4 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Memristor-based Offset Cancellation Technique in Analog CrossbarsabstractAnalog computing platforms have been a popular and promising research area that suggest efficient ways of computation compared to its digital counterparts. Memristor based crossbars drew attention by computing the vector-matrix calculation intensive tasks such as Artificial Intelligence (AI) and Machine Learning (ML) in one time step. Although they provide an energy efficient way of computing these tasks, analog computation in general suffers from non-idealities and systematic errors in the circuitry, which could degrade the performance and accuracy significantly. One of the issues is the random offset associated with the op-amps in the system resulting from the process and mismatch variations. In this paper, a novel technique is offered to reduce the negative effects of the random offset and increase the output accuracy. This newly proposed system uses minimum extra circuitry and additional power consumption and only requires the crossbar to be enlarged by two extra rows. The intrinsic issue of the analog crossbars, interconnect parasitics, must be incorporated into the problem, and a way to separate the offset and wire resistance issues from each other is offered. The functionality of the system has been shown with a case study in the results section where the op-amps have$\sigma_{offset}=3mV$. The effectiveness of the offered technique demonstrates a 6× better accuracy with the mitigation of the offset problem. The proposed method can be used in memristor and other analog crossbars to achieve a greater performance and thus improve their competitiveness. Anil Korkmaz, Gianluca Zoppo, Francesco Marrone, Fernando Corinto, Su-In Yi, R. Stanley Williams, Samuel Palermo |
ISCAS | 6 |
| 2023 | Gaussian Process for Nonlinear Regression via Memristive CrossbarsabstractOver the last decade, Gaussian processes (GPs) have become popular in the area of machine learning and data analysis for their flexibility and robustness. Despite their attractive formulation, practical use in large-scale problems remains out of reach due to computational complexity. Existing direct computational methods for manipulations involving large-scale$n\times n$covariance matrices require$O(n^{3})$calculations. In this work, we present the design and evaluation of a simulated computing platform for exact GP inference, that achieves true model parallelism using memristive crossbars. To achieve a one-shot solution, a linear equation solver and a vector-matrix multiplication solver crossbar configurations are used together, reducing the number of operations from$O(n^{3})$to$O(n)$. The transistor level op-amps, ADC models for quantization, circuit and interconnect parasitics, together with the finite memristor precision are incorporated into the system simulation. The analog system resulted in %1.51 mean error and %2.93 average variance error in solving a nonlinear regression problem. The proposed method achieved 9× to 144× better energy efficiency compared to TPU and 7× compared to a custom analog linear regression solver. Gianluca Zoppo, Anil Korkmaz, Francesco Marrone, Su-In Yi, Samuel Palermo, Fernando Corinto, R. Stanley Williams |
ISCAS | 7 |
| 2022 | Analog Acceleration of the Power Method using Memristor CrossbarsabstractDetermining the dominant eigenvector of matrices and graphs is one of the most fundamental tasks in many machine learning problems, including spectral clustering, Hyperlink Induced Topic Search (HITS), Markov Chains, PageRank and eigenvector centralities. Among the several algorithms used, the Power Method is one of the simplest iterative approaches. It relies on multiple vector-matrix multiplications (VMMs) and a normalization step to prevent divergent behaviours. Recently, efficiency of the memristor crossbars in solving VMMs have been demonstrated using fundamental laws of the circuit theory. In this work, we propose a circuit to accelerate the Power iteration algorithm including current-mode termination for the memristor crossbars and a normalization circuit. The normalization step together with the feedback loop of the complete circuit ensure stability and convergence of the dominant eigenvector. The system allows the observation of the evolution of the outputs. We implement a transistor level peripheral circuitry around the memristor crossbar and take non-idealities such as wire parasitics, source driver resistance and finite memristor precision into account. We compute the eigenvector centrality to demonstrate the performance of the proposed system. We compare our results to the ones coming from the conventional digital computers and observe significant energy savings while maintaining a competitive accuracy. Anil Korkmaz, Gianluca Zoppo, Francesco Marrone, Fernando Corinto, R. Stanley Williams, Samuel Palermo |
ISCAS | 5 |
| 2022 | Entrenching Decision Trees in a Robust Molecular Circuit ElementabstractBy mapping logic complexities onto molecular redox transitions, we embed decision trees in a nanoscale memristive element. Our circuit element is robust, and the switching events are deterministic. These molecular elements may offer a substantial advancement in stateful in-memory computing technology. T. Venkatesan, Sreebrata Goswami, R. Stanley Williams, Sreetosh Goswami |
ISCAS | 3 |
| 2022 | Molecular building blocks for non-linear circuitsabstractWe have discovered molecular circuit elements that resolve decade long challenges in organic electronics. Our devices are robust, reliable, uniform and we have developed a deterministic understanding of their molecular redox mechanisms. The films are compatible with existing circuit fabrication protocols. In addition, via molecular engineering routes, we are able to achieve unprecedented characteristic tunability in voltage, current, number of states and also in the mode of operation. These elements can function both as a nonvolatile memory for in-memory computing and also, on the edge of chaos depending on how they are operated. T. Venkatesan, R. Stanley Williams, Sreebrata Goswami, Sreetosh Goswami |
ISCAS | 2 |
| 2022 | Combinatorial Optimization in Hopfield Networks with Noise and Diagonal PerturbationsabstractWe demonstrate via simulations that transient perturbations introduced by non-zero diagonal elements in a Hopfield network can improve NP-hard graph optimization efficiency by more than a factor of two. Such perturbations enhance the known effects of circuit noise typical of memristor-based networks in escaping local minima (incorrect solutions) and finding the global minimum (correct solution) of the Hopfield energy. We provide systematic simulations of NP-hard graph problems with controlled nonidealities in memristor arrays modeled as noise amplitude and diagonal perturbations. Furthermore, our approach improves Hopfield network optimization efficiencies to solve NP-hard problems regardless of the graph size (30 × 30, 60 × 60, and 80 × 80) and connectivity (30%, 50%, and 70%). Su-In Yi, Suhas Kumar, R. Stanley Williams |
ISCAS | 3 |
| 2021 | Design of Tunable Analog Filters Using Memristive CrossbarsabstractTunable front-end filters are necessary in wireless systems that support multiple frequency bands. N-path filters have the potential for wide tuning ranges, but require multiple high-frequency mixing clocks and suffer from harmonic responses. Another option is digital filtering. However, this requires high-speed analog-to-digital converters and the time complexity to perform Discrete Fourier Transform (DFT) and Inverse Discrete Fourier Transform (IDFT) operations using conventional digital computers is O(N2). Conversely, memristor crossbars have experimentally demonstrated the ability to perform various signal processing tasks, including Discrete Cosine Transform (DCT), in one time-step. In this work, this has been taken a step further by presenting full DFT and IDFT operations using memristive crossbars and proposes the implementation of novel continuous-time tunable analog filters by applying filter coefficients in between these DFT and IDFT crossbars. This highly scalable new method of filtering leverages advantages found both in conventional digital and analog filters and allows any digital filter to be implemented in the analog domain with a filter order as large as half the utilized crossbar size. The proposed architecture allows for the generation of arbitrary filter functions with tunable corner (or center) frequencies and bandwidths from 0.2-20GHz. Stopband attenuation greater than 40dB is achieved with as low as 2-bit memristor precision, while close to 80dB is possible with 8-bit precision. The filter system consumes 106mW to support the 20GHz frequency range, resulting in a 5.3mW/GHz energy metric. Anil Korkmaz, Chaoyi He, Linda Katehi, R. Stanley Williams, Samuel Palermo |
ISCAS | 4 |
| 2021 | NbO2-Mott Memristor: A Circuit- Theoretic InvestigationabstractThis paper presents a circuit-theoretic analysis of a NbO2-Mott memristor fabricated at Hewlett-Packard Labs. It investigates mechanisms behind the origin of complexity based on local activity, which characterizes the behavior of this outstanding nanodevice. We propose an accurate, particularly simplified version of a recently introduced physical model suitable for large-scale circuit simulations. Following the concept of local activity, we then conduct a small-signal circuit-theoretic derivation of the impedance and associated small-signal equivalent circuit elements to analyze device stability and frequency response. Finally, our analysis reveals locally active operating regions, as well as regions where the device dynamics are positioned on the edge of chaos. The latter regions are crucial for designing bio-inspired computing systems. Ioannis Messaris, Timothy D. Brown, Ahmet Samil Demirkol, Alon Ascoli, Mohamad Moner Al Chawa, R. Stanley Williams, Ronald Tetzlaff, Leon O. Chua |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2021 | Improved Hopfield Network Optimization Using Manufacturable Three-Terminal Electronic SynapsesabstractWe illustrate novel optimization techniques via simulations for Hopfield networks constructed from manufacturable three-terminal Silicon-Oxide-Nitride-Oxide-Silicon (SONOS) synaptic circuit elements. We first present a computationally-light, memristor-based, highly accurate static compact model for the SONOS synapses used in our simulations. We then show how to exploit analog errors in programming resistances and current leakage, and the continuous tunability of the SONOS synapses to enable transient chaotic group dynamics, to accelerate the convergence of a Hopfield network. We project improvements in energy consumption and time to solution relative to existing CPUs and GPUs by at least 4 orders of magnitude, and also exceed the projected performance of two-terminal memristor-based crossbars in addition to a 100-fold increase in error-resilient array size (i.e. problem size). Su-In Yi, Suhas Kumar, R. Stanley Williams |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2021 | Analog Solutions of Discrete Markov Chains via Memristor CrossbarsabstractProblems involving discrete Markov Chains are solved mathematically using matrix methods. Recently, several research groups have demonstrated that matrix-vector multiplication can be performed analytically in a single time step with an electronic circuit that incorporates an open-loop memristor crossbar that is effectively a resistive random-access memory. Ielmini and co-workers have taken this a step further by demonstrating that linear algebraic systems can also be solved in a single time step using similar hardware with feedback. These two approaches can both be applied to Markov chains, in the first case using matrix-vector multiplication to compute successive updates to a discrete Markov process and in the second directly calculating the stationary distribution by solving a constrained eigenvector problem. We present circuit models for open-loop and feedback configurations, and perform detailed analyses that include memristor programming errors, thermal noise sources and element nonidealities in realistic circuit simulations to determine both the precision and accuracy of the analog solutions. We provide mathematical tools to formally describe the trade-offs in the circuit model between power consumption and the magnitude of errors. We compare the two approaches by analyzing Markov chains that lead to two different types of matrices, essentially random and ill-conditioned, and observe that ill-conditioned matrices suffer from significantly larger errors. We compare our analog results to those from digital computations and find a significant power efficiency advantage for the crossbar approach for similar precision results. Gianluca Zoppo, Anil Korkmaz, Francesco Marrone, Samuel Palermo, Fernando Corinto, R. Stanley Williams |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2020 | A Simplified Model for a NbO2 Mott Memristor Physical RealizationabstractIn this paper, we propose a new model for practical, nano-scale, NbO2-based Mott memristors, which is based on a thorough analysis performed on a recently presented physics-based model for these devices. Our investigations revealed that the 3D Poole-Frenkel conduction mechanism adopted in the aforementioned model, can be well-approximated by a transport equation in which: a) memristor current is expressed as a linear function of memristor voltage and b) the device memductance is solely dependent on the device temperature which represents the memristor state. The resulting simplified mathematical form of the original differential algebraic equation set is not only more suitable for simulating large-scale, nano-scale NbO2-based memristor circuits, but is also ideal for circuit-theoretic investigations which may allow an in depth understanding of the peculiar nonlinear behaviors of these devices. Ioannis Messaris, Ronald Tetzlaff, Alon Ascoli, R. Stanley Williams, Suhas Kumar, Leon O. Chua |
ISCAS | 4 |
| 2019 | PUMA: A Programmable Ultra-efficient Memristor-based Accelerator for Machine Learning InferenceabstractMemristor crossbars are circuits capable of performing analog matrix-vector multiplications, overcoming the fundamental energy efficiency limitations of digital logic. They have been shown to be effective in special-purpose accelerators for a limited set of neural network applications. We present the Programmable Ultra-efficient Memristor-based Accelerator (PUMA) which enhances memristor crossbars with general purpose execution units to enable the acceleration of a wide variety of Machine Learning (ML) inference workloads. PUMA's microarchitecture techniques exposed through a specialized Instruction Set Architecture (ISA) retain the efficiency of in-memory computing and analog circuitry, without compromising programmability. We also present the PUMA compiler which translates high-level code to PUMA ISA. The compiler partitions the computational graph and optimizes instruction scheduling and register allocation to generate code for large and complex workloads to run on thousands of spatial cores. We have developed a detailed architecture simulator that incorporates the functionality, timing, and power models of PUMA's components to evaluate performance and energy consumption. A PUMA accelerator running at 1 GHz can reach area and power efficiency of 577 GOPS/s/mm 2 and 837~GOPS/s/W, respectively. Our evaluation of diverse ML applications from image recognition, machine translation, and language modelling (5M-800M synapses) shows that PUMA achieves up to 2,446× energy and 66× latency improvement for inference compared to state-of-the-art GPUs. Compared to an application-specific memristor-based accelerator, PUMA incurs small energy overheads at similar inference latency and added programmability. Aayush Ankit, Izzat El Hajj, Sai Rahul Chalamalasetti, Geoffrey Ndu, Martin Foltin, R. Stanley Williams, Paolo Faraboschi, Wen-Mei W. Hwu, John Paul Strachan, Kaushik Roy 0001, Dejan S. Milojicic |
ASPLOS | 6 |
| 2018 | Large Memristor Crossbars for Analog ComputingabstractMemristor with tunable non-volatile resistance offers in-memory computing capability that avoids the von-Neumann bottleneck. However, large-scale experimental demonstration to this end is yet to be implemented due to the immaturity of the device and integration technologies. Here in this paper we report our recent process in analog computing using analog-voltage-amplitude-vector input and analog-memristor-conductance matrix, with applications in signal and image processing. The vector matrix multiplication is processed in the memristor crossbars in one step, with 5-8 bit precision depending on the array size. The demonstration is made possible by high memristor yield (99.8%), stable multilevel memresistance states, linear current-voltage (IV) relation in the operation range, and low wire resistance between the cells. Can Li 0024, Yunning Li, Hao Jiang 0017, J. Joshua Yang, Qiangfei Xia, Miao Hu 0002, Eric Montgomery, Noraica Dávila, Catherine Graves, John Paul Strachan, R. Stanley Williams, Ning Ge 0001, Mark Barnell, Qing Wu 0002 |
ISCAS | 16 |
| 2016 | Brain Inspired ComputingabstractNo abstract available. R. Stanley Williams |
ASPLOS | 1 |
| 2016 | Dot-product engine for neuromorphic computing: programming 1T1M crossbar to accelerate matrix-vector multiplicationabstractVector-matrix multiplication dominates the computation time and energy for many workloads, particularly neural network algorithms and linear transforms (e.g, the Discrete Fourier Transform). Utilizing the natural current accumulation feature of memristor crossbar, we developed the Dot-Product Engine (DPE) as a high density, high power efficiency accelerator for approximate matrix-vector multiplication. We firstly invented a conversion algorithm to map arbitrary matrix values appropriately to memristor conductances in a realistic crossbar array, accounting for device physics and circuit issues to reduce computational errors. The accurate device resistance programming in large arrays is enabled by close-loop pulse tuning and access transistors. To validate our approach, we simulated and benchmarked one of the state-of-the-art neural networks for pattern recognition on the DPEs. The result shows no accuracy degradation compared to software approach (99 % pattern recognition accuracy for MNIST data set) with only 4 Bit DAC/ADC requirement, while the DPE can achieve a speed-efficiency product of 1,000× to 10,000× compared to a custom digital ASIC. Miao Hu 0002, John Paul Strachan, Emmanuelle M. Grafals, Noraica Dávila, Catherine Graves, Sity Lam, Ning Ge 0001, J. Joshua Yang, R. Stanley Williams |
DAC | 10 |
| 2016 | Fading memory effects in a memristor for Cellular Nanoscale Network applications
Alon Ascoli, Ronald Tetzlaff, Leon O. Chua, John Paul Strachan, R. Stanley Williams |
DATE | 5 |
| 2016 | ISAAC: A Convolutional Neural Network Accelerator with In-Situ Analog Arithmetic in CrossbarsabstractA number of recent efforts have attempted to design accelerators for popular machine learning algorithms, such as those involving convolutional and deep neural networks (CNNs and DNNs). These algorithms typically involve a large number of multiply-accumulate (dot-product) operations. A recent project, DaDianNao, adopts a near data processing approach, where a specialized neural functional unit performs all the digital arithmetic operations and receives input weights from adjacent eDRAM banks. This work explores an in-situ processing approach, where memristor crossbar arrays not only store input weights, but are also used to perform dot-product operations in an analog manner. While the use of crossbar memory as an analog dot-product engine is well known, no prior work has designed or characterized a full-fledged accelerator based on crossbars. In particular, our work makes the following contributions: (i) We design a pipelined architecture, with some crossbars dedicated for each neural network layer, and eDRAM buffers that aggregate data between pipeline stages. (ii) We define new data encoding techniques that are amenable to analog computations and that can reduce the high overheads of analog-to-digital conversion (ADC). (iii) We define the many supporting digital components required in an analog CNN accelerator and carry out a design space exploration to identify the best balance of memristor storage/compute, ADCs, and eDRAM storage on a chip. On a suite of CNN and DNN workloads, the proposed ISAAC architecture yields improvements of 14.8×, 5.5×, and 7.5× in throughput, energy, and computational density (respectively), relative to the state-of-the-art DaDianNao architecture. Ali Shafiee, Anirban Nag, Naveen Muralimanohar, Rajeev Balasubramonian, John Paul Strachan, Miao Hu 0002, R. Stanley Williams, Vivek Srikumar |
ISCA | 7 |
| 2014 | New materials for memristive switchingabstractMaterials play a critical role in memristive devices and the research community is aggressively searching for the most applicable material systems for memristive switching. Two representative examples of newly developed switching materials are presented in this paper, including nitride memristors and Pt doped SiO2nanometallic memristors. The former represents nonoxide systems that might be more compatible with nitride electrodes preferred in a fab and the latter represents engineered materials that exhibit a better controllability over the formation of switching channel(s). Byung Joon Choi, Ning Ge 0001, J. Joshua Yang, Min-Xian Zhang, R. Stanley Williams, Kate J. Norris, Nobuhiko P. Kobayashi |
ISCAS | 5 |
| 2013 | Physics-based memristor modelsabstractIn order to utilize memristors in circuits, one needs high-quality predictive models that can be used for simulations to act as a design aid. Whenever possible, we base our models on the known physics of the memristors we are using. To that end, we perform a wide range of materials characterizations and electronic measurements on which to base the model. However, given the complexity of the physical processes that occur in the devices, such as drift-diffusion-thermophoresis in ion-migration based memristors and Mott transitions in locally active memristors, the corresponding detailed mathematical descriptions are far too complex to solve analytically and numerical solutions are too time consuming to include in a simulation. We thus need to find simpler, analytical approximations that can match the measured behavior of the memristors over many orders of magnitude in time and a wide range of applied voltage. We present some new models that we have developed and describe how they are derived. R. Stanley Williams, Matthew D. Pickett, John Paul Strachan |
ISCAS | 1 |
| 2013 | Memristive devices in computing system: Promises and challengesabstractMemristive devices with a simple structure are not only very small but also very versatile, which makes them an ideal candidate used for the next generation computing system in the post-Si era. The working mechanism of the devices and a family of nanodevices built based on this working mechanism are introduced first followed by some proposed applications of these novel devices. The promises and challenges of these devices are then discussed, together with the significant progresses made recently in dealing with these challenges. J. Joshua Yang, R. Stanley Williams |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2012 | Designing memristors: Physics, materials science and engineeringabstractRecently, memory and storage have taken a front seat in computer hardware as it experiences an explosive growth at a rate faster than Moore's law for the past 10 years. With the upcoming challenges for further FLASH scaling into the next generations, emerging technologies have appeared portraying perspectives with the potential to shift computer architecture concepts. Here we present a brief overview and progress in our quest to create a device with competitive attributes, with the ultimate goal of achieving a universal, non-volatile data storage solution. Gilberto Medeiros-Ribeiro, J. Joshua Yang, Janice H. Nickel, Antonio Torrezan, John Paul Strachan, R. Stanley Williams |
ISCAS | 6 |
| 2010 | Hybrid CMOS/memristor circuitsabstractThis is a brief review of recent work on the prospective hybrid CMOS/memristor circuits. Such hybrids combine the flexibility, reliability and high functionality of the CMOS subsystem with very high density of nanoscale thin film resistance switching devices operating on different physical principles. Simulation and initial experimental results demonstrate that performance of CMOS/memristor circuits for several important applications is well beyond scaling limits of conventional VLSI paradigm. Dmitri B. Strukov, Duncan R. Stewart, Julien Borghetti, Xuema Li, Matthew D. Pickett, Gilberto Medeiros-Ribeiro, Warren Robinett, Gregory S. Snider, John Paul Strachan, Qiangfei Xia, J. Joshua Yang, R. Stanley Williams |
ISCAS | 13 |
| 2008 | Nanoelectronic and Nanophotonic InterconnectabstractA significant performance limitation in integrated circuits has become the metal interconnect, which is responsible for depressing the on-chip data bandwidth while consuming an increasing percentage of power. These problems will grow as wire diameters scale down and the resistance-capacitance product of the interconnect wires increases hyperbolically, which threatens to choke off the computational performance increases of chips that we have come to expect over time. We examine some of the quantitative implications of these trends by analyzing the International Technology Roadmap for Semiconductors. We compare the potential of replacing the global electronic interconnect of future chips with a photonic interconnect and see that there is in principle a four order of magnitude bandwidth-to-power ratio advantage for the latter. This indicates that it could be possible to dramatically improve chip performance without scaling transistors but rather utilize the capability of existing transistors much more efficiently. However, at this time it is not clear if these advantages can be realized. We discuss various issues related to the architecture and components necessary to implement on-chip photonic interconnect. Raymond G. Beausoleil, Philip Kuekes, Gregory S. Snider, Shih-Yuan Wang, R. Stanley Williams |
Proc. IEEE | 5 |
| 2000 | Molecular nanoelectronicsabstractBoth molecular switching and nanoscale wires have been recently demonstrated. We discuss their possible applications in circuits with emphasis on cross-point memories and programmable logic arrays. R. Stanley Williams, Philip Kuekes |
ISCAS | 1 |
| 1993 | The nanomanipulator: a virtual-reality interface for a scanning tunneling microscopeabstractWe present an atomic-scale teleoperation system that uses a head-mounted display and force-feedback manipulator arm for a user interface and a Scanning Tunneling Microscope (STM) as a sensor and effector.The system approximates presence at the atomic scale, placing the scientist on the surface, in control, w h i l e the experiment is happening.A scientist using the Nanomanipulator can view incoming STM data, feel the surface, and modify the surface (using voltage pulses) in real time.The Nanomanipulator has been used to study the effects of bias pulse duration on the creation of gold mounds.We intend to use the system to make controlled modifications to silicon surfaces. Russell M. Taylor II, Warren Robinett, Vernon L. Chi, Frederick P. Brooks Jr., William V. Wright, R. Stanley Williams, Erik J. Snyder |
SIGGRAPH | 6 |