Brian Crafton

dblp:237/9870 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 15 · 7 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2026 CryoBoost: A 40nm Cryogenic-CMOS Matrix Multiplication Accelerator for Energy Efficient Computing
abstract
This paper proposes a cryogenic Matrix Multiplication (MATMUL) accelerator chip to address the exponential increase in energy consumption in AI model training. The accelerator, designed in a 40nm CMOS process, leverages a Liquid Nitrogen-based cooling system, operating from 300K to 77K. The design comprises 10 processing elements (PEs) operating on 4x4 matrices at INT8 precision, interconnected by a core data ring. The PEs use a load-store architecture with a 128-bit very long instruction word (VLIW)-based controller. The paper presents a detailed characterization of the supply voltage versus frequency, performance, and power across different temperatures. The results indicate a significant reduction in power at cryogenic temperatures, with up to 45.7% reduction in power at 77K compared to 300K at iso performance. The maximum energy efficiency increases from 1.975GHz/W at 300K to 2.497GHz/W at 77K, yielding a 26.4% gain. This translates to up to 20% reduction in training energy for Large Language Models.
Rakshith Saligram, Samuel Spetalnick, Brian Crafton, Muya Chang, Alec Nordlund, Joshua Gess, Ruslan Nagimov, Arijit Raychowdhury
ACM Great Lakes Symposium on VLSI3
2026 Ultrafast Generative AI by Ultradense 3D Integration: A Case Study on LLM-based Edge Inference
abstract
Generative AI (GenAI) is one of the most critical applications today, continually challenging the limits of semiconductor technology. We introduce a very fine-grained 3D memory-on-logic architecture along with a novel data mapping strategy to support Large Language Model (LLM)-based GenAI, including both prefill and generation stages. Our conceptual analysis shows how ultradense 3D connectivity can enhance text generation speed and energy-efficiency well-beyond current limits. Preliminary findings from a basic analytical model indicate that the single batch autoregressive generation rate for Llama 3.2 1B could surpass 5K tokens/sec by maximizing weight locality and enhancing memory bandwidth through massively parallel 3D links between Multiply-Accumulate (MAC) units in the logic tier and their dedicated memory partitions in the 3D stack. We also explore the impact of advanced logic nodes and quantify their benefits in reducing prefill latency. Finally, we examine the challenges associated with memory access power and power density under extreme bandwidth conditions and present pipelined access strategies to address them.
Kerem Akarvardar, Xiaoyu Sun 0001, Brian Crafton, Xiaochen Peng, Haruki Mori, Abhiroop Bhattacharjee, Hidehiro Fujiwara, H.-S. Philip Wong
ACM Trans. Design Autom. Electr. Syst.3
2025 Finding the Pareto Frontier of Low-Precision Data Formats and MAC Architecture for LLM Inference
abstract
To accelerate AI applications, numerous data formats and physical implementations of matrix multiplication have been proposed, creating a complex design space. This paper studies the efficient MAC implementation of the integer, floating-point, posit, and logarithmic number system (LNS) data formats and Microscaling (MX) and VectorScaled Quantization (VSQ) block data formats. We evaluate the area, power, and numerical accuracy (evaluated as signal-to-quantization noise ratio) of $\mathbf{3 5, 0 0 0}$ MAC designs spanning each data format and several key design parameters such as the inner product size and accumulation width. We find that for the same numerical accuracy, pareto optimal MAC designs with emerging data formats (LNS16, MXINT8, VSQINT4) achieve $1.8 \times 2.2 \times$, and $1.9 \times$ TOPs/W improvement compared to FP16, FP8, and FP4 dot product implementations.
Brian Crafton, Xiaochen Peng, Xiaoyu Sun 0001, Ashwin Sanjay Lele, Win-San Khwa, Kerem Akarvardar
DAC1
2025 Characterization and Mitigation of ADC Noise by Reference Tuning in RRAM-Based Compute-In-Memory
abstract
With the escalating demand for power-efficient neural network architectures, non-volatile compute-in-memory de-signs have garnered significant attention. However, owing to the nature of analog computation, susceptibility to noise remains a critical concern. This study confronts this challenge by introducing a detailed model that incorporates noise factors arising from both ADCs and RRAM devices. The experimental data is derived from a 40nm foundry RRAM test-chip, wherein different reference voltage configurations are applied, each tailored to its respective module. The mean and standard deviation values of HRS and LRS cells are derived through a randomized vector, forming the foundation for noise simulation within our analytical framework. Additionally, the study examines the read-disturb effects, shedding light on the potential for accuracy deterioration in neural networks due to extended exposure to high-voltage stress. This phenomenon is mitigated through the proposed low-voltage read mode. Leveraging our derived comprehensive fault model from the RRAM test-chip, we evaluate CIM noise impact on both supervised learning (time-independent) and reinforcement learning (time-dependent) tasks, and demonstrate the effectiveness of reference tuning to mitigate noise impacts.
Ying-Hao Wei, Zishen Wan, Brian Crafton, Samuel Spetalnick, Arijit Raychowdhury
ISCAS3
2024 Efficient Processing of MLPerf Mobile Workloads Using Digital Compute-In-Memory Macros
abstract
Compute-in-memory (CIM) has recently emerged as a promising design paradigm to accelerate deep neural network (DNN) processing. Continuously better energy and area efficiency at the macrolevel had been reported through many testchips over the last few years. However, in those macro design-oriented studies, accelerator-level considerations, such as memory accesses and processing of entire DNN workloads have not been investigated in-depth. In this article, we aim to fill this gap starting with the characteristics of our latest CIM macro fabricated with cutting-edge FinFET CMOS technology at 4-nm node. We then study, through an accelerator simulator developed in-house, three key items that would determine the efficiency of our CIM macro in the accelerator context while running MLPerf Mobile suite: 1) dataflow optimization; 2) optimal selection of CIM macro dimensions to further improve macro utilization; and 3) optimal combination of multiple CIM macros. Although there is typically a stark contrast between macro-level peak and accelerator-level average throughput and energy efficiency, the aforementioned optimizations are shown to improve the macro utilization by$3.04\times $and reduce the energy-delay product (EDP) to$0.34\times $compared to the original macro on MLPerf Mobile inference workloads. While we exploit a digital CIM macro in this study, the findings and proposed methods remain valid for other types of CIM (such as analog CIM and analog–digital–hybrid CIM) as well.
Xiaoyu Sun 0001, Weidong Cao 0001, Brian Crafton, Kerem Akarvardar, Haruki Mori, Hidehiro Fujiwara, Hiroki Noguchi, Yu-Der Chih, Meng-Fan Chang, Yih Wang, Tsung-Yung Jonathan Chang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2023 Neuromorphic Swarm on RRAM Compute-in-Memory Processor for Solving QUBO Problem
abstract
Combinatorial optimization problems prevail in engineering and industry. Some are NP-hard and thus become difficult to solve on edge devices due to limited power and computing resources. Quadratic Unconstrained Binary Optimization (QUBO) problem is a valuable emerging model that can formulate numerous combinatorial problems, such as Max-Cut, traveling salesman problems, and graphic coloring. QUBO model also reconciles with two emerging computation models, quantum computing and neuromorphic computing, which can potentially boost the speed and energy efficiency in solving combinatorial problems. In this work, we design a neuromorphic QUBO solver composed of a swarm of spiking neural networks (SNN) that conduct a population-based meta-heuristic search for solutions. The proposed model can achieve about x20 40 speedup on large QUBO problems in terms of time steps compared to a traditional neural network solver. As a codesign, we evaluate the neuromorphic swarm solver on a 40nm 25mW Resistive RAM (RRAM) Compute-in-Memory (CIM) SoC with a 2.25MB RRAM-based accelerator and an embedded Cortex M3 core. The collaborative SNN swarm can fully exploit the specialty of CIM accelerator in matrix and vector multiplications. Compared to previous works, such an algorithm-hardware synergized solver exhibits advantageous speed and energy efficiency for edge devices.
Ashwin Sanjay Lele, Muya Chang, Samuel Spetalnick, Brian Crafton, Arijit Raychowdhury, Yan Fang 0002
DAC4
2023 Live Demonstration: Hybrid RRAM and SRAM SoC for Fused Frame and Event Target Tracking
abstract
Event and frame cameras capture the complemen-tary spatial and temporal details of a scene providing an accuracy vs. latency trade-off. Fusing these processing modalities using convolutional (CNN) and spiking neural networks (SNN) respectively has been shown for target tracking. We present our heterogeneous RRAM compute-in-memory (CIM) and SRAM compute-near-memory (CNM) SoC for simultaneous processing of CNN and SNN. We will show the advantage of using fused vision over frame-only vision and demonstrate python programmable data streaming. The visitors will be able to see the processing-dependent dynamic power gating of non-volatile RRAM and in-memory error correction capability.
Ashwin Sanjay Lele, Muya Chang, Samuel Spetalnick, Yan Fang 0002, Brian Crafton, Shota Konno, Arijit Raychowdhury
ISCAS5
2022 Improving compute in-memory ECC reliability with successive correction
abstract
Compute in-memory (CIM) is an exciting technique that minimizes data transport, maximizes memory throughput, and performs computation on the bitline of memory sub-arrays. This is especially interesting for machine learning applications, where increased memory bandwidth and analog domain computation offer improved area and energy efficiency. Unfortunately, CIM faces new challenges traditional CMOS architectures have avoided. In this work, we explore the impact of device variation (calibrated with measured data on foundry RRAM arrays) and propose a new class of error correcting codes (ECC) for hard and soft errors in CIM. We demonstrate single, double, and triple error correction offering over 16,000× reduction in bit error rate over a design without ECC and over 427× over prior work, while consuming only 29.1% area and 26.3% power overhead.
Brian Crafton, Zishen Wan, Samuel Spetalnick, Jong-Hyeok Yoon, Carlos Tokunaga, Vivek De, Arijit Raychowdhury
DAC1
2022 Characterization and Mitigation of IR-Drop in RRAM-based Compute In-Memory
abstract
Compute in-memory (CIM) is an exciting circuit innovation that promises to increase effective memory bandwidth and perform computation on the bitlines of memory sub-arrays. Utilizing embedded non-volatile memories (eNVM) such as resistive random access memory (RRAM), various forms of neural networks can be implemented. Unfortunately, CIM faces new challenges traditional CMOS architectures have avoided. In this work, we characterize the impact of IR-drop and device variation (calibrated with measured data on foundry RRAM) and evaluate different approaches to write verify. Using various voltages and pulse widths we program cells to offset IR-drop and demonstrate a $136.4 \times $ reduction in BER during CIM.
Brian Crafton, Connor Talley, Samuel Spetalnick, Jong-Hyeok Yoon, Arijit Raychowdhury
ISCAS1
2021 Merged Logic and Memory Fabrics for AI Workloads
abstract
As we approach the end of the silicon roadmap, we observe a steady increase in both the research effort toward and quality of embedded non-volatile memories (eNVM). Integrated in a dense array, eNVM such as resistive random access memory (RRAM), spin transfer torque based random access memory, or phase change random access memory (PCRAM) can perform compute in-memory (CIM) using the physical properties of the device. The combination of eNVM and CIM seeks to minimize both data transport and leakage power while offering density up to 10x that of traditional 6T SRAM. Despite these exciting new properties, these devices introduce problems that were not faced by traditional CMOS and SRAM based designs. While some of these problems will be solved by further research and development, properties such as significant cell-to-cell variance and high write power will persist due to the physical limitations of the devices. As a result, circuit and system level designs must account for and mitigate the problems that arise. In this work we introduce these problems from the system level and propose solutions that improve performance while mitigating the impact of the non-ideal properties of eNVM. Using statistics from the application and known properties of the eNVM, we can configure a CIM accelerator to minimize error from cell-to-cell variance and maximize throughput while minimizing write energy.
Brian Crafton, Samuel Spetalnick, Arijit Raychowdhury
ASP-DAC1
2021 Statistical Optimization of Compute In-Memory Performance Under Device Variation
abstract
Compute in-memory (CIM) is a promising technique that minimizes data transport, maximizes memory throughput, and performs computation on the bitline of memory sub-arrays. Utilizing embedded non-volatile memories (eNVM) such as resistive random access memory (RRAM), various forms of neural networks can be implemented. Unfortunately, CIM faces new challenges traditional CMOS architectures have avoided. In this work, we explore the impact of device variation (calibrated with measured data on foundry RRAM arrays) and propose a new algorithm based on device variation to increase both performance and accuracy for CIM designs. We demonstrate up to 36% power improvement and 44% performance improvement, while satisfying any error constraint.
Brian Crafton, Samuel Spetalnick, Jong-Hyeok Yoon, Arijit Raychowdhury
ISLPED1
2021 A Hardware-Friendly Approach Towards Sparse Neural Networks Based on LFSR-Generated Pseudo-Random Sequences
abstract
The increase in the number of edge devices has led to the emergence of edge computing where the computations are performed on the device. In recent years, deep neural networks (DNNs) have become the state-of-the-art method in a broad range of applications, from image recognition, to cognitive tasks to control. However, neural network models are typically large and computationally expensive and therefore not deployable on power and memory constrained edge devices. Sparsification techniques have been proposed to reduce the memory foot-print of neural network models. However, they typically lead to substantial hardware and memory overhead. In this article, we propose a hardware-aware pruning method using linear feedback shift register (LFSRs) to generate the locations of non-zero weights in real-time during inference. We call this LFSR-generated pseudorandom sequence based sparsity (LGPS) technique. We explore two different architectures for our hardware-friendly LGPS technique, based on (1) row/column indexing with LFSRs and (2) column-wise indexing with nested LFSRs, respectively. Using the proposed method, we present a total saving of energy and area up to 37.47% and 49.93% respectively and speed up of 1.53× w.r.t the baseline pruning method, for the VGG-16 network on down-sampled ImageNet.
Foroozan Karimzadeh, Ningyuan Cao, Brian Crafton, Justin K. Romberg, Arijit Raychowdhury
IEEE Trans. Circuits Syst. I Regul. Pap.3
2020 Hardware-Aware Pruning of DNNs using LFSR-Generated Pseudo-Random Indices
abstract
Deep neural networks (DNNs) have been emerged as the state-of-the-art algorithms in broad range of applications. To reduce the memory foot-print of DNNs, in particular for embedded applications, sparsification techniques have been proposed. Unfortunately, these techniques come with a large hardware overhead. In this paper, we present a hardware-aware pruning method where the locations of non-zero weights are derived in real-time from a Linear Feedback Shift Registers (LFSRs). Using the proposed method, we demonstrate a total saving of energy and area up to 63.96% and 64.23% for VGG-16 network on down-sampled ImageNet, respectively for iso-compression-rate and iso-accuracy.
Foroozan Karimzadeh, Ningyuan Cao, Brian Crafton, Justin K. Romberg, Arijit Raychowdhury
ISCAS3
2020 Breaking Barriers: Maximizing Array Utilization for Compute in-Memory Fabrics
abstract
Compute in-memory (CIM) is a promising technique that minimizes data transport, the primary performance bottleneck and energy cost of most data intensive applications. This has found wide-spread adoption in accelerating neural networks for machine learning applications. Utilizing a crossbar architecture with emerging non-volatile memories (eNVM) such as dense resistive random access memory (RRAM) or phase change random access memory (PCRAM), various forms of neural networks can be implemented to greatly reduce power and increase on chip memory capacity. However, compute in-memory faces its own limitations at both the circuit and the device levels. Although compute in-memory using the crossbar architecture can greatly reduce data transport, the rigid nature of these large fixed weight matrices forfeits the flexibility of traditional CMOS and SRAM based designs. In this work, we explore the different synchronization barriers that occur from the CIM constraints. Furthermore, we propose a new allocation algorithm and data flow based on input data distributions to maximize utilization and performance for compute-in memory based designs. We demonstrate a$\boldsymbol{7.47}\times$performance improvement over a naive allocation method for CIM accelerators on ResNet18.
Brian Crafton, Samuel Spetalnick, Gauthaman Murali, Tushar Krishna, Sung Kyu Lim, Arijit Raychowdhury
VLSI-SOC1
2019 Local Learning in RRAM Neural Networks with Sparse Direct Feedback Alignment
abstract
Neural networks utilizing non-volatile random access memory (NVM) exhibit excellent power reduction over traditional CMOS implementations. RRAM (resistive random access memory) is one such emerging memory technology offering low energy, good endurance, and a large analog conductance window. When implemented in a crossbar architecture, these networks are able to bypass the von-Neumann bottleneck by performing compute in-memory. This architecture works well for inference; however, training the network is far more challenging. Networks built using RRAM can be trained on-chip with gradient descent or off-chip where weights are transferred. Backpropagation, while effective in training von-Neumann architectures, is inefficient when memory and compute are partitioned together. Commonly referred to as the weight transport problem, each neuron's dependence on the weights and errors located deeper in the network requires reading the weights in each layer before computing and applying the error. This presents a key challenge in performing efficient on chip training for non von-Neumann architectures. In this work we demonstrate an alternative to backpropagation called sparse direct feedback alignment which bypasses the weight transport problem. We simulate crossbars of HfOx RRAM based on experimental data to explore the performance, area, and energy trade-offs of using bio-plausible algorithms on the MNIST and EMNIST datasets.
Brian Crafton, Matt West 0002, Padip Basnet, Eric Vogel, Arijit Raychowdhury
ISLPED1