VLDB 2026 Research / reviewers in the wild / expert
Samuel Spetalnick
dblp:266/9691 · also Samuel D. Spetalnick
· DBLP profile ↗
11ranked-venue papers
1as first author
10since 2021 · last 2026
0000-0003-1627-9002ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 1 first-author · 10 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CryoBoost: A 40nm Cryogenic-CMOS Matrix Multiplication Accelerator for Energy Efficient ComputingabstractThis paper proposes a cryogenic Matrix Multiplication (MATMUL) accelerator chip to address the exponential increase in energy consumption in AI model training. The accelerator, designed in a 40nm CMOS process, leverages a Liquid Nitrogen-based cooling system, operating from 300K to 77K. The design comprises 10 processing elements (PEs) operating on 4x4 matrices at INT8 precision, interconnected by a core data ring. The PEs use a load-store architecture with a 128-bit very long instruction word (VLIW)-based controller. The paper presents a detailed characterization of the supply voltage versus frequency, performance, and power across different temperatures. The results indicate a significant reduction in power at cryogenic temperatures, with up to 45.7% reduction in power at 77K compared to 300K at iso performance. The maximum energy efficiency increases from 1.975GHz/W at 300K to 2.497GHz/W at 77K, yielding a 26.4% gain. This translates to up to 20% reduction in training energy for Large Language Models. Rakshith Saligram, Samuel Spetalnick, Brian Crafton, Muya Chang, Alec Nordlund, Joshua Gess, Ruslan Nagimov, Arijit Raychowdhury |
ACM Great Lakes Symposium on VLSI | 2 |
| 2025 | Characterization and Mitigation of ADC Noise by Reference Tuning in RRAM-Based Compute-In-MemoryabstractWith the escalating demand for power-efficient neural network architectures, non-volatile compute-in-memory de-signs have garnered significant attention. However, owing to the nature of analog computation, susceptibility to noise remains a critical concern. This study confronts this challenge by introducing a detailed model that incorporates noise factors arising from both ADCs and RRAM devices. The experimental data is derived from a 40nm foundry RRAM test-chip, wherein different reference voltage configurations are applied, each tailored to its respective module. The mean and standard deviation values of HRS and LRS cells are derived through a randomized vector, forming the foundation for noise simulation within our analytical framework. Additionally, the study examines the read-disturb effects, shedding light on the potential for accuracy deterioration in neural networks due to extended exposure to high-voltage stress. This phenomenon is mitigated through the proposed low-voltage read mode. Leveraging our derived comprehensive fault model from the RRAM test-chip, we evaluate CIM noise impact on both supervised learning (time-independent) and reinforcement learning (time-dependent) tasks, and demonstrate the effectiveness of reference tuning to mitigate noise impacts. Ying-Hao Wei, Zishen Wan, Brian Crafton, Samuel Spetalnick, Arijit Raychowdhury |
ISCAS | 4 |
| 2024 | H3DFact: Heterogeneous 3D Integrated CIM for Factorization with Holographic Perceptual RepresentationsabstractDisentangling attributes of various sensory signals is central to human-like perception and reasoning and a critical task for higher-order cognitive and neuro-symbolic AI systems. An elegant approach to represent this intricate factorization is via high-dimensional holographic vectors drawing on brain-inspired vector symbolic architectures. However, holographic factorization involves iterative computation with high-dimensional matrix-vector multiplications and suffers from non-convergence problems. In this paper, we present H3DFact, a heterogeneous 3D integrated in-memory compute engine capable of efficiently factorizing high-dimensional holographic representations. H3DFact exploits the computation-in-superposition capability of holographic vectors and the intrinsic stochasticity associated with memristive-based 3D compute-in-memory. Evaluated on large-scale factorization and perceptual problems, H3DFact demonstrates superior capability in factorization accuracy and operational capacity by up to five orders of magnitude, with 5.5 x compute density, 1.2 x energy efficiency improvements, and 5.9 x less silicon footprint compared to iso-capacity 2D designs. Zishen Wan, Che-Kai Liu, Mohamed Ibrahim 0002, Hanchen Yang 0001, Samuel Spetalnick, Tushar Krishna, Arijit Raychowdhury |
DATE | 5 |
| 2023 | Neuromorphic Swarm on RRAM Compute-in-Memory Processor for Solving QUBO ProblemabstractCombinatorial optimization problems prevail in engineering and industry. Some are NP-hard and thus become difficult to solve on edge devices due to limited power and computing resources. Quadratic Unconstrained Binary Optimization (QUBO) problem is a valuable emerging model that can formulate numerous combinatorial problems, such as Max-Cut, traveling salesman problems, and graphic coloring. QUBO model also reconciles with two emerging computation models, quantum computing and neuromorphic computing, which can potentially boost the speed and energy efficiency in solving combinatorial problems. In this work, we design a neuromorphic QUBO solver composed of a swarm of spiking neural networks (SNN) that conduct a population-based meta-heuristic search for solutions. The proposed model can achieve about x20 40 speedup on large QUBO problems in terms of time steps compared to a traditional neural network solver. As a codesign, we evaluate the neuromorphic swarm solver on a 40nm 25mW Resistive RAM (RRAM) Compute-in-Memory (CIM) SoC with a 2.25MB RRAM-based accelerator and an embedded Cortex M3 core. The collaborative SNN swarm can fully exploit the specialty of CIM accelerator in matrix and vector multiplications. Compared to previous works, such an algorithm-hardware synergized solver exhibits advantageous speed and energy efficiency for edge devices. Ashwin Sanjay Lele, Muya Chang, Samuel Spetalnick, Brian Crafton, Arijit Raychowdhury, Yan Fang 0002 |
DAC | 3 |
| 2023 | Live Demonstration: Hybrid RRAM and SRAM SoC for Fused Frame and Event Target TrackingabstractEvent and frame cameras capture the complemen-tary spatial and temporal details of a scene providing an accuracy vs. latency trade-off. Fusing these processing modalities using convolutional (CNN) and spiking neural networks (SNN) respectively has been shown for target tracking. We present our heterogeneous RRAM compute-in-memory (CIM) and SRAM compute-near-memory (CNM) SoC for simultaneous processing of CNN and SNN. We will show the advantage of using fused vision over frame-only vision and demonstrate python programmable data streaming. The visitors will be able to see the processing-dependent dynamic power gating of non-volatile RRAM and in-memory error correction capability. Ashwin Sanjay Lele, Muya Chang, Samuel Spetalnick, Yan Fang 0002, Brian Crafton, Shota Konno, Arijit Raychowdhury |
ISCAS | 3 |
| 2022 | Improving compute in-memory ECC reliability with successive correctionabstractCompute in-memory (CIM) is an exciting technique that minimizes data transport, maximizes memory throughput, and performs computation on the bitline of memory sub-arrays. This is especially interesting for machine learning applications, where increased memory bandwidth and analog domain computation offer improved area and energy efficiency. Unfortunately, CIM faces new challenges traditional CMOS architectures have avoided. In this work, we explore the impact of device variation (calibrated with measured data on foundry RRAM arrays) and propose a new class of error correcting codes (ECC) for hard and soft errors in CIM. We demonstrate single, double, and triple error correction offering over 16,000× reduction in bit error rate over a design without ECC and over 427× over prior work, while consuming only 29.1% area and 26.3% power overhead. Brian Crafton, Zishen Wan, Samuel Spetalnick, Jong-Hyeok Yoon, Carlos Tokunaga, Vivek De, Arijit Raychowdhury |
DAC | 3 |
| 2022 | Characterization and Mitigation of IR-Drop in RRAM-based Compute In-MemoryabstractCompute in-memory (CIM) is an exciting circuit innovation that promises to increase effective memory bandwidth and perform computation on the bitlines of memory sub-arrays. Utilizing embedded non-volatile memories (eNVM) such as resistive random access memory (RRAM), various forms of neural networks can be implemented. Unfortunately, CIM faces new challenges traditional CMOS architectures have avoided. In this work, we characterize the impact of IR-drop and device variation (calibrated with measured data on foundry RRAM) and evaluate different approaches to write verify. Using various voltages and pulse widths we program cells to offset IR-drop and demonstrate a $136.4 \times $ reduction in BER during CIM. Brian Crafton, Connor Talley, Samuel Spetalnick, Jong-Hyeok Yoon, Arijit Raychowdhury |
ISCAS | 3 |
| 2022 | A Practical Design-Space Analysis of Compute-in-Memory With SRAMabstractAnalog-domain compute-in-memory (CIM) is a technique that has emerged in part as a response to the memory-intensive vector-matrix-multiplications (VMMs) required to implement important emerging applications, notably machine learning inference. Implemented CIM systems have demonstrated good energy efficiency for lower-precision systems and/or with loosened compute-level accuracy requirements.A prioriit is unclear exactly how the efficiency advantages of CIM emerge and therefore the generalizability of these advantages, beyond the specific demonstrated examples, is unclear. Noting that not all VMM-heavy workloads can tolerate imperfect accuracy and/or reduced precision, this work combines high-level models with circuit models and simulations to examine the efficiency gains and penalties associated with CIM in static random-access memory (SRAM) arrays. Extracted models which are needed to make assertive statements about CIM are developed and discussed. An energy comparison to standard SRAM is made, and the issues of accuracy loss and area are contextualized. Finally, a few example models comparing the energy efficiency of CIM to that of SRAM are shown to verify that CIM is most effective for error-tolerant, low-precision applications. Samuel Spetalnick, Arijit Raychowdhury |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2021 | Merged Logic and Memory Fabrics for AI WorkloadsabstractAs we approach the end of the silicon roadmap, we observe a steady increase in both the research effort toward and quality of embedded non-volatile memories (eNVM). Integrated in a dense array, eNVM such as resistive random access memory (RRAM), spin transfer torque based random access memory, or phase change random access memory (PCRAM) can perform compute in-memory (CIM) using the physical properties of the device. The combination of eNVM and CIM seeks to minimize both data transport and leakage power while offering density up to 10x that of traditional 6T SRAM. Despite these exciting new properties, these devices introduce problems that were not faced by traditional CMOS and SRAM based designs. While some of these problems will be solved by further research and development, properties such as significant cell-to-cell variance and high write power will persist due to the physical limitations of the devices. As a result, circuit and system level designs must account for and mitigate the problems that arise. In this work we introduce these problems from the system level and propose solutions that improve performance while mitigating the impact of the non-ideal properties of eNVM. Using statistics from the application and known properties of the eNVM, we can configure a CIM accelerator to minimize error from cell-to-cell variance and maximize throughput while minimizing write energy. Brian Crafton, Samuel Spetalnick, Arijit Raychowdhury |
ASP-DAC | 2 |
| 2021 | Statistical Optimization of Compute In-Memory Performance Under Device VariationabstractCompute in-memory (CIM) is a promising technique that minimizes data transport, maximizes memory throughput, and performs computation on the bitline of memory sub-arrays. Utilizing embedded non-volatile memories (eNVM) such as resistive random access memory (RRAM), various forms of neural networks can be implemented. Unfortunately, CIM faces new challenges traditional CMOS architectures have avoided. In this work, we explore the impact of device variation (calibrated with measured data on foundry RRAM arrays) and propose a new algorithm based on device variation to increase both performance and accuracy for CIM designs. We demonstrate up to 36% power improvement and 44% performance improvement, while satisfying any error constraint. Brian Crafton, Samuel Spetalnick, Jong-Hyeok Yoon, Arijit Raychowdhury |
ISLPED | 2 |
| 2020 | Breaking Barriers: Maximizing Array Utilization for Compute in-Memory FabricsabstractCompute in-memory (CIM) is a promising technique that minimizes data transport, the primary performance bottleneck and energy cost of most data intensive applications. This has found wide-spread adoption in accelerating neural networks for machine learning applications. Utilizing a crossbar architecture with emerging non-volatile memories (eNVM) such as dense resistive random access memory (RRAM) or phase change random access memory (PCRAM), various forms of neural networks can be implemented to greatly reduce power and increase on chip memory capacity. However, compute in-memory faces its own limitations at both the circuit and the device levels. Although compute in-memory using the crossbar architecture can greatly reduce data transport, the rigid nature of these large fixed weight matrices forfeits the flexibility of traditional CMOS and SRAM based designs. In this work, we explore the different synchronization barriers that occur from the CIM constraints. Furthermore, we propose a new allocation algorithm and data flow based on input data distributions to maximize utilization and performance for compute-in memory based designs. We demonstrate a$\boldsymbol{7.47}\times$performance improvement over a naive allocation method for CIM accelerators on ResNet18. Brian Crafton, Samuel Spetalnick, Gauthaman Murali, Tushar Krishna, Sung Kyu Lim, Arijit Raychowdhury |
VLSI-SOC | 2 |