Kyler R. Scott

dblp:321/0260 · DBLP profile ↗
← Back
5ranked-venue papers
5as first author
5since 2021 · last 2026
0000-0002-8872-6177ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 5 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 A Flash-Based Reconfigurable QCNN Inference Accelerator for Edge Applications
abstract
This paper presents a novel in-memory flash-based Quantized Convolutional Neural Network (QCNN) inference accelerator Integrated Circuit (IC), which achieves extremely high throughput and extremely low power, energy, and memory requirements. In this manner, our design significantly improves upon state-of-the-art inference accelerators operating at the edge. Our accelerator IC is designed as a field-programmable, reconfigurable dataflow architecture, which allows a wide variety of trained QCNN models to be programmed to our IC. Once programmed with a model, our IC performs inference withzeroaccess to off-chip memory. In our accelerator, the core computing structures are arrays of flash Field Effect Transistors (FETs) integrated in three dimensions (3D). We use 3D NOR flash stacks, thus achieving an area efficient design with the promise of further improvements as the depth of 3D NOR flash process technology advances. Our flash arrays perform dot product operations in the analog voltage domain, using an extremely low supply voltage (V DDlo, which is 100 mV in our design), resulting in a very high energy efficiency. The digital I/Os of our arrays are in the nominal supply voltage domain for our process (V DDhi, which is 0.8 V in our design), which allows them to be easily stored in registers without the need for level shifters. Neuron weights in our QCNN accelerator are stored in-memory in our flash arrays. These arrays perform dot product operations between a vector of all neuron inputs (in theV DDhidomain) and a vector of corresponding weights (stored in the flash devices as programmedVTHvalues), resulting in an output in theV DDlodomain. We show that for inference on ImageNet with a quantized AlexNet model, our IC achieves extremely low inference latency (28.7 ms at most), power consumption (5.82 mW at most), and energy consumption (4.35 µJ at most), making it especially suited for use in resource-constrained edge devices, such as portable medical devices used for disease diagnosis and ECG signal classification. We prove that our design is robust to Process, Voltage, and Temperature (PVT) variations, and report a drop in Top-1 (Top-5) ImageNet validation accuracy of at most 0.90% (1.07%) across PVT variations. We compare our QCNN accelerator against recent state-of-the-art in-memory neural network accelerators and report an energy efficiency that is at least 15.2µ higher and a parameter density that is at least 3.32µ higher than the best existing design.
Kyler R. Scott, Sunil P. Khatri
IEEE Trans. Computers1
2024 A Mixed-Signal Quantized Neural Network Accelerator Using Flash Transistors
abstract
This paper presents a mixed-signal architecture for implementing Quantized Neural Networks (QNNs) using flash transistors, to achieve extremely high throughput with extremely low power, energy and memory requirements. Its low resource utilization makes our design especially suited for use in edge devices. The network weights are stored in-memory using flash transistors, and neurons perform operations in the analog current domain. Our design can be programmed with any QNN whose hyperparameters (the number of layers, filters, or filter size, etc) do not exceed the maximum provisioned. Once the flash devices are programmed with a trained model and our IC is given an input, our architecture performs inference with zero access to off-chip memory. We demonstrate the robustness of our design under current-mode non-linearities arising from process, voltage, and temperature (PVT) variations. We test validation accuracy on the ImageNet dataset, and show that our IC suffers only 0.71% and 0.92% reduction in classification accuracy for Top-1 and Top-5 outputs, respectively. Our implementation achieves between$2.1\times $and$125\times $better energy efficiency than previous NVM-based QNN accelerators. Our approach provides layer partitioning and neuron sharing options, which allow us to trade off latency, power, and area amongst each other.
Kyler R. Scott, Cheng-Yen Lee, Sunil P. Khatri, Sarma B. K. Vrudhula
IEEE Trans. Circuits Syst. I Regul. Pap.1
2023 An Extremely Low-voltage Floating Gate Artificial Neuron
abstract
This paper presents an artificial neuron design, based on a floating gate (flash) transistor array, that operates with an extremely low supply voltage domain$VDD_{lo}$(which we chose to be 100 mV). Neuron inputs and outputs are in another domain$VDD_{hi}$(which we chose to be 0.8 V). Since$VDD_{hi}$is the nominal supply voltage for our process, input and output data can be stored in registers without the need for level shifters. Neuron weights are stored in-memory in a flash transistor array, using a novel differential conductance encoding scheme. Our neuron performs operations in the analog voltage domain. We compare our design against a recent flash-based mixed-signal neuron design, which has reported the best results (in terms of power and energy, with insignificant error) thus far. We demonstrate that our neuron design beats the previous design in static power consumption$\mathbf{(52.9}\times$less), static energy consumption$\mathbf{(1.41}\times$less), and layout area (44% less). The goal of this work is to design a low-power, low-energy neuron which could be used to implement an entire Neural Network (NN). Our focus is on the neuron, and not the whole NN.
Kyler R. Scott, Sunil P. Khatri
ISCAS1
2022 A Flash-based Current-mode IC to Realize Quantized Neural Networks
abstract
This paper presents a mixed-signal architecture for implementing Quantized Neural Networks (QNNs) using flash transistors to achieve extremely high throughput with extremely low power, energy and memory requirements. Its low resource consumption makes our design especially suited for use in edge devices. The network weights are stored in-memory using flash transistors, and nodes perform operations in the analog current domain. Our design can be programmed with any QNN whose hyperparameters (the number of layers, filters, or filter size, etc) do not exceed the maximum provisioned. Once the flash devices are programmed with a trained model and the IC is given an input, our architecture performs inference with zero access to off-chip memory. We demonstrate the robustness of our design under current-mode non-linearities arising from process and voltage variations. We test validation accuracy on the ImageNet dataset, and show that our IC suffers only 0.6% and 1.0% reduction in classification accuracy for Top-1 and Top-5 outputs, respectively. Our implementation results in a$\sim \boldsymbol {50}\times$reduction in latency and energy when compared to a recently published mixed-signal ASIC implementation, with similar power characteristics. Our approach provides layer partitioning and node sharing possibilities, which allow us to trade off latency, power, and area amongst each other.
Kyler R. Scott, Cheng-Yen Lee, Sunil P. Khatri, Sarma B. K. Vrudhula
DATE1
2022 A Flash-based Digital to Analog Converter for Low Power Applications
abstract
This paper presents a novel technique for digital to analog conversion, which uses flash transistors embedded in a traditional current-steering digital to analog converter (DAC) architecture. Our design utilizes flash transistors as programmable, tunable current sources. Our DAC achieves extremely low latency, area, and power, making it especially suited for the internet of things (IoT) and other applications with a highly constrained resource budget. In addition, the use of flash transistors allows a user to cancel errors due to process/voltage variations and chip aging. This tuning can be performed in the fabrication facility, or in-field. We perform a Monte-Carlo analysis to demonstrate the performance of our DAC in a real-world context. Compared to recent DACs intended for use in IoT devices (or for applications that are highly resource constrained), we significantly improve error metrics. We reduce chip area by 4.3× compared with the smallest DAC intended for IoT. Additionally, we improve throughput by 55 ×and energy per conversion by 33×, compared with the fastest and the most energy efficient DAC intended for IoT, respectively. The proposed 12-bit DAC achieves a throughput of 100 MS/s, an energy per conversion of 36.6 fJ. Based on Monte Carlo analysis, we report a maximum INL (DNL) of 1.242 LSB (0.757 LSB), an average INL (DNL) of 0.286 LSB (0.088 LSB), an ENOB of 11.80 bits, and an SFDR of 82.89 dB.
Kyler R. Scott, Sunil P. Khatri
ICCD1