Kevin Siu

dblp:220/9008 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5Software engineering, systems software and programming languages · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Hardware accelerators and domain-specific architectures · 88% Memory systems · 12%
Artificial intelligence
4 papers
Efficient and distributed learning · 100%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures
machine learning accelerator
1.032019
ShapeShifter: Enabling Fine-Grain Data Width Adaptation in Deep Learning · MICRO 2019
Diffy: a Déjà vu-Free Differential Deep Neural Network Accelerator · MICRO 2018
Loom: exploiting weight and activation precisions to accelerate convolutional neural networks · DAC 2018
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator
0.822019
Laconic deep learning inference acceleration · ISCA 2019
Bit-Tactical: A Software/Hardware Approach to Exploiting Value and Bit Sparsity in Neural Networks · ASPLOS 2019
Machine learning › Efficient and distributed learning
model compression
0.422018
Diffy: a Déjà vu-Free Differential Deep Neural Network Accelerator · MICRO 2018
Loom: exploiting weight and activation precisions to accelerate convolutional neural networks · DAC 2018
Machine learning › Efficient and distributed learning
inference acceleration
0.412019
Laconic deep learning inference acceleration · ISCA 2019
Hardware accelerators and domain-specific architectures › sparsity exploitation
bit-level sparsity exploitation
0.412019
Bit-Tactical: A Software/Hardware Approach to Exploiting Value and Bit Sparsity in Neural Networks · ASPLOS 2019
Memory systems
memory compression
0.412019
ShapeShifter: Enabling Fine-Grain Data Width Adaptation in Deep Learning · MICRO 2019
Machine learning › Efficient and distributed learning › model compression › neural network compression
activation compression
0.312018
Diffy: a Déjà vu-Free Differential Deep Neural Network Accelerator · MICRO 2018
Hardware accelerators and domain-specific architectures › machine learning accelerator › CNN accelerator
CNN inference accelerator
0.312018
Loom: exploiting weight and activation precisions to accelerate convolutional neural networks · DAC 2018
Hardware accelerators and domain-specific architectures › machine learning accelerator › DNN accelerator
precision-scalable accelerator
0.312018
Loom: exploiting weight and activation precisions to accelerate convolutional neural networks · DAC 2018
Machine learning › Efficient and distributed learning › model compression
quantization
0.222019
ShapeShifter: Enabling Fine-Grain Data Width Adaptation in Deep Learning · MICRO 2019
Loom: exploiting weight and activation precisions to accelerate convolutional neural networks · DAC 2018

Methods — techniques the papers use, named apart from their topics

fine-grain data width encoding · 0.8dynamic width selection · 0.8per-layer precision · 0.7bit-parallel multiply-accumulate · 0.7static scheduling · 0.4sparse shuffling network · 0.4
YearPublicationVenuePosition
2019 Bit-Tactical: A Software/Hardware Approach to Exploiting Value and Bit Sparsity in Neural Networks
abstract
Weight and activation sparsity can be leveraged in hardware to boost the performance and energy efficiency of Deep Neural Networks during inference. Fully capitalizing on sparsity requires re-scheduling and mapping the execution stream to deliver non-zero weight/activation pairs to multiplier units for maximal utilization and reuse. However, permitting arbitrary value re-scheduling in memory space and in time places a considerable burden on hardware to perform dynamic at-runtime routing and matching of values, and incurs significant energy inefficiencies. Bit-Tactical (TCL) is a neural network accelerator where the responsibility for exploiting weight sparsity is shared between a novel static scheduling middleware, and a co-designed hardware front-end with a lightweight sparse shuffling network comprising two (2- to 8-input) multiplexers per activation input. We empirically motivate two back-end designs chosen to target bit-sparsity in activations, rather than value-sparsity, with two benefits: a) we avoid handling the dynamically sparse whole-value activation stream, and b) we uncover more ineffectual work. TCL outperforms other state-of-the-art accelerators that target sparsity for weights and activations, the dynamic precision requirements of activations, or their bit-level sparsity for a variety of neural networks.
Alberto Delmas Lascorz, Patrick Judd, Dylan Malone Stuart, Zissis Poulos, Mostafa Mahmoud, Sayeh Sharify, Milos Nikolic 0002, Kevin Siu, Andreas Moshovos
ASPLOS8
2019 Laconic deep learning inference acceleration
abstract
We present a method for transparently identifying ineffectual computations during inference with Deep Learning models. Specifically, by decomposing multiplications down to the bit level, the amount of work needed by multiplications during inference can be potentially reduced by at least 40× across a wide selection of neural networks (8b and 16b). This method produces numerically identical results and does not affect overall accuracy. We present Laconic, a hardware accelerator that implements this approach to boost energy efficiency for inference with Deep Learning Networks. Laconic judiciously gives up some of the work reduction potential to yield a low-cost, simple, and energy efficient design that outperforms other state-of-the-art accelerators: an optimized DaDianNao-like design [13], Eyeriss [15], SCNN [71], Pragmatic [3], and BitFusion [83]. We study 16b, 8b, and 1b/2b fixed-point quantized models.
Sayeh Sharify, Alberto Delmas Lascorz, Mostafa Mahmoud, Milos Nikolic 0002, Kevin Siu, Dylan Malone Stuart, Zissis Poulos, Andreas Moshovos
ISCA5
2019 ShapeShifter: Enabling Fine-Grain Data Width Adaptation in Deep Learning
abstract
We show that selecting a data width for all values in Deep Neural Networks, quantized or not and even if that width is different per layer, amounts to worst-case design. Much shorter data widths can be used if we target the common case by adjusting the data type width at a much finer granularity. We propose ShapeShifter, where we group weights and activations and encode them using a width specific to each group and where typical group sizes vary from 16 to 256 values. The per group widths are selected statically for the weights and dynamically by hardware for the activations. We present two applications of ShapeShifter. In the first, that is applicable to any system, ShapeShifter reduces off- and on-chip storage and communication. This ShapeShifter-based memory compression is simple and low cost yet reduces off-chip traffic to 33% and 36% for 8-bit and 16-bit models respectively. This makes it possible to sustain higher performance for a given off-chip memory interface while also boosting energy efficiency. In the second application, we show how ShapeShifter can be implemented as a surgical extension over designs that exploit variable precision in time.
Alberto Delmas Lascorz, Sayeh Sharify, Isak Edo Vivancos, Dylan Malone Stuart, Omar Mohamed Awad, Patrick Judd, Mostafa Mahmoud, Milos Nikolic 0002, Kevin Siu, Zissis Poulos, Andreas Moshovos
MICRO9
2018 Loom: exploiting weight and activation precisions to accelerate convolutional neural networks
abstract
Loom (LM), a hardware inference accelerator for Convolutional Neural Networks (CNNs) is presented. In LM every bit of data precision that can be saved translates to proportional performance gains. For both weights and activations LM exploits profile-derived per layer precisions. However, at runtime LM further trims activation precisions at a much smaller than a layer granularity. On average, across several image classification CNNs and for a configuration that can perform the equivalent of 128 16b × 16b multiply-accumulate operations per cycle LM outperforms a state-of-the-art bit-parallel accelerator [3] by 3.19× without any loss in accuracy while being 2.59× more energy efficient. LM can trade-off accuracy for additional improvements in execution performance and energy efficiency and compares favorably to an accelerator that targeted only activation precisions.
Sayeh Sharify, Alberto Delmas Lascorz, Kevin Siu, Patrick Judd, Andreas Moshovos
DAC3
2018 Diffy: a Déjà vu-Free Differential Deep Neural Network Accelerator
abstract
We show that Deep Convolutional Neural Network (CNN) implementations of computational imaging tasks exhibit spatially correlated values. We exploit this correlation to reduce the amount of computation, communication, and storage needed to execute such CNNs by introducing Diffy, a hardware accelerator that performs Differential Convolution. Diffy stores, communicates, and processes the bulk of the activation values as deltas. Experiments show that, over five state-of-the-art CNN models and for HD resolution inputs, Diffy boosts the average performance by 7.1× over a baseline value-agnostic accelerator [1] and by 1.41× over a state-of-the-art accelerator that processes only the effectual content of the raw activation values [2]. Further, Diffy is respectively 1.83× and 1.36× more energy efficient when considering only the on-chip energy. However, Diffy requires 55% less on-chip storage and 2.5× less off-chip bandwidth compared to storing the raw values using profiled per-layer precisions [3]. Compared to using dynamic per group precisions [4], Diffy requires 32% less storage and 1.43× less off-chip memory bandwidth. More importantly, Diffy provides the performance necessary to achieve real-time processing of HD resolution images with practical configurations. Finally, Diffy is robust and can serve as a general CNN accelerator as it improves performance even for image classification models.
Mostafa Mahmoud, Kevin Siu, Andreas Moshovos
MICRO2