Alina Vasilciuc

dblp:301/6378 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Performance modeling and evaluation · 67% Hardware accelerators and domain-specific architectures · 33%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
model compression
0.512021
Evaluation of Optimized CNNs on Heterogeneous Accelerators Using a Novel Benchmarking Approach · IEEE Trans. Computers 2021
Machine learning › Efficient and distributed learning › model compression
pruning and quantization
0.512021
Evaluation of Optimized CNNs on Heterogeneous Accelerators Using a Novel Benchmarking Approach · IEEE Trans. Computers 2021
Performance modeling and evaluation › benchmarking › computer architecture benchmarking
accelerator benchmarking
0.512021
Evaluation of Optimized CNNs on Heterogeneous Accelerators Using a Novel Benchmarking Approach · IEEE Trans. Computers 2021
Performance modeling and evaluation
benchmarking
0.512021
Evaluation of Optimized CNNs on Heterogeneous Accelerators Using a Novel Benchmarking Approach · IEEE Trans. Computers 2021
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator
0.512021
Evaluation of Optimized CNNs on Heterogeneous Accelerators Using a Novel Benchmarking Approach · IEEE Trans. Computers 2021

Methods — techniques the papers use, named apart from their topics

quantization · 1.0pruning · 1.0benchmarking · 1.0
YearPublicationVenuePosition
2021 Evaluation of Optimized CNNs on Heterogeneous Accelerators Using a Novel Benchmarking Approach
abstract
Numerous algorithmic optimization techniques have been proposed to alleviate the computational complexity of convolutional neural networks. Given the broad selection of AI accelerators, it is not obvious which approach benefits from which optimization most. The design space includes a large number of deployment settings (batch sizes, power modes, etc.) and unclear measurement methods. This research provides clarity into this design space, leveraging a novel benchmarking approach. We provide a theoretical evaluation of different CNNs and hardware platforms, focusing on understanding the impact of pruning and quantization as primary optimization techniques. We benchmark across a spectrum of FPGA, GPU, TPU, and VLIW processors for systematically pruned and quantized neural networks (ResNet50, GoogLeNetv1, MobileNetv1, a VGG derivative, a multilayer perceptron) over many deployment options, considering power, latency, and throughput at a specific accuracy. Our findings show that channel pruning is most effective and works across most hardware platforms, with speedups directly correlated to the reduction in compute load, while FPGAs benefit the most from quantization. Pruning and quantization are orthogonal, and yield optimal design points when combined. Further in-depth results can be found at our web portal, where we share all experimental data, provide data analytics, and invite the community to contribute.
Michaela Blott, Nicholas J. Fraser, Giulio Gambardella, Lisa Halder, Johannes Kath, Zachary Neveu, Yaman Umuroglu, Alina Vasilciuc, Miriam Leeser, Linda Doyle
IEEE Trans. Computers8