Ankur Deshwal

dblp:156/8639 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
3since 2021 · last 2021
0000-0002-3046-1292ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2021 An Automated Approach to Accelerate DNNs on Edge Devices
abstract
Deployment of Deep Neural Networks (DNNs) on edge devices can significantly increase the utility of DNNs for a variety of applications. However executing DNN models on the edge device is still a major challenge, as heavy computation and memory bandwidth requirements of such models limit their adoption. Employing a highly optimized code for DNN model execution can easily enable many more use-cases than currently possible. However, current strategies are still based on manual optimization for efficient resource utilization. This is not only cumbersome but also requires a high level of expert intervention in the rapidly changing DNN Model landscape. In this work, we provide an automated way of optimizing Convolutional Neural Network (CNN) models using Deep Reinforcement Learning (DRL) algorithm. The experiments with our DRL technique demonstrate 1.85×, 1.58×, 1.64× speedup in execution time for MobileNetV1, MobileNetV2 and Efficientnet-lite0 CNN models respectively on Mobile CPU devices.
Ujjawal Chugh, Arnab Mitra 0004, Ankur Deshwal, N. P. Swaroop, Aditi Saluja, Joonho Song
ISCAS3
2021 Transport Triggered near Memory Accelerator for Deep Learning
abstract
As throughput of neural network accelerator datapaths have grown, memory has consistently fallen behind. Although attempts have been made to improve performance of neural networks via approaches such as batching, several layers often starve for memory bandwidth when used for tasks such as online inferencing. In order to mitigate memory bandwidth limitations, we propose a near memory accelerator for mobile devices based on Transport-Triggered Architecture (TTA) and evaluate its performance benefits compared to the existing approaches. Through experiments we demonstrate that our proposed accelerator achieves up to 4.3× speedup with respect to a non-near memory accelerator and the proposed TTA data-path achieves area and energy efficiency close to fixed-function datapaths while offering programmability similar to VLIW cores.
Kavitha T. Madhu, Saptarsi Das, Abhishek Tyagi, Ankur Deshwal, Joonho Song
ISCAS4
2021 Autotuning LSTM for Accelerated Execution on Edge
abstract
Deployment of Deep Neural Networks (DNNs) on edge devices is highly desirable to address user privacy concerns and minimize the turnaround time of AI applications. However, the execution of DNN models on a battery-operated device requires a highly optimized implementation specific to the target hardware. Moreover, as different layers of a DNN exhibit distinct computation and memory characteristics, it is imperative to optimize each layer separately. This is in contrast to the widely deployed library-based approach where all the configurations of DNN operations share the same implementation. In this paper, we address this issue by auto-tuning the implementation of Long Short Term Memory (LSTM) operations which are widely used in sequence based AI applications. To exhaustively search through the space of optimizations and its parameters, we develop a high-level autotuning framework based on Halide. We use grid search to find the parameters that lead to minimum runtime and further present TPE based search method to find the near-optimal runtime in a limited number of trials. We observe 2.2× -3.1× speedup in execution time for LSTM layers used in widely deployed GNMT and DeepSpeech2 models.
Aditi Saluja, Arnab Mitra 0004, Ankur Deshwal, Kavitha T. Madhu, Ujjawal Chugh, Joonho Song
ISCAS3
2020 Sparse CNN Architecture Search (Scas)
abstract
Advent of deep neural networks has revolutionized Computer Vision. However, designing of such models with high accuracy and low computation requirements is a difficult task and needs extensive human expertise. Recent advances in Neural Architecture Search use various methods like Deep Reinforcement Learning, Evolutionary methods, Gradient Descent, HyperNetworks etc. to automatically generate neural networks with high level of accuracy. However, large size of such generated models limit their practical use. Recent findings about lottery ticket hypothesis suggest the existence of sparse subnetworks (winning tickets) which can reach the accuracy comparable to that of original dense network. In this paper, we present a method for leveraging redundancies inherent to deep Convolutional Neural Networks (CNN) to guide the generation of sparse CNN models (to find the architectures with winning tickets) without significant loss in accuracy. We evaluate our proposed method with different NAS methods on CFAR-10,CIFAR-100 and MNIST datasets. Our results show a reduction ranging from $2\times \mathrm{to} 12\times$ in terms of model size and $2\times \mathrm{to}19\times$ in terms of number of MAC operations with less than 1% drop in accuracy.
Yeshwanth V, Ankur Deshwal, Sundeep Krishnadasan, Joonho Song
ICME2
2020 A Systolic Dataflow Based Accelerator for CNNs
abstract
Modern Artificial Intelligence (AI) systems deploy Convolution Neural Networks (CNN) as they offer very high accuracy. Computational complexity of CNNs necessitates hardware acceleration, especially in mobile phones and other hand held devices due to stringent power and area budget. In this paper, we propose a systolic dataflow based accelerator architecture with the aim of improving energy efficiency as well as scalability. We demonstrate that a prototype of our proposed accelerator comprising 4096 Multiply-Accumulate (MAC) units is capable of achieving 8.89-9.54 TOPs/W power efficiency when executing a number of state of the art CNN models. We also observe that the throughput of the accelerator scales almost linearly with power dissipated. This is attributed to low power consumption due to minimal non-compute overhead. Our proposed accelerator outperforms state of the art accelerators power efficiency by factors of 2.6 to 3.8× while executing various layers of Inception V3 model.
Saptarsi Das, Arnab Roy 0010, Kiran Kolar Chandrasekharan, Ankur Deshwal, Sehwan Lee
ISCAS4