Imlijungla Longchar

dblp:261/6377 · DBLP profile ↗
← Back
3ranked-venue papers
3as first author
3since 2021 · last 2026
0009-0009-8063-441XORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 first-author · 3 since 2021
YearPublicationVenuePosition
2026 SpALEn: Sparsity Aware Load Balancing Inference Engine for Neural Network
abstract
In the field of computer science, Convolutional Neural Network (CNN) algorithms are a crucial tool that contributes to the advancement of Computer Vision. CNNs are composed of an enormous number of multiplication and addition operations performed on the input data to calculate the probability and predict the output. In multiplication, if any of the operands is zero-valued, then it is irrelevant in that particular output, and hence, these computations can be omitted to avoid unnecessary computation. In this study, we propose an architecture that performs the computations through parallel Processing Elements(PEs) and is also capable of skipping the ineffectual zero-valued computation to improve PE utilisation. Our proposed work, SpALEn, adopts the channel-first dataflow and is designed to perform the inference function with zero-skipping for enhanced performance. Moreover, due to the skipped computation, load imbalance occurs as the computation workload varies among the PEs. A dynamic logic is designed to mitigate this and ensure the hardware resources are utilised thoroughly. SpALEn achieves a speedup of 12 \(\times\) in comparison to a dense architecture.
Imlijungla Longchar, Hemangee K. Kapoor
ACM J. Emerg. Technol. Comput. Syst.1
2023 ADaMaT: Towards an Adaptive Dataflow for Maximising Throughput in Neural Network Inference
abstract
With the development of research in hardware for Convolutional Neural Network(CNNs) Algorithms, it becomes crucial to examine the different aspects of hardware design. CNNs are mainly used in computer vision applications, and translating these algorithms into hardware calls for adopting appropriate dataflow to improve the utilisation of hardware resources resulting in higher throughput. In particular, the inference task at each neuron position can be assigned to a compute unit in the hardware accelerator, and several such neuron positions can be completed in parallel. We observe that adopting a static dataflow for an architecture can result in the under-utilisation of resources because of the different dimensions of the data in the network. The motivation of this paper is built upon the need for adaptive dataflow for the design to improve the multiply-and-accumulate (MAC) utilisation in CNNs. We propose a method, ADaMaT, which adapts the dataflow at runtime by appropriately assigning tasks to the MAC units depending on the dimensions of the layers instead of a pre-determined assignment. The adaptive assignment tries to maximise the MAC utilisation and improve the throughput. We have performed a comparative analysis among different static dataflows and our proposed ADaMaT dataflow.
Imlijungla Longchar, Hemangee K. Kapoor
VLSI-SoC1
2022 ZaLoBI: Zero avoiding Load Balanced Inference accelerator
abstract
Convolutional neural networks are prevalent machine learning tools used in computer vision. Their ubiquitous use and high compute requirement have given rise to the design and development of accelerators for the same. Among several approaches to improve the performance of these accelerators, exploiting data sparsity has become very popular. Along similar lines, this paper proposes a design that skips the computation of zero-valued data operands and achieves better speedup. The savings in zero-valued computations also results in energy savings. The proposed accelerator exploits two levels of data parallelism to distribute work across multiple processing elements (PEs). The random distribution of zero values results in certain PEs getting idle due to the skipping of computations, thus creating load imbalance in the system. To address this issue, we extend our contribution in performing load balancing by dynamically scheduling tasks to the idle PEs. Our zero avoiding load-balanced accelerator (ZaLoBI) achieves around 76% and 5.57% speedup over the respective baselines and also outperforms the state-of-the-art works while saving energy.
Imlijungla Longchar, Palash Das 0001, Hemangee K. Kapoor
VLSI-SoC1