Neelesh Gupta

dblp:118/7101 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
4since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 Net2Tab: Tabularizing neural networks with applications to data prefetching
Pengmiao Zhang, Neelesh Gupta, Rajgopal Kannan, Viktor Prasanna 0001
J. Parallel Distributed Comput.2
2025 Context-Driven Performance Modeling for Causal Inference Operators on Neural Processing Units
abstract
The proliferation of large language models has driven demand for long-context inference on resourceconstrained edge platforms. However, deploying these models on Neural Processing Units (NPUs) presents significant challenges due to architectural mismatch: the quadratic complexity of standard attention conflicts with NPU memory and compute patterns. This paper presents a comprehensive performance analysis of causal inference operators on a modern NPU, benchmarking quadratic attention against sub-quadratic alternatives including structured state-space models and causal convolutions. Our analysis reveals a spectrum of critical bottlenecks: quadratic attention becomes severely memory-bound with catastrophic cache inefficiency, while sub-quadratic variants span from computebound on programmable vector cores to memory-bound by data movement. These findings provide essential insights for codesigning hardware-aware models and optimization strategies to enable efficient long-context inference on edge platforms.
Neelesh Gupta, Rakshith Jayanth, Dhruv Parikh, Viktor Prasanna 0001
HiPC1
2024 TabConv: Low-Computation CNN Inference via Table Lookups
abstract
Convolutional Neural Networks (CNNs) have demonstrated remarkable ability throughout the field of computer vision. However, CNN inference requires a large number of arithmetic operations making them expensive to deploy in hardware. Current approaches alleviate this issue by developing hardware-supported, algorithmic processes to simplify spatial convolution functions. However, these methods still heavily rely on matrix multiplication, leading to significant computational overhead. To bridge the gap between hardware, algorithmic acceleration, and approximate matrix multiplication, we propose TabConv, a novel, table-based approximation for convolution to significantly reduce arithmetic operations during inference. Additionally, we introduce a priority masking technique based on cosine similarity to select layers for table-based approximation, thereby maintaining the model performance. We evaluate our approach on popular CNNs: ResNet-18, ResNet-34, and NetworkIn-Network (NIN). TabConv preserves over 93% of the original model's performance while reducing arithmetic operations by 36.5%, 25.8%, and 99.4% for ResNet-18 on CIFAR-10, CIFAR-100, and MNIST, respectively, 35.6% and 99.3% for ResNet-34 on CIFAR-10 and MNIST, and 98.9% for NIN on MNIST, achieving low-computation inference.
Neelesh Gupta, Narayanan Kannan, Pengmiao Zhang, Viktor Prasanna 0001
CF1
2024 Attention, Distillation, and Tabularization: Towards Practical Neural Network-Based Prefetching
abstract
Attention-based Neural Networks (NN) have demonstrated their effectiveness in accurate memory access prediction, an essential step in data prefetching. However, the substantial computational overheads associated with these models result in high inference latency, limiting their feasibility as practical prefetchers. To close the gap, we propose a new approach based on tabularization that significantly reduces model complexity and inference latency without sacrificing prediction accuracy. Our novel tabularization methodology takes input as a distilled, yet highly accurate attention-based model for memory access prediction and efficiently converts its expensive matrix multiplications into a hierarchy of fast table lookups. As an exemplar of the above approach, we develop DART, a prefetcher comprised of a simple hierarchy of tables. With a modest 0.09 drop in F1-score, DART reduces 99.99% of arithmetic operations from the original attention-based model and 91.83% from the distilled model. DART accelerates the large model inference by 170× and the distilled model by 9.4×. DART has comparable latency and storage costs as state-of-the-art rule-based prefetcher BO but surpasses it by 6.1% in IPC improvement. DART outperforms state-of-the-art NN-based prefetchers TransFetch by 33.1% and Voyager by 37.2% in terms of IPC improvement, primarily due to its low prefetching latency.
Pengmiao Zhang, Neelesh Gupta, Rajgopal Kannan, Viktor Prasanna 0001
IPDPS2
2009 E2-SCAN: an extended credit strategy-based energy-efficient security scheme for wireless ad hoc networks
abstract
Utilising the battery life and the limited bandwidth available in mobile ad hoc networks (MANETs) in the most efficient manner is an important issue, along with providing security at the network layer. The authors propose, design and describe E2-SCAN, an energy-efficient network layered security solution for MANETs, which protects both routing and packet forwarding functionalities in the context of the on demand distance vector protocol. E2-SCAN is an advanced approach that builds on and improves upon some of the state-of-the-art results available in the literature. The proposed E2-SCAN algorithm protects the routing and data forwarding operations through the same reactive approach, as is provided by the SCAN algorithm. It also enhances the security of the network by detecting and reacting to the malicious nodes. In E2-SCAN, the immediate one-hop neighbour nodes collaboratively monitor. E2-SCAN adopts a modified novel credit strategy to decrease its overhead as the time evolves. Through both analysis and simulation results, the authors demonstrate the effectiveness of E2-SCAN over SCAN in a hostile environment.
Sanjay K. Dhurandher, Sudip Misra, Sombir Ahlawat, Neelesh Gupta, Nitesh Gupta
IET Commun.4