Arghadip Das

dblp:291/4716 · DBLP profile ↗
← Back
6ranked-venue papers
5as first author
6since 2021 · last 2026
0000-0001-6043-4685ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 COSMOS: Designing Energy-Efficient Context-Aware Multimodal Cognitive Systems
Arghadip Das, Yatharth Agarwal, Soumendu Kumar Ghosh, Arnab Raha, Vijay Raghunathan
ISLPED1
2025 GraNNite: Enabling High-Performance Execution of Graph Neural Networks on Resource-Constrained Neural Processing Units
abstract
Graph Neural Networks (GNNs) are crucial for learning and reasoning over graph-structured data, with applications in network analysis, recommendation systems, and speech analytics. Deploying them on edge devices, such as client PCs and laptops, enables real-time processing, enhances privacy, and reduces cloud dependency. For instance, GNNs can augment Retrieval-Augmented Generation (RAG) for Large Language Models (LLMs) and enable event-based vision tasks. However, irregular memory access, sparse graphs, and dynamic structures lead to high latency and energy consumption on resource-constrained devices. Modern edge processors combine CPUs, GPUs, and NPUs, where NPUs excel at data-parallel tasks but face challenges with irregular GNN computations. To address these gaps, we present GraNNite, the first hardware-aware framework tailored to optimize GNN deployment on commercial-off-the-shelf (COTS) state-of-the-art (SOTA) DNN accelerators using a systematic three-step methodology: (1) enabling GNN execution on NPUs, (2) optimizing performance, and (3) trading accuracy for further performance and energy efficiency gains. Towards that end, the first category includes techniques such as GraphSplit for workload distribution and StaGr for static graph aggregation, while GrAd and NodePad handle real-time updates for dynamic graphs. Next, performance improvement is acquired through techniques such as EffOp for control-heavy operations and GraSp for sparsity exploitation. For Graph Convolution layers, PreG, SymG, and CacheG reduce redundancy and memory transfers. The final class of techniques deals with quality vs efficiency tradeoffs – QuantGr applies INT8 quantization to lower memory usage and computation time, while GrAx1, GrAx2, and GrAx3 optimize graph attention, broadcast-add, and sample-and-aggregate (SAGE)-max aggregation for higher throughput with minimal quality loss. Experimental evaluations on Intel® Core™ Ultra Series 1 and 2 AI PCs demonstrate that GraNNite achieves speedups of 2.6× to 7.6× over default NPU mappings, with energy efficiency improvements up to 8.6× compared to CPUs and GPUs. Across various GNN models, GraNNite delivers up to 10.8× and 6.7× higher performance than CPUs and GPUs, respectively. Our code implementation is available at this link.
Arghadip Das, Shamik Kundu, Arnab Raha, Soumendu Kumar Ghosh, Deepak Mathaikutty, Vijay Raghunathan
IJCNN1
2025 Demo Abstract: ECO: Low Power Context-Aware Multimodal AI on NPUs
abstract
We present ECO, the first system enabling efficient multimodal AI deployment on commercial Neural Processing Units (NPUs) through context-aware sensor and compute optimizations. ECO introduces runtime-tunable, NPU-architecture-aware knobs—approximate interpolation, quantization, and model scaling—that adapt to system conditions such as energy availability and sensor reliability. Deployed on an Intel Core Ultra Series 2 NPU with RGB and LiDAR inputs for a semantic segmentation application, ECO achieves up to 4.9× performance and 11.3× energy-efficiency improvement over CPU. Compared to systems lacking runtime context adaptability, ECO preserves higher segmentation quality (48.1 mean IoU in %, referred to as IoU hereafter) vs. 37.9 IoU under energy constraints and restores accuracy from 30.6 IoU to 40.0 IoU in sensor failure scenarios. The demo video and the ECO codebase are available at https://github.com/arghadippurdue/ECO%5FDemo.
Arghadip Das, Yatharth Agarwal, Soumendu Kumar Ghosh, Arnab Raha, Vijay Raghunathan
ISLPED1
2025 Demo Abstract: A Low-Power Real-Time Hardware Accelerator for Edge Detection Using Stochastic Computing
abstract
We present a low-power, stochastic computing-based method for real-time video edge detection. Traditional Sobel-based pipelines are often resource-intensive and consume substantial power. In this work, we simplify the Sobel operator within a stochastic computing framework to achieve significant reductions in hardware complexity and energy consumption. We implement the proposed design on a Basys 3 FPGA interfaced with an OV7670 camera, demonstrating real-time performance. Experimental results show up to 17% energy savings, 84% reduction in LUT utilization, and a 68% decrease in RAM storage and a substantial reduction in hardware footprint compared to a traditional implementation. Demo video link: https://github.com/arghadippurdue/StoBelDemo.
Priyajit Ghosh, Rajarshi Mukherjee, Auro Anand Saha, Sutirtha Naha, Arghadip Das, Arnab Raha, Mrinal K. Naskar
ISLPED5
2024 Toward Energy-Efficient Collaborative Inference Using Multisystem Approximations
abstract
Cooperative inference applications have seen considerable potential with distributed deep neural networks (DDNNs). One use for DDNNs is the classification of 3-D objects from a set of 2-D images or views. This approach is also known as multiview convolutional neural networks (MVCNNs). However, due to the intensive computational demands, substantial communication overhead, high-inference delay, and energy limits, it is difficult to deploy MVCNN on resource-constrained edge devices. This article proposes for the first time the concept of distributed approximate systems (DRAX), which employs a multidevice approach to approximate computing and uses synergistic approximations of various edge computing systems to enable energy-efficient collaborative DDNN inference.DRAXperforms a significance-aware approximation of multiple nodes and prunes the large design space using the nonuniform contribution of various perspectives/views to the final inference to achieve optimal quality-energy tradeoff. In addition, we also propose a novel remaining energy-aware heuristic, which dynamically chooses the approximation degree based on the user-provided quality bounds and further increases the system lifetime. The experimental results obtained from a prototype of a 12-view 3-D object classification system implemented on an Intel Stratix IV FPGA development board demonstrate substantial energy savings ($2.6 \times$to$8\times$) for minimal (<1%) application-level quality loss.
Arghadip Das, Soumendu Kumar Ghosh, Arnab Raha, Vijay Raghunathan
IEEE Internet Things J.1
2023 HIPEDAP: Energy-Efficient Hardware Accelerators for Hidden Periodicity Detection
abstract
Hidden periodicity detection (HPD) forms the basis of various emerging and complex applications such as detecting tandem repeats in DNA, absence seizure detection in EEG signals,etc.. The solutions to the period estimation problem were not satisfactorily accurate until Ramanujan sums (RS) were used to explore the periodic decomposition of signals. Its use in hidden periodicity detection was streamlined to form Ramanujan Filter Bank (RFB), but its usage in the applications proved to be computationally expensive. This paper proposes HIPEDAP, an efficient set of hardware accelerators for hidden periodicity detection applications using Ramanujan Filter Bank. HIPEDAPis developed by proposing several incrementally efficient microarchitectures, from Arch-A to E targeting improvements in different aspects of the design such as area, power, and performance. Further, the inherent error resilience exhibited by HPD applications enables us to propose an approximate architecture Arch-F, that synergistically applies multiple approximation techniques such as approximate adder, multiplier, and precision scaling on top of Arch-E, resulting in significant performance and energy benefits. Experimental results obtained after synthesizing the microarchitectures on 45 nm technology demonstrate that the optimized Arch-E design is able to achieve 4.7X, 8.2X, and 1.7X improvements in terms of area, frequency of execution, and power, respectively. Further, Arch-F demonstrate additional power savings of 14.6% on average (max 32.2%) over Arch-E for almost no loss in application-level quality. Finally, across a suite of practical applications, HIPEDAPexhibited a speed-up in the range of$5.1 \;{\times }\; 10^{2}$X to$3.2 \;{\times }\; 10^{4}$compared to its software implementations.
Arghadip Das, Chandrachur Majumder, Debaprasad De, Arnab Raha, Mrinal K. Naskar
IEEE Trans. Computers1