Soumendu Kumar Ghosh

dblp:160/3442 · also Soumendu Ghosh 0003 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0001-6776-1427ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1
YearPublicationVenuePosition
2026 COSMOS: Designing Energy-Efficient Context-Aware Multimodal Cognitive Systems
Arghadip Das, Yatharth Agarwal, Soumendu Kumar Ghosh, Arnab Raha, Vijay Raghunathan
ISLPED3
2025 GraNNite: Enabling High-Performance Execution of Graph Neural Networks on Resource-Constrained Neural Processing Units
abstract
Graph Neural Networks (GNNs) are crucial for learning and reasoning over graph-structured data, with applications in network analysis, recommendation systems, and speech analytics. Deploying them on edge devices, such as client PCs and laptops, enables real-time processing, enhances privacy, and reduces cloud dependency. For instance, GNNs can augment Retrieval-Augmented Generation (RAG) for Large Language Models (LLMs) and enable event-based vision tasks. However, irregular memory access, sparse graphs, and dynamic structures lead to high latency and energy consumption on resource-constrained devices. Modern edge processors combine CPUs, GPUs, and NPUs, where NPUs excel at data-parallel tasks but face challenges with irregular GNN computations. To address these gaps, we present GraNNite, the first hardware-aware framework tailored to optimize GNN deployment on commercial-off-the-shelf (COTS) state-of-the-art (SOTA) DNN accelerators using a systematic three-step methodology: (1) enabling GNN execution on NPUs, (2) optimizing performance, and (3) trading accuracy for further performance and energy efficiency gains. Towards that end, the first category includes techniques such as GraphSplit for workload distribution and StaGr for static graph aggregation, while GrAd and NodePad handle real-time updates for dynamic graphs. Next, performance improvement is acquired through techniques such as EffOp for control-heavy operations and GraSp for sparsity exploitation. For Graph Convolution layers, PreG, SymG, and CacheG reduce redundancy and memory transfers. The final class of techniques deals with quality vs efficiency tradeoffs – QuantGr applies INT8 quantization to lower memory usage and computation time, while GrAx1, GrAx2, and GrAx3 optimize graph attention, broadcast-add, and sample-and-aggregate (SAGE)-max aggregation for higher throughput with minimal quality loss. Experimental evaluations on Intel® Core™ Ultra Series 1 and 2 AI PCs demonstrate that GraNNite achieves speedups of 2.6× to 7.6× over default NPU mappings, with energy efficiency improvements up to 8.6× compared to CPUs and GPUs. Across various GNN models, GraNNite delivers up to 10.8× and 6.7× higher performance than CPUs and GPUs, respectively. Our code implementation is available at this link.
Arghadip Das, Shamik Kundu, Arnab Raha, Soumendu Kumar Ghosh, Deepak Mathaikutty, Vijay Raghunathan
IJCNN4
2025 Demo Abstract: ECO: Low Power Context-Aware Multimodal AI on NPUs
abstract
We present ECO, the first system enabling efficient multimodal AI deployment on commercial Neural Processing Units (NPUs) through context-aware sensor and compute optimizations. ECO introduces runtime-tunable, NPU-architecture-aware knobs—approximate interpolation, quantization, and model scaling—that adapt to system conditions such as energy availability and sensor reliability. Deployed on an Intel Core Ultra Series 2 NPU with RGB and LiDAR inputs for a semantic segmentation application, ECO achieves up to 4.9× performance and 11.3× energy-efficiency improvement over CPU. Compared to systems lacking runtime context adaptability, ECO preserves higher segmentation quality (48.1 mean IoU in %, referred to as IoU hereafter) vs. 37.9 IoU under energy constraints and restores accuracy from 30.6 IoU to 40.0 IoU in sensor failure scenarios. The demo video and the ECO codebase are available at https://github.com/arghadippurdue/ECO%5FDemo.
Arghadip Das, Yatharth Agarwal, Soumendu Kumar Ghosh, Arnab Raha, Vijay Raghunathan
ISLPED3
2024 SwiSS: Switchable Single-Sided Sparsity-based DNN Accelerators
abstract
Deep Neural Networks (DNNs) exhibit sparsity in both activation and weight tensors, but certain layers have higher weight sparsity, while others have higher activation sparsity. This challenges the conventional approach of fixing sparsity acceleration to either weights or activations alone. Conversely, harnessing both-sided sparsity necessitates complex design logic for identifying participating non-zero weights and activation pairs during a multiply-accumulate operation, leading to a significant impact on energy efficiency and area overhead in the edge accelerator. In this paper, we, for the first time, exploit the unbalanced sparsity in DNNs to propose the concept of dynamically Switchable Single-sided Sparsity, SwiSS, to improve energy efficiency in edge DNN accelerators. Through a novel self-adaptive dynamic sparsity selection algorithm, SwiSS can determine whether to enable one-sided weight or one-sided activation sparsity for a sparsity-enabled DNN accelerator. This capability allows SwiSS to dynamically exploit both sides of sparsity while maximizing the associated power and area benefits in the accelerator. Evaluation on state-of-the-art network-dataset configurations conducted on FlexNN [12] accelerator architecture demonstrates that SwiSS yields up to 30.76% and 8.29% improvements in power and area overheads, respectively (which translates to 1.42X and 1.08X improvement in TOPS/W and TOPS/mm2, respectively), compared to a combined two-sided sparsity scenario, with a negligible drop in sparsity acceleration.
Shamik Kundu, Soumendu Kumar Ghosh, Arnab Raha, Deepak Mathaikutty
ISLPED2
2024 Toward Energy-Efficient Collaborative Inference Using Multisystem Approximations
abstract
Cooperative inference applications have seen considerable potential with distributed deep neural networks (DDNNs). One use for DDNNs is the classification of 3-D objects from a set of 2-D images or views. This approach is also known as multiview convolutional neural networks (MVCNNs). However, due to the intensive computational demands, substantial communication overhead, high-inference delay, and energy limits, it is difficult to deploy MVCNN on resource-constrained edge devices. This article proposes for the first time the concept of distributed approximate systems (DRAX), which employs a multidevice approach to approximate computing and uses synergistic approximations of various edge computing systems to enable energy-efficient collaborative DDNN inference.DRAXperforms a significance-aware approximation of multiple nodes and prunes the large design space using the nonuniform contribution of various perspectives/views to the final inference to achieve optimal quality-energy tradeoff. In addition, we also propose a novel remaining energy-aware heuristic, which dynamically chooses the approximation degree based on the user-provided quality bounds and further increases the system lifetime. The experimental results obtained from a prototype of a 12-view 3-D object classification system implemented on an Intel Stratix IV FPGA development board demonstrate substantial energy savings ($2.6 \times$to$8\times$) for minimal (<1%) application-level quality loss.
Arghadip Das, Soumendu Kumar Ghosh, Arnab Raha, Vijay Raghunathan
IEEE Internet Things J.2
2024 PArtNNer: Platform-Agnostic Adaptive Edge-Cloud DNN Partitioning for Minimizing End-to-End Latency
abstract
The last decade has seen the emergence of Deep Neural Networks (DNNs) as the de facto algorithm for various computer vision applications. In intelligent edge devices, sensor data streams acquired by the device are processed by a DNN application running on either the edge device itself or in the cloud. However, “edge-only” and “cloud-only” execution of State-of-the-Art DNNs may not meet an application’s latency requirements due to the limited compute, memory, and energy resources in edge devices, dynamically varying bandwidth of edge-cloud connectivity networks, and temporal variations in the computational load of cloud servers. This work investigates distributed (partitioned) inference across edge devices (mobile/end device) and cloud servers to minimize end-to-end DNN inference latency. We study the impact of temporally varying operating conditions and the underlying compute and communication architecture on the decision of whether to run the inference solely on the edge, entirely in the cloud, or by partitioning the DNN model execution among the two. Leveraging the insights gained from this study and the wide variation in the capabilities of various edge platforms that run DNN inference, we propose PArtNNer , a platform-agnostic adaptive DNN partitioning algorithm that finds the optimal partitioning point in DNNs to minimize inference latency. PArtNNer can adapt to dynamic variations in communication bandwidth and cloud server load without requiring pre-characterization of underlying platforms. Experimental results for six image classification and object detection DNNs on a set of five commercial off-the-shelf compute platforms and three communication standards indicate that PArtNNer results in 10.2× and 3.2× (on average) and up to 21.1× and 6.7× improvements in end-to-end inference latency compared to execution of the DNN entirely on the edge device or entirely on a cloud server, respectively. Compared to pre-characterization-based partitioning approaches, PArtNNer converges to the optimal partitioning point 17.6× faster.
Soumendu Kumar Ghosh, Arnab Raha, Vijay Raghunathan, Anand Raghunathan
ACM Trans. Embed. Comput. Syst.1
2023 Energy-Efficient Approximate Edge Inference Systems
abstract
The rapid proliferation of the Internet of Things and the dramatic resurgence of artificial intelligence based application workloads have led to immense interest in performing inference on energy-constrained edge devices. Approximate computing (a design paradigm that trades off a small degradation in application quality for disproportionate energy savings) is a promising technique to enable energy-efficient inference at the edge. This article introduces the concept of an approximate edge inference system ( AxIS ) and proposes a systematic methodology to perform joint approximations between different subsystems in a deep neural network (DNN)-based edge inference system, leading to significant energy benefits compared to approximating individual subsystems in isolation. We use a smart camera system that executes various DNN-based image classification and object detection applications to illustrate how the sensor, memory, compute, and communication subsystems can all be approximated synergistically. We demonstrate our proposed methodology using two variants of a smart camera system: (a) Cam Edge , where the DNN is executed locally on the edge device, and (b) Cam Cloud , where the edge device sends the captured image to a remote cloud server that executes the DNN. We have prototyped such an approximate inference system using an Intel Stratix IV GX-based Terasic TR4-230 FPGA development board. Experimental results obtained using six large DNNs and four compact DNNs running image classification applications demonstrate significant energy savings (≈ 1.6× -4.7× for large DNNs and ≈ 1.5× -3.6× for small DNNs), for minimal (<1%) loss in application-level quality. Furthermore, results using four object detection DNNs exhibit energy savings of ≈ 1.5× -5.2× for similar quality loss. Compared to approximating a single subsystem in isolation, AxIS achieves 1.05× -3.25× gains in energy savings for image classification and 1.35× -4.2× gains for object detection on average, for minimal (<1%) application-level quality loss.
Soumendu Kumar Ghosh, Arnab Raha, Vijay Raghunathan
ACM Trans. Embed. Comput. Syst.1
2021 Special Session: Approximate TinyML Systems: Full System Approximations for Extreme Energy-Efficiency in Intelligent Edge Devices
abstract
Approximate computing (AxC) has advanced from being an emerging design paradigm to becoming one of the most popular and effective methods of energy optimization for applications in the domains of computer vision, image/video processing, data mining, analytics, and search. The simultaneous rise of artificial intelligence (AI) has provided an additional thrust to the adoption of various AxC techniques in intelligent edge platforms where energy-efficiency is not only desirable but necessary. In spite of the big rise in interest for AxC, the adoption of approximate hardware has mostly been limited to only one component of the system (usually the processing subsystem) which often contributes only a fraction of the overall system-level power. A full system approach to AxC enables us to extend approximations to other subsystems, such as the memory, sensor, and communications subsystems. This paper presents the foundational concepts of an approximate TinyML system that applies approximations synergistically to multiple subsystems in an edge inference device. These approximations are applied intelligently to significantly reduce energy while incurring a negligible loss in application-level quality. We demonstrate multiple versions of an approximate smart camera system that can execute state-of-the-art deep neural networks (DNNs) while consuming only a fraction of the total energy in a typical system.
Arnab Raha, Soumendu Kumar Ghosh, Debabrata Mohapatra, Deepak Mathaikutty, Raymond Sung, Cormac Brick, Vijay Raghunathan
ICCD2
2020 Approximate inference systems (AxIS): end-to-end approximations for energy-efficient inference at the edge
abstract
The rapid proliferation of the Internet-of-Things (IoT) and the dramatic resurgence of artificial intelligence (AI) based application workloads has led to immense interest in performing inference on energy-constrained edge devices. Approximate computing (a design paradigm that yields large energy savings at the cost of a small degradation in application quality) is a promising technique to enable energy-efficient inference at the edge. This paper introduces the concept of an approximate inference system (AxIS) and proposes a systematic methodology to perform joint approximations across different subsystems in a deep neural network-based inference system, leading to significant energy benefits compared to approximating individual subsystems in isolation. We use a smart camera system that executes various convolutional neural network (CNN) based image recognition applications to illustrate how the sensor, memory, compute, and communication subsystems can all be approximated synergistically. We demonstrate our proposed methodology using two variants of a smart camera system: (a) Camedge, where the CNN executes locally on the edge device, and (b) Camcloud, where the edge device sends the captured image to a remote cloud server that executes the CNN. We have prototyped such an approximate inference system using an Altera Stratix IV GX-based Terasic TR4-230 FPGA development board. Experimental results obtained using six CNNs demonstrate significant energy savings (around 1.7× for Camedge and 3.5× for Camcloud) for minimal (< 1%) loss in application quality. Compared to approximating a single subsystem in isolation, AxIS achieves additional energy benefits of 1.6×--1.7× (Camedge) and 1.4×--3.4× (Camcloud) on average for minimal application-level quality loss.
Soumendu Kumar Ghosh, Arnab Raha, Vijay Raghunathan
ISLPED1
2017 Railway bridge health monitoring system using smart wireless sensor network: demo
abstract
Railway Bridge Health Monitoring is of prime importance as damages in bridges can lead to heavy casualties. Hence, monitoring is necessary to provide safety services to the millions of people around the world. Recently, wireless sensor network (WSN) has come up as a promising technology for health monitoring. However, memory and energy constraints of WSN, and extracting intelligible information from signals obtained from complex bridge structures are the technical challenges in this domain. In this paper, we give a brief overview of a novel railway bridge health monitoring system using smart WSN.
Sathvik Dev Velagandula, Nirjhar Dhang, Raja Datta, Soumendu Kumar Ghosh, Maroju Suman
WISEC4