Cyan Subhra Mishra

dblp:263/7470 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
11since 2021 · last 2025
0000-0002-5532-9757ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 3 first-author · 9 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Salient Store: Enabling Smart Storage for Continuous Learning Edge Servers
abstract
Global IP-video traffic is projected to exceed 4.8 Zettabytes annually by 2025, driven largely by edge applications like autonomous driving and urban mobility that generate hundreds of terabytes per device. While significant research has optimized edge inference and training architectures, the critical challenges of data archival and storage have been largely overlooked. Our analysis reveals that archival operations, not inference, dominate system resources, consuming up to 80% of memory bandwidth and one-third of CPU cycles in edge servers. We present Salient Store, a novel computational storage architecture that transforms passive storage devices into active participants in the video analytics pipeline. Salient Store integrates motion-aware layered neural compression directly within storage FPGAs, reusing feature maps from inference to eliminate redundant computation while employing anchor-delta temporal encoding to maximize compression efficiency. By executing these operations entirely within the storage plane through peer-to-peer communication, our system bypasses host memory bottlenecks that plague traditional architectures. Comprehensive evaluation across five real-world datasets demonstrates that Salient Store reduces end-to-end archival latency by $6.18 \times$, decreases host-side data movement by $5.63 \times$, and maintains up to 47 dB PSNR while reducing system power consumption by 65%. This work fundamentally re-imagines storage for continuous learning systems, transforming it from a passive bottleneck into an acceleration layer that operates symbiotically with neural inference and training.
Cyan Subhra Mishra, Deeksha Chaudhary, Mahmut T. Kandemir, Chita R. Das
PACT1
2025 NExUME: Adaptive Training and Inference for DNNs under Intermittent Power Environments
abstract
The deployment of Deep Neural Networks (DNNs) in energy-constrained environments, such as Energy Harvesting Wireless Sensor Networks (EH-WSNs), introduces significant challenges due to the intermittent nature of power availability. This study introduces NExUME, a novel training methodology designed specifically for DNNs operating under such constraints. We propose a dynamic adjustment of training parameters—dropout rates and quantization levels—that adapt in real-time to the available energy, which varies in energy harvesting scenarios. This approach utilizes a model that integrates the characteristics of the network architecture and the specific energy harvesting profile. It dynamically adjusts training strategies, such as the intensity and timing of dropout and quantization, based on predictions of energy availability. This method not only conserves energy but also enhances the network’s adaptability, ensuring robust learning and inference capabilities even under stringent power constraints. Our results show a 6% to 22% improvement in accuracy over current methods, with an increase of less than 5% in computational overhead. This paper details the development of the adaptive training framework, describes the integration of energy profiles with dropout and quantization adjustments, and presents a comprehensive evaluation using real-world data. Additionally, we introduce a novel dataset aimed at furthering the application of energy harvesting in computational settings.
Cyan Subhra Mishra, Deeksha Chaudhary, Jack Sampson, Mahmut T. Kandemir, Chita R. Das
ICLR1
2025 CORD: Parallelizing Query Processing Across Multiple Computational Storage Devices
abstract
Query processing on large-scale scientific datasets often suffers from performance bottlenecks due to significant data transfers between storage nodes and applications in decoupled distributed storage environments. This issue is particularly pronounced in high-selectivity queries where unnecessary data is transferred between the storage plane and the compute plane. To tackle this challenge, we introduce the integration of SmartSSDs, functioning as Computational Storage Devices (CSDs), into the storage layer. By offloading simple filter-projection operations to these CSDs, we significantly reduce data transfer bottlenecks, leading to lower query latency and higher throughput. Our novel framework, CORD (parallelizing query processing across multiple Computational stORage Devices), facilitates parallel query execution across multiple CSDs while considering data locality. CORD is compatible with any decoupled storage system equipped with CSDs. Our extensive empirical evaluation demonstrates that CORD achieves up to$93 \times$speedup for high-selectivity queries compared to traditional (compute plane) execution strategy and offers a further$1.64 \times$speedup in cases of uneven data distribution. Additionally, we present two optimizations for batch query processing. Results from our experiments with 4 CSDs reveal substantial performance improvements provided by the optimizations embedded in CORD.
Wahid Uz Zaman, Cyan Subhra Mishra, Saleh AlSaleh, Abutalib Aghayev, Mahmut T. Kandemir
IPDPS2
2024 Usas: A Sustainable Continuous-Learning' Framework for Edge Servers
abstract
Edge servers have recently become very popular for performing localized analytics, especially on video, as they reduce data traffic and protect privacy. However, due to their resource constraints, these servers often employ compressed models, which are typically prone to data drift. Consequently, for edge servers to provide cloud-comparable quality, they must also perform continuous learning to mitigate this drift. However, at expected deployment scales, performing continuous training on every edge server is not sustainable due to their aggregate power demands on grid supply and associated sustainability footprints. To address these challenges, we propose Us.as,´ an approach combining algorithmic adjustments, hardware-software co-design, and morphable acceleration hardware to enable the training of workloads on these edge servers to be powered by renewable, but intermittent, solar power that can sustainably scale alongside data sources. Our evaluation of Us.as on a real-world´ traffic dataset indicates that our continuous learning approach simultaneously improves both accuracy and efficiency: Us.as´ offers a 4.96% greater mean accuracy than prior approaches while our morphable accelerator that adapts to solar variance can save up to {234.95kWH, 2.63MWH}/year/edge-server compared to a {DNN accelerator, data center scale GPU}, respectively.
Cyan Subhra Mishra, Jack Sampson, Mahmut T. Kandemir, Narayanan Vijaykrishnan, Chita R. Das
HPCA1
2022 Exploiting Frame Similarity for Efficient Inference on Edge Devices
abstract
Deep neural networks (DNNs) are being widely used in various computer vision tasks as they can achieve very high accuracy. However, the large number of parameters employed in DNNs can result in long inference times for vision tasks, thus making it even more challenging to deploy them in the compute- and memory-constrained mobile/edge devices. To boost the inference of DNNs, some existing works employ compression (model pruning or quantization) or enhanced hardware. How-ever, most prior works focus on improving model structure and implementing custom accelerators. As opposed to the prior work, in this paper, we target the video data that are processed by edge devices, and study the similarity between frames. Based on that, we propose two runtime approaches to boost the performance of the inference process, while achieving high accuracy.Specifically, considering the similarities between successive video frames, we propose a frame-level compute reuse algorithm based on the motion vectors of each frame. With frame-level reuse, we are able to skip 53% of frames in inference with negligible overhead and remain within less than 1% mAP (accuracy) drop for the object detection task. Additionally, we implement a partial inference scheme to enable region/tile-level reuse. Our experiments on a representative mobile device (Pixel 3 Phone) show that the proposed partial inference scheme achieves 2 × speedup over the baseline approach that performs full inference on every frame. We integrate these two data reuse algorithms to accelerate the neural network inference and improve its energy efficiency. More specifically, for each frame in the video, we can dynamically select between (i) performing a full inference, (ii) performing a partial inference, or (iii) skipping the inference altogether. Our experimental evaluations using six different videos reveal that the proposed schemes are up to 80% (56% on average) energy efficient and 2.2× performance efficient compared to the conventional scheme, which performs full inference, while losing less than 2% accuracy. Additionally, the experimental analysis indicates that our approach outperforms the state-of-the-art work with respect to accuracy and/or performance/energy savings.
Ziyu Ying 0001, Shulin Zhao 0001, Haibo Zhang 0005, Cyan Subhra Mishra, Sandeepa Bhuyan, Mahmut T. Kandemir, Anand Sivasubramaniam, Chita R. Das
ICDCS4
2022 Pushing Point Cloud Compression to the Edge
abstract
As Point Clouds (PCs) gain popularity in processing millions of data points for 3D rendering in many applications, efficient data compression becomes a critical issue. This is because compression is the primary bottleneck in minimizing the latency and energy consumption of existing PC pipelines. Data compression becomes even more critical as PC processing is pushed to edge devices with limited compute and power budgets. In this paper, we propose and evaluate two complementary schemes, intra-frame compression and inter-frame compression, to speed up the PC compression, without losing much quality or compression efficiency. Unlike existing techniques that use sequential algorithms, our first design, intra-frame compression, exploits parallelism for boosting the performance of both geometry and attribute compression. The proposed parallelism brings around $43.7 \times$ performance improvement and 96.6% energy savings at a cost of $1.01 \times$ larger compressed data size. To further improve the compression efficiency, our second scheme, inter-frame compression, considers the temporal similarity among the video frames and reuses the attribute data from the previous frame for the current frame. We implement our designs on an NVIDIA Jetson AGX Xavier edge GPU board. Experimental results with six videos show that the combined compression schemes provide $34.0 \times$ speedup compared to a state-of-the-art scheme, with minimal impact on quality and compression ratio.
Ziyu Ying 0001, Shulin Zhao 0001, Sandeepa Bhuyan, Cyan Subhra Mishra, Mahmut T. Kandemir, Chita R. Das
MICRO4
2022 Cocktail: A Multidimensional Optimization for Model Serving in Cloud
Jashwant Raj Gunasekaran, Cyan Subhra Mishra, Prashanth Thinakaran, Bikash Sharma, Mahmut T. Kandemir, Chita R. Das
NSDI2
2021 Kraken: Adaptive Container Provisioning for Deploying Dynamic DAGs in Serverless Platforms
abstract
The growing popularity of microservices has led to the proliferation of online cloud service-based applications, which are typically modelled as Directed Acyclic Graphs (DAGs) comprising of tens to hundreds of microservices. The vast majority of these applications are user-facing, and hence, have stringent SLO requirements. Serverless functions, having short resource provisioning times and instant scalability, are suitable candidates for developing such latency-critical applications. However, existing serverless providers are unaware of the workflow characteristics of application DAGs, leading to container over-provisioning in many cases. This is further exacerbated in the case of dynamic DAGs, where the function chain for an application is not known a priori. Motivated by these observations, we propose Kraken, a workflow-aware resource management framework that minimizes the number of containers provisioned for an application DAG while ensuring SLO-compliance. We design and implement Kraken on OpenFaaS and evaluate it on a multi-node Kubernetes-managed cluster. Our extensive experimental evaluation using DeathStarbench workload suite and real-world traces demonstrates that Kraken spawns up to 76% fewer containers, thereby improving container utilization and saving cluster-wide energy by up to 4x and 48%, respectively, when compared to state-of-the art schedulers employed in serverless platforms.
Vivek M. Bhasi, Jashwant Raj Gunasekaran, Prashanth Thinakaran, Cyan Subhra Mishra, Mahmut T. Kandemir, Chita R. Das
SoCC4
2021 Origin: Enabling On-Device Intelligence for Human Activity Recognition Using Energy Harvesting Wireless Sensor Networks
abstract
There is an increasing demand for performing machine learning tasks, such as human activity recognition (HAR) on emerging ultra-low-power internet of things (IoT) platforms. Recent works show substantial efficiency boosts from performing inference tasks directly on the IoT nodes rather than merely transmitting raw sensor data. However, the computation and power demands of deep neural network (DNN) based inference pose significant challenges when executed on the nodes of an energy-harvesting wireless sensor network (EH-WSN). Moreover, managing inferences requiring responses from multiple energy-harvesting nodes imposes challenges at the system level in addition to the constraints at each node. This paper presents a novel scheduling policy along with an adaptive ensemble learner to efficiently perform HAR on a distributed energy-harvesting body area network. Our proposed policy, Origin, strategically ensures efficient and accurate individual inference execution at each sensor node by using a novel activity-aware scheduling approach. It also leverages the continuous nature of human activity when coordinating and aggregating results from all the sensor nodes to improve final classification accuracy. Further, Origin proposes an adaptive ensemble learner to personalize the optimizations based on each individual user. Experimental results using two different HAR data-sets show Origin, while running on harvested energy, to be at least 2.5% more accurate than a classical battery-powered energy aware HAR classifier continuously operating at the same average power.
Cyan Subhra Mishra, Jack Sampson, Mahmut T. Kandemir, Narayanan Vijaykrishnan
DATE1
2021 HoloAR: On-the-fly Optimization of 3D Holographic Processing for Augmented Reality
abstract
Hologram processing is the primary bottleneck and contributes to more than 50% of energy consumption in battery-operated augmented reality (AR) headsets. Thus, improving the computational efficiency of the holographic pipeline is critical. The objective of this paper is to maximize its energy efficiency without jeopardizing the hologram quality for AR applications. Towards this, we take the approach of analyzing the workloads to identify approximation opportunities. We show that, by considering various parameters like region of interest and depth of view, we can approximate the rendering of the virtual object to minimize the amount of computation without affecting the user experience. Furthermore, by optimizing the software design flow, we propose HoloAR, which intelligently renders the most important object in sight to the clearest detail, while approximating the computations for the others, thereby significantly reducing the amount of computation, saving energy, and gaining performance at the same time. We implement our design in an edge GPU platform to demonstrate the real-world applicability of our research. Our experimental results show that, compared to the baseline, HoloAR achieves, on average, 2.7 × speedup and 73% energy savings.
Shulin Zhao 0001, Haibo Zhang 0005, Cyan Subhra Mishra, Sandeepa Bhuyan, Ziyu Ying 0001, Mahmut T. Kandemir, Anand Sivasubramaniam, Chita R. Das
MICRO3
2021 MLPP: Exploring Transfer Learning and Model Distillation for Predicting Application Performance
abstract
Performance prediction for applications is quintessential towards detecting malicious hardware and software vulnerabilities. Typically application performance is predicted using the profiling data generated from hardware tools such as linux perf. By leveraging the data, prediction models, both machine learning (ML) based and non ML-based have been proposed. However a majority of these models suffer from either loss in prediction accuracy, very large model sizes, and/or lack of general applicability to different hardware types such as wearables, handhelds, desktops etc. To address the aforementioned inefficiencies, in this paper we proposed MLPP, a machine learning based performance prediction model which can accurately predict application performance, and at the same time be easily transferable to a wide both mobile and desktop hardware platforms by leveraging transfer learning technique. Furthermore, MLPP incorporates model distillation techniques to significantly reduce the model size. Through our extensive experimentation and evaluation we show that MLPP can achieve up to 92.5% prediction accuracy while reducing the model size by up to 3.5 ×.
Jashwant Raj Gunasekaran, Cyan Subhra Mishra
NAS2
2020 ResiRCA: A Resilient Energy Harvesting ReRAM Crossbar-Based Accelerator for Intelligent Embedded Processors
abstract
Many recent works have shown substantial efficiency boosts from performing inference tasks on Internet of Things (IoT) nodes rather than merely transmitting raw sensor data. However, such tasks, e.g., convolutional neural networks (CNNs), are very compute intensive. They are therefore challenging to complete at sensing-matched latencies in ultra-low-power and energy-harvesting IoT nodes. ReRAM crossbar-based accelerators (RCAs) are an ideal candidate to perform the dominant multiplication-and-accumulation (MAC) operations in CNNs efficiently, but conventional, performance-oriented RCAs, while energy-efficient, are power hungry and ill-optimized for the intermittent and unstable power supply of energy-harvesting IoT nodes. This paper presents the ResiRCA architecture that integrates a new, lightweight, and configurable RCA suitable for energy harvesting environments as an opportunistically executing augmentation to a baseline sense-and-transmit battery-powered IoT node. To maximize ResiRCA throughput under different power levels, we develop the ResiSchedule approach for dynamic RCA reconfiguration. The proposed approach uses loop tiling-based computation decomposition, model duplication within the RCA, and inter-layer pipelining to reduce RCA activation thresholds and more closely track execution costs with dynamic power income. Experimental results show that ResiRCA together with ResiSchedule achieve average speedups and energy efficiency improvements of 8× and 14× respectively compared to a baseline RCA with intermittency-unaware scheduling.
Keni Qiu, Nicholas Jao, Mengying Zhao, Cyan Subhra Mishra, Gulsum Gudukbay Akbulut, Sethu Jose, Jack Sampson, Mahmut T. Kandemir, Narayanan Vijaykrishnan
HPCA4
2020 Déjà View: Spatio-Temporal Compute Reuse for' Energy-Efficient 360° VR Video Streaming
abstract
The emergence of virtual reality (VR) and augmented reality (AR) has revolutionized our lives by enabling a 360° artificial sensory stimulation across diverse domains, including, but not limited to, sports, media, healthcare, and gaming. Unlike the conventional planar video processing, where memory access is the main bottleneck, in 360° VR videos the compute is the primary bottleneck and contributes to more than 50% energy consumption in battery-operated VR headsets. Thus, improving the computational efficiency of the video processing pipeline in a VR is critical. While prior efforts have attempted to address this problem through acceleration using a GPU or FPGA, none of them has analyzed the 360° VR pipeline to examine if there is any scope to optimize the computation with known techniques such as memoization.Thus, in this paper, we analyze the VR computation pipeline and observe that there is significant scope to skip computations by leveraging the temporal and spatial locality in head orientation and eye correlations, respectively, resulting in computation reduction and energy efficiency. The proposed Déjà View design takes advantage of temporal reuse by memoizing head orientation and spatial reuse by establishing a relationship between left and right eye projection, and can be implemented either on a GPU or an FPGA. We propose both software modifications for existing compute pipeline and microarchitectural additions for further enhancement. We evaluate our design by implementing the software enhancements on an NVIDIA Jetson TX2 GPU board and our microarchitectural additions on a Xilinx Zynq-7000 FPGA model using five video workloads. Experimental results show that Déjà View can provide 34% computation reduction and 17% energy saving, compared to the state-of-the-art design.
Shulin Zhao 0001, Haibo Zhang 0005, Sandeepa Bhuyan, Cyan Subhra Mishra, Ziyu Ying 0001, Mahmut T. Kandemir, Anand Sivasubramaniam, Chita R. Das
ISCA4