Sandeepa Bhuyan

dblp:252/1368 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0002-0679-9058ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 2 first-author · 7 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 GSCoder: Enabling Fast and Efficient Encoding for Game Streaming Applications
abstract
The recent proliferation of cloud gaming (also referred to as game streaming), enabling high-fidelity gaming quality on edge devices without high-end hardware, promises a transformative and democratized gaming experience across diverse populations. Yet, streaming high definition (4K UHD) game frames to thin-client devices, especially mobile, requires substantially higher bandwidth than traditional video streaming and often results in frame drops, degrading user experience. We identified that this high bandwidth demand arises from inefficient compression of game frames with real-time compression requirement (60 frames per second (FPS)) because the rapid, irregular motion in game frames violates the predictability assumptions baked into standard video motion estimation algorithms.To address this issue, we propose and evaluate GSCoder, a realtime efficient game frame compression framework that utilizes motion cues from the game’s rendering pipeline to directly acquire and optimize encoding-compliant motion vectors precisely, bypassing the costly motion estimation step inherent in standard encoders. Our evaluation, conducted in five open-source games with varying motion complexities, demonstrates that GSCoder achieves on average 49% and 19% higher compression efficiency than state-of-the-art (SOTA) game and video encoders, respectively, while maintaining real-time performance and delivering high-quality streams. Moreover, GSCoder delivers at least $3.5 \times$ encoding speedup compared to SOTA video encoders.
Sandeepa Bhuyan, Ziyu Ying 0001, Vivek M. Bhasi, Mahmut T. Kandemir, Chita R. Das
MASCOTS1
2024 Foveated HDR: Efficient HDR Content Generation on Edge Devices Leveraging User's Visual Attention
abstract
In recent years, high dynamic range (HDR) content has become increasingly popular for its ability to represent a broader brightness range, enhancing the realism and immersion in applications like augmented reality/virtual reality (AR/VR) on edge devices. While DNN-based solutions are effective for reconstructing high-fidelity HDR content, due to their high computational demands and memory usage, the DNN-based HDR reconstruction takes up to several seconds to generate one HDR image, making it very challenging to deploy such techniques onto the edge devices.
Ziyu Ying 0001, Sandeepa Bhuyan, Yingtian Zhang, Mahmut T. Kandemir, Chita R. Das
ICCAD2
2024 GameStreamSR: Enabling Neural-Augmented Game Streaming on Commodity Mobile Platforms
abstract
Cloud gaming (also referred to as Game Streaming) is a rapidly emerging application that is changing the way people enjoy video games. However, if the user demands a high-resolution (e.g., 2 K or 4 K) stream, the game frames require high bandwidth and the stream often suffers from a significant number of frame drops due to network congestion degrading the Quality of Experience (QoE). Recently, the DNN-based Super Resolution (SR) technique has gained prominence as a practical alternative for streaming low-resolution frames and upscaling them at the client for enhanced video quality. However, performing such DNN-based tasks on resource-constrained and battery-operated mobile platforms is very expensive and also fails to meet the real-time requirement (60 frames per second (FPS)). Unlike traditional video streaming, where the frames can be downloaded and buffered, and then upscaled by their playback turn, Game Streaming is real-time and interactive, where the frames are generated on the fly and cannot tolerate high latency/lags for frame upscaling. Thus, state-of-the-art (SOTA) DNN-based SR cannot satisfy the mobile Game Streaming requirements. Towards this, we propose GameStreamSR, a framework for enabling real-time Super Resolution for Game Streaming applications on mobile platforms. We take visual perception nature into consideration and propose to only apply DNN-based SR to the regions with high visual importance and upscale the remaining regions using traditional solutions such as bilinear interpolation. Especially, we leverage the depth data from the game rendering pipeline to intelligently localize the important regions, called regions of importance (RoI), in the rendered game frames. Our evaluation of ten popular games on commodity mobile platforms shows that our proposal can enable realtime (60 FPS) neurally-augmented SR. Our design achieves a $13 \times$ frame rate speedup (and $\approx 4 \times$ Motion-to-Photon latency improvement) for the reference frames and a $1.6 \times$ frame rate speedup for the non-reference frames, which translates to, on average $2 \times$ FPS performance improvement and 26-33% energy savings over the SOTA DNN-based SR execution, while achieving about 2dB PSNR gain and better perceptual quality than the current SOTA.
Sandeepa Bhuyan, Ziyu Ying 0001, Mahmut T. Kandemir, Mahanth Gowda, Chita R. Das
ISCA1
2023 EdgePC: Efficient Deep Learning Analytics for Point Clouds on Edge Devices
abstract
Recently, point cloud (PC) has gained popularity in modeling various 3D objects (including both synthetic and real-life) and has been extensively utilized in a wide range of applications such as AR/VR, 3D reconstruction, and autonomous driving. For such applications, it is critical to analyze/understand the surrounding scenes properly. To achieve this, deep learning based methods (e.g., convolutional neural networks (CNNs)) have been widely employed for higher accuracy. Unlike the deep learning on conventional 2D images/videos, where the feature computation (matrix multiplication) is the major bottleneck, in point cloud-based CNNs, the sample and neighbor search stages are the primary bottlenecks, and collectively contribute to 54% (up to 80%) of the overall execution latency on a typical edge device. While prior efforts have attempted to solve this issue by designing custom ASICs or pipelining the neighbor search with other stages, to our knowledge, none of them has tried to "structurize" the unstructured PC data for improving computational efficiency.
Ziyu Ying 0001, Sandeepa Bhuyan, Yingtian Zhang, Mahmut T. Kandemir, Chita R. Das
ISCA2
2022 Exploiting Frame Similarity for Efficient Inference on Edge Devices
abstract
Deep neural networks (DNNs) are being widely used in various computer vision tasks as they can achieve very high accuracy. However, the large number of parameters employed in DNNs can result in long inference times for vision tasks, thus making it even more challenging to deploy them in the compute- and memory-constrained mobile/edge devices. To boost the inference of DNNs, some existing works employ compression (model pruning or quantization) or enhanced hardware. How-ever, most prior works focus on improving model structure and implementing custom accelerators. As opposed to the prior work, in this paper, we target the video data that are processed by edge devices, and study the similarity between frames. Based on that, we propose two runtime approaches to boost the performance of the inference process, while achieving high accuracy.Specifically, considering the similarities between successive video frames, we propose a frame-level compute reuse algorithm based on the motion vectors of each frame. With frame-level reuse, we are able to skip 53% of frames in inference with negligible overhead and remain within less than 1% mAP (accuracy) drop for the object detection task. Additionally, we implement a partial inference scheme to enable region/tile-level reuse. Our experiments on a representative mobile device (Pixel 3 Phone) show that the proposed partial inference scheme achieves 2 × speedup over the baseline approach that performs full inference on every frame. We integrate these two data reuse algorithms to accelerate the neural network inference and improve its energy efficiency. More specifically, for each frame in the video, we can dynamically select between (i) performing a full inference, (ii) performing a partial inference, or (iii) skipping the inference altogether. Our experimental evaluations using six different videos reveal that the proposed schemes are up to 80% (56% on average) energy efficient and 2.2× performance efficient compared to the conventional scheme, which performs full inference, while losing less than 2% accuracy. Additionally, the experimental analysis indicates that our approach outperforms the state-of-the-art work with respect to accuracy and/or performance/energy savings.
Ziyu Ying 0001, Shulin Zhao 0001, Haibo Zhang 0005, Cyan Subhra Mishra, Sandeepa Bhuyan, Mahmut T. Kandemir, Anand Sivasubramaniam, Chita R. Das
ICDCS5
2022 Pushing Point Cloud Compression to the Edge
abstract
As Point Clouds (PCs) gain popularity in processing millions of data points for 3D rendering in many applications, efficient data compression becomes a critical issue. This is because compression is the primary bottleneck in minimizing the latency and energy consumption of existing PC pipelines. Data compression becomes even more critical as PC processing is pushed to edge devices with limited compute and power budgets. In this paper, we propose and evaluate two complementary schemes, intra-frame compression and inter-frame compression, to speed up the PC compression, without losing much quality or compression efficiency. Unlike existing techniques that use sequential algorithms, our first design, intra-frame compression, exploits parallelism for boosting the performance of both geometry and attribute compression. The proposed parallelism brings around $43.7 \times$ performance improvement and 96.6% energy savings at a cost of $1.01 \times$ larger compressed data size. To further improve the compression efficiency, our second scheme, inter-frame compression, considers the temporal similarity among the video frames and reuses the attribute data from the previous frame for the current frame. We implement our designs on an NVIDIA Jetson AGX Xavier edge GPU board. Experimental results with six videos show that the combined compression schemes provide $34.0 \times$ speedup compared to a state-of-the-art scheme, with minimal impact on quality and compression ratio.
Ziyu Ying 0001, Shulin Zhao 0001, Sandeepa Bhuyan, Cyan Subhra Mishra, Mahmut T. Kandemir, Chita R. Das
MICRO3
2021 HoloAR: On-the-fly Optimization of 3D Holographic Processing for Augmented Reality
abstract
Hologram processing is the primary bottleneck and contributes to more than 50% of energy consumption in battery-operated augmented reality (AR) headsets. Thus, improving the computational efficiency of the holographic pipeline is critical. The objective of this paper is to maximize its energy efficiency without jeopardizing the hologram quality for AR applications. Towards this, we take the approach of analyzing the workloads to identify approximation opportunities. We show that, by considering various parameters like region of interest and depth of view, we can approximate the rendering of the virtual object to minimize the amount of computation without affecting the user experience. Furthermore, by optimizing the software design flow, we propose HoloAR, which intelligently renders the most important object in sight to the clearest detail, while approximating the computations for the others, thereby significantly reducing the amount of computation, saving energy, and gaining performance at the same time. We implement our design in an edge GPU platform to demonstrate the real-world applicability of our research. Our experimental results show that, compared to the baseline, HoloAR achieves, on average, 2.7 × speedup and 73% energy savings.
Shulin Zhao 0001, Haibo Zhang 0005, Cyan Subhra Mishra, Sandeepa Bhuyan, Ziyu Ying 0001, Mahmut T. Kandemir, Anand Sivasubramaniam, Chita R. Das
MICRO4
2020 Déjà View: Spatio-Temporal Compute Reuse for' Energy-Efficient 360° VR Video Streaming
abstract
The emergence of virtual reality (VR) and augmented reality (AR) has revolutionized our lives by enabling a 360° artificial sensory stimulation across diverse domains, including, but not limited to, sports, media, healthcare, and gaming. Unlike the conventional planar video processing, where memory access is the main bottleneck, in 360° VR videos the compute is the primary bottleneck and contributes to more than 50% energy consumption in battery-operated VR headsets. Thus, improving the computational efficiency of the video processing pipeline in a VR is critical. While prior efforts have attempted to address this problem through acceleration using a GPU or FPGA, none of them has analyzed the 360° VR pipeline to examine if there is any scope to optimize the computation with known techniques such as memoization.Thus, in this paper, we analyze the VR computation pipeline and observe that there is significant scope to skip computations by leveraging the temporal and spatial locality in head orientation and eye correlations, respectively, resulting in computation reduction and energy efficiency. The proposed Déjà View design takes advantage of temporal reuse by memoizing head orientation and spatial reuse by establishing a relationship between left and right eye projection, and can be implemented either on a GPU or an FPGA. We propose both software modifications for existing compute pipeline and microarchitectural additions for further enhancement. We evaluate our design by implementing the software enhancements on an NVIDIA Jetson TX2 GPU board and our microarchitectural additions on a Xilinx Zynq-7000 FPGA model using five video workloads. Experimental results show that Déjà View can provide 34% computation reduction and 17% energy saving, compared to the state-of-the-art design.
Shulin Zhao 0001, Haibo Zhang 0005, Sandeepa Bhuyan, Cyan Subhra Mishra, Ziyu Ying 0001, Mahmut T. Kandemir, Anand Sivasubramaniam, Chita R. Das
ISCA3
2019 Understanding Energy Efficiency in IoT App Executions
abstract
Billions of Internet-of-Things (IoT) devices such as sensors, actuators, computing units, etc., are connected to form IoT platforms. However, it is observed that such hardware platforms today spend a significant proportion of their energy in communication between the CPU and the sensors (which are controlled by micro-controller unit (MCU)). Motivated by this observation, two simple, yet effective, optimizations are proposed to minimize the energy consumption. The first optimization, called Batching, interrupts the CPU after collecting multiple sensor data points at the MCU (instead of only 1), and thus, minimizes the interrupt overheads. The second optimization, called Computation Offloading to MCU (COM), offloads app-specific computations to the MCU to minimize data transfer overheads, and makes use of the relatively low energy footprint, and low-compute capabilities of the MCU in place of the CPU in the hub. However, questions such as why these two schemes are needed, where the energy benefit comes from, which IoT apps are suitable for these optimizations, etc., remain unclear. To better understand the Batching and COM approaches towards energy efficiency in IoT app executions, we characterize ten representative workloads on a Raspberry Pi and ESP8266 MCU platform, and evaluate the energy savings using these two optimizations and illustrate that for light-weight workloads (where COM is applicable), Batching and COM reduce the energy consumption by 52% and 85%, respectively when compared to the baseline. And for heavy-weight apps (where COM is not possible due to limited capacity of MCU), by offloading the light-weight apps and batching for the heavy-weight, Batching + COM (BCOM) benefits 10% energy savings compared to the baseline.
Shulin Zhao 0001, Prasanna Venkatesh Rengasamy, Haibo Zhang 0005, Sandeepa Bhuyan, Nachiappan Chidambaram Nachiappan, Anand Sivasubramaniam, Mahmut T. Kandemir, Chita R. Das
ICDCS4