VLDB 2026 Research / reviewers in the wild / expert
Ziyu Ying 0001
dblp:236/6777-1
· DBLP profile ↗
9ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0002-0854-9933ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 4 first-author · 8 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Pirate: No Compromise Low-Bandwidth VR Streaming for Edge DevicesabstractDue to the limited compute power and storage capabilities of edge platforms, ''streaming'' often provides a better VR experience compared to ''rendering''. Yet, achieving high-quality VR streaming faces two significant challenges, namely, bandwidth limitations and the need for real-time operation with high frames per second (FPS). Previous efforts have tended to prioritize either conserving bandwidth without real-time performance or ensuring real-time operation without substantial bandwidth savings. In this work, we incorporate the concept of ''stereo similarity'' to develop a novel real-time stereo video compression framework for streaming, called Pirate. Unlike the previously proposed approaches that rely on large machine learning-based models for synthesizing stereo pairs from both eyes with disparity maps (which can be impractical for most edge platforms due to their high computational cost), Pirate iteratively synthesizes the target eye view using only a single eye view and its corresponding disparity and optical flow information, with alternating left or right eye transmission. This enables us to generate target view at an extremely low computational cost, even under bandwidth constraints as low as 0.1 bits per pixel (bpp), while maintaining a high frame rate of 90 FPS. Our evaluations also reveal that, the proposed approach not only achieves real-time VR streaming with a 20%-40% reduction in bandwidth usage, but also maintains similar superior quality standards. Yingtian Zhang, Ziyu Ying 0001, Wanhang Lu, Sijie Lan, Huijuan Xu 0001, Kiwan Maeng, Anand Sivasubramaniam, Mahmut T. Kandemir, Chita R. Das |
ASPLOS (2) | 3 |
| 2025 | GSCoder: Enabling Fast and Efficient Encoding for Game Streaming ApplicationsabstractThe recent proliferation of cloud gaming (also referred to as game streaming), enabling high-fidelity gaming quality on edge devices without high-end hardware, promises a transformative and democratized gaming experience across diverse populations. Yet, streaming high definition (4K UHD) game frames to thin-client devices, especially mobile, requires substantially higher bandwidth than traditional video streaming and often results in frame drops, degrading user experience. We identified that this high bandwidth demand arises from inefficient compression of game frames with real-time compression requirement (60 frames per second (FPS)) because the rapid, irregular motion in game frames violates the predictability assumptions baked into standard video motion estimation algorithms.To address this issue, we propose and evaluate GSCoder, a realtime efficient game frame compression framework that utilizes motion cues from the game’s rendering pipeline to directly acquire and optimize encoding-compliant motion vectors precisely, bypassing the costly motion estimation step inherent in standard encoders. Our evaluation, conducted in five open-source games with varying motion complexities, demonstrates that GSCoder achieves on average 49% and 19% higher compression efficiency than state-of-the-art (SOTA) game and video encoders, respectively, while maintaining real-time performance and delivering high-quality streams. Moreover, GSCoder delivers at least $3.5 \times$ encoding speedup compared to SOTA video encoders. Sandeepa Bhuyan, Ziyu Ying 0001, Vivek M. Bhasi, Mahmut T. Kandemir, Chita R. Das |
MASCOTS | 2 |
| 2024 | Foveated HDR: Efficient HDR Content Generation on Edge Devices Leveraging User's Visual AttentionabstractIn recent years, high dynamic range (HDR) content has become increasingly popular for its ability to represent a broader brightness range, enhancing the realism and immersion in applications like augmented reality/virtual reality (AR/VR) on edge devices. While DNN-based solutions are effective for reconstructing high-fidelity HDR content, due to their high computational demands and memory usage, the DNN-based HDR reconstruction takes up to several seconds to generate one HDR image, making it very challenging to deploy such techniques onto the edge devices. Ziyu Ying 0001, Sandeepa Bhuyan, Yingtian Zhang, Mahmut T. Kandemir, Chita R. Das |
ICCAD | 1 |
| 2024 | GameStreamSR: Enabling Neural-Augmented Game Streaming on Commodity Mobile PlatformsabstractCloud gaming (also referred to as Game Streaming) is a rapidly emerging application that is changing the way people enjoy video games. However, if the user demands a high-resolution (e.g., 2 K or 4 K) stream, the game frames require high bandwidth and the stream often suffers from a significant number of frame drops due to network congestion degrading the Quality of Experience (QoE). Recently, the DNN-based Super Resolution (SR) technique has gained prominence as a practical alternative for streaming low-resolution frames and upscaling them at the client for enhanced video quality. However, performing such DNN-based tasks on resource-constrained and battery-operated mobile platforms is very expensive and also fails to meet the real-time requirement (60 frames per second (FPS)). Unlike traditional video streaming, where the frames can be downloaded and buffered, and then upscaled by their playback turn, Game Streaming is real-time and interactive, where the frames are generated on the fly and cannot tolerate high latency/lags for frame upscaling. Thus, state-of-the-art (SOTA) DNN-based SR cannot satisfy the mobile Game Streaming requirements. Towards this, we propose GameStreamSR, a framework for enabling real-time Super Resolution for Game Streaming applications on mobile platforms. We take visual perception nature into consideration and propose to only apply DNN-based SR to the regions with high visual importance and upscale the remaining regions using traditional solutions such as bilinear interpolation. Especially, we leverage the depth data from the game rendering pipeline to intelligently localize the important regions, called regions of importance (RoI), in the rendered game frames. Our evaluation of ten popular games on commodity mobile platforms shows that our proposal can enable realtime (60 FPS) neurally-augmented SR. Our design achieves a $13 \times$ frame rate speedup (and $\approx 4 \times$ Motion-to-Photon latency improvement) for the reference frames and a $1.6 \times$ frame rate speedup for the non-reference frames, which translates to, on average $2 \times$ FPS performance improvement and 26-33% energy savings over the SOTA DNN-based SR execution, while achieving about 2dB PSNR gain and better perceptual quality than the current SOTA. Sandeepa Bhuyan, Ziyu Ying 0001, Mahmut T. Kandemir, Mahanth Gowda, Chita R. Das |
ISCA | 2 |
| 2023 | EdgePC: Efficient Deep Learning Analytics for Point Clouds on Edge DevicesabstractRecently, point cloud (PC) has gained popularity in modeling various 3D objects (including both synthetic and real-life) and has been extensively utilized in a wide range of applications such as AR/VR, 3D reconstruction, and autonomous driving. For such applications, it is critical to analyze/understand the surrounding scenes properly. To achieve this, deep learning based methods (e.g., convolutional neural networks (CNNs)) have been widely employed for higher accuracy. Unlike the deep learning on conventional 2D images/videos, where the feature computation (matrix multiplication) is the major bottleneck, in point cloud-based CNNs, the sample and neighbor search stages are the primary bottlenecks, and collectively contribute to 54% (up to 80%) of the overall execution latency on a typical edge device. While prior efforts have attempted to solve this issue by designing custom ASICs or pipelining the neighbor search with other stages, to our knowledge, none of them has tried to "structurize" the unstructured PC data for improving computational efficiency. Ziyu Ying 0001, Sandeepa Bhuyan, Yingtian Zhang, Mahmut T. Kandemir, Chita R. Das |
ISCA | 1 |
| 2022 | Exploiting Frame Similarity for Efficient Inference on Edge DevicesabstractDeep neural networks (DNNs) are being widely used in various computer vision tasks as they can achieve very high accuracy. However, the large number of parameters employed in DNNs can result in long inference times for vision tasks, thus making it even more challenging to deploy them in the compute- and memory-constrained mobile/edge devices. To boost the inference of DNNs, some existing works employ compression (model pruning or quantization) or enhanced hardware. How-ever, most prior works focus on improving model structure and implementing custom accelerators. As opposed to the prior work, in this paper, we target the video data that are processed by edge devices, and study the similarity between frames. Based on that, we propose two runtime approaches to boost the performance of the inference process, while achieving high accuracy.Specifically, considering the similarities between successive video frames, we propose a frame-level compute reuse algorithm based on the motion vectors of each frame. With frame-level reuse, we are able to skip 53% of frames in inference with negligible overhead and remain within less than 1% mAP (accuracy) drop for the object detection task. Additionally, we implement a partial inference scheme to enable region/tile-level reuse. Our experiments on a representative mobile device (Pixel 3 Phone) show that the proposed partial inference scheme achieves 2 × speedup over the baseline approach that performs full inference on every frame. We integrate these two data reuse algorithms to accelerate the neural network inference and improve its energy efficiency. More specifically, for each frame in the video, we can dynamically select between (i) performing a full inference, (ii) performing a partial inference, or (iii) skipping the inference altogether. Our experimental evaluations using six different videos reveal that the proposed schemes are up to 80% (56% on average) energy efficient and 2.2× performance efficient compared to the conventional scheme, which performs full inference, while losing less than 2% accuracy. Additionally, the experimental analysis indicates that our approach outperforms the state-of-the-art work with respect to accuracy and/or performance/energy savings. Ziyu Ying 0001, Shulin Zhao 0001, Haibo Zhang 0005, Cyan Subhra Mishra, Sandeepa Bhuyan, Mahmut T. Kandemir, Anand Sivasubramaniam, Chita R. Das |
ICDCS | 1 |
| 2022 | Pushing Point Cloud Compression to the EdgeabstractAs Point Clouds (PCs) gain popularity in processing millions of data points for 3D rendering in many applications, efficient data compression becomes a critical issue. This is because compression is the primary bottleneck in minimizing the latency and energy consumption of existing PC pipelines. Data compression becomes even more critical as PC processing is pushed to edge devices with limited compute and power budgets. In this paper, we propose and evaluate two complementary schemes, intra-frame compression and inter-frame compression, to speed up the PC compression, without losing much quality or compression efficiency. Unlike existing techniques that use sequential algorithms, our first design, intra-frame compression, exploits parallelism for boosting the performance of both geometry and attribute compression. The proposed parallelism brings around $43.7 \times$ performance improvement and 96.6% energy savings at a cost of $1.01 \times$ larger compressed data size. To further improve the compression efficiency, our second scheme, inter-frame compression, considers the temporal similarity among the video frames and reuses the attribute data from the previous frame for the current frame. We implement our designs on an NVIDIA Jetson AGX Xavier edge GPU board. Experimental results with six videos show that the combined compression schemes provide $34.0 \times$ speedup compared to a state-of-the-art scheme, with minimal impact on quality and compression ratio. Ziyu Ying 0001, Shulin Zhao 0001, Sandeepa Bhuyan, Cyan Subhra Mishra, Mahmut T. Kandemir, Chita R. Das |
MICRO | 1 |
| 2021 | HoloAR: On-the-fly Optimization of 3D Holographic Processing for Augmented RealityabstractHologram processing is the primary bottleneck and contributes to more than 50% of energy consumption in battery-operated augmented reality (AR) headsets. Thus, improving the computational efficiency of the holographic pipeline is critical. The objective of this paper is to maximize its energy efficiency without jeopardizing the hologram quality for AR applications. Towards this, we take the approach of analyzing the workloads to identify approximation opportunities. We show that, by considering various parameters like region of interest and depth of view, we can approximate the rendering of the virtual object to minimize the amount of computation without affecting the user experience. Furthermore, by optimizing the software design flow, we propose HoloAR, which intelligently renders the most important object in sight to the clearest detail, while approximating the computations for the others, thereby significantly reducing the amount of computation, saving energy, and gaining performance at the same time. We implement our design in an edge GPU platform to demonstrate the real-world applicability of our research. Our experimental results show that, compared to the baseline, HoloAR achieves, on average, 2.7 × speedup and 73% energy savings. Shulin Zhao 0001, Haibo Zhang 0005, Cyan Subhra Mishra, Sandeepa Bhuyan, Ziyu Ying 0001, Mahmut T. Kandemir, Anand Sivasubramaniam, Chita R. Das |
MICRO | 5 |
| 2020 | Déjà View: Spatio-Temporal Compute Reuse for' Energy-Efficient 360° VR Video StreamingabstractThe emergence of virtual reality (VR) and augmented reality (AR) has revolutionized our lives by enabling a 360° artificial sensory stimulation across diverse domains, including, but not limited to, sports, media, healthcare, and gaming. Unlike the conventional planar video processing, where memory access is the main bottleneck, in 360° VR videos the compute is the primary bottleneck and contributes to more than 50% energy consumption in battery-operated VR headsets. Thus, improving the computational efficiency of the video processing pipeline in a VR is critical. While prior efforts have attempted to address this problem through acceleration using a GPU or FPGA, none of them has analyzed the 360° VR pipeline to examine if there is any scope to optimize the computation with known techniques such as memoization.Thus, in this paper, we analyze the VR computation pipeline and observe that there is significant scope to skip computations by leveraging the temporal and spatial locality in head orientation and eye correlations, respectively, resulting in computation reduction and energy efficiency. The proposed Déjà View design takes advantage of temporal reuse by memoizing head orientation and spatial reuse by establishing a relationship between left and right eye projection, and can be implemented either on a GPU or an FPGA. We propose both software modifications for existing compute pipeline and microarchitectural additions for further enhancement. We evaluate our design by implementing the software enhancements on an NVIDIA Jetson TX2 GPU board and our microarchitectural additions on a Xilinx Zynq-7000 FPGA model using five video workloads. Experimental results show that Déjà View can provide 34% computation reduction and 17% energy saving, compared to the state-of-the-art design. Shulin Zhao 0001, Haibo Zhang 0005, Sandeepa Bhuyan, Cyan Subhra Mishra, Ziyu Ying 0001, Mahmut T. Kandemir, Anand Sivasubramaniam, Chita R. Das |
ISCA | 5 |