VLDB 2026 Research / reviewers in the wild / expert
Shu Shi
dblp:57/3883
· DBLP profile ↗
29ranked-venue papers
13as first author
10since 2021 · last 2026
0009-0000-9893-612XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 11 first-author · 3 since 2021Computer networks · 9 · 3 first-author · 6 since 2021Systems, architecture and hardware · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hermit: A Flow-Collaborative Transport Scheme for Multi-Source Video On-Demand StreamingabstractToday's fast-growing Video-on-Demand (VoD) service needs efficient content delivery to guarantee the user experience. To reduce costs, the industry has been exploring the adoption of unstable, heterogeneous, low-performance edge nodes as cost-efficient alternatives to expensive CDN servers. To compensate for the resulting degradation in user experience, Multi-source Parallel Downloading (MPD) is becoming a new VoD transport paradigm. However, existing transport optimization solutions face performance obstacles when applied to the MPD scenarios. They cannot handle the contention between MPD flows of the same download task, which is likely to occur at the shared last-hop, and lack the ability to quickly adapt to the unstable network environments brought by dynamic, heterogeneous, and low-performance edge nodes. To fill this gap, we propose Hermit, a VoD-oriented MPD transport algorithm. Hermit (1) continuously monitors the state of the flows and makes timely scheduling decisions, and (2) efficiently coordinates across the flows to mitigate self-contention at the shared last hop. As a client-driven scheme, Hermit does not require cumbersome coordination among edge nodes, nor does it increase server complexity. Through extensive experiments on real-world large-scale testbed and locally emulated network conditions, we demonstrate that Hermit can improve the consistent downloading rate by 9.2% to 21.3%. Shaorui Ren, Enhuan Dong, Haiping Wang 0002, Jia Zhang 0010, Zili Meng, Mingwei Xu 0001, Shu Shi, Hebin Yu, Zhichen Xue, Yajie Peng, Xiaofei Pang |
ICC | 8 |
| 2026 | Trinity: Exploiting Latency Sensitivity to Improve Quality of Experience on Cloud VR Gaming
Yongqiang Gui, Yanyan Suo, Sandesh Dhawaskar Sathyanarayana, Klara Nahrstedt, Shu Shi |
MMSys | 7 |
| 2026 | HCDN: Coordinated Stream Scheduling for Cost-Effective Live Video Delivery
Liying Wang 0011, Chengke Wang, Mingming Lu, Qingyue Li, Song Geng, Linsen Wang, Kaida Hu, Haoyuan Huang, Shimao Tian, Ri Lu, Mingfei Hao, Chenren Xu, Shu Shi |
NSDI | 15 |
| 2026 | Medley: Optimizing Midgress Bandwidth for Commercial Live Streaming CDNs
Haiping Wang 0002, Wanxin Shi, Sandesh Dhawaskar Sathyanarayana, Shu Shi, Yinghao Yu, La Zuo, Hebin Yu, Ruoshi Sun, Yajie Peng, Xiaofei Pang, Ruili Fang, Zhenpeng Zhu, Yang Xu 0010 |
NSDI | 4 |
| 2025 | ACE: Sending Burstiness Control for High-Quality Real-time CommunicationabstractModern real-time communication (RTC) demands both ultra-low latency and consistently high visual quality. Yet, as content becomes more dynamic and RTTs shrink, we reveal a previously overlooked problem: long-tail queuing latency in the sender's pacing queue between encoder and network. This phenomenon is rooted in a mismatch between the bursty frame stream produced by the encoder and the smooth traffic expected by the network. Existing approaches trying to smoothen the bitrate inevitably force an undesirable trade-off between latency and video quality. To address this, we propose a dual-control approach that manages both the encoding and transmission burstiness. At the sender, we dynamically adjust the bucket size of a token-based pacer to control burstiness at the granularity of frame level. Within the encoder, we introduce an adaptive complexity mechanism that smoothens frame sizes without sacrificing quality. Trace-driven emulation and real-world experiments show our solution ACE reduces end-to-end 95th percentile latency by up to 43% while maintaining superior visual quality versus the state of the art. Xiangjie Huang, Haiping Wang 0002, Hebin Yu, Sandesh Dhawaskar Sathyanarayana, Shu Shi, Zili Meng |
SIGCOMM | 6 |
| 2024 | Magpie: Improving the Efficiency of A/B Tests for Large Scale Video-on-Demand SystemsabstractWith the exponential rise in video traffic, researchers and developers require more effective tools to validate the efficacy of designed algorithms for Video-on-Demand (VoD) system. However, traditional experimental platforms face two main challenges: a lack of realistic testing and the need for longer and significant effort. To overcome these limitations, we propose Magpie, an efficient experimental platform tailored for VoD systems. Magpie leverages a realistic operational setting, rapid testing, and high reproducibility to closely simulate online user environments without impacting production systems. Compared to conventional simulations, our evaluation demonstrates that Magpie reduces the disparity with online experiments by 85.6%. Deployed within our company-a leading video content provider in China-Magpie has efficiently validated over tens of algorithms, with 80% demonstrating enhanced performance in subsequent online tests. Hebin Yu, Haiping Wang 0002, Chenfei Tian, Sandesh Dhawaskar Sathyanarayana, Shu Shi, Zhichen Xue, Shuaixin Yu, Yajie Peng, Xiaofei Pang |
IMC | 5 |
| 2024 | AGiLE: Enhancing Adaptive GOP in Live Video StreamingabstractAs live streaming video continues to gain popularity, encoding efficiency remains a critical challenge. Current commercial systems limit the Group of Picture (GOP) length to optimize for spontaneous viewer access, but this often compromises encoding efficiency, especially for popular live streams with intricate background textures and minimal global motion. This paper introduces AGiLE (Adaptive GOP in Live video Encoding), an innovative solution that employs 'pseudo-GOP' to separate encoding efficiency from transmission needs. AGiLE is designed for easy industry adoption and includes a supervised-learning based algorithm for adaptive GOP selection. Our experiments on Douyin's popular live content indicate that AGiLE can reduce bandwidth usage by up to 3.48%, making it a promising solution for the future of live streaming. Cheng Chen 0061, Wenpei Yin, Zhexiong Huang, Shu Shi |
MMSys | 4 |
| 2024 | Enhancing Resource Management of the World's Largest PCDN System for On-Demand Video Streaming
Haiping Wang 0002, Shu Shi, Xiaofei Pang, Yajie Peng, Zhichen Xue, Jiangchuan Liu |
USENIX ATC | 3 |
| 2023 | Poster: E3PO - An Open Platform for 360° Video Streaming Simulation and EvaluationabstractThe simulation and evaluation of diverse 360° video streaming systems present inherent challenges due to the varied design objectives, streaming strategies, and evaluation metrics. This poster introduces E3PO, a versatile and extensible simulation and evaluation platform for 360° video streaming systems. E3PO excels in simulating all proposed variations of 360° video streaming methods, while also generating the precise view that users experience. We have implemented E3PO and developed a variety of examples that simulate different streaming approaches. The promising results from our evaluation of these simulated scenarios demonstrate E3PO's substantial potential to assist researchers in testing new designs, fine-tuning parameters, and comparing performance with their counterparts. Yongqiang Gui, Yanyan Suo, Tian Zhang 0010, Shu Shi |
IMC | 5 |
| 2023 | TwinStar: A Practical Multi-path Transmission Framework for Ultra-Low Latency Video DeliveryabstractUltra-low latency video streaming has received explosive growth in the past few years. However, existing methods all focus on single-path transmission, which is ineffective in dealing with really poor network conditions. To tackle their problems, we propose TwinStar, a novel multi-path framework to improve the experience quality of ultra-low latency video. The core idea of TwinStar is to concurrently leverage multiple paths to mitigate the negative impacts of network jitter on a single path. In particular, by carefully designing the video encoding, data allocation and loss recovery, TwinStar is very robust to handle network dynamics and deliver high-quality video services. We have deployed TwinStar in a commercial cloud gaming platform and evaluated it with real-world networks. The extensive experiments demonstrate that TwinStar significantly outperforms the single-path transmission methods, with 91% reduction in stall ratio and 11% improvement in PSNR across all regions. Haiping Wang 0002, Siping Tao, Hebin Yu, Shu Shi |
ACM Multimedia | 6 |
| 2020 | SiEVE: Semantically Encoded Video Analytics on Edge and CloudabstractRecent advances in computer vision and neural networks have made it possible for more surveillance videos to be automatically searched and analyzed by algorithms rather than humans. This happened in parallel with advances in edge computing where videos are analyzed over hierarchical clusters that contain edge devices, close to the video source. However, the current video analysis pipeline has several disadvantages when dealing with such advances. For example, video encoders have been designed for a long time to please human viewers and be agnostic of the downstream analysis task (e.g., object detection). Moreover, most of the video analytics systems leverage 2-tier architecture where the encoded video is sent to either a remote cloud or a private edge server but does not efficiently leverage both of them. In response to these advances, we present SIEVE, a 3-tier video analytics system to reduce the latency and increase the throughput of analytics over video streams. In SIEVE, we present a novel technique to detect objects in compressed video streams. We refer to this technique as semantic video encoding because it allows video encoders to be aware of the semantics of the downstream task (e.g., object detection). Our results show that by leveraging semantic video encoding, we achieve close to 100% object detection accuracy with decompressing only 3.5% of the video frames which results in more than 100x speedup compared to classical approaches that decompress every video frame. Tarek Elgamal, Shu Shi, Rittwik Jana, Klara Nahrstedt |
ICDCS | 2 |
| 2019 | Mobile VR on edge cloud: a latency-driven designabstractIn this paper we design and implement MEC-VR, a mobile VR system that uses a Mobile Edge Cloud (MEC) to deliver high quality VR content to today's mobile devices using 4G/LTE cellular networks. Our main contribution is in realizing a low latency control loop that streams VR scenes containing only the user's Field of View (FoV) and a latency-adaptive margin area around the FoV. This allows the clients to render locally at a high refresh rate to accommodate and compensate for the head movements before the next motion update arrives. Compared with prior approaches, our MEC-VR design requires no viewpoint prediction, supports dynamic and live VR content, and adapts to the real-world latency experienced in cellular networks between the MEC and mobile devices. We implement a prototype of MEC-VR and evaluate its performance on a MEC node connected to an LTE testbed. We demonstrate that MEC-VR can effectively stream live VR content up to 8K resolution over 4G/LTE networks and achieve more than 80% of bandwidth savings. Shu Shi, Michael Hwang, Rittwik Jana |
MMSys | 1 |
| 2019 | Real time streaming of 8K 360 degree video to mobile VR headsetsabstractIn this demo, we showcase how to stream 8K 360° video to commodity mobile devices without pre-processing or viewpoint prediction using our MEC-VR system. Users can freely select 360° video from YouTube and immediately become immersed in the video of 8K resolution using a Samsung GearVR compatible smartphone. Our system can support zoom in and zoom out while saving up to 80% bandwidth. Shu Shi, Michael Hwang, Bo Han 0001, Vijay Gopalakrishnan, Rittwik Jana |
MMSys | 1 |
| 2019 | Freedom: Fast Recovery Enhanced VR Delivery Over Mobile NetworksabstractIn this paper we design and implement Freedom, a mobile VR system that deliver high quality VR content on today's mobile devices using 4G/LTE cellular networks. Compared to existing state-of-the-art, Freedom does not rely on any video frame pre- rendering or viewpoint prediction. We send a latency-adaptive VAM frame that contains pixels around the FoV. This allows the clients to render locally at a high refresh rate of 60 Hz to accommodate and compensate for the user's head movements before the next server update arrives. We demonstrate that Freedom is the first system in the world that can support dynamic and live 8K resolution VR content, while adapting to the real-world latency variations experienced in cellular networks. Compared to streaming the whole 360° panoramic VR content, we show that Freedom achieves up to 80% bandwidth savings. Finally, we provide detailed end to end latency measurements of actual VR systems by running extensive experiments in a private LTE testbed using a Mobile Edge Cloud (MEC). Shu Shi, Rittwik Jana |
MobiSys | 1 |
| 2019 | Latency Adaptive Streaming of 8K 360 Degree Video to Mobile VR HeadsetsabstractIn this demo, we showcase how to stream 8K 360° video to commod- ity mobile devices without pre-processing or viewpoint prediction using our Freedom system. Users can freely select the viewpoint, zoom in and zoom out to enjoy the full quality of the 8K resolu- tion using any Samsung GearVR compatible smartphone. We also present how our system works internally to dynamically adjust margin size to accommodate to network latency and how our ap- proach can save up to 80% bandwidth compared to streaming full video to mobile devices. Shu Shi, Michael Hwang, Rittwik Jana |
MobiSys | 1 |
| 2017 | LiveJack: Integrating CDNs and Edge Clouds for Live Content BroadcastingabstractEmerging commercial live content broadcasting platforms are facing great challenges to accommodate large scale dynamic viewer populations. Existing solutions constantly suffer from balancing the cost of deploying at the edge close to the viewers and the quality of content delivery. We propose LiveJack, a novel network service to allow CDN servers to seamlessly leverage ISP edge cloud resources. LiveJack can elastically scale the serving capacity of CDN servers by integrating Virtual Media Functions (VMF) in the edge cloud to accommodate flash crowds for very popular contents. LiveJack introduces minor application layer changes for streaming service providers and is completely transparent to end users. We have prototyped LiveJack in both LAN and WAN environments. Evaluations demonstrate that LiveJack can increase CDN server capacity by more than six times, and can effectively accommodate highly dynamic workloads with an improved service quality. Bo Yan 0004, Shu Shi, Yong Liu 0013, Weizhe Yuan, Haoqin He, Rittwik Jana, Yang Xu 0010, H. Jonathan Chao |
ACM Multimedia | 2 |
| 2014 | A Real-Time Smart Display Detection SystemabstractA smart display detection system is proposed that allows users to connect with displays using mobile cameras. The smart displays dynamically update a server with information about the current screen content and the system matches captured images from mobile devices with the screen information. A synchronized timestamped matching strategy is employed to achieve high performance in detecting screens playing motion intensive video and an aggressive feature selection method is used to minimize bandwidth requirements. Shu Shi, John W. Barrus |
ACM Multimedia | 1 |
| 2012 | A real-time remote rendering system for interactive mobile graphicsabstractMobile devices are gradually changing people's computing behaviors. However, due to the limitations of physical size and power consumption, they are not capable of delivering a 3D graphics rendering experience comparable to desktops. Many applications with intensive graphics rendering workloads are unable to run on mobile platforms directly. This issue can be addressed with the idea of remote rendering: the heavy 3D graphics rendering computation runs on a powerful server and the rendering results are transmitted to the mobile client for display. However, the simple remote rendering solution inevitably suffers from the large interaction latency caused by wireless networks, and is not acceptable for many applications that have very strict latency requirements. In this article, we present an advanced low-latency remote rendering system that assists mobile devices to render interactive 3D graphics in real-time. Our design takes advantage of an image based rendering technique: 3D image warping, to synthesize the mobile display from the depth images generated on the server. The research indicates that the system can successfully reduce the interaction latency while maintaining the high rendering quality by generating multiple depth images at the carefully selected viewpoints. We study the problem of viewpoint selection, propose a real-time reference viewpoint prediction algorithm, and evaluate the algorithm performance with real-device experiments. Shu Shi, Klara Nahrstedt, Roy H. Campbell |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2011 | Distortion over latency: Novel metric for measuring interactive performance in remote rendering systemsabstractA new metric distortion over latency (DOL) is proposed in this paper to overcome the deficiency of the traditional metric interaction latency in measuring the interactive performance of the modern remote rendering systems, which are enhanced with different latency reduction techniques. The proposed metric is novel in combining both latency and rendering quality into one score for measurement. Our experiments validate that in many scenarios, our new metric can effectively distinguish the performance difference between systems while interaction latency can not. The paper also introduces how DOL can be efficiently calculated at runtime. Shu Shi, Klara Nahrstedt, Roy H. Campbell |
ICME | 1 |
| 2011 | Tele-immersive gaming for everybodyabstractIn this demonstration, we present two 3D tele-immersive games: light-saber dual and block fencing that merge 3D video representations of participants in real-time to enable remote interactions in a virtual world. The light-saber dual arranges participants in a symmetric setup where both participants interact with each other in a virtual world with similar goals. On the other hand, the block fencing creates an asymmetric setup where participants interact with virtual objects having different goals. Using these two setups, we address the challenges and novelty of our solutions in portable environment setup, data acquisition, multi-stream synchronization, multi-stream session management, mobile device rendering, and overlay communication in the design and implementation of advanced 3D tele-immersive systems. Ahsan Arefin, Zixia Huang, Raoul Rivas, Shu Shi, Wanmin Wu, Klara Nahrstedt |
ACM Multimedia | 4 |
| 2011 | Building low-latency remote rendering systems for interactive 3D graphics rendering on mobile devicesabstractThe recent explosion of mobile devices is changing people's computing behaviors and more and more applications are ported to mobile platforms. However, some applications, such as 3D video tele-immersion and 3D video gaming that require intensive computation or network bandwidth are not capable of running on mobile devices yet. Remote rendering is a simple but effective solution. A workstation with enough computation and network bandwidth resources (e.g., cloud server) is served as the rendering server. It receives and renders the source contents (e.g., 3D graphics or 3D video), and sends the rendering results (2D images) to one or multiple clients. The client simply receives and displays the result images. Using remote rendering for 3D game or 3D video rendering on mobile devices can solve both computation and bandwidth problems. Shu Shi |
ACM Multimedia | 1 |
| 2011 | Using graphics rendering contexts to enhance the real-time video coding for mobile cloud gamingabstractThe emerging cloud gaming service has been growing rapidly, but not yet able to reach mobile customers due to many limitations, such as bandwidth and latency. We introduce a 3D image warping assisted real-time video coding method that can potentially meet all the requirements of mobile cloud gaming. The proposed video encoder selects a set of key frames in the video sequence, uses the 3D image warping algorithm to interpolate other non-key frames, and encodes the key frames and the residues frames with an H.264/AVC encoder. Our approach is novel in taking advantage of the run-time graphics rendering contexts (rendering viewpoint, pixel depth, camera motion, etc.) from the 3D game engine to enhance the performance of video encoding for the cloud gaming service. The experiments indicate that our proposed video encoder has the potential to beat the state-of-art x264 encoder in the scenario of real-time cloud gaming. For example, by implementing the proposed method in a 3D tank battle game, we experimentally show that more than 2 dB quality improvement is possible. Shu Shi, Cheng-Hsin Hsu, Klara Nahrstedt, Roy H. Campbell |
ACM Multimedia | 1 |
| 2011 | ViewMark: An interactive videoconferencing system for mobile devicesabstractViewMark, a server-client based interactive mobile videoconferencing system is proposed in this paper to enhance the remote meeting experience for mobile users. Compared with the state-of-the-art mobile videoconferencing technology, ViewMark is novel in allowing a mobile user to interactively change the viewpoint of the remote video, create viewmarks, and hear with spatial audio. In addition, ViewMark also streams the screen of the presentation slides to mobile devices. In this paper, we introduce the system design of ViewMark in details, compare the devices that can be used to implement interactive videoconferencing, and demonstrate the prototype system we have built on Windows Mobile platform. Shu Shi, Zhengyou Zhang |
MMSP | 1 |
| 2010 | Real-time parallel remote rendering for mobile devices using graphics processing unitsabstractDemand for 3D visualization is increasing in mobile devices as users have come to expect more realistic immersive experiences. However, limited networking and computing resources on mobile devices remain challenges. A solution is to have a proxy-based framework that offloads the burden of rendering computation from mobile devices to more powerful servers. We present the implementation of a framework for parallel remote rendering using commodity Graphics Processing Units (GPUs) in the proxy servers. Experiments show that this framework substantially improves the performance of rendering computation of 3D video. Wucherl Yoo, Shu Shi, Won Jong Jeon, Klara Nahrstedt, Roy H. Campbell |
ICME | 2 |
| 2010 | "I'm the Jedi!" - A Case Study of User Experience in 3D Tele-immersive GamingabstractIn this paper, we present the results from a quantitative and qualitative study of distributed gaming in 3D tele-immersive (3DTI) environments. We explore the Quality of Experience (QoE) of users in the new cyber-physical gaming environment. Guided by a theoretical QoE model, we conducted a case study and evaluated the impact of various Quality of Service (QoS) metrics (e.g., end-to-end delay, visual quality, etc.) on 3DTI gaming experience. We also identified a number of non-technical factors that are not captured by the original theoretical model, such as age, social interaction, and physical setup. Our analysis highlights new implications for the next-generation gaming system design, as well as a more comprehensive conceptual framework that captures non-technical influences for user experience in such environments. Wanmin Wu, Ahsan Arefin, Zixia Huang, Pooja Agarwal, Shu Shi, Raoul Rivas, Klara Nahrstedt |
ISM | 5 |
| 2010 | A high-quality low-delay remote rendering system for 3D videoabstractAs an emerging technology, 3D video has shown a great potential to become the next generation media for tele-immersion. However, streaming and rendering this dynamic 3D data in real-time requires tremendous network bandwidth and computing resources. In this paper, we build a remote rendering model to better study different remote rendering designs and define 3D video rendering as an optimization problem. Moreover, we design a 3D video remote rendering system that significantly reduces the delay while maintaining high rendering quality. We also propose a reference viewpoint prediction algorithm with super sampling support that requires much less computation resources but provides better performance than the search-based algorithms proposed in the related work. Shu Shi, Mahsa Kamali, Klara Nahrstedt, John C. Hart, Roy H. Campbell |
ACM Multimedia | 1 |
| 2009 | Real-time remote rendering of 3D video for mobile devicesabstractAt the convergence of computer vision, graphics, and multimedia, the emerging 3D video technology promises immersive experiences in a truly seamless environment. However, the requirements of huge network bandwidth and computing resources make it still a big challenge to render 3D video on mobile devices at real-time. In this paper, we present how remote rendering framework can be used to solve the problem. The differences between dynamic 3D video and static graphic models are analyzed. A general proxy-based framework is presented to render 3D video streams on the proxy and transmit the rendered scene to mobile devices over a wireless network. An image-based approach is proposed to enhance 3D interactivity and reduce the interaction delay. Experiments prove that the remote rendering framework can be effectively used for quality 3D video rendering on mobile devices in real time. Shu Shi, Won Jong Jeon, Klara Nahrstedt, Roy H. Campbell |
ACM Multimedia | 1 |
| 2009 | MobileTI: a portable tele-immersive systemabstractWe present MobileTI, a portable tele-immersive system that merges 3D video representations of users in real time to enable remote collaboration across geographical distances. With portability as a main goal, we address the challenges in the camera setup, time synchronization, video acquisition, and networking in the design and implementation of the system. Having been deployed in public performances, MobileTI proves to be effective, efficient, and user-friendly. Our experimental findings in terms of technical performance and user feedback are presented. Wanmin Wu, Raoul Rivas, Ahsan Arefin, Shu Shi, Renata M. Sheppard, Bach D. Bui, Klara Nahrstedt |
ACM Multimedia | 4 |
| 2008 | View-dependent real-time 3d video compression for mobile devicesabstract3D video is an emerging technology that promises immersive experiences in a truly seamless environment. Currently, 3D video systems still require excessive bandwidth and computation power provided by gigabit switches and multi-core workstations machines. In order to extend the experience to mobile devices, we present a view-dependent compression methodology that shows great promise in making 3D video a reality on resource-constrained mobile devices. Using our technology, we are able to achieve a software-only rendering on a Nokia N800 PDA with only wireless network transmission. We believe that with the use of newer handhelds and improvements to our compression techniques, we will be able to deliver full-motion 3D video soon. Shu Shi, Klara Nahrstedt, Roy H. Campbell |
ACM Multimedia | 1 |