Yinghao Yu

dblp:147/5299 · DBLP profile ↗
← Back
26ranked-venue papers
8as first author
17since 2021 · last 2026
0000-0002-2744-845XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 15 · 5 first-author · 11 since 2021Computer networks · 10 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 GFS: A Preemption-aware Scheduling Framework for GPU Clusters with Predictive Spot Instance Management
abstract
The surge in large language models (LLMs) has fundamentally reshaped the landscape of GPU usage patterns, creating an urgent need for more efficient management strategies. While cloud providers employ spot instances to reduce costs for low-priority (LP) tasks, existing schedulers still grapple with high eviction rates and lengthy queuing times. To address these limitations, we present GFS, a novel preemptive scheduling framework that enhances service-level objective (SLO) compliance for high-priority (HP) tasks while minimizing preemptions to LP tasks. Firstly, GFS utilizes a lightweight forecasting model that predicts GPU demand among different tenants, enabling proactive resource management. Secondly, GFS employs a dynamic allocation mechanism to adjust the spot quota for LP tasks with guaranteed durations. Lastly, GFS incorporates a preemptive scheduling policy that prioritizes HP tasks while minimizing the impact on LP tasks. We demonstrate the effectiveness of GFS through both real-world implementation and simulations. The results show that GFS reduces eviction rates by 33.0%, and cuts queuing delays by 44.1% for LP tasks. Furthermore, GFS enhances the GPU allocation rate by up to 22.8% in real production clusters. In a production cluster of more than 10,000 GPUs, GFS yields roughly $459,715 in monthly benefits.
Jiaang Duan, Shenglin Xu, Shiyou Qian, Dingyu Yang, Kangjin Wang, Chenzhi Liao, Yinghao Yu, Qin Hua, Hanwen Hu, Dongqing Bao, Tianyu Lu, Jian Cao 0001, Guangtao Xue, Liping Zhang 0013, Gang Chen 0001
ASPLOS (1)7
2026 FlashPS: Efficient Generative Image Editing with Mask-aware Caching and Scheduling
abstract
Generative image editing using diffusion models has become a prevalent application in today's AI cloud services. In production environments, image editing typically involves a mask that specifies the regions of an image template to be edited. The use of mask provides direct control over the editing process and introduces sparsity in the model inference. In this paper, we present FlashPS, a system that efficiently serves image editing requests. The key insight behind FlashPS is that image editing only modifies the masked regions of image templates, while preserving the original content in the unmasked areas. Driven by this insight, FlashPS judiciously skips redundant computations associated with the unmask areas by reusing cached intermediate activations from previous inferences. To mitigate the high cache loading overhead, FlashPS employs a bubble-free pipeline scheme that overlaps computation with cache loading. Additionally, to reduce queuing latency in online serving while improving the GPU utilization, FlashPS proposes a novel continuous batching strategy for diffusion model serving, allowing newly arrived requests to join the running batch in just one step of denoising computation, without waiting for the entire batch to complete. As heterogenous masks induce imbalanced load, FlashPS also develops a load balancing strategy that takes into account the loads of both computation and cache loading. Collectively, FlashPS outperforms state-of-the-art diffusion serving systems for image editing, achieving up to 3× higher throughput and reducing average request latency by up to 14.7× while ensuring image quality.
Xiaoxiao Jiang, Suyi Li 0002, Lingyun Yang, Tianyu Feng, Zhipeng Di, Weiyi Lu, Guoxuan Zhu, Xiu Lin, Yinghao Yu, Tao Lan, Lin Qu, Liping Zhang 0013, Wei Wang 0030
EuroSys10
2026 eGPU: Production-Scale Elastic Sharing Over 10,000 GPUs
abstract
As the cost of GPUs continues to rise, GPU-sharing solutions have become increasingly important for improving efficiency and maximizing resource utilization. At the same time, large-scale operational deployments of such solutions remain relatively less explored, especially in heterogeneous production environments where workload dynamics and orchestration complexity introduce new practical considerations. In this paper, we introduce eGPU, an elastic, efficient, and scalable GPU-sharing framework tailored for production-scale concurrent machine learning (ML) training and inference. eGPU enables fine-grained, runtime-adjustable sharing of GPUs across multiple jobs, while preserving high resource utilization and fault isolation. To address communication bottlenecks, eGPU supports native NVLink/NCCL-based communication between shared GPU instances, capabilities that are limited or unavailable in many existing designs. Built with production deployment in mind, eGPU integrates with Kubernetes (K8s) to support large-scale orchestration. It has been deployed and running stably in production clusters with over$\text{1 0, 0 0 0 ~ G P U s}$for five years. Our evaluation results show that eGPU achieves elastic and precise control over instance sizes, improves job efficiency by 21 % to 31% than SOTA sharing solutions, saves the number of GPUs required by up to$8 \times$, and improves cluster GPU utilization by more than$\mathrm{3} \times$.
Xiaochuan Tang, Hao Qi 0008, Jianbo Dong, Yinghao Yu, Zhennan Xue, Daocheng Ying, Zheng Cao 0003, Xiaoyi Lu 0001
HPCA4
2026 Medley: Optimizing Midgress Bandwidth for Commercial Live Streaming CDNs
Haiping Wang 0002, Wanxin Shi, Sandesh Dhawaskar Sathyanarayana, Shu Shi, Yinghao Yu, La Zuo, Hebin Yu, Ruoshi Sun, Yajie Peng, Xiaofei Pang, Ruili Fang, Zhenpeng Zhu, Yang Xu 0010
NSDI6
2026 Attack of the Bubbles: Straggler-Resilient Pipeline Parallelism for Large Model Training
Tianyuan Wu, Lunxi Cao, Hanfeng Lu, Xiaoxiao Jiang, Yinghao Yu, Siran Yang, Jiamang Wang, Lin Qu, Liping Zhang 0013, Wei Wang 0030
NSDI5
2026 Defrag: Reducing Resource Fragmentation in Large-Scale Heterogeneous GPU Clusters
abstract
Technology companies have built large-scale heterogeneous GPU clusters to support various workloads. However, their cluster machines are found underutilized with severe resource fragmentation. The main causes are myopic online scheduling of incoming tasks and complex placement constraints specified by users or systems. In this paper, we propose to use task migration as a measure to alleviate resource fragmentation. Our trace-driven analysis on a production cluster with 12kmachines reveals that almost at any random snapshot, most of the concurrent tasks have long run-times and small migration times, thus justifying the feasibility of task migration. By making the complex constraints mathematically tractable, we formulate an integer linear programming problem. An efficient heuristic algorithm called Iterative Partitioned Defragmentation (IPD) is presented to perform task migration in multiple iterations of computation. We design and implementDefrag, an operational resource defragmentation system on Kubernetes, and deploy it on the production clusters. Trace-driven experiments show thatDefragcan reduce up to 80% idle CPUs and 29% idle GPUs on average. Furthermore,Defragcan refine the performance of online scheduling strategies by reducing up to 57% idle CPUs and 35% idle GPUs at the time of execution. Our real-world experiment also demonstrates the effectiveness ofDefrag.
Yuedong Xu 0001, Jun Wu 0006, Yinghao Yu
IEEE Trans. Netw.4
2025 Reducing the End-to-End Latency of DNN-Based Recommendation Systems in GPU Pools
abstract
While intelligent applications (e.g., recommendation systems) prefer different CPU-GPU ratios, GPU pooling technique that decouples the GPU and CPU resources yields substantial flexibility when serving diverse applications. With such architecture, DNN-based recommendation services often offload the compute-intensive neural network layers to the remote GPU pool for high resource utilization. However, such a paradigm results in the long end-to-end latency due to two causes: 1) the intermediate data is copied for multiple times during the entire process in current GPU pooling practices, incurring heavy overheads; 2) the content transferred to the GPU pool involves multiple small tensors, suffering from poor bandwidth efficiency. To solve these problems, we design Zero, a runtime system that incorporates a zero-copy transmission mechanism as well as a dynamic tensor merging policy. The zero-copy transmission mechanism unifies memory management across the inference framework and the RPC framework, accompanied by an elaborated serialization protocol to fully eliminate redundant data copying. Meanwhile, the tensor merging policy deliberately organizes small tensors into larger data blocks, so as to transfer them with higher efficiency. Experimental results show that, compared with prior work, Zero reduces the latency of typical recommendation models by up to 15.1% (10.1% on average).
Guangqiang Luan, Pu Pang, Quan Chen 0002, Chen Chen 0067, Guoyao Xu, Chi Zhang 0005, Yanyi Zi, Yinghao Yu, Liping Zhang 0013, Minyi Guo
IPDPS8
2025 GPU-Disaggregated Serving for Deep Learning Recommendation Models at Scale
Lingyun Yang, Yongchen Wang, Yinghao Yu, Qizhen Weng 0001, Jianbo Dong, Chi Zhang 0005, Yanyi Zi, Zechao Zhang, Menglei Zheng, Lanlan Xi, Binzhang Fu, Tao Lan, Liping Zhang 0013, Lin Qu, Wei Wang 0030
NSDI3
2025 Katz: Efficient Workflow Serving for Diffusion Models with Many Adapters
Suyi Li 0002, Lingyun Yang, Xiaoxiao Jiang, Hanfeng Lu, Dakai An, Zhipeng Di, Weiyi Lu, Yinghao Yu, Tao Lan, Lin Qu, Liping Zhang 0013, Wei Wang 0030
USENIX ATC10
2025 GREYHOUND: Hunting Fail-Slows in Hybrid-Parallel Training at Scale
Tianyuan Wu, Wei Wang 0030, Yinghao Yu, Siran Yang, Qinkai Duan, Jiamang Wang, Lin Qu, Liping Zhang 0013
USENIX ATC3
2023 Beware of Fragmentation: Scheduling GPU-Sharing Workloads with Fragmentation Gradient Descent
Qizhen Weng 0001, Lingyun Yang, Yinghao Yu, Wei Wang 0030, Xiaochuan Tang, Liping Zhang 0013
USENIX ATC3
2022 Workload consolidation in alibaba clusters: the good, the bad, and the ugly
abstract
Web companies typically run latency-critical long-running services and resource-intensive, throughput-hungry batch jobs in a shared cluster for improved utilization and reduced cost. Despite many recent studies on workload consolidation, the production practice remains largely unknown. This paper describes our efforts to efficiently consolidate the two types of workloads in Alibaba clusters to support the company's e-commerce businesses.
Yongkang Zhang 0003, Yinghao Yu, Wei Wang 0030, Qiukai Chen, Tianchen Ding, Qizhen Weng 0001, Lingyun Yang, Jian He 0004, Liping Zhang 0013
SoCC2
2022 MLaaS in the Wild: Workload Analysis and Scheduling in Large-Scale Heterogeneous GPU Clusters
Qizhen Weng 0001, Wencong Xiao, Yinghao Yu, Wei Wang 0030, Jian He 0004, Yong Li 0045, Liping Zhang 0013, Wei Lin 0016
NSDI3
2022 Towards Dependency-Aware Cache Management for Data Analytics Applications
abstract
Memory caches are being used aggressively in today's data analytics systems such as Spark, Tez, and Piccolo. The significant performance impact of caches and their limited sizes call for efficient cache management in data analytics clusters. However, prevalent data analytics systems employ rather simple cache management policies—notably Least Recently Used (LRU) and Least Frequently Used (LFU)—that areobliviousto the application semantics of data dependency, expressed as directed acyclic graphs (DAGs). Without this knowledge, cache management can, at best, be performed by “guessing” the future data access patterns based on history, which frequently results in inefficient, erroneous caching with a low hit rate and a long response time. Worse still, the lack of data dependency knowledge makes it impossible to retain theall-or-nothingcache property of cluster applications, in that a compute task cannot be sped up unless all the dependent data has been kept in the main memory. In this paper, we propose a novel cache replacement policy, named Least Reference Count (LRC), which exploits the application's data dependency information to optimize the cache management. LRC keeps track of thereference countof each data block, defined as the number of dependent child blocks that have not been computed yet, and always evicts the block with the smallest reference count. Furthermore, we incorporate the all-or-nothing requirement into LRC by coordinately managing the reference counts of all the input data blocks for the same computation. We demonstrate the efficacy of LRC through both empirical analysis and cluster deployments against popular benchmarking workloads. Our Spark implementation shows that, the proposed policies well address the all-or-nothing requirement and significantly improve the cache performance. Compared with LRU and a recently proposed caching policy called MEMTUNE, LRC improves the caching performance of typical workloads in production clusters by 22 and 284 percent, respectively.
Yinghao Yu, Chengliang Zhang, Wei Wang 0030, Jun Zhang 0004, Khaled Ben Letaief
IEEE Trans. Cloud Comput.1
2022 Missing Data Repairs for Traffic Flow With Self-Attention Generative Adversarial Imputation Net
abstract
With the rapid development of sensor technologies, time series data collected by multiple and spatially distributed sensors have been widely used in different research fields. Examples of such data include geo-tagged temperature data collected by temperature sensors, air pollutant monitoring data, and traffic data collected by road traffic sensors. Due to sensor failure, communication errors and storage loss, etc., data collected by sensors inevitably includes missing data. However, models commonly used in the analysis of such large-scale data often rely on complete data sets. This paper proposes a model for the imputation of missing data of traffic flow, which combines a self-attention mechanism, an auto-encoder, and a generative adversarial network, into a self-attention generative adversarial imputation net (SA-GAIN). The introduction of the self-attention mechanism can help the proposed model to effectively capture correlations between spatially-distributed sensors at different time points. Adversarial training through two neural networks, called generators and discriminators, allows the proposed model to generate imputed data close to the real data. In comparison with different imputation models, the proposed model shows the best performance in imputing missing data.
Pulin Zhang, Yinghao Yu, Salvatore Antonio Biancardo
IEEE Trans. Intell. Transp. Syst.3
2021 George: Learning to Place Long-Lived Containers in Large Clusters with Operation Constraints
abstract
Online cloud services are widely deployed as Long-Running Applications (LRAs) hosted in containers. Placing LRA containers turns out to be particularly challenging due to the complex interference between co-located containers and the operation constraints in production clusters such as fault tolerance, disaster avoidance and incremental deployment. Existing schedulers typically provide APIs for operators to manually specify the container scheduling requirements and offer only qualitative scheduling guidelines for container placement. Such schedulers, do not perform well in terms of both performance and scale, while also requiring manual intervention.
Suyi Li 0002, Wei Wang 0030, Yinghao Yu, Bo Li 0001
SoCC4
2021 Morphling: Fast, Near-Optimal Auto-Configuration for Cloud-Native Model Serving
abstract
Machine learning models are widely deployed in production cloud to provide online inference services. Efficiently deploying inference services requires careful tuning of hardware and runtime configurations (e.g., GPU type, GPU memory, batch size), which can significantly improve the model serving performance and reduce cost. However, existing autoconfiguration approaches for general workloads, such as Bayesian optimization and white-box prediction, are inefficient in navigating the high-dimensional configuration space of model serving, incurring high sampling cost.
Lingyun Yang, Yinghao Yu, Wei Wang 0030, Bo Li 0001, Xianchao Sun, Jian He 0004, Liping Zhang 0013
SoCC3
2020 RepBun: Load-Balanced, Shuffle-Free Cluster Caching for Structured Data
abstract
Cluster caching systems increasingly store structured data objects in the columnar format. However, these systems routinely face the imbalanced load that significantly impairs the I/O performance. Existing load-balancing solutions, while effective for reading unstructured data objects, fall short in handling columnar data. Unlike unstructured data that can only be read through a full-object scan, columnar data supports direct query of specific columns with two distinct access patterns: (1) columns have the heavily skewed popularity, and (2) hot columns are likely accessed together in a query job. Based on these two access patterns, we propose an effective load-balancing solution for structured data. Our solution, which we call RepBun, groups hot columns into a bundle. It then copies multiple replicas of the column bundle and stores them uniformly across servers. We show that RepBun achieves improved load balancing with reduced memory overhead, while avoiding data shuffling between cache servers. We implemented RepBun atop Alluxio, a popular in-memory distributed storage, and evaluate its performance through EC2 deployment against the TPC-H benchmark work-load. Experimental results show that RepBun outperforms the existing load-balancing solutions with significantly shorter read latency and faster query completion.
Minchen Yu, Yinghao Yu, Yunchuan Zheng, Baichen Yang, Wei Wang 0030
INFOCOM2
2020 Achieving Load-Balanced, Redundancy-Free Cluster Caching with Selective Partition
abstract
Data-intensive clusters increasingly rely on in-memory storages to improve I/O performance. However, the routinely observed file popularity skew and load imbalance create hot spots, which significantly degrade the benefits of in- memory caching. Common approaches to tame load imbalance include copying multiple replicas of hot files and creating parity chunks using storage codes. Yet, these techniques either suffer from high memory overhead due to cache redundancy or incur non-trivial encoding/decoding complexity. In this paper, we propose an effective approach to achieve load balancing without cache redundancy or encoding/decoding overhead. Our solution, termed SP-Cache,selectively partitionsfiles based on the loads they contribute and evenly caches those partitions across the cluster. We develop an efficient algorithm to determine the optimal number of partitions for a hot file—too few partitions are incapable of mitigating hot spots, while too many are susceptible to stragglers. We have implemented SP-Cache atop Alluxio, a popular in-memory distributed storage system, and evaluated its performance through EC2 deployment and trace-driven simulations. SP-Cache can quickly react to the changing load by dynamically re-balancing cache servers. Compared to the state-of-the-art solution, SP-Cache reduces the file access latency by up to 40 percent in both the mean and the tail, using 40 percent less memory.
Yinghao Yu, Wei Wang 0030, Renfei Huang, Jun Zhang 0004, Khaled Ben Letaief
IEEE Trans. Parallel Distributed Syst.1
2019 LACS: Load-Aware Cache Sharing with Isolation Guarantee
abstract
Cluster caching has been increasingly deployed in front of cloud storage to improve I/O performance. In shared, multi-tenant environments such as cloud datacenters, cluster caches are constantly contended by many users. Enforcing performance isolation between users hence becomes imperative to cluster caching. A user's caching performance critically depends on two factors: (1) the amount of cache allocation and (2) the load of servers in which its files are cached. However, existing cache sharing policies only provide guarantees on the amount of cache allocation, while remaining agnostic to the load of cache servers. Consequently, "mice" users having files co-located with "elephants" contributing heavy data accesses may experience extremely long latency, hence receiving no isolation. In this paper, we propose a Load-Aware Cache Sharing scheme (LACS) to enforce isolation between users. LACS keeps track of the load contributed by each user and reins back the congestions caused by elephant users by throttling their cache usage and network bandwidth. We have implemented LACS atop Alluxio, a popular cluster caching system. EC2 deployment shows that LACS achieves performance isolation in the presence of elephants, while improving the mean read latency by up to 80.4% (25.3% on average) over the state-of-the-art load balancing technique.
Yinghao Yu, Wei Wang 0030, Jun Zhang 0004, Khaled Ben Letaief
ICDCS1
2018 OpuS: Fair and Efficient Cache Sharing for In-Memory Data Analytics
abstract
We study the fair cache allocation problem in shared cloud environments, where many users and applications contend for the main memory to cache shared datasets or files. Unlike other resources such as CPUs and networks, in-memory caches can be non-exclusively shared across many users, e.g., a cached columnar dataset queried by many Spark SQL jobs. This results in a unique challenge of the "free-riding" problem, where a user lies about its caching preferences to trick other users to cache files for it, using their allocated cache space. We show that existing cache allocation policies either suffer from such manipulations or result in poor efficiency. To address this problem, we propose a new cache allocation algorithm, termed OpuS, or Opportunistic Sharing for high efficiency. We show that OpuS provides performance isolation between users and is strategy-proof against "free-riding" manipulations. We have implemented OpuS as a pluggable cache manager in Alluxio, a popular memory-centric filesystem. Cluster deployment and trace-driven simulations demonstrate that OpuS allocates each user a fair share of caches while achieving near-optimal efficiency in cache utilization.
Yinghao Yu, Wei Wang 0030, Jun Zhang 0004, Qizhen Weng 0001, Khaled Ben Letaief
ICDCS1
2018 SP-cache: load-balanced, redundancy-free cluster caching with selective partition
Yinghao Yu, Renfei Huang, Wei Wang 0030, Jun Zhang 0004, Khaled Ben Letaief
SC1
2017 LERC: Coordinated Cache Management for Data-Parallel Systems
abstract
Memory caches are being aggressively used in today's data- parallel frameworks such as Spark, Tez and Storm. By caching input and intermediate data in memory, compute tasks can witness speedup by orders of magnitude. To maximize the chance of in-memory data access, existing cache algorithms, be it recency- or frequency-based, settle on cache hit ratio as the optimization objective. However, unlike the conventional belief, we show in this paper that simply pursuing a higher cache hit ratio of individual data blocks does not necessarily translate into faster task completion in data-parallel environments. A data-parallel task typically depends on multiple input data blocks. Unless all of these blocks are cached in memory, no speedup will result. To capture this all-or-nothing property, we propose a more relevant metric, called effective cache hit ratio. Specifically, a cache hit of a data block is said to be effective if it can speed up a compute task. In order to optimize the effective cache hit ratio, we propose the Least Effective Reference Count (LERC) policy that persists the dependent blocks of a compute task as a whole in memory. We have implemented the LERC policy as a memory manager in Spark and evaluated its performance through Amazon EC2 deployment. Evaluation results demonstrate that LERC helps speed up data-parallel jobs by up to 37% compared with the widely employed least-recently-used (LRU) policy.
Yinghao Yu, Wei Wang 0030, Jun Zhang 0004, Khaled Ben Letaief
GLOBECOM1
2017 LRC: Dependency-aware cache management for data analytics clusters
abstract
Memory caches are being aggressively used in today's data-parallel systems such as Spark, Tez, and Piccolo. However, prevalent systems employ rather simple cache management policies — notably the Least Recently Used (LRU) policy — that are oblivious to the application semantics of data dependency, expressed as a directed acyclic graph (DAG). Without this knowledge, memory caching can at best be performed by “guessing” the future data access patterns based on historical information (e.g., the access recency and/or frequency), which frequently results in inefficient, erroneous caching with low hit ratio and a long response time. In this paper, we propose a novel cache replacement policy, Least Reference Count (LRC), which exploits the application-specific DAG information to optimize the cache management. LRC evicts the cached data blocks whose reference count is the smallest. The reference count is defined, for each data block, as the number of dependent child blocks that have not been computed yet. We demonstrate the efficacy of LRC through both empirical analysis and cluster deployments against popular benchmarking workloads. Our Spark implementation shows that, compared with LRU, LRC speeds up typical applications by 60%.
Yinghao Yu, Wei Wang 0030, Jun Zhang 0004, Khaled Ben Letaief
INFOCOM1
2016 Joint Subcarrier and CPU Time Allocation for Mobile Edge Computing
abstract
In mobile edge computing systems, mobile devices can offload compute-intensive tasks to a nearby cloudlet, so as to save energy and extend battery life. Unlike a fully-fledged cloud, a cloudlet is a small-scale datacenter deployed at a wireless access point, and thus is highly constrained by both radio and compute resources. We show in this paper that separately optimizing the allocation of either compute or radio resource - as most existing works did - is highly suboptimal: the congestion of compute resource leads to the waste of radio resource, and vice versa. To address this problem, we propose a joint scheduling algorithm that allocates both radio and compute resources coordinately. Specifically, we consider a cloudlet in an Orthogonal Frequency-Division Multiplexing Access (OFDMA) system with multiple mobile devices, where we study subcarrier allocation for task offloading and CPU time allocation for task execution in the cloudlet. Simulation results show that the proposed algorithm significantly outperforms per-resource optimization, accommodating more offloading requests while achieving salient energy saving.
Yinghao Yu, Jun Zhang 0004, Khaled Ben Letaief
GLOBECOM1
2016 Flow-Level QoE of Video Streaming in Wireless Networks
abstract
The Quality of Experience (QoE) of streaming service is often degraded by frequent playback interruptions. To mitigate the interruptions, the media player prefetches streaming contents before starting playback, at a cost of initial delay. We study the QoE of streaming from the perspective of flow dynamics. First, a framework is developed for QoE when streaming users join the network randomly and leave after downloading completion. We model the distribution of prefetching delay using partial differential equations (PDEs), and the probability generating function of playout buffer starvations using ordinary differential equations (ODEs) for constant bit-rate (CBR) streaming. The explicit form starvation probabilities and mean start-up delay are obtained by use of a matrix function approach. Second, we extend our framework to characterize the throughput variation caused by opportunistic scheduling at the base station, and the playback variation of variable bit-rate (VBR) streaming. Our study reveals that the flow dynamics is the fundamental reason of playback starvation. The QoE of streaming service is dominated by the first moments such as the average throughput of opportunistic scheduling and the mean playback rate. While the variances of throughput and playback rate have very limited impact on starvation behavior in practice.
Yuedong Xu 0001, Salah-Eddine Elayoubi, Eitan Altman, Rachid El Azouzi, Yinghao Yu
IEEE Trans. Mob. Comput.5