Chen Wang 0039

dblp:82/4206-39 · DBLP profile ↗
← Back
30ranked-venue papers
5as first author
20since 2021 · last 2026
0000-0003-0204-2362ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 4 since 2021Computer networks · 5 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 4 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Who Watches the Watchers? On the Reliability of Softwarizing Cloud Application Management
Jiawei Tyler Gu, Yiming Su, Bogdan Alexandru Stoica, Xudong Sun 0013, William X. Zheng, Akond Ashfaque Ur Rahman, Chen Wang 0039, Tianyin Xu
NSDI9
2026 Sakkara: Intelligent Topology-Aware Scheduling for Kubernetes in the Age of AI
abstract
The rapid growth of Artificial Intelligence (AI) workloads has introduced unprecedented challenges to modern cloud-native systems, particularly in Kubernetes (K8s)-based environments. These workloads often demand low-latency communication, high resource locality, and efficient utilization of heterogeneous hardware devices such as Graphics Processing Units (GPUs) and specialized accelerators. However, the existing scheduling mechanisms in K8s are typically unaware of the underlying physical topology, leading to performance degradation and inefficient resource usage. This paper presents Sakkara, a novel topology-aware scheduling framework designed to optimize the placement of AI workloads in K8s clusters. Sakkara incorporates a hierarchical model of the Data Center (DC), including nodes and racks, enabling flexible scheduling strategies that account for resource availability and risk-aware metrics that mitigate performance interference and constraint violations caused by topology-unaware placement. Sakkara extends existing scheduling logic in K8s with placement strategies that guide pod allocation using configurable topology constraints, aiming to minimize communication costs and maximize workload performance. We evaluated Sakkara on a representative AI workload, a distributed training application under different cluster configurations. Experimental results show that Sakkara improves job completion time, throughput, and memory utilization compared to available K8s schedulers, achieving improvements of up to 10%. Sakkara, available as open-source, offers a promising pathway toward topology-conscious orchestration of AI workloads in next-generation cloud environments.
José Santos 0001, Asser N. Tantawi, Pavlos Maniotis, Chen Wang 0039, Olivier Tardieu, Tim Wauters, Filip De Turck
IEEE Trans. Netw. Serv. Manag.4
2025 Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference
abstract
Large language models have been widely adopted across different tasks, but their auto-regressive generation nature often leads to inefficient resource utilization during inference. While batching is commonly used to increase throughput, per-formance gains plateau beyond a certain batch size, especially with smaller models, a phenomenon that existing literature typically explains as a shift to the compute-bound regime. In this paper, through an in-depth GPU-level analysis, we reveal that large-batch inference remains memory-bound, with most GPU compute capabilities underutilized due to DRAM bandwidth saturation as the primary bottleneck. To address this, we propose a Batching Configuration Advisor (BCA) that optimizes memory allocation, reducing GPU memory requirements with minimal impact on throughput. The freed memory and underutilized GPU compute capabilities can then be leveraged by concurrent workloads. Specifically, we use model replication to improve serving throughput and GPU utilization. Our findings challenge conventional assumptions about LLM inference, offering new in-sights and practical strategies for improving resource utilization, particularly for smaller language models.
Pol G. Recasens, Ferran Agullo, Chen Wang 0039, Olivier Tardieu, Jordi Torres, Josep Lluís Berral
CLOUD4
2025 Towards Efficient Key-Value Cache Management for Prefix Prefilling in LLM Inference
abstract
The increasing adoption of large language models (LLMs) with extended context windows necessitates efficient Key-Value Cache (KVC) management to optimize inference performance. Inference workloads like Retrieval-Augmented Generation (RAG) and agents exhibit high cache reusability, making efficient caching critical to reducing redundancy and improving speed. We analyze real-world KVC access patterns using publicly available traces and evaluate commercial key-value stores like Redis and state-of-the-art RDMA-based systems (CHIME [1] and Sherman [2]) for KVC metadata management. Our work demonstrates the lack of tailored storage solution for KVC prefilling, underscores the need for an efficient distributed caching system with optimized metadata management for LLM workloads, and provides insights into designing improved KVC management systems for scalable, low-latency inference.
Chen Wang 0039
CLOUD3
2025 A House United Within Itself: SLO-Awareness for On-Premises Containerized ML Inference Clusters via Faro
abstract
This paper tackles the challenge of running multiple ML inference jobs (models) under time-varying workloads, on a constrained on-premises production cluster. Our system Faro takes in latency Service Level Objectives (SLOs) for each job, auto-distills them into utility functions, "sloppifies" these utility functions to make them amenable to mathematical optimization, automatically predicts workload via probabilistic prediction, and dynamically makes implicit cross-job resource allocations, in order to satisfy cluster-wide objectives, e.g., total utility, fairness, and other hybrid variants. A major challenge Faro tackles is that using precise utilities and high-fidelity predictors, can be too slow (and in a sense too precise!) for the fast adaptation we require. Faro's solution is to "sloppify" (relax) its multiple design components to achieve fast adaptation without overly degrading solution quality. Faro is implemented in a stack consisting of Ray Serve running atop a Kubernetes cluster. Trace-driven cluster deployments show that Faro achieves 2.3×-23× lower SLO violations compared to state-of-the-art systems.
Beomyeol Jeon, Chen Wang 0039, Diana Arroyo, Alaa Youssef, Indranil Gupta
EuroSys2
2025 Evaluating the Network Effects of Orchestration Strategies for AI Workloads in Modern Data Centers
abstract
The exponential growth in Artificial Intelligence (AI) adoption presents unique challenges and opportunities for deploying AI workloads in modern Data Center (DC) networks, particularly in terms of performance, scalability, and reliability. AI workloads, such as inference and distributed training, impose different network demands: inference is primarily computebound and typically requires low network latency, while distributed training is network-bound and requires high bandwidth, placing significant strain on the network. This paper focuses on the network requirements of widely known AI communication patterns, and studies their impact on modern DC architectures by analyzing the effects of different orchestration strategies-specifically packing and spreading-on throughput, response time, and network congestion. The results show that packing strategies generally deliver higher performance for most covered AI collectives. However, spreading strategies can be beneficial in certain scenarios, such as when larger workloads span across higher number of racks, as they can help mitigate network congestion between the switches of leaf-spine network configurations. This paper offers valuable insights into optimizing the orchestration of popular AI collectives in data center networks, presenting informed strategies to improve performance in response to growing AI demands, with findings demonstrating completion time reductions of up to 30 %.
José Santos 0001, Pavlos Maniotis, Chen Wang 0039, Asser N. Tantawi, Olivier Tardieu, Tim Wauters, Filip De Turck
NetSoft3
2024 Optimizing Simultaneous Autoscaling for Serverless Cloud Computing
abstract
This paper explores resource allocation in server-less cloud computing platforms and proposes an optimization approach for autoscaling systems. Serverless computing relieves users from resource management tasks, enabling focus on application functions. However, dynamic resource allocation and function replication based on changing loads remain crucial. Typically, autoscalers in these platforms utilize threshold-based mechanisms to adjust function replicas independently. We model applications as interconnected graphs of functions, where requests probabilistically traverse the graph, triggering associated function execution. Our objective is to develop a control policy that optimally allocates resources on servers, minimizing failed requests and response time in reaction to load changes. Using a fluid approximation model and Separated Continuous Linear Programming (SCLP), we derive an optimal control policy that determines the number of resources per replica and the required number of replicas over time. We evaluate our approach using a simulation framework built with Python and simpy. Comparing against threshold-based autoscaling, our approach demonstrates significant improvements in average response times and failed requests, ranging from 15% to over 300% in most cases. We also explore the impact of system and workload parameters on performance, providing insights into the behavior of our optimization approach under different conditions. Overall, our study contributes to advancing resource allocation strategies, enhancing efficiency and reliability in serverless cloud computing platforms.
Harold J. Ship, Evgeny Shindin, Chen Wang 0039, Diana Arroyo, Asser N. Tantawi
CLOUD3
2024 Cloud-native Workflow Scheduling using a Hybrid Priority Rule, Dynamic Resource Allocation, and Dynamic Task Partition
abstract
As cloud-native workflow orchestration tools become increasingly important for complex data science workloads, there is a growing need for more efficient scheduling. Existing cloud schedulers rely on basic heuristics and user choice for task partitioning for parallel computing, leading to under-utilization of cluster resources and prolonged job completion times. To address this, we propose a novel workflow scheduling algorithm that leverages workflow characteristics to enhance resource utilization and reduce weighted job completion time. The algorithm combines three sub-algorithms, each reflecting a distinct aspect of the scheduling strategy: 1) Hybrid Maximum Children (MC) -Weighted Shortest Critical Path Time (WSCPT) rule alternates between two heuristics, MC and WSCPT, which prioritize jobs based on workflow structure and critical path, respectively. The choice between these heuristics is dynamically adjusted according to the cluster queue size. 2) Dynamic Resource Allocation (DRA), which dynamically adjusts the number of executors assigned to each workflow, and 3) Dynamic Task Partition (DTP), which autonomously determines the task parallelism level. We tested our algorithm with extensive experiments on various workflow types using Spark-imitated simulation. Our algorithm outperformed other schedulers, including learning-based models, by reducing 21-47% of the combined performance of average job completion time and makespan for unweighted workflows and reducing at least 50% of weighted job completion time for weighted workflows.
Jungeun Shin, Diana Arroyo, Asser N. Tantawi, Chen Wang 0039, Alaa Youssef, Rakesh Nagi
SoCC4
2024 Optimizing GPU Multiplexing for Efficient and Cost-Effective Access to Diverse Large Language Models in GPU Clusters
abstract
Large Language Models (LLMs) are a cornerstone of modern artificial intelligence research, gaining popularity and encouraging adoption in varying domains. The burgeoning interest among researchers drives the need for building a dedicated GPU cluster serving an extensive and varied collection of LLMs, specifically designed to facilitate experimental exploration and innovation in the field of AI research. However, this requirement poses significant challenges in optimizing GPU multiplexing. The primary problem is ensuring that a wide variety of fast-changing LLM models remain accessible to researchers while minimizing GPU costs and reducing idle time. This is further complicated by the fluctuating popularity of models; at any given time, only a few models would be in high demand, leading to underutilization of expensive GPU resources. To address this challenge, the paper presents a methodology that involves profiling the performance of various LLMs across different Multi-Instance GPU (MIG) slice configurations under various loads. By analyzing the trade-offs between GPU utilization and LLM inference performance on vLLM server, we identify the optimal MIG partitioning strategies that can dynamically adapt to the changing landscape of model popularity and usage patterns with latency SLA retained ($50 \mathrm{~ms} /$ token), thereby reducing the GPU cost by 50% and enhancing the energy efficiency by 35% in our GPU cluster.
Chen Wang 0039, Max Calman, Rina Nakazawa
MASCOTS2
2024 Dexter: A Performance-Cost Efficient Resource Allocation Manager for Serverless Data Analytics
abstract
Leveraging serverless platforms for the efficient execution of distributed data analytics frameworks, such as Apache Spark [3], has gained substantial interest since early 2022. The elasticity, free-of-management, and on-demand scalability of serverless have motivated the effort in deploying distributed data analytics applications to serverless platforms. However, effectively auto-scaling resources for such complex workloads so that we can fully benefit from the resource elasticity of serverless remains challenging. Mis-configuration can result in severe performance and cost issues arising from resource under- and over-provisioning.
Anna Maria Nestorov, Diego Marron, Alberto Gutierrez-Torre, Chen Wang 0039, Claudia Misale, Alaa Youssef, David Carrera 0001, Josep Lluís Berral
Middleware4
2024 Power-aware Deep Learning Model Serving with μ-Serve
Haoran Qiu, Weichao Mao, Archit Patke, Shengkun Cui, Saurabh Jha, Chen Wang 0039, Hubertus Franke, Zbigniew T. Kalbarczyk, Tamer Basar, Ravishankar K. Iyer
USENIX ATC6
2023 Advancing Cloud Sustainability: A Versatile Framework for Container Power Model Training
abstract
Estimating power consumption in modern Cloud is important to account for the power consumed by each container. The challenge is that multiple customers are sharing the same hardware platform, where physical information is mostly obscured. In addition, there is the overhead in power consumption that the Cloud control plane induces. This paper addresses these challenges and introduces a pipeline framework for container power model training on the basis of available performance counters and other metrics. The proposed model utilizes machine learning techniques to predict the power consumed by the control plane and associated processes when running together with the user containers, and uses it for isolating the dynamic power consumed by the user-inducing workload. Applying the proposed power model does not require online power measurements, nor does it need machine information, or information on other tenants sharing the same machine. The results of cross-workload, cross-platform experiments demonstrated the higher accuracy of the model when predicting power consumption of unseen containers on unknown platforms, including on virtual machines.
Sunyanan Choochotkaew, Chen Wang 0039, Tatsuhiro Chiba, Marcelo Amaral, Tamar Eilam
MASCOTS2
2023 Multi-Agent Meta-Reinforcement Learning: Sharper Convergence Rates with Task Similarity
abstract
Multi-agent reinforcement learning (MARL) has primarily focused on solving a single task in isolation, while in practice the environment is often evolving, leaving many related tasks to be solved. In this paper, we investigate the benefits of meta-learning in solving multiple MARL tasks collectively. We establish the first line of theoretical results for meta-learning in a wide range of fundamental MARL settings, including learning Nash equilibria in two-player zero-sum Markov games and Markov potential games, as well as learning coarse correlated equilibria in general-sum Markov games. Under natural notions of task similarity, we show that meta-learning achieves provable sharper convergence to various game-theoretical solution concepts than learning each task separately. As an important intermediate step, we develop multiple MARL algorithms with initialization-dependent convergence guarantees. Such algorithms integrate optimistic policy mirror descents with stage-based value updates, and their refined convergence guarantees (nearly) recover the best known results even when a good initialization is unknown. To our best knowledge, such results are also new and might be of independent interest. We further provide numerical simulations to corroborate our theoretical findings.
Weichao Mao, Haoran Qiu, Chen Wang 0039, Hubertus Franke, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Tamer Basar
NeurIPS3
2023 Acto: Automatic End-to-End Testing for Operation Correctness of Cloud System Management
abstract
Cloud systems are increasingly being managed by operation programs termed operators, which automate tedious, human-based operations. Operators of modern management platforms like Kubernetes, Twine, and ECS implement declarative interfaces based on the state-reconciliation principle. An operation declares a desired system state and the operator automatically reconciles the system to that declared state.
Jiawei Tyler Gu, Xudong Sun 0013, Yuxuan Jiang 0016, Chen Wang 0039, Mandana Vaziri, Owolabi Legunsen, Tianyin Xu
SOSP5
2023 AWARE: Automate Workload Autoscaling with Reinforcement Learning in Production Cloud Systems
Haoran Qiu, Weichao Mao, Chen Wang 0039, Hubertus Franke, Alaa Youssef, Zbigniew T. Kalbarczyk, Tamer Basar, Ravishankar K. Iyer
USENIX ATC3
2023 Diktyo: Network-Aware Scheduling in Container-Based Clouds
abstract
Containers have revolutionized application deployment and life-cycle management in current cloud platforms. Applications have evolved from single monoliths to complex graphs of loosely-coupled microservices. However, the efficient allocation of microservice-based applications is challenging due to their complex inter-dependencies. Further, recent applications are becoming even more delay-sensitive, demanding lower latency between dependent microservices. Scheduling policies in popular container orchestration platforms mainly aim to increase the resource efficiency of the infrastructure, insufficient for latency-sensitive applications. Application domains such as the Internet of Things and multi-tier Web services would benefit from network-aware policies that consider network latency and bandwidth in the scheduling process. Previous works have studied network-aware scheduling via theoretical formulations or heuristic-based methods evaluated via simulations or small testbeds, making their full applicability in popular platforms difficult. This paper proposes a novel network-aware framework for the popular Kubernetes (K8s) platform named Diktyo that determines the placement of dependent microservices in long-running applications focused on reducing the application’s end-to-end latency and guaranteeing bandwidth reservations. Simulations show that Diktyo can significantly reduce the network latency for various applications across different infrastructure topologies compared to default K8s scheduling plugins. Also, experiments in a K8s cluster with microservice benchmark applications show that Diktyo can increase database throughput by 22% and reduce application response time by 45%.
José Santos 0001, Chen Wang 0039, Tim Wauters, Filip De Turck
IEEE Trans. Netw. Serv. Manag.2
2022 SIMPPO: a scalable and incremental online learning framework for serverless resource management
abstract
Serverless Function-as-a-Service (FaaS) offers improved programmability for customers, yet it is not server-"less" and comes at the cost of more complex infrastructure management (e.g., resource provisioning and scheduling) for cloud providers. To maintain service-level objectives (SLOs) and improve resource utilization efficiency, recent research has been focused on applying online learning algorithms such as reinforcement learning (RL) to manage resources. Despite the initial success of applying RL, we first show in this paper that the state-of-the-art single-agent RL algorithm (S-RL) suffers up to 4.8x higher p99 function latency degradation on multi-tenant serverless FaaS platforms compared to isolated environments and is unable to converge during training. We then design and implement a scalable and incremental multi-agent RL framework based on Proximal Policy Optimization (SIMPPO). Our experiments demonstrate that in multi-tenant environments, SIMPPO enables each RL agent to efficiently converge during training and provides online function latency performance comparable to that of S-RL trained in isolation with minor degradation (<9.2%). In addition, SIMPPO reduces the p99 function latency by 4.5x compared to S-RL in multi-tenant cases.
Haoran Qiu, Weichao Mao, Archit Patke, Chen Wang 0039, Hubertus Franke, Zbigniew T. Kalbarczyk, Tamer Basar, Ravishankar K. Iyer
SoCC4
2022 Cloud-native workflow scheduling using a hybrid priority rule and dynamic task parallelism
abstract
Demand for efficient cloud-native workflow scheduling is growing as many data science workloads are composed of several tasks with dependencies. As container technology becomes more prevalent in cloud communities, containerized workflow orchestration tools are introduced and become standard for scheduling workflows. However, current schedulers use simple heuristics and rely on the user's choice on priority and parallelism level of tasks without accounting for workflow-specific information.
Jungeun Shin, Diana Arroyo, Asser N. Tantawi, Chen Wang 0039, Alaa Youssef, Rakesh Nagi
SoCC4
2022 A Mean-Field Game Approach to Cloud Resource Management with Function Approximation
abstract
Reinforcement learning (RL) has gained increasing popularity for resource management in cloud services such as serverless computing. As self-interested users compete for shared resources in a cluster, the multi-tenancy nature of serverless platforms necessitates multi-agent reinforcement learning (MARL) solutions, which often suffer from severe scalability issues. In this paper, we propose a mean-field game (MFG) approach to cloud resource management that is scalable to a large number of users and applications and incorporates function approximation to deal with the large state-action spaces in real-world serverless platforms. Specifically, we present an online natural actor-critic algorithm for learning in MFGs compatible with various forms of function approximation. We theoretically establish its finite-time convergence to the regularized Nash equilibrium under linear function approximation and softmax parameterization. We further implement our algorithm using both linear and neural-network function approximations, and evaluate our solution on an open-source serverless platform, OpenWhisk, with real-world workloads from production traces. Experimental results demonstrate that our approach is scalable to a large number of users and significantly outperforms various baselines in terms of function latency and resource utilization efficiency.
Weichao Mao, Haoran Qiu, Chen Wang 0039, Hubertus Franke, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer, Tamer Basar
NeurIPS3
2021 Theta-Scan: Leveraging Behavior-Driven Forecasting for Vertical Auto-Scaling in Container Cloud
abstract
Detection of behavior patterns on resource usage in containerized Cloud applications is necessary for proper resource provisioning. Applications can use CPU/Memory with repetitive patterns, following a trend over time independently. By identifying such patterns, resource forecasting models can be fit better, reducing over/under-provisioning via fewer resizing operations. Here we present ThetaScan, a time-series analysis method for vertical auto-scaling of containers in the Cloud, based on the detection of stationarity/trending and periodicity on resource consumption. Our method leverages the Theta Forecaster algorithm with deseasonalization that, in our provisioning scenario, only requires the estimated periodicity for resource consumption as principal hyper-parameter. Commonly used behavior detection methods require manual hyper-parameter tuning, making them infeasible for automation. Besides, it can be used at multi-scales (minute/hour/day), detecting hourly and daily patterns to improve resource usage prediction. Experiments show that we can detect behaviors in resource consumption that common methods miss, without requiring extensive manual tuning. We can reduce the resizing triggers compared to fixed-size scheduling around ~ 10% – 15%, reduce over-provisioning of CPU and Memory through periodic-based provisioning. Also a ~ 60% on multiscale resource forecasting for traces showing periodicity at different levels in respect to single-scale.
Josep Lluís Berral, David Buchaca Prats, Claudia Herron, Chen Wang 0039, Alaa Youssef
CLOUD4
2020 Proactive Container Auto-scaling for Cloud Native Machine Learning Services
abstract
Understanding the resource usage behaviors of the ever-increasing machine learning workloads are critical to cloud providers offering Machine Learning (ML) services. Capable of auto-scaling resources for customer workloads can significantly improve resource utilization, thus greatly reducing the cost. Here we leverage the AI4DL framework [1] to characterize workload and discover resource consumption phases. We advance the existing technology to an incremental phase discovery method that applies to more general types of ML workload for both training and inference. We use a time-window MultiLayer Perceptron (MLP) to predict phases in containers with different types of workload. Then, we propose a predictive vertical auto-scaling policy to resize the container dynamically according to phase predictions. We evaluate our predictive auto-scaling policies on 561 long-running containers with multiple types of ML workloads. The predictive policy can reduce up to 38% of allocated CPU compared to the default resource provisioning policies by developers. By comparing our predictive policies with commonly used reactive auto-scaling policies, we find that they can accurately predict sudden phase transitions (with an F1-score of 0.92) and significantly reduce the number of out-of-memory errors (350 vs. 20). Besides, we show that the predictive auto-scaling policy maintains the number of resizing operations close to the best reactive policies.
David Buchaca Prats, Josep Lluís Berral, Chen Wang 0039, Alaa Youssef
CLOUD3
2020 Using GANs for Sharing Networked Time Series Data: Challenges, Initial Promise, and Open Questions
abstract
Limited data access is a longstanding barrier to data-driven research and development in the networked systems community. In this work, we explore if and how generative adversarial networks (GANs) can be used to incentivize data sharing by enabling a generic framework for sharing synthetic datasets with minimal expert knowledge. As a specific target, our focus in this paper is on time series datasets with metadata (e.g., packet loss rate measurements with corresponding ISPs). We identify key challenges of existing GAN approaches for such workloads with respect to fidelity (e.g., long-term dependencies, complex multidimensional relationships, mode collapse) and privacy (i.e., existing guarantees are poorly understood and can sacrifice fidelity). To improve fidelity, we design a custom workflow called DoppelGANger (DG) and demonstrate that across diverse real-world datasets (e.g., bandwidth measurements, cluster requests, web sessions) and use cases (e.g., structural characterization, predictive modeling, algorithm comparison), DG achieves up to 43% better fidelity than baseline models. Although we do not resolve the privacy problem in this work, we identify fundamental challenges with both classical notions of privacy and recent advances to improve the privacy properties of GANs, and suggest a potential roadmap for addressing these challenges. By shedding light on the promise and challenges, we hope our work can rekindle the conversation on workflows for data sharing.
Zinan Lin 0001, Alankar Jain, Chen Wang 0039, Giulia Fanti, Vyas Sekar
Internet Measurement Conference3
2019 FfDL: A Flexible Multi-tenant Deep Learning Platform
abstract
Deep learning (DL) is becoming increasingly popular in several application domains and has made several new application features involving computer vision, speech recognition and synthesis, self-driving automobiles, drug design, etc. feasible and accurate. As a result, large scale "on-premise" and "cloud-hosted" deep learning platforms have become essential infrastructure in many organizations. These systems accept, schedule, manage and execute DL training jobs at scale.
K. R. Jayaram, Vinod Muthusamy, Parijat Dube, Vatche Isahagian, Chen Wang 0039, Benjamin Herta, Scott Boag, Diana Arroyo, Asser N. Tantawi, Archit Verma, Falk Pollok, Rania Khalaf
Middleware5
2019 AI Gauge: Runtime Estimation for Deep Learning in the Cloud
abstract
Major cloud providers, including IBM Cloud, Amazon Web Services, Microsoft Azure, and Google Cloud, offer services to train, debug, store, and deploy machine learning models at scale. For enhanced user experience in SLA-driven control, cost effective budgeting, elastic scaling, and efficient operations, estimating the runtime of training a machine learning model is important. We present AI Gauge, a cloud service to estimate runtime and cost for training deep learning models under different configuration options on the cloud. AI Gauge is designed using micro-service architecture and performs estimations based on machine learning models calibrated by an extensive and continuously populated job trace data-set. We show that AI Gauge can accurately predict the remaining time of running jobs based on its runtime progress (<; 10% relative error) and can accurately predict the total runtime for a job before it starts with 7-8% relative error on average.
Parijat Dube, Tonghoon Suk, Chen Wang 0039
SBAC-PAD3
2018 Comparing Cloud Content Delivery Networks for Adaptive Video Streaming
abstract
Cloud vendors offer content delivery network (CDN) services to compete for the video market. The user experience and the costs of providing the same video streaming service can vary when using different cloud CDNs. We emulate video streaming users in PlanetLab cloud to measure cloud CDNs including Amazon Web Service (AWS) CloudFront, Microsoft Azure Verizon CDN, and Google Cloud CDN. We leverage an approximated Quality of Experience (QoE) as a metric for evaluation. Our study finds that: 1) cloud vendors vary in providing QoE across regions; the video provider should assign a user to the CDN offering the best QoE at his location; 2) the QoE provided by one CDN can change over time; the video provider should adapt the CDN selection according to the real time QoE measurement; 3) cloud CDNs vary in scalability; streaming sessions may crash when there is bursty user demand; video providers should choose among the cloud CDNs that can properly scale; 4) regarding the cost, some cloud CDN is more economical than others given certain cache hit rate; video providers can minimize their costs by forcing free trial users to stream from the cheapest one.
Chen Wang 0039, Andal Jayaseelan, Hyong S. Kim 0001
IEEE CLOUD1
2018 Resource Profile Advisor for Containers in Cognitive Platform
abstract
Containers have transformed the cluster management into an application oriented endeavor, thus being widely used as the deployment units (i.e., micro-services) of large scale cloud services. As opposed to VMs, containers allow for resource provisioning with fine granularity and their resource usage directly reflects the micro-service behaviors. Container management systems like Kubernetes and Mesos provision resources to containers according to the capacity requested by the developers. Resource usages estimated by the developers are grossly inaccurate. They tend to be risk-averse and over provision resources, as under-provisioning would cause poor runtime performance or failures.
Mehmet Fatih Aktas, Chen Wang 0039, Alaa Youssef, Malgorzata Steinder
SoCC2
2017 Identifying Persistent and Recurrent QoE Anomalies for DASH Streaming in the Cloud
abstract
Quality of Experience (QoE) anomalies widely exist in all types of video services. As video services migrate to the Cloud, unique challenges occur to deploy video services in the Cloud environment. We study the QoE anomalies for users in a video service deployed in a production Cloud CDN. We use a QoE anomaly identification system, QRank, to identify anomalous systems. We consider Cloud CDN servers, Cloud CDN networks, transit networks, user access networks and different types of user devices. Our extensive experiments in production Cloud find several interesting insights about QoE anomalies of video streaming in the Cloud. 91.4% of QoE anomalies are detected on 15.32% of users. These users experience QoE anomalies persistently and recurrently. The Cloud servers and networks seldom cause QoE anomalies. More than 99.98% of QoE anomalies are identified in anomalous systems including the transit networks, the access networks and user devices. We infer that transit networks are the actual bottleneck systems for QoE anomalies in production Cloud. More than 95% of persistent and recurrent QoE anomalies are identified in less than 10 transit networks. We collect latency measurements to anomalous networks and the analysis indicates that the limited capacity in transit networks are the major cause of QoE anomalies. Resulting anomalies impair user QoEs persistently or recurrently. In order to provide good user QoE, the Cloud provider should identify transit networks that may become bottlenecks for high quality video streaming and appropriate peering with Internet Service Providers (ISPs) to bypass these bottlenecks.
Chen Wang 0039, Hyong S. Kim 0001, Ricardo Morla
CloudCom1
2016 QWatch: Detecting and Locating QoE Anomaly for VoD in the Cloud
abstract
Commercial large-scale VoD systems such as Netflix and Hulu rely on CDNs to deliver videos to users around the world. Various anomalies occur often and degrade users' Quality of Experience (QoE). Detecting and locating such anomalies are highly complex due to a large number of different entities involved in the end-to-end video delivery. These entities include VoD provider, CDN/Cloud providers, transit ISPs, access ISPs, and end user devices. QoE perceived by the users is a critical metric for VoD providers. We propose QWatch, a scalable monitoring system, which detects and locates anomalies based on the end user QoE in real-time. We evaluate QWatch in a controlled VoD system and production Microsoft Azure Cloud and CDN. QWatch effectively detects and locates QoE anomalies in our extensive experiments. We discuss insights obtained from running VoD system with 200 worldwide users in production Cloud.
Chen Wang 0039, Hyong S. Kim 0001, Ricardo Morla
CloudCom1
2015 QoE Driven Server Selection for VoD in the Cloud
abstract
In commercial Video-on-Demand (VoD) systems, user's Quality of Experience (QoE) is the key factor for user satisfaction. In order to improve user's QoE, VoD providers replicate popular videos in geo-distributed Cloud and deploy cache servers close to users. Generally, the VoD provider selects a server for the user request according to the user's location. Usually geographically closely located servers would provide lower network delay. However, the performance of VoD servers deployed in cloud virtual machines (VM) depends not only on the network delay but also resource contention due to other VMs and highly dynamic user demands. Thus, QoE offered by the server varies greatly over time as user demands and network traffic fluctuate regardless of the location. Selecting a server close to users sometimes reduces the network delay but cannot guarantee QoE in general. We believe that end users have the best perception of server performance in terms of their QoE rather than the servers themselves. What user perceives incorporate performance of all elements, such as network delay and server response time in VoD service. We propose VoD server selection schemes that dynamically select servers according to user's QoE feedback. We integrate our server selection schemes with Dynamic Adaptive Streaming over HTTP (DASH) clients and evaluate our system both in simulation and in Google Cloud. Results show our system improves user QoE up to 20% compared to existing solutions.
Chen Wang 0039, Hyong S. Kim 0001, Ricardo Morla
CLOUD1
2015 Users Know Better: A QoE Based Adaptive Control System for VoD in the Cloud
abstract
As VoD systems migrate to the Cloud, new challenges emerge in managing user Quality-of- Experience (QoE). The complexity of the cloud system due to virtualization and resource sharing complicates the QoE management. Operational failures in the Cloud could be challenging for QoE as well. We believe that end users have the best perception of system performance in terms of their QoE. We propose a QoE based adaptive control system for VoD in the Cloud. The system learns server performance from the user QoE and then adaptively selects servers for users accordingly. We deploy our proposed system in Google Cloud and evaluate it with hundreds of clients deployed all over the world. Results show that given the same amount of resources, our system provides 9% to 30% more users with QoE above the Mean Opinion Score (MOS) "good" level than the existing measurement based server selection systems. The system guarantees a better QoE (above 6% better) for 90% users. Additionally, our system discovers operational failures by monitoring QoE and prevents streaming session crashes. A computational overhead analysis shows that our system can easily scale to large VoD systems containing thousands of servers.
Chen Wang 0039, Hyong S. Kim 0001, Ricardo Morla
GLOBECOM1