EDBT 2026 Demo / reviewers in the wild / expert
Jia Rao
dblp:38/618
· DBLP profile ↗
70ranked-venue papers
9as first author
17since 2021 · last 2026
0000-0002-2133-4363ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 52 · 7 first-author · 14 since 2021Software engineering, systems software and programming languages · 8 · 1 first-author · 3 since 2021Computer networks · 6 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 4Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AEP: Achieving Hierarchical Fault Tolerance in DSM Through Atomic Execution ProtectionabstractRecent advances in CXL 3.0 have revitalized Distributed Shared Memory (DSM), enabling zero-copy sharing across machine nodes at near-NUMA latency and making shared-memory communication practical for distributed applications. However, DSM memory allocators remain vulnerable to partial failures: client crashes during critical, non-atomic memory operations can corrupt metadata and stall the entire system. Existing solutions either block all clients during recovery or impose heavy runtime overhead due to distributed fault-tolerance protocols. We propose Atomic Execution Protection (AEP), a kernel-assisted fault tolerance mechanism that guarantees atomic completion of user-space critical operations by deferring termination until protected regions finish. AEP tolerates intra-node partial failures without costly distributed coordination. To extend beyond a single node, we further design AEP-DSM, a hierarchical DSM allocator that combines AEP-protected node-level allocators with cross-node coordination. Our Linux AEP prototype integrated with a cross-node DSM consistency protocol (CXL-SHM) achieves near-native local performance and improves throughput by up to 13.3x over state-of-the-art DSM allocators. Zixuan Wang 0030, Hang Huang, Jia Rao, Hui Lu 0001, Hao Fan 0006, Song Wu 0001, Hai Jin 0001 |
EuroSys | 4 |
| 2026 | Scaling Attention Beyond GPUs for LLM InferenceabstractScaling inference for large language models is increasingly constrained by limited GPU memory, primarily due to the expanding intermediate states (KV caches) required for long-context generation and multi-user workloads. Once the KV cache exceeds the capacity of high-bandwidth memory, it must be offloaded to host memory and reloaded on demand, a workflow severely bottlenecked by the CPU–GPU interconnect, typically PCIe. Existing approaches exploiting offload KV caches to CPU memory and selectively reload partial segments for attention computation often underutilize CPU compute resources and suffer from accuracy degradation. We present Beyond, a drop-in runtime that integrates a smart offloading scheme to selectively identify and retain salient KV entries across continuous decoding sessions, together with a hybrid CPU–GPU attention mechanism for scalable inference. Beyond executes dense attention over recent KV entries stored in GPU memory while performing parallel, per-head sparse attention on salient contextual KV entries residing in CPU memory. The outputs are fused efficiently through a log-sum-exp scheme. During the bandwidth-constrained decoding phase, oversized KV caches are processed cooperatively by the aggregated CPU and GPU memory bandwidth, with only minimal PCIe data movement. Experiments across diverse models and workloads demonstrate that Beyond improves scalability, supports longer sequences and larger batch sizes, and outperforms existing sparse attention baselines in both efficiency and accuracy—all on commodity GPU hardware. Weishu Deng, Peiran Du, Lingfeng Xiang, Chen Zhong 0002, Faraz Ahmed, Lianjie Cao, Puneet Sharma 0001, Song Jiang 0001, Hui Lu 0001, Jia Rao |
HPDC | 12 |
| 2026 | KPT-Fork: Enhancing User Data Isolation via Kernel Page Table ForkabstractModern operating systems (OS) utilize a shared kernel address design, enabling the OS kernel to access physical memory addresses conveniently and efficiently. However, this design introduces vulnerabilities that attackers can exploit to arbitrarily read/write content from any kernel-space address, known as data-oriented attacks. Existing MMU-based isolation approaches face inherent challenges in scalability, compatibility, and context switch overhead. Inspired by the reference kernel page table, we propose to use user-private reference kernel page tables (UP-KPTs) as the foundation for data isolation in a shared kernel address space. A UP-KPT is a customized kernel page table tailored for a specific container, encompassing address mappings of private user data that are not visible to other containers. To protect sensitive system-wide data, an administrator can unmap such data from UP-KPTs and make it inaccessible to potentially malicious containers. UP-KPT provides strong data isolation while maintaining full compatibility with legacy applications and existing kernel components. We develop KPT-fork, a kernel isolation framework that enables the creation and management of UP-KPTs. KPT-fork entails two important designs: 1) a UP-KPT structure that enables data isolation while preserving the simplicity, convenience, and efficiency of a shared kernel page table; 2) a private memory allocator that efficiently transforms static and direct address mappings of private data in shared kernel addresses into dynamic user-private mappings. We demonstrate through three case studies that KPT-fork can be extended to protect various types of data. Our evaluation shows that KPT fork can defend against a wide range of data-oriented attacks while incurring negligible overhead. Zixuan Wang 0030, Lisong Pan, Jia Rao, Hao Fan 0006, Song Wu 0001 |
IEEE Trans. Cloud Comput. | 3 |
| 2025 | Can Hardware Outsmart Software in Tiered Memory Management? A CMM-H Case StudyabstractWith the advent of Compute Express Link (CXL), hardware-managed memory tiering has become a reality. In this paper, we investigate Samsung's CXL Memory Module-Hybrid (CMM-H), a CXL Type 3 device integrating DRAM and NAND flash managed by an FPGA-based controller and providing byte-addressable memory interface via the cxl.mem protocol. We perform a detailed evaluation of CMM-H and compare its performance with OS-level and block-level tiering solutions. Our results highlight the performance benefits of CMM-H for cache-hit scenarios and identify key limitations for cache-miss situations, offering insights into the trade-offs involved in adopting hardware-managed memory tiering in emerging CXL-based systems. Lingfeng Xiang, Lianjie Cao, Faraz Ahmed, Jia Rao, Hui Lu 0001, Puneet Sharma 0001 |
SYSTOR | 6 |
| 2025 | DSA-2LM: A CPU-Free Tiered Memory Architecture with Intel DSA
Ruili Liu, Teng Ma 0006, Yingdi Shan, Zheng Liu 0022, Lingfeng Xiang, Hui Lu 0001, Jia Rao, Kang Chen 0001, Yongwei Wu 0001 |
USENIX ATC | 10 |
| 2024 | Mega: More Efficient Graph Attention for GNNsabstractGraph neural networks (GNNs) have demonstrated effectiveness across diverse application domains by leveraging graph information to uncover intrinsic correlations alongside feature representation. This enables GNNs to explore richer information compared to conventional neural networks, resulting in enhanced predictive performance. However, the integration of graphstructured data into the learning process poses two challenges. First, the sparsity and irregularity of the graph representation result in inefficient and expensive memory accesses on throughput-oriented accelerators, such as GPUs. Second, as GNN training involves interleaved graph operations to extract topological information and neural operations to update node or edge embeddings, the joint optimization of these two operations on accelerators is challenging due to their distinct resource requirements. Among GNN graph operations, graph attention which helps focus GNN training on highly correlated nodes is critical to training performance and model accuracy. However, our profiling of representative GNNs reveals that irregular memory access during graph attention accounts for the dominating overhead in GNN training. To address this issue, this paper proposes a more efficient graph attention method Megato accelerate GNN training. MegA converts the original graph representation into one that regularizes memory access patterns for graph attention. Specifically, during preprocessing, Megatraverses a graph to derive a schedule for graph attention and uses the schedule to reorganize the graph representation for optimized memory access. MegA explores several techniques to balance memory access efficiency and preserve the original graph properties to avoid the loss of model accuracy. Experimental results with representative GNNs and graph data sets show that Megaconsistently outperforms conventional graph attention methods with up to 3x speedup. Weishu Deng, Jia Rao |
ICDCS | 2 |
| 2024 | Nomad: Non-Exclusive Memory Tiering via Transactional Page Migration
Lingfeng Xiang, Weishu Deng, Hui Lu 0001, Jia Rao, Ren Wang 0001 |
OSDI | 5 |
| 2024 | vKernel: Enhancing Container Isolation via Private Code and DataabstractContainer technology is increasingly adopted in cloud environments. However, the lack of isolation in the shared kernel becomes a significant barrier to the wide adoption of containers. The challenges lie in how to simultaneously attain high performance and isolation. On the one hand, kernel-level isolation mechanisms, such asseccomp,capabilities, andapparmor, achieve good performance without much overhead, but lack the support for per-container customization. On the other hand, user-level and VM-based isolation offer superior security guarantees and allow for customization since a container is assigned a dedicated kernel, however, at the cost of high overhead. We presentvKernel, a kernel isolation framework. It maintains a minimal set of code and data that are either sensitive or are prone to interference in a virtual kernel instance (vKI). vKernel relies on inline hooks to intercept and redirect requests sent to the host kernel to a vKI, where container-specific security rules, functions, and data are implemented. Through case studies, we demonstrate that under vKernel user-defined data isolation and kernel customization can be supported with a reasonable engineering effort. An evaluation of vKernel with micro-benchmarks, cloud services, real-world applications show that vKernel achieves good security guarantees, but with much less overhead. Hang Huang, Jia Rao, Song Wu 0001, Hao Fan 0006, Chen Yu 0003, Hai Jin 0001, Kun Suo, Lisong Pan |
IEEE Trans. Computers | 3 |
| 2023 | Accelerating Packet Processing in Container Overlay Networks via Packet-level ParallelismabstractOverlay networks serve as the de facto network virtualization technique for providing connectivity among distributed containers. Despite the flexibility in building customized private container networks, overlay networks incur significant performance loss compared to physical networks (i.e., the native). The culprit lies in the inclusion of multiple network processing stages in overlay networks, which prolongs the network processing path and overloads CPU cores. In this paper, we propose mFlow, a novel packet steering approach to parallelize the in-kernel data path of network flows. mFlow exploits packet-level parallelism in the kernel network stack by splitting the packets of the same flow into multiple micro-flows, which can be processed in parallel on multiple cores. mFlow devises new, generic mechanisms for flow splitting while preserving in-order packet delivery with little overhead. Our evaluation with both micro-benchmarks and real-world applications demonstrates the effectiveness of mFlow, with significantly improved performance – e.g., by 81% in TCP throughput and 139% in UDP compared to vanilla overlay networks. mFlow even achieved higher TCP throughput than the native (e.g., 29.8 vs. 26.6 Gbps). Jiaxin Lei, Manish Munikar, Hui Lu 0001, Jia Rao |
IPDPS | 4 |
| 2023 | PVM: Efficient Shadow Paging for Deploying Secure Containers in Cloud-native EnvironmentabstractIn cloud-native environments, containers are often deployed within lightweight virtual machines (VMs) to ensure strong security isolation and privacy protection. With the growing demand for customized cloud services, third-party vendors are turning to infrastructure-as-a-service (IaaS) cloud providers to build their own cloud-native platforms, necessitating the need to run a VM or a guest that hosts containers inside another VM instance leased from an IaaS cloud. State-of-the-art nested virtualization in the x86 architecture relies heavily on the host hypervisor to expose hardware virtualization support to the guest hypervisor, not only complicating cloud management but also raising concerns about an increased attack surface at the host hypervisor. Hang Huang, Jiangshan Lai, Jia Rao, Hui Lu 0001, Wenlong Hou, Zhengyu He, Weidong Han 0003, Tao Ma 0006, Song Wu 0001 |
SOSP | 3 |
| 2023 | P2CACHE: Exploring Tiered Memory for In-Kernel File Systems Caching
Lingfeng Xiang, Jia Rao, Hui Lu 0001 |
USENIX ATC | 3 |
| 2023 | Adapt Burstable Containers to Variable CPU ResourcesabstractIn the age of the cloud-native, container technology, referred as OS-level virtualization, is increasingly adopted to deploy cloud applications. Compared with virtual machines, containers are lightweight and flexible in resource management. An important quality-of-service (QoS) class in container management is burstable container, whose resource limits are higher than the actual requests allowing a container to expand whenever demands ramp up and additional resources become available. However, efficiently managing burstable containers is challenging, especially for CPU resources. On the one hand, burstable containers should maintain sufficient concurrency, in the form of threads, to utilize extendable CPU resources. On the other hand, the degree of concurrency necessary for utilizing peak CPU resources leads to suboptimal performance when a container's CPU allocation is constrained. In this paper, we recommend that the number of threads in burstable containers should always be set to the CPU limit to guarantee extensibility. However, modern operating systems (OSes) fall short of efficiently managing thread oversubscription. First, the OS CPU scheduler is inefficient for scheduling excessive threads and lacks container awareness. Second, the existing blocking synchronization supported by the OS kernel is inefficient in handling the sleep and wakeup of excessive threads. Finally, the non-blocking synchronization may waste CPUs performing busy waiting when more than one thread in the run queue. To this end, we present a user-level adaptive container scheduler and two OS mechanisms,virtual blockingandbusy-waiting detection, to avoid inefficiency in managing burstable containers without requiring program code changes. Experimental results show that our approaches can keep burstable containers efficient while allowing the applications in containers to take advantage of additional CPUs. The performance gain under high system load is up to 29.7×. Hang Huang, Jia Rao, Song Wu 0001, Hai Jin 0001, Duoqiang Wang, Kun Suo, Lisong Pan |
IEEE Trans. Computers | 3 |
| 2022 | Characterizing the performance of intel optane persistent memory: a close look at its on-DIMM bufferingabstractWe present a comprehensive and in-depth study of Intel Optane DC persistent memory (DCPMM). Our focus is on exploring the internal design of Optane's on-DIMM read-write buffering and its impacts on application-perceived performance, read and write amplifications, the overhead of different types of persists, and the tradeoffs between persistency models. While our measurements confirm the results of the existing profiling studies, we have new discoveries and offer new insights. Notably, we find that read and write are managed differently in separate on-DIMM read and write buffers. Comparable in size, the two buffers serve distinct purposes. The read buffer offers higher concurrency and effective on-DIMM prefetching, leading to high read bandwidth and superior sequential performance. However, it does not help hide media access latency. In contrast, the write buffer offers limited concurrency but is a critical stage in a pipeline that supports asynchronous write in the DDR-T protocol. Surprisingly, in addition to write coalescing, the write buffer delivers lower than read and consistent write latency regardless of the working set size, the type of write, the access pattern, or the persistency model. Furthermore, we discover that the mismatch between cacheline access granularity and the 3D-Xpoint media access granularity negatively impacts the effectiveness of CPU cache prefetching and leads to wasted persistent memory bandwidth. Lingfeng Xiang, Xingsheng Zhao, Jia Rao, Song Jiang 0001, Hong Jiang 0001 |
EuroSys | 3 |
| 2022 | Prism: Streamlined Packet Processing for Containers with Flow PrioritizationabstractAdvanced high-speed network cards have made packet processing in host operating systems a major performance bottleneck. The kernel network stack gives rise to various sources of overheads that limit the throughput and lengthen the per-packet processing latency. The problem is further exacerbated for short-lived, latency-sensitive network flows such as control packets, online gaming, database requests, etc. — in a highly utilized system, especially in virtualized (containerized) cloud environments, short flows can experience excessively long in-kernel queuing delays. As a consequence, recent research works propose to bypass the kernel network stack to enable lightweight, custom userspace network stacks for improved performance, but at a heavy cost of compatibility and security. In this paper, we take a different approach: We first analyze various sources of inefficiencies in the kernel network stack and propose ways to mitigate them without compromising systems compatibility, security, or flexibility. Further, we propose Prism, a novel mechanism in the kernel network stack to differentiate incoming packets based on their performance requirements and streamline the processing stages of multi-stage packet processing pipelines (e.g., in container overlay networks). Our evaluation demonstrates that Prism can significantly improve the latency of high-priority flows in container overly networks in the presence of heavy low-priority background traffic. Manish Munikar, Jiaxin Lei, Hui Lu 0001, Jia Rao |
ICDCS | 4 |
| 2021 | Parallelizing packet processing in container overlay networksabstractContainer networking, which provides connectivity among containers on multiple hosts, is crucial to building and scaling container-based microservices. While overlay networks are widely adopted in production systems, they cause significant performance degradation in both throughput and latency compared to physical networks. This paper seeks to understand the bottlenecks of in-kernel networking when running container overlay networks. Through profiling and code analysis, we find that a prolonged data path, due to packet transformation in overlay networks, is the culprit of performance loss. Furthermore, existing scaling techniques in the Linux network stack are ineffective for parallelizing the prolonged data path of a single network flow. Jiaxin Lei, Manish Munikar, Kun Suo, Hui Lu 0001, Jia Rao |
EuroSys | 5 |
| 2021 | Towards Exploiting CPU Elasticity via Efficient Thread OversubscriptionabstractElasticity is an essential feature of cloud computing, which allows users to dynamically add or remove resources in response to workload changes. However, building applications that truly exploit elasticity is non-trivial. Traditional applications need to be modified to efficiently utilize variable resources. This paper explores thread oversubscription, i.e., provisioning more threads than the available cores, to exploit CPU elasticity in the cloud. While maintaining sufficient concurrency allows applications to utilize additional CPUs when more are made available, it is widely believed that thread oversubscription introduces prohibitive overheads due to excessive context switches, loss of locality, and contention on shared resources. Hang Huang, Jia Rao, Song Wu 0001, Hai Jin 0001, Hong Jiang 0001, Hao Che, Xiaofeng Wu 0002 |
HPDC | 2 |
| 2021 | SwitchFlow: preemptive multitasking for deep learningabstractAccelerators, such as GPU, are a scarce resource in deep learning (DL). Effectively and efficiently sharing GPU leads to improved hardware utilization as well as user experiences, who may need to wait for hours to access GPU before a long training job is done. Spatial and temporal multitasking on GPU have been studied in the literature, but popular deep learning frameworks, such as Tensor-Flow and PyTorch, lack the support of GPU sharing among multiple DL models, which are typically represented as computation graphs, heavily optimized by underlying DL libraries, and run on a complex pipeline spanning CPU and GPU. Our study shows that GPU kernels, spawned from computation graphs, can barely execute simultaneously on a single GPU and time slicing may lead to low GPU utilization. Xiaofeng Wu 0002, Jia Rao, Wei Chen 0038, Hang Huang, Chris Ding, Heng Huang 0001 |
Middleware | 2 |
| 2020 | Sledge: Towards Efficient Live Migration of Docker ContainersabstractModern large-scale cloud platforms require live migration technique on Docker containers with stateful workload to support load balancing, host maintenance, and Quality of Service (QoS) improvement. Efficient and scalable Docker live migration is expected to guarantee the component-integrity (image, runtime, and management context) with negligible downtime. In this paper, we present a highly efficient live migration system called Sledge, which ensures the component-integrity by integrating both images and management context during runtime migration. The key insight is that the layered image can be leveraged to reduce the migration overhead, and appropriately selective migration of management context will effectively improve QoS with negligible downtime. To achieve good scalability, a lightweight container registry mechanism for end-to-end image migration is designed to avoid the redundant layers transmission. In addition, a dynamic context loading scheme is proposed to precisely load the management context into the running daemon, which can significantly reduce downtime. Experiments show that, compared with the state-of-the-art, Sledge reduces 57% of total migration time, 55% of image migration time, and 70% downtime. Song Wu 0001, Jiang Xiao 0001, Hai Jin 0001, Yingxi Zhang, Guoqiang Shi, Tingyu Lin 0001, Jia Rao, Jizhong Jiang |
CLOUD | 8 |
| 2020 | Towards Lightweight Serverless Computing via Unikernel as a FunctionabstractServerless computing, also known as “Function as a Service (FaaS)”, is emerging as an event-driven paradigm of cloud computing. In the FaaS model, applications are programmed in the form of functions that are executed and managed separately. Functions are triggered by cloud users and are provisioned dynamically through containers or virtual machines (VMs). The startup delays of containers or VMs usually lead to rather high latency of response to cloud users. Moreover, the communication between different functions generally relies on virtual net devices or shared memory, and may cause extremely high performance overhead. In this paper, we propose Unikernel-as-a-Function (UaaF), a much more lightweight approach to serverless computing. Applications are abstracted as a combination of different functions, and each function are built as an unikernel in which the function is linked with a specified minimum-sized library operating system (LibOS). UaaF offers extremely low startup latency to execute functions, and an efficient communication model to speed up inter-functions interactions. We exploit an new hardware technique (namely VMFUNC) to invoke functions in other unikernels seamlessly (mostly like inter-process communications), without suffering performance penalty of VM Exits. We implement our proof-of-concept prototype based on KVM and deploy UaaF in three unikernels (MirageOS, IncludeOS, and Solo5). Experimental results show that U aaF can significantly reduce the startup latency and memory usage of serverless cloud applications. Moreover, the VMFUNC-based communication model can also significantly improve the performance of function invocations between different unikernels. Haikun Liu, Jia Rao, Xiaofei Liao, Hai Jin 0001, Yu Zhang 0027 |
IWQoS | 3 |
| 2020 | Preemptive and Low Latency Datacenter Scheduling via Lightweight ContainersabstractDatacenters are evolving to host heterogeneous workloads on shared clusters to reduce the operational cost and achieve higher resource utilization. However, it is challenging to schedule heterogeneous workloads with diverse resource requirements and QoS constraints. On one hand, latency-critical jobs need to be scheduled as soon as they are submitted to avoid any queuing delays. On the other hand, best-effort long jobs should be allowed to occupy the cluster when there are idle resources to improve cluster utilization. The challenge lies in how to minimize the queuing delays of short jobs while maximizing cluster utilization. In this article, we propose and develop BIG-C, a container-based resource management framework for data-intensive cluster computing. The key design is to leverage lightweight virtualization, a.k.a, containers, to make tasks preemptable in cluster scheduling. We devise two types of preemption strategies: immediate and graceful preemptions and show their effectiveness and tradeoffs with loosely-coupled MapReduce workloads as well as iterative, in-memory Spark workloads. Based on the mechanisms for task preemption, we further develop job-level and task-level preemptive policies as well as a preemptive fair share cluster scheduler. Our implementation on Yarn and evaluation with synthetic and production workloads show that low job latency and high resource utilization can be both attained when scheduling heterogeneous workloads on a contended cluster. Wei Chen 0038, Xiaobo Zhou 0002, Jia Rao |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2019 | Pigeon: an Effective Distributed, Hierarchical Datacenter Job SchedulerabstractIn today's datacenters, job heterogeneity makes it difficult for schedulers to simultaneously meet latency requirements and maintain high resource utilization. The state-of-the-art datacenter schedulers, including centralized, distributed, and hybrid schedulers, fail to ensure low latency for short jobs in large-scale and highly loaded systems. The key issues are the scalability in centralized schedulers, ineffective and inefficient probing and resource sharing in both distributed and hybrid schedulers. Zhijun Wang 0001, Huiyang Li, Xiaocui Sun, Jia Rao, Hao Che, Hong Jiang 0001 |
SoCC | 5 |
| 2019 | CNTC: A Container Aware Network Traffic Control Framework
Lin Gu 0002, Junjian Guan, Song Wu 0001, Hai Jin 0001, Jia Rao, Kun Suo, Deze Zeng |
GPC | 5 |
| 2019 | Adaptive Resource Views for ContainersabstractAs OS-level virtualization advances, containers have become a viable alternative to virtual machines in deploying applications in the cloud. Unlike virtual machines, which allow guest OSes to run atop virtual hardware, containers have direct access to physical hardware and share one OS kernel. While the absence of virtual hardware abstractions eliminates most virtualization overhead, it presents unique challenges for containerized applications to efficiently utilize the underlying hardware. The lack of hardware abstraction exposes the total amount of resources that are shared among all containers to each individual container. Parallel runtimes (e.g., OpenMP) and managed programming languages (e.g., Java) that rely on OS-exported information for resource management could suffer from suboptimal performance. In this paper, we develop a per-container view of resources to export information on the actual resource allocation to containerized applications. The central design of the resource view is a per-container sys\_namespace that calculates the effective capacity of CPU and memory in the presence of resource sharing among containers. We further create a virtual sysfs to seamlessly interface user space applications with sys\_namespace. We use two case studies to demonstrate how to leverage the continuously updated resource view to enable elasticity in the HotSpot JVM and OpenMP. Experimental results show that an accurate view of resource allocation leads to more appropriate configurations and improved performance in a variety of containerized applications. Hang Huang, Jia Rao, Song Wu 0001, Hai Jin 0001, Kun Suo, Xiaofeng Wu 0002 |
HPDC | 2 |
| 2019 | Preemptive Multi-Queue Fair QueuingabstractFair queuing (FQ) algorithms have been widely adopted in computer systems to share resources among multiple users. Modern operating systems and hypervisors use variants of FQ algorithms to implement the critical OS resource management -- the thread scheduler. While the existing FQ algorithms enforce fair CPU allocation on a per-core basis, there lacks an algorithm to fairly allocate CPU on multiple cores. This common deficiency in state-of-the-art multicore schedulers causes unfair CPU allocations to parallel programs using blocking synchronization, leading to severe performance degradation. Parallel threads that frequently block due to synchronization exhibit deceptive idleness and are penalized by the thread scheduler. To this end, we propose a preemptive multi-queue fair queuing (P-MQFQ) algorithm that uses a centralized queue to fairly dispatch threads from different programs based on their received CPU bandwidth from multiple cores. We demonstrate that P-MQFQ can be approximated by augmenting the existing load balancing in the OS without requiring to implement the centralized queue or undermining scalability. We implement P-MQFQ in Linux and Xen, respectively, and show significantly improved utilization and performance for parallel programs. Kun Suo, Xiaofeng Wu 0002, Jia Rao, Song Wu 0001, Hai Jin 0001 |
HPDC | 4 |
| 2019 | N-Docker: A NVM-HDD Hybrid Docker Storage Framework to Improve Docker Performance
Lin Gu 0002, Qizhi Tang, Song Wu 0001, Hai Jin 0001, Yingxi Zhang, Guoqiang Shi, Tingyu Lin 0001, Jia Rao |
NPC | 8 |
| 2018 | Characterizing and optimizing hotspot parallel garbage collection on multicore systemsabstractThe proliferation of applications, frameworks, and services built on Java have led to an ecosystem critically dependent on the underlying runtime system, the Java virtual machine (JVM). However, many applications running on the JVM, e.g., big data analytics, suffer from long garbage collection (GC) time. The long pause time due to GC not only degrades application throughput and causes long latency, but also hurts overall system efficiency and scalability. Kun Suo, Jia Rao, Hong Jiang 0001, Witawas Srisa-an |
EuroSys | 2 |
| 2018 | vNetTracer: Efficient and Programmable Packet Tracing in Virtualized NetworksabstractAs the scale of cloud systems continues to grow, virtualized networks that provide connectivity between services within and across data centers, are becoming increasingly important to the performance and reliability of the cloud. Despite many advantages, including fast deployment, ease of management, and programmability, virtualized networks require additional layers of abstraction and complicate monitoring and diagnosis of performance issues compared to traditional networks on physical hardware. Virtualized networks usually connect components in multiple protection domains, such as a guest OS, the hypervisor, network bridges, and separate virtualized network functions. There is no efficient means to trace packet transmission across the boundaries. Furthermore, it is challenging to reason about the performance of dynamic virtualized networks. Therefore, fine-grained, user customizable, and reconfigurable network tracing becomes a great need. To address these challenges, we built vNetTracer, an efficient and programmable packet profiler for virtualized networks. vNetTracer relies on the extended Berkeley Packet Filter (eBPF) to dynamically insert user-defined trace programs into a live virtualized network without any changes to the applications or restarts of the monitored network. Through three case studies, we demonstrate the effectiveness of vNetTracer in diagnosing various virtualized networking problems. Kun Suo, Wei Chen 0038, Jia Rao |
ICDCS | 4 |
| 2018 | eBrowser: Making Human-Mobile Web Interactions Energy Efficient with Event Rate LearningabstractDue to the limited screen size of mobile devices, finger movements on touchscreen, such as scrolling and pinching (i.e., zooming in or out), are frequently used on mobile Web browsers and WebView-based apps, consuming considerable energy on mobile devices. While existing works on mobile Web browsers focus on reducing the power consumption or optimizing the performance of webpage loading, the power consumption of mobile Web interactions, especially after webpage loading, has received comparatively little attention. Motivated by an empirical study of the power consumption and user experience survey of human-mobile interactions, we design and implement eBrowser, an energy-efficient mobile Web interaction framework. It leverages a cloud-based machine learning model to enable personalized interaction event rate for individual users according to the interaction speed of their finger movement and the content of rendered webpages. To adapt to user behavior changes, eBrowser continuously monitors the interaction experience on each mobile device and periodically updates the personalized event rate model with incremental learning in the cloud. We implement eBrowser in Chromium and deploy the event rate model in a remote Aliyun cloud instance. Experimental results show that eBrowser reduces the energy consumption of mobile Web interactions by up to 43.8% with negligible runtime overhead, while guaranteeing user satisfaction on both mobile browsers and WebView-based apps. Fei Xu 0009, Zhi Zhou 0006, Jia Rao |
ICDCS | 4 |
| 2018 | Improving Resource Utilization through Demand Aware Process SchedulingabstractTraditional process scheduling in the operating system focuses on high CPU utilization while achieving fairness among the processes. However, this can lead to an inefficient usage of other hardware resources, e.g., the caches, which have limited capacity and is a scarce resource on most systems. This paper extends a traditional operating system scheduler to schedule processes more efficiently against hardware resources. Through the introduction of a new concept, a progress period, which models the variation of resource access characteristics during application execution, our scheduling extension dynamically monitors the changes in resource access behavior of each process being scheduled, tracks their collective usage of hardware resources, and schedules the processes to decrease overall system power consumption without compromising performance. Testing this scheduling system on programs on an Intel(R) Xeon(R) E5-2420 CPU with twelve kernels from the BLAS suite and five applications from the SPLASH-2 benchmark suite yielded a 48% maximum decrease in system energy consumption (average 12%), and a 1.88x maximum increase in application performance (average 1.16x). Brandon Nesterenko, Qing Yi, Jia Rao |
ICPP | 3 |
| 2018 | An Analysis and Empirical Study of Container NetworksabstractContainers, a form of lightweight virtualization, provide an alternative means to partition hardware resources among users and expedite application deployment. Compared to virtual machines (VMs), containers incur less overhead and allow a much higher consolidation ratio. Container networking, a vital component in container-based virtualization, is still not well understood. Many techniques have been developed to provide connectivity between containers on a single host or across multiple machines. However, there lacks an in-depth analysis of their respective advantages, limitations, and performance in a cloud environment. In this paper, we perform a comprehensive study of representative container networks. We first conduct a qualitative comparison of their applicable scenarios, levels of security isolation, and overhead. Then we quantitatively evaluate the throughput, latency, scalability, and startup cost of various container networks in a realistic cloud environment. We find that virtualized network in containers incurs non-negligible overhead compared to physical networks. Performance degradation varies depending on the type of network protocol and packet size. Our experiments show that there is no clear winner in performance and users need to select an appropriate container network based on the requirements and characteristics of their workloads. Kun Suo, Wei Chen 0038, Jia Rao |
INFOCOM | 4 |
| 2018 | Dynamic vertical memory scalability for OpenJDK cloud applicationsabstractThe cloud is an increasingly popular platform to deploy applications as it lets cloud users to provide resources to their applications as needed. Furthermore, cloud providers are now starting to offer a "pay-as-you-use" model in which users are only charged for the resources that are really used instead of paying for a statically sized instance. This new model allows cloud users to save money, and cloud providers to better utilize their hardware. Rodrigo Bruno, Paulo Ferreira 0001, Ruslan Synytsky, Tetiana Fydorenchyk, Jia Rao, Hang Huang, Song Wu 0001 |
ISMM | 5 |
| 2017 | mBalloon: enabling elastic memory management for big data processingabstractBig Data processing often suffers from significant memory pressure, resulting in excessive garbage collection (GC) and out-of-memory (OOM) errors, harming system performance and reliability. Therefore, users tend to give an excessive heap size to applications to avoid job failure, causing low cluster utilization. Wei Chen 0038, Aidi Pi, Jia Rao, Xiaobo Zhou 0002 |
SoCC | 3 |
| 2017 | Preserving I/O prioritization in virtualized OSesabstractWhile virtualization helps to enable multi-tenancy in data centers, it introduces new challenges to the resource management in traditional OSes. We find that one important design in an OS, prioritizing interactive and I/O-bound workloads, can become ineffective in a virtualized OS. Resource multiplexing between multiple tenants breaks the assumption of continuous CPU availability in physical systems and causes two types of priority inversions in virtualized OSes. In this paper, we present xBalloon, a lightweight approach to preserving I/O prioritization. It uses a balloon process in the virtualized OS to avoid priority inversion in both short-term and long-term scheduling. Experiments in a local Xen environment and Amazon EC2 show that xBalloon improves I/O performance in a recent Linux kernel by as much as 136% on network throughput, 95% on disk throughput, and 125x on network tail latency. Kun Suo, Jia Rao, Luwei Cheng, Xiaobo Zhou 0002, Francis C. M. Lau 0001 |
SoCC | 3 |
| 2017 | Addressing Performance Heterogeneity in MapReduce Clusters with Elastic TasksabstractMapReduce applications, which require access to a large number of computing nodes, are commonly deployed in heterogeneous environments. The performance discrepancy between individual nodes in a heterogeneous cluster present significant challenges to attain good performance in MapReduce jobs. MapReduce implementations designed and optimized for homogeneous environments perform poorly on heterogeneous clusters. We attribute suboptimal performance in heterogeneous clusters to significant load imbalance between map tasks. We identify two MapReduce designs that hinder load balancing: (1) static binding between mappers and their data makes it difficult to exploit data redundancy for load balancing; (2) uniform map sizes is not optimal for nodes with heterogeneous performance. To address these issues, we propose FlexMap, a user-transparent approach that dynamically provisions map tasks to match distinct machine capacity in heterogeneous environments. We implemented FlexMap in Hadoop-2.6.0. Experimental results show that it reduces job completion time by as much as 40% compared to stock Hadoop and 30% to SkewTune. Wei Chen 0038, Jia Rao, Xiaobo Zhou 0002 |
IPDPS | 2 |
| 2017 | Container-Based Cloud Platform for Mobile Computation OffloadingabstractWith the explosive growth of smartphones and cloud computing, mobile cloud, which leverages cloud resource to boost the performance of mobile applications, becomes attrac- tive. Many efforts have been made to improve the performance and reduce energy consumption of mobile devices by offloading computational codes to the cloud. However, the offloading cost caused by the cloud platform has been ignored for many years. In this paper, we propose Rattrap, a lightweight cloud platform which improves the offloading performance from cloud side. To achieve such goals, we analyze the characteristics of typical of- floading workloads and design our platform solution accordingly. Rattrap develops a new runtime environment, Cloud Android Container, for mobile computation offloading, replacing heavy- weight virtual machines (VMs). Our design exploits the idea of running operating systems with differential kernel features inside containers with driver extensions, which partially breaks the limitation of OS-level virtualization. With proposed resource sharing and code cache mechanism, Rattrap fundamentally improves the offloading performance. Our evaluation shows that Rattrap not only reduces the startup time of runtime environments and shows an average speedup of 16x, but also saves a large amount of system resources such as 75% memory footprint and at least 79% disk capacity. Moreover, Rattrap improves offloading response by as high as 63% over the cloud platform based on VM, and thus saving the battery life. Song Wu 0001, Jia Rao, Hai Jin 0001, Xiaohai Dai |
IPDPS | 3 |
| 2017 | Scheduler activations for interference-resilient SMP virtual machine schedulingabstractThe wide adoption of SMP virtual machines (VMs) and resource consolidation present challenges to efficiently executing multi-threaded programs in the cloud. An important problem is the semantic gaps between the guest OS and the hypervisor. The well-known lock-holder preemption (LHP) and lock-waiter preemption (LWP) problems are examples of such semantic gaps, in which the hypervisor is unaware of the activities in the guest OS and adversely deschedules virtual CPUs (vCPUs) that are executing in critical sections. Existing studies have focused on inferring a high-level semantic state of the guest OS to aid hypervisor-level scheduling so as to avoid the LHP and LWP problems. Kun Suo, Luwei Cheng, Jia Rao |
Middleware | 4 |
| 2017 | Preemptive, Low Latency Datacenter Scheduling via Lightweight Virtualization
Wei Chen 0038, Jia Rao, Xiaobo Zhou 0002 |
USENIX ATC | 2 |
| 2017 | Improving Performance of Heterogeneous MapReduce Clusters with Adaptive Task TuningabstractDatacenter-scale clusters are evolving toward heterogeneous hardware architectures due to continuous server replacement. Meanwhile, datacenters are commonly shared by many users for quite different uses. It often exhibits significant performance heterogeneity due to multi-tenant interferences. The deployment of MapReduce on such heterogeneous clusters presents significant challenges in achieving good application performance compared to in-house dedicated clusters. As most MapReduce implementations are originally designed for homogeneous environments, heterogeneity can cause significant performance deterioration in job execution despite existing optimizations on task scheduling and load balancing. In this paper, we observe that the homogeneous configuration of tasks on heterogeneous nodes can be an important source of load imbalance and thus cause poor performance. Tasks should be customized with different configurations to match the capabilities of heterogeneous nodes. To this end, we propose a self-adaptive task tuning approach, Ant, that automatically searches the optimal configurations for individual tasks running on different nodes. In a heterogeneous cluster, Ant first divides nodes into a number of homogeneous subclusters based on their hardware configurations. It then treats each subcluster as a homogeneous cluster and independently applies the self-tuning algorithm to them. Ant finally configures tasks with randomly selected configurations and gradually improves tasks configurations by reproducing the configurations from best performing tasks and discarding poor performing configurations. To accelerate task tuning and avoid trapping in local optimum, Ant uses genetic algorithm during adaptive task configuration. Experimental results on a heterogeneous physical cluster with varying hardware capabilities show that Ant improves the average job completion time by 31, 20, and 14 percent compared to stock Hadoop (Stock), customized Hadoop with industry recommendations (Heuristic), and a profilingbased configuration approach (Starfish), respectively. Furthermore, we extend Ant to virtual MapReduce clusters in a multi-tenant private cloud. Specifically, Ant characterizes a virtual node based on two measured performance statistics: I/O rate and CPU steal time. It uses k-means clustering algorithm to classify virtual nodes into configuration groups based on the measured dynamic interference. Experimental results on virtual clusters with varying interferences show that Ant improves the average job completion time by 20, 15, and 11 percent compared to Stock, Heuristic and Starfish, respectively. Dazhao Cheng, Jia Rao, Yanfei Guo, Changjun Jiang 0002, Xiaobo Zhou 0002 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2017 | iShuffle: Improving Hadoop Performance with Shuffle-on-WriteabstractHadoop is a popular implementation of the MapReduce framework for running data-intensive jobs on clusters of commodity servers.Shuffle, the all-to-all input data fetching phase between the map and reduce phase can significantly affect job performance. However, the shuffle phase and reduce phase are coupled together in Hadoop and the shuffle can only be performed by running the reduce tasks. This leaves the potential parallelism between multiple waves of map and reduce unexploited and resource wastage in multi-tenant Hadoop clusters, which significantly delays the completion of jobs in a multi-tenant Hadoop cluster. More importantly, Hadoop lacks the ability to schedule task efficiently and mitigate the data distribution skew among reduce tasks, which leads to further degradation of job performance. In this work, we propose to decouple shuffle from reduce tasks and convert it into a platform service provided by Hadoop. We presentiShuffle, a user-transparent shuffle service that pro-actively pushes map output data to nodes via a novelshuffle-on-writeoperation and flexibly schedules reduce tasks considering workload balance. Experimental results with representative workloads and Facebook workload trace show that iShuffle reduces job completion time by as much as 29.6 and 34 percent in single-user and multi-user clusters, respectively. Yanfei Guo, Jia Rao, Dazhao Cheng, Xiaobo Zhou 0002 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2017 | Moving Hadoop into the Cloud with Flexible Slot Management and Speculative ExecutionabstractLoad imbalance is a major source of overhead in parallel programs such as MapReduce. Due to the uneven distribution of input data, tasks with more data become stragglers and delay the overall job completion. Running Hadoop in a private cloud opens up opportunities for expediting stragglers with more resources but also introduces problems that often outweigh the performance gain: (1) performance interference from co-running jobs may create new stragglers; (2) there exists a semantic gap between the Hadoop task management and resource pool-based virtual cluster management preventing tasks from using resources efficiently. In this paper, we strive to make Hadoop more resilient to data skew and more efficient in cloud environments. We presentFlexSlot, a user-transparent task slot management scheme that automatically identifies map stragglers and resizes their slots accordingly to accelerate task execution. FlexSlot adaptively changes the number of slots on each virtual node to balance the resource usage so that the pool of resources can be efficiently utilized. FlexSlot further improves mitigation of data skew with an adaptive speculative execution strategy. Experimental results show that FlexSlot effectively reduces job completion time up to$47.2$percent compared to stock Hadoop and two recently proposed skew mitigation and speculative execution approaches. Yanfei Guo, Jia Rao, Changjun Jiang 0002, Xiaobo Zhou 0002 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2016 | Towards Application-centric Fairness in Multi-tenant Clouds with Adaptive CPU Sharing ModelabstractThe performance of cloud application is often quite disappointing due to unmanaged consolidation. Therefore, efforts are required to reduce co-tenants interference and provide predictable application performance in multi-tenant cloud environments. In this paper, we examined the complex interplay among cloud tenants as they compete for CPU time, and shared hardware resources. We propose Adaptive CPU Sharing (ACS) approach that reduces co-tenants interference and provides predictable application performance. Our approach is to monitor the progress of submitted applications at runtime, tracks the slowdown of individual application and applies adjustment until convergence. Thus, when an application suffered more slowdown, we allocate more CPU to reduce unfairness. In establishing system support for fine-grained profiling, we report system level activities at sub-second granularity. We predicted application performance degradation by creating a mathematical relationship between high-level application performance and low-level machine events (i.e., CPU steal time and L2 caches miss rate). We validate the added value of our approach by comparing application performance slowdowns (average) with various datasets. Based on our experimental results, our approach helps mitigate co-tenant interference and reduces unfairness by minimizing the overall application slowdowns. Anthony O. Ayodele, Jia Rao, Terrance E. Boult |
CLOUD | 2 |
| 2016 | Characterizing and Optimizing the Performance of Multithreaded Programs Under InterferenceabstractAs virtualization becomes ubiquitous in datacenters, there is a growing interest in characterizing application performance in multi-tenant environments to improve datacenter resource management. The performance of parallel programs is notoriously difficult to reason about in virtualized environments. Although performance degradations caused by virtualization and interferences have been extensively studied, there still lacks a comprehensive understanding why parallel programs have unpredictable slowdowns when co-located with different types of workloads. Jia Rao, Qing Yi |
PACT | 2 |
| 2016 | vScale: automatic and efficient processor scaling for SMP virtual machinesabstractSMP virtual machines (VMs) have been deployed extensively in clouds to host multithreaded applications. A widely known problem is that when CPUs are oversubscribed, the scheduling delays due to VM preemption give rise to many performance problems because of the impact of these delays on thread synchronization and I/O efficiency. Dynamically changing the number of virtual CPUs (vCPUs) by considering the available physical CPU (pCPU) cycles has been shown to be a promising approach. Unfortunately, there are currently no efficient mechanisms to support such vCPU-level elasticity. Luwei Cheng, Jia Rao, Francis C. M. Lau 0001 |
EuroSys | 2 |
| 2016 | Elastic Power-Aware Resource Provisioning of Heterogeneous Workloads in Self-Sustainable DatacentersabstractWhile major Cloud service operators have taken various initiatives to operate their datacenters with renewable energy partially or completely, it is challenging to effectively utilize the renewable energy since its generation depends on dynamic natural conditions. In this paper, we propose and develop an elastic power-aware resource provisioning approach (ePower) for heterogeneous workloads in self-sustainable datacenters that completely rely on renewable energy. We aim to maximize the system goodput and control the system power consumption with respect to green power supply. ePower takes challenges and advantages of dynamic power supply, heterogeneous workload characteristics and QoS requirements, and automatically optimizes elastic resource allocations to workloads. The core of ePower design is a novel power-aware simulated annealing algorithm with fuzzy performance modeling for the efficient search of an optimal resource allocation. We have implemented ePower in a university cloud testbed hosting Gridmix2 and RUBiS benchmark applications. We utilize real weather data traces to simulate the green power generation and supply in the experiments. Experimental results demonstrate ePower can achieve near-to-optimal system performance while being resilient to dynamic power availability. It outperforms a representative resource provisioning approach for heterogeneous workloads by at least 24% in improving system goodput and 35 percent in reducing QoS violations. Dazhao Cheng, Jia Rao, Changjun Jiang 0002, Xiaobo Zhou 0002 |
IEEE Trans. Computers | 2 |
| 2015 | Performance Measurement and Interference Profiling in Multi-tenant CloudsabstractThe ongoing rush for cloud-based services by small, medium, and large-scale organizations to reduce operational cost and to have more flexibility in the deployment and management of business applications cannot be overemphasized. However, the performance of in-cloud applications is often quite disappointing and unpredictable. Cloud users often perceive the sub optimal and unpredictable performance as anomalies as it is hard to conduct capacity planning based on such unreliable measurements. Performance interference due to resource sharing has been well studied in literature. Representative work includes the study of shared CPU caches, memory bandwidth, hard disks, network bandwidth, and the fair allocation CPU time. There lacks a comprehensive understanding of the complex interplay for shared resource under contention such as CPU. In this research work, we focus on predicting application performance by establishing a mathematical relationship between the high-level performance and the low-level CPU multiplexing. We design a synthetic workload with controllable CPU demands to emulate interference workloads in the cloud. We begin our measurements in a controlled environment to study the impact of CPU allocation on application performance. Based on the results from our experiments, we established the interdependency between CPU steal time, and application performance, and confirms that the percentage of CPU steal time influence application performance, even when workloads of equal parameters were submitted for processing on the same system platform. Our experimental results were evaluated against similar experimental results we performed on Amazon EC2 m3.medium model instance. We confirmed that the workload runtime duration on Amazon EC2 m3. medium model instance are significantly been impacted by high CPU steal time percentage due to interference from co-tenants. Therefore, the workload runtime slowdown percentage on submitted workloads on Amazon EC2.medium model instance is the hidden cost incurred by Cloud subscribers in term of time lost. Cloud service providers should pay close attention to CPU steal time percentage as part of system optimization efforts on Xen based cloud platform. We present Multi-tenant Performance Measurement and Interference Profiling system, a Xen based multi-tenant cloud environment designed for performance measurement and profiling. Anthony O. Ayodele, Jia Rao, Terrance E. Boult |
CLOUD | 2 |
| 2015 | StoreApp: A shared storage appliance for efficient and scalable virtualized Hadoop clustersabstractVirtualizing Hadoop clusters provides many benefits, including rapid deployment, on-demand elasticity and secure multi-tenancy. However, a simple migration of Hadoop to a virtualized environment does not fully exploit these benefits. The dual role of a Hadoop worker, acting as both a compute node and a data node, makes it difficult to achieve efficient IO processing, maintain data locality, and exploit resource elasticity in the cloud. We find that decoupling per-node storage from its computation opens up opportunities for IO acceleration, locality improvement, and on-the-fly cluster resizing. To fully exploit these opportunities, we propose StoreApp, a shared storage appliance for virtual Hadoop worker nodes co-located on the same physical host. To completely separate storage from computation and prioritize IO processing, StoreApp pro-actively pushes intermediate data generated by map tasks to the storage node. StoreApp also implements late-binding task creation to take the advantage of prefetched data due to mis-aligned records. Experimental results show that StoreApp achieves up to 61% performance improvement compared to stock Hadoop and resizes the cluster to the (near) optimal degree of parallelism. Yanfei Guo, Jia Rao, Dazhao Cheng, Changjun Jiang 0002, Cheng-Zhong Xu 0001, Xiaobo Zhou 0002 |
INFOCOM | 2 |
| 2015 | Resource and Deadline-Aware Job Scheduling in Dynamic Hadoop ClustersabstractAs Hadoop is becoming increasingly popular in large-scale data analysis, there is a growing need for providing predictable services to users who have strict requirements on job completion times. While earliest deadline first scheduling (EDF) like algorithms are popular in guaranteeing job deadlines in real-time systems, they are not effective in a dynamic Hadoop environment, i.e., a Hadoop cluster with dynamically available resources. As there is a growing number of Hadoop clusters deployed on hybrid systems, e.g., infrastructure powered by mix of traditional and renewable energy, and cloud platforms hosting heterogeneous workloads, variable resource availability becomes common when running Hadoop jobs. In this paper, we propose, RDS, a Resource and Deadline-aware Hadoop job Scheduler that takes future resource availability into consideration when minimizing job deadline misses. We formulate the job scheduling problem as an online optimization problem and solve it using an efficient receding horizon control algorithm. To aid the control, we design a self-learning model to estimate job completion times and use a simple but effective model to predict future resource availability. We have implemented RDS in the open source Hadoop implementation and performed evaluations with various benchmark workloads. Experimental results show that RDS substantially reduces the penalty of deadline misses by at least 36% and 10% compared with Fair Scheduler and EDF scheduler, respectively. Dazhao Cheng, Jia Rao, Changjun Jiang 0002, Xiaobo Zhou 0002 |
IPDPS | 2 |
| 2015 | Self-Boosted Co-scheduling for SMP Virtual MachinesabstractIn this paper, we propose a self-boosted co-scheduling(SBCO) algorithm to reduce synchronization latency among consolidated virtual machines. Different from conventional co-scheduling which requires all runnable sibling vCPUs that are from the same VM to be scheduled at precisely the same time, SBCO reorders all these sibling vCPUs threads coarsely at the same level in their respective run queue, then schedules them at the same time window, and maintains global fairness between consolidated VMs. SBCO minimizes costly pCPU preemption and preserves the flexibility of the dynamic mapping between vCPUs and pCPUs. We have implemented SBCO in KVM and conducted comprehensive evaluations with various workloads. Results shows that SBCO is able to reduce the number of context switches significantly and achieve overall performance improve up to 10% compared with other competitors and improve up to 60% compared with the default scheduler. Yudi Wei, Cheng-Zhong Xu 0001, Jia Rao |
MASCOTS | 4 |
| 2015 | Understanding Parallel Performance Under Interferences in Multi-tenant CloudsabstractThe performance of parallel programs is notoriously difficult to reason in virtualized environments. Although performance degradations caused by virtualization and interferences have been well studied, there is little understanding why different parallel programs have unpredictable slow- downs. We find that unpredictable performance is the result of complex interplays between the design of the program, the memory hierarchy of the hosting system, and the CPU scheduling in the hypervisor. We develop a profiling tool, vProfile, to decompose parallel runtime into three parts: compute, steal and synchronization. With the help of time breakdown, we devise two optimizations at the hypervisor to reduce slowdowns. Jia Rao, Xiaobo Zhou 0002, Qing Yi |
SIGMETRICS | 2 |
| 2014 | Improving MapReduce performance in heterogeneous environments with adaptive task tuningabstractThe deployment of MapReduce in datacenters and clouds present several challenges in achieving good job performance. Compared to in-house dedicated clusters, datacenters and clouds often exhibit significant hardware and performance heterogeneity due to continuous server replacement and multi-tenant interferences. As most Mapreduce implementations assume homogeneous clusters, heterogeneity can cause significant load imbalance in task execution, leading to poor performance and low cluster utilizations. Despite existing optimizations on task scheduling and load balancing, MapReduce still performs poorly on heterogeneous clusters. Dazhao Cheng, Jia Rao, Yanfei Guo, Xiaobo Zhou 0002 |
Middleware | 2 |
| 2014 | Towards fair and efficient SMP virtual machine schedulingabstractAs multicore processors become prevalent in modern computer systems, there is a growing need for increasing hardware utilization and exploiting the parallelism of such platforms. With virtualization technology, hardware utilization is improved by encapsulating independent workloads into virtual machines (VMs) and consolidating them onto the same machine. SMP virtual machines have been widely adopted to exploit parallelism. For virtualized systems, such as a public cloud, fairness between tenants and the efficiency of running their applications are keys to success. However, we find that existing virtualization platforms fail to enforce fairness between VMs with different number of virtual CPUs (vCPU) that run on multiple CPUs. We attribute the unfairness to the use of per-CPU schedulers and the load imbalance on these CPUs that incur inaccurate CPU allocations. Unfortunately, existing approaches to reduce unfairness, e.g., dynamic load balancing and CPU capping, introduce significant inefficiencies to parallel workloads. Jia Rao, Xiaobo Zhou 0002 |
PPoPP | 1 |
| 2014 | FlexSlot: Moving Hadoop Into the Cloud with Flexible Slot ManagementabstractLoad imbalance is a major source of overhead in Hadoop where the uneven distribution of input data among tasks can significantly delays the job completion. Running Hadoop in a private cloud opens up opportunities for mitigating data skew with elastic resource allocation, where stragglers are expedited with more resources, yet introduces problems that often cancel out the performance gain: (1) performance interference from co running jobs may create new stragglers, (2) there exist a semantic gap between Hadoop task management and resource pool-based virtual cluster management preventing efficient resource usage. We present Flex Slot, a user-transparent task slot management scheme that automatically identifies map stragglers and resizes their slots accordingly to accelerate task execution. Flex Slot adaptively changes the number of slots on each virtual node to promote efficient usage of resource pool. Experimental results with representative benchmarks show that Flex Slot effectively reduces job completion time by 46% and achieves better resource utilization. Yanfei Guo, Jia Rao, Changjun Jiang 0002, Xiaobo Zhou 0002 |
SC | 2 |
| 2013 | Optimizing virtual machine scheduling in NUMA multicore systemsabstractAn increasing number of new multicore systems use the Non-Uniform Memory Access architecture due to its scalable memory performance. However, the complex interplay among data locality, contention on shared on-chip memory resources, and cross-node data sharing overhead, makes the delivery of an optimal and predictable program performance difficult. Virtualization further complicates the scheduling problem. Due to abstract and inaccurate mappings from virtual hardware to machine hardware, program and system-level optimizations are often not effective within virtual machines. We find that the penalty to access the “uncore” memory subsystem is an effective metric to predict program performance in NUMA multicore systems. Based on this metric, we add NUMA awareness to the virtual machine scheduling. We propose a Bias Random vCPU Migration (BRM) algorithm that dynamically migrates vCPUs to minimize the system-wide uncore penalty. We have implemented the scheme in the Xen virtual machine monitor. Experiment results on a two-way Intel NUMA multicore system with various workloads show that BRM is able to improve application performance by up to 31.7% compared with the default Xen credit scheduler. Moreover, BRM achieves predictable performance with, on average, no more than 2% runtime variations. Jia Rao, Xiaobo Zhou 0002, Cheng-Zhong Xu 0001 |
HPCA | 1 |
| 2013 | Interference and locality-aware task scheduling for MapReduce applications in virtual clusters
Xiangping Bu, Jia Rao, Cheng-Zhong Xu 0001 |
HPDC | 2 |
| 2013 | eBase: A baseband unit cluster testbed to improve energy-efficiency for cloud radio access networkabstractRecently, power consumption of radio access networks (RAN) has attracted a lot of attention since energy cost takes up a vast portion of operational expenditure (OPEX). Specifically, more than half of the energy is consumed by base stations (BS). Baseband Unit (BBU), which is responsible for baseband signal processing, is an important part of BS, and BBU consolidation in the form of a cluster is an promising solution to enhance energy efficiency. Although there are several existing work addressing similar ideas from theoretical perspective, there is still a gap between theoretical and practical solutions. Hence, by leveraging cloud computing and software-defined radio technologies, we develop a BBU cluster testbed, eBase, to consolidates multiple isolated BBUs into a virtualized BBU and then provide unified baseband resource pool for RAN. We present the architecture and implementation of eBase in detail. We also design a distributed resource management middleware to realize comprehensive resource management strategy for BBU cluster in order to improve energy efficiency. Through experiments, we find eBase can reduce total energy consumption by about 20% while satisfying about 95% QoS. These results not only demonstrate BBU cluster's significant capability to reduce energy cost and improve resource utilization, but also show eBase is a capable prototyping testbed for cloud radio access network research. Zhen Kong, Jiayu Gong, Cheng-Zhong Xu 0001, Jia Rao |
ICC | 5 |
| 2013 | V-Cache: Towards Flexible Resource Provisioning for Multi-tier Applications in IaaS CloudsabstractAlthough the resource elasticity offered by Infrastructure-as-a-Service (IaaS) clouds opens up opportunities for elastic application performance, it also poses challenges to application management. Cluster applications, such as multi-tier websites, further complicates the management requiring not only accurate capacity planning but also proper partitioning of the resources into a number of virtual machines. Instead of burdening cloud users with complex management, we move the task of determining the optimal resource configuration for cluster applications to cloud providers. We find that a structural reorganization of multi-tier websites, by adding a caching tier which runs on resources debited from the original resource budget, significantly boosts application performance and reduces resource usage. We propose V-Cache, a machine learning based approach to flexible provisioning of resources for multi-tier applications in clouds. V-Cache transparently places a caching proxy in front of the application. It uses a genetic algorithm to identify the incoming requests that benefit most from caching and dynamically resizes the cache space to accommodate these requests. We develop a reinforcement learning algorithm to optimally allocate the remaining capacity to other tiers. We have implemented V-Cache on a VMware-based cloud testbed. Experiment results with the RUBiS and WikiBench benchmarks show that V-Cache outperforms a representative capacity management scheme and a cloud-cache based resource provisioning approach by at least 15% in performance, and achieves at least 11% and 21% savings on CPU and memory resources, respectively. Yanfei Guo, Palden Lama, Jia Rao, Xiaobo Zhou 0002 |
IPDPS | 3 |
| 2013 | QoS Guarantees and Service Differentiation for Dynamic Cloud ApplicationsabstractCloud elasticity allows dynamic resource provisioning in concert with actual application demands. Feedback control approaches have been applied with success to resource allocation in physical servers. However, cloud dynamics make the design of an accurate and stable resource controller challenging, especially when application-level performance is considered as the measured output. Application-level performance is highly dependent on the characteristics of workload and sensitive to cloud dynamics. To address these challenges, we extend a self-tuning fuzzy control (STFC) approach, originally developed for response time assurance in web servers to resource allocation in virtualized environments. We introduce mechanisms for adaptive output amplification and flexible rule selection in the STFC approach for better adaptability and stability. Based on the STFC, we further design a two-layer QoS provisioning framework, DynaQoS, that supports adaptive multi-objective resource allocation and service differentiation. We implement a prototype of DynaQoS on a Xen-based cloud testbed. Experimental results on representative server workloads show that STFC outperforms popular controllers such as Kalman filter, ARMA and, Adaptive PI in the control of CPU, memory, and disk bandwidth resources under both static and dynamic workloads. Further results with multiple control objectives and service classes demonstrate the effectiveness of DynaQoS in performance-power control and service differentiation. Jia Rao, Yudi Wei, Jiayu Gong, Cheng-Zhong Xu 0001 |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2013 | Coordinated Self-Configuration of Virtual Machines and Appliances Using a Model-Free Learning ApproachabstractCloud computing has a key requirement for resource configuration in a real-time manner. In such virtualized environments, both virtual machines (VMs) and hosted applications need to be configured on-the-fly to adapt to system dynamics. The interplay between the layers of VMs and applications further complicates the problem of cloud configuration. Independent tuning of each aspect may not lead to optimal system wide performance. In this paper, we propose a framework, namely CoTuner, for coordinated configuration of VMs and resident applications. At the heart of the framework is a model-free hybrid reinforcement learning (RL) approach, which combines the advantages of Simplex method and RL method and is further enhanced by the use of system knowledge guided exploration policies. Experimental results on Xen-based virtualized environments with TPC-W and TPC-C benchmarks demonstrate that CoTuner is able to drive a virtual server cluster into an optimal or near-optimal configuration state on the fly, in response to the change of workload. It improves the systems throughput by more than 30 percent over independent tuning strategies. In comparison with the coordinated tuning strategies based on basic RL or Simplex algorithm, the hybrid RL algorithm gains 25 to 40 percent throughput improvement. Xiangping Bu, Jia Rao, Cheng-Zhong Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2012 | An efficient index for massive IOT data in cloud environmentabstractThe Internet of Things (IOT) has been widely applied in many fields, while the IOT data are always large volume, update frequently and inherently multi-dimensional, these characteristics bring big challenges to the traditional DBMSs. The traditional DBMSs have rich functionality and can deal with multi-attributes access efficiently, they can not scale good enough to deal with large volume data and can not support high insert throughput. The cloud-based database systems have good scalability, but they don't support multi-dimensional access natively.In order to deal with the large volume of IOT data, we propose an update and query efficient index framework (UQE-Index) based on key-value store that can support high insert throughput and provide efficient multi-dimensional query simultaneously. We implemented a prototype based on HBase and did comprehensive experiments to test our solution's scalability and efficiency. Youzhong Ma, Jia Rao, Weisong Hu, Xiaofeng Meng 0001, Yu Zhang 0083, Yunpeng Chai, Chunqiu Liu |
CIKM | 2 |
| 2012 | URL: A unified reinforcement learning approach for autonomic cloud management
Cheng-Zhong Xu 0001, Jia Rao, Xiangping Bu |
J. Parallel Distributed Comput. | 2 |
| 2011 | DynaQoS: Model-free self-tuning fuzzy control of virtualized resources for QoS provisioningabstractCloud elasticity allows dynamic resource provisioning in concert with actual application demands. Feedback control approaches have been applied with success to resource allocation in physical servers. However, cloud dynamics make the design of an accurate and stable resource controller more challenging, especially when response time is considered as the measured output. Response time is highly dependent on the characteristics of workload and sensitive to cloud dynamics. To address the challenges, we extend a self-tuning fuzzy control (STFC) approach, originally developed for response time assurance in web servers to resource allocation in virtualized environments. We introduce mechanisms for adaptive output amplification and flexible rule selection in the STFC approach for better adaptability and stability. Based on the STFC, we further design a two-layer QoS provisioning framework, DynaQoS, that supports adaptive multi-objective resource allocation and service differentiation. We implement a prototype of DynaQoS on a Xen-based cloud testbed. Experimental results on an E-Commerce benchmark show that STFC outperforms popular controllers such as Kalman filter, ARMA and adaptive PI by at least 16% and 37% under both static and dynamic workloads, respectively. Further results with multiple control objectives and service classes demonstrate the effectiveness of DynaQoS in performance-power control and service differentiation. Jia Rao, Yudi Wei, Jiayu Gong, Cheng-Zhong Xu 0001 |
IWQoS | 1 |
| 2011 | A Model-free Learning Approach for Coordinated Configuration of Virtual Machines and AppliancesabstractCloud computing has a key requirement for resource configuration in a real-time manner. In such virtualized environments, both virtual machines (VMs) and hosted applications need to be configured on-the-fly to adapt to system dynamics. The interplay between the layers of VMs and applications further complicates the problem of cloud configuration. Independent tuning of each aspect may not lead to optimal system wide performance. In this paper, we propose a framework, namely CoTuner, for coordinated configuration of VMs and resident applications. At the heart of the framework is a model-free hybrid reinforcement learning (RL) approach, which combines the advantages of Simplex and RL methods and is further enhanced by the use of system knowledge guided exploration policies. Experimental results on Xen-based virtualized environments with TPC-W and TPC-C benchmarks demonstrate that CoTuner is able to drive a virtual server system into an optimal or near optimal configuration state dynamically, in response to the change of workload. It improves the systems throughput by more than 30% over independent tuning strategies. In comparison with the coordinated tuning strategies based solely on Simplex or basic RL algorithm, the hybrid RL algorithm gains 30% to 40% throughput improvement. Moreover, the algorithm is able to reduce SLA violation of the applications by more than 80%. Xiangping Bu, Jia Rao, Cheng-Zhong Xu 0001 |
MASCOTS | 2 |
| 2011 | A Distributed Self-Learning Approach for Elastic Provisioning of Virtualized Cloud ResourcesabstractAlthough cloud computing has gained sufficient popularity recently, there are still some key impediments to enterprise adoption. Cloud management is one of the top challenges. The ability of on-the-fly partitioning hardware resources into virtual machine(VM) instances facilitates elastic computing environment to users. But the extra layer of resource virtualization poses challenges on effective cloud management. The factors of time-varying user demand, complicated interplay between co-hosted VMs and the arbitrary deployment of multitier applications make it difficult for administrators to plan good VM configurations. In this paper, we propose a distributed learning mechanism that facilitates self-adaptive virtual machines resource provisioning. We treat cloud resource allocation as a distributed learning task, in which each VM being a highly autonomous agent submits resource requests according to its own benefit. The mechanism evaluates the requests and replies with feedback. We develop a reinforcement learning algorithm with a highly efficient representation of experiences as the heart of the VM side learning engine. We prototype the mechanism and the distributed learning algorithm in an iBalloon system. Experiment results on an Xen-based cloud test bed demonstrate the effectiveness of iBalloon. The distributed VM agents are able to reach near-optimal configuration decisions in 7 iteration step sat no more than 5% performance cost. Most importantly, iBalloon shows good scalability on resource allocation by scaling to 128 correlated VMs. Jia Rao, Xiangping Bu, Cheng-Zhong Xu 0001 |
MASCOTS | 1 |
| 2011 | Self-adaptive provisioning of virtualized resources in cloud computingabstractIn this paper, we propose a distributed learning mechanism that facilitates self-adaptive virtual machines resource provisioning. We treat cloud resource allocation as a distributed learning task, in which each VM being a highly autonomous agent submits resource requests according to its own benefit. The mechanism evaluates the requests and replies with feedback. We develop a reinforcement learning algorithm with a highly efficient representation of experiences as the heart of the VM side learning engine. We prototype the mechanism and the distributed learning algorithm in an iBalloon system. Experiment results on a Xen-based cloud testbed demonstrate the effectiveness of iBalloon. Jia Rao, Xiangping Bu, Cheng-Zhong Xu 0001 |
SIGMETRICS | 1 |
| 2011 | Rethink the virtual machine templateabstractServer virtualization technology facilitates the creation of an elastic computing infrastructure on demand. There are cloud applications like server-based computing and virtual desktop that concern startup latency and require impromptu requests for VM creation in a real-time manner. Conventional template-based VM creation is a time consuming process and lacks flexibility for the deployment of stateful VMs. In this paper, we present an abstraction of VM substrate to represent generic VM instances in miniature. Unlike templates that are stored as an image file in disk, VM substrates are docked in memory in a designated VM pool. They can be activated into stateful VMs without machine booting and application initialization. The abstraction leverages an arrange of techniques, including VM miniaturization, generalization, clone and migration, storage copy-on-write, and on-the-fly resource configuration, for rapid deployment of VMs and VM clusters on demand. We implement a prototype on a Xen platform and show that a server with typical configuration of TB disk and GB memory can accommodate more substrates in memory than templates in disk and stateful VMs can be created from the same or different substrates and deployed on to the same or different physical hosts in a cluster without causing any configuration conflicts. Experimental results show that general purpose VMs or a VM cluster for parallel computing can be deployed in a few seconds. We demonstrate the usage of VM substrates in a mobile gaming application. Jia Rao, Cheng-Zhong Xu 0001 |
VEE | 2 |
| 2011 | Online Capacity Identification of Multitier Websites Using Hardware Performance CountersabstractUnderstanding server capacity is crucial to system capacity planning, configuration, and QoS-aware resource management. Conventional stress testing approaches measure server capacity offline in terms of application-level performance metrics like response time and throughput. They are limited in measurement accuracy and timeliness. In a multitier website, resource bottleneck often shifts between tiers as client access pattern changes. This makes the problem of online capacity measurement even more challenge. This paper presents an online measurement approach based on low-level hardware performance metrics such as instructions execution rate and cache access behavior. Such metrics together define a system internal running state. The measurement approach uses machine learning techniques to infer application-level performance at each tier from a set of selected hardware performance counters. A coordinated predictor is induced over individual tier-wide models to make global system performance prediction and identify the bottleneck when the system becomes overloaded. Experiments were conducted on a two-tier Tomcat/MySQL-configured website using TPC-W benchmarks. Experimental results demonstrated that this approach was able to achieve an overload prediction accuracy of higher than 90 percent for a priori known input traffic mix and over 85 percent accuracy even for traffic causing frequent bottleneck shifting. It costs less than 0.5 percent runtime overhead for data collection and no more than 50 ms for each online decision making. Jia Rao, Cheng-Zhong Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2009 | A Reinforcement Learning Approach to Online Web Systems Auto-configurationabstractIn a web system, configuration is crucial to the performance and service availability. It is a challenge, not only because of the dynamics of Internet traffic, but also the dynamic virtual machine environment the system tends to be run on. In this paper, we propose a reinforcement learning approach for autonomic configuration and reconfiguration of multi-tier web systems. It is able to adapt performance parameter settings not only to the change of workload, but also to the change of virtual machine configurations. The RL approach is enhanced with an efficient initialization policy to reduce the learning time for online decision. The approach is evaluated using TPC-W benchmark on a three-tier website hosted on a Xen-based virtual machine environment. Experiment results demonstrate that the approach can auto-configure the web system dynamically in response to the change in both workload and VM resource. It can drive the system into a near-optimal configuration setting in less than 25 trial-and-error iterations. Xiangping Bu, Jia Rao, Cheng-Zhong Xu 0001 |
ICDCS | 2 |
| 2008 | Online Measurement of the Capacity of Multi-Tier Websites Using Hardware Performance CountersabstractUnderstanding server capacity is crucial for system capacity planning, configuration, and QoS-aware resource management. Conventional stress testing approaches measure the server capacity in terms of application-level performance metrics like response time and throughput. They are limited in measurement accuracy and timeliness. In a multitier website, resource bottleneck often shifts between tiers as client access pattern changes. This makes the capacity measurement even more challenging. This paper presents a measurement approach based on hardware performance counter metrics. The approach uses machine learning techniques to infer application-level performance at each tier. A coordinated predictor is induced over individual tier models to estimate system-wide performance and identify the bottleneck when the system becomes overloaded. Experimental results demonstrate that this approach is able to achieve an overload prediction accuracy of higher than 90% for a priori known input traffic patterns and over 85% accuracy even for traffic causing frequent bottleneck shifting. It costs less than 0.5% runtime overhead for data collection and no more than 50 ms for each on-line decision. Jia Rao, Cheng-Zhong Xu 0001 |
ICDCS | 1 |
| 2008 | CoSL: A coordinated statistical learning approach to measuring the capacity of multi-tier websitesabstractWebsite capacity determination is crucial to measurement-based access control, because it determines when to turn away excessive client requests to guarantee consistent service quality under overloaded conditions. Conventional capacity measurement approaches based on high-level performance metrics like response time and throughput may result in either resource over-provisioning or lack of responsiveness. It is because a website may have different capacities in terms of the maximum concurrent level when the characteristic of workload changes. Moreover, bottleneck in a multi-tier website may shift among tiers as client access pattern changes. In this paper, we present an online robust measurement approach based on statistical machine learning techniques. It uses a Bayesian network to correlate low level instrumentation data like system and user cpu time, available memory size, and I/O status that are collected at run-time to high level system states in each tier. A decision tree is induced over a group of coordinated Bayesian models in different tiers to identify the bottleneck dynamically when the system is overloaded. Experimental results demonstrate its accuracy and robustness in different traffic loads. Jia Rao, Cheng-Zhong Xu 0001 |
IPDPS | 1 |
| 2006 | Annotated Logic Reasoning Based On XML
Fuxi Zhu, Duanzhi Chen, Jia Rao |
WEBIST (1) | 3 |