Ming Liu 0006

dblp:20/2039-6 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
0since 2021 · last 2015
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Cloud and datacenter computing · 49% Energy-efficient computing · 23% GPUs and heterogeneous computing · 13%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing
virtualization
0.422015
Understanding the virtualization "Tax" of scale-out pass-through GPUs in GaaS clouds: An empirical study · HPCA 2015
Optimizing virtual machine consolidation performance on NUMA server architecture for cloud workloads · ISCA 2014
GPUs and heterogeneous computing › GPU resource management
GPU virtualization
0.212015
Understanding the virtualization "Tax" of scale-out pass-through GPUs in GaaS clouds: An empirical study · HPCA 2015
Memory systems › memory management
NUMA memory management
0.212014
Optimizing virtual machine consolidation performance on NUMA server architecture for cloud workloads · ISCA 2014
Cloud and datacenter computing › virtualization › virtual machine management
virtual machine consolidation
0.212014
Optimizing virtual machine consolidation performance on NUMA server architecture for cloud workloads · ISCA 2014
Energy-efficient computing
datacenter power management
0.212013
Enabling datacenter servers to scale out economically and sustainably · MICRO 2013
Energy-efficient computing › datacenter power management
power oversubscription
0.212013
Enabling datacenter servers to scale out economically and sustainably · MICRO 2013
Performance modeling and evaluation
workload characterization
0.112015
Understanding the virtualization "Tax" of scale-out pass-through GPUs in GaaS clouds: An empirical study · HPCA 2015
Energy-efficient computing › datacenter power management
power provisioning
0.012013
Enabling datacenter servers to scale out economically and sustainably · MICRO 2013

Methods — techniques the papers use, named apart from their topics

empirical measurement · 0.2VMEXIT analysis · 0.2page fault handling · 0.2buddy allocator · 0.2P2M swap FIFO · 0.2
YearPublicationVenuePosition
2015 Understanding the virtualization "Tax" of scale-out pass-through GPUs in GaaS clouds: An empirical study
abstract
Pass-through techniques enable virtual machines to directly access hardware GPU resources in an exclusive mode, delivering extraordinary graphics performance for client users in GaaS clouds. However, the virtualization overheads of pass-through GPUs may decrease the frame rate of graphics workloads by reducing the occupancy rate of the GPU working queue. In this work, we make the first attempt to characterize pass-through GPUs running in different consolidation scenarios and uncover the root causes of these overheads. Towards this end, we set up state-of-the-art empirical platforms equipped with NVIDIA GRID GPUs and execute graphics intensive workloads running in GaaS clouds. We first demonstrate the existence of virtualization overheads, which can slow down the GPU command generation rate. Compared with a bare-metal system, the performance of pass-through GPUs degrades 9.0% and 21.5% under a single VM and 8-VMs respectively. We analyze the workflow of Windows display driver model and VMEXIT events distribution and identify four factors (i.e. HLT instruction and idle domain, external interrupt delivery, IOMMU, and memory subsystem) that contribute to the performance degradation. Our evaluation results show that: (1) the VM-VMM context switch caused by a HLT instruction and wake-up interrupt injection of an idle domain result in 66. 7% idle time for a single pass-through GPU; (2) the external interrupt delivery and tasklet processing cause additional overheads. When 8 VMs are consolidated, the interrupt delivery processing time and interrupt frequency rise 30.7% and 127.3%, respectively; (3) the existing IOMMU design scales well with pass-through GPUs; and (4) interactions of domain guest's software stacks impact the hardware prefetching mechanism so that it fails to compensate the rapidly growing LLC miss rate when more pass-through GPU VMs are added. To the best of our knowledge, this is the first work that characterizes pass-through GPU virtualization overheads and underlying reasons. This study highlights valuable insights for improving the performance of future virtualized GPU systems.
Ming Liu 0006, Tao Li 0006, Neo Jia, Andy Currid, Vladimir Troy
HPCA1
2015 Optimization of Resource Allocation and Energy Efficiency in Heterogeneous Cloud Data Centers
abstract
Performance and energy efficiency are major concerns in cloud computing data centers. More often, they carry conflicting requirements making optimization a challenge. Further complications arise when heterogeneous hardware and data center management technologies are combined. For example, heterogeneous hardware such as General Purpose Graphics Processing Units (GPGPUs) improve performance at the cost of greater power consumption while virtualization technologies improve resource management and utilization at the cost of degraded performance. In this paper, we focus on exploiting heterogeneity introduced by GPUs to reduce power budget requirements for servers while maintaining performance. To maintain or improve overall server performance at reduced power budget, we propose two enhancements: (a) We borrow power from co-located multithreaded virtual machines (VMs) and reallocate it to GPU VMs. (b) To compensate multi-threaded VMs and re-boost their performance, we propose to borrow virtual computing resources from GPU VMs and reallocate them to CPU VMs. Combining the two techniques minimizes server power budget while maintaining overall server performance. Our results show that server power budget can be reduced by almost 18% at the average cost of 13% performance degradation per virtual machine. In addition, reallocating virtual resources improves the performance of multi-threaded applications by 30% without affecting GPU applications. Combining both techniques reduces server energy consumption by 47 % with minimum performance degradation.
Amer Qouneh, Ming Liu 0006, Tao Li 0006
ICPP2
2014 Optimizing virtual machine consolidation performance on NUMA server architecture for cloud workloads
abstract
Server virtualization and workload consolidation enable multiple workloads to share a single physical server, resulting in significant energy savings and utilization improvements. The shift of physical server architectures to NUMA and the increasing popularity of scale-out cloud applications undermine workload consolidation efficiency and result in overall system degradation. In this work, we characterize the consolidation of cloud workloads on NUMA virtualized systems, estimate four different sources of architecture overhead, and explore optimization opportunities beyond the default NUMA-aware hypervisor memory management. Motivated by the observed architectural impact on cloud workload consolidation performance, we propose three optimization techniques incorporating NUMA access overhead into the hypervisor's virtual machine memory allocation and page fault handling routines. Among these, estimation of the memory zone access overhead serves as a foundation for the other two techniques: a NUMA overhead aware buddy allocator and a P2M swap FIFO. Cache hit rate, cycle loss due to cache miss, and IPC serve as indicators to estimate the access cost of each memory node. Our optimized buddy allocator dynamically selects low-overhead memory zones and “proportionally” distributes memory pages across target nodes. The P2M swap FIFO records recently unusedlists for mapping exchanges to rebalance memory access pressure within one domain. Our real system based evaluations show a 41.1% performance improvement when consolidating 16-VMs on a 4-socket server (the proposed allocator contributes 22.8% of the performance gain and the P2M swap FIFO accounts for the rest). Furthermore, our techniques can cooperate well with other methods (i.e. vCPU migration) and scale well when varying VM memory size and the number of sockets in a physical host.
Ming Liu 0006, Tao Li 0006
ISCA1
2014 Understanding the Impact of vCPU Scheduling on DVFS-Based Power Management in Virtualized Cloud Environment
abstract
Virtualized platform has emerged as a prominent environment for cloud computing, especially in today's power-constrained data centers. However, due to a lack of coordination between runtime power management and a virtual CPU (vCPU) scheduler, existing virtualized cloud platform is far from efficient. First, current frequency control mechanism is unable to satisfy the fast-changing vCPU frequency requirement imposed by vCPU scheduler, which we refer to as demand imbalance problem. In addition, newly created vCPUs, if scheduled solely based on fairness, can cause inefficient frequency rise and drop on an unmatched physical core, which we refer to as utilization mismatch problem. In both cases, the system incurs degraded power efficiency and sub-optimal workload performance. In this study we perform a comprehensive analysis on the interplay between vCPU scheduling and processor-centric power control in virtualized cloud environment. Using representative workloads from Cloud Suite and real server deployment, we examine the energy/performance implications of frequency scaling and vCPU scheduling on both single-VM and multi-VM cloud host. We show that existing virtualized platform has the potential to improve energy efficiency and workload performance by 32% and 25%, respectively, if vCPUs are balanced and appropriately scheduled. We also show that dirty page rate, virtual block device processing rate, virtual network packets arrival rate, and network I/O buffer availability are important efficiency indicators for energy-efficient virtualized cloud system design.
Ming Liu 0006, Chao Li 0009, Tao Li 0006
MASCOTS1
2014 Towards Automated Provisioning and Emergency Handling in Renewable Energy Powered Datacenters
Chao Li 0009, Rui Wang 0014, Yang Hu 0001, Ruijin Zhou, Ming Liu 0006, Longjun Liu, Jingling Yuan, Tao Li 0006, Depei Qian 0001
J. Comput. Sci. Technol.5
2013 Enabling datacenter servers to scale out economically and sustainably
abstract
As cloud applications proliferate and data-processing demands increase, server resources must grow to unleash the performance of emerging workloads that scale well with large number of compute nodes. Nevertheless, power has become a crucial bottleneck that restricts horizontal scaling (scale out) of server systems, especially in datacenters that employ power over-subscription. When a datacenter hits the maximum capacity of its power provisioning equipment, the owner has to either build another facility or upgrade existing utility power infrastructure -- both approaches add huge capital expenditure, require significant construction lead time, and can further increase the owner's carbon footprint.
Chao Li 0009, Yang Hu 0001, Ruijin Zhou, Ming Liu 0006, Longjun Liu, Jingling Yuan, Tao Li 0006
MICRO4