Yaozu Dong

dblp:99/566 · DBLP profile ↗
← Back
34ranked-venue papers
9as first author
2since 2021 · last 2023
0000-0002-9623-5933ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 31 · 9 first-author · 2 since 2021Computer networks · 2Software engineering, systems software and programming languages · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
15 papers
Cloud and datacenter computing · 51% Memory systems · 17% GPUs and heterogeneous computing · 14%
Software engineering, system software, and programming languages
5 papers
Operating systems · 100%

Topics — the 27 heaviest of 31, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing
virtualization
2.8112022
MDev-NVMe: Mediated Pass-Through NVMe Virtualization Solution With Adaptive Polling · IEEE Trans. Computers 2022
gMig: Efficient vGPU Live Migration with Overlapped Software-Based Dirty Page Verification · IEEE Trans. Parallel Distributed Syst. 2020
coIOMMU: A Virtual IOMMU with Cooperative DMA Buffer Tracking for Efficient Memory Management in Direct I/O · USENIX ATC 2020
GPUs and heterogeneous computing › GPU resource management
GPU virtualization
1.452020
gMig: Efficient vGPU Live Migration with Overlapped Software-Based Dirty Page Verification · IEEE Trans. Parallel Distributed Syst. 2020
Scalable GPU Virtualization with Dynamic Sharing of Graphics Memory Space · IEEE Trans. Parallel Distributed Syst. 2018
gScale: Scaling up GPU Virtualization with Dynamic Sharing of Graphics Memory Space · USENIX ATC 2016
Memory systems
hybrid memory
0.712023
FlexHM: A Practical System for Heterogeneous Memory with Flexible and Efficient Performance Optimizations · ACM Trans. Archit. Code Optim. 2023
Cloud and datacenter computing › resource management
memory resource management
0.712023
FlexHM: A Practical System for Heterogeneous Memory with Flexible and Efficient Performance Optimizations · ACM Trans. Archit. Code Optim. 2023
Memory systems › hybrid memory
non-volatile memory and DRAM
0.712023
FlexHM: A Practical System for Heterogeneous Memory with Flexible and Efficient Performance Optimizations · ACM Trans. Archit. Code Optim. 2023
Storage systems
flash and SSD
0.612022
MDev-NVMe: Mediated Pass-Through NVMe Virtualization Solution With Adaptive Polling · IEEE Trans. Computers 2022
Storage systems › i/o architecture › i/o subsystem
direct i/o
0.412020
coIOMMU: A Virtual IOMMU with Cooperative DMA Buffer Tracking for Efficient Memory Management in Direct I/O · USENIX ATC 2020
Memory systems › memory management › memory management unit
IOMMU
0.412020
coIOMMU: A Virtual IOMMU with Cooperative DMA Buffer Tracking for Efficient Memory Management in Direct I/O · USENIX ATC 2020
Cloud and datacenter computing › virtualization › virtual machine migration
live migration
0.412020
gMig: Efficient vGPU Live Migration with Overlapped Software-Based Dirty Page Verification · IEEE Trans. Parallel Distributed Syst. 2020
Cloud and datacenter computing › virtualization
i/o virtualization
0.432013
SR-IOV Based Network Interrupt-Free Virtualization with Event Based Polling · IEEE J. Sel. Areas Commun. 2013
ReNIC: Architectural extension to SR-IOV I/O virtualization for efficient replication · ACM Trans. Archit. Code Optim. 2012
High performance network virtualization with SR-IOV · HPCA 2010
Storage systems
storage virtualization
0.312018
MDev-NVMe: A NVMe Storage Virtualization Solution with Mediated Pass-Through · USENIX ATC 2018
Operating systems › resource management › memory management
virtual memory
0.212023
FlexHM: A Practical System for Heterogeneous Memory with Flexible and Efficient Performance Optimizations · ACM Trans. Archit. Code Optim. 2023
Cloud and datacenter computing › virtualization
hardware-assisted virtualization
0.212014
HYVI: A HYbrid VIrtualization Solution Balancing Performance and Manageability · IEEE Trans. Parallel Distributed Syst. 2014
Cloud and datacenter computing › virtualization
device pass-through
0.212022
MDev-NVMe: Mediated Pass-Through NVMe Virtualization Solution With Adaptive Polling · IEEE Trans. Computers 2022
Cloud and datacenter computing › cloud storage
multi-tenant cloud storage
0.212022
MDev-NVMe: Mediated Pass-Through NVMe Virtualization Solution With Adaptive Polling · IEEE Trans. Computers 2022
Hardware reliability and fault tolerance
memory mirroring
0.212013
kMemvisor: flexible system wide memory mirroring in virtual environments · HPDC 2013
Cloud and datacenter computing › virtualization › i/o virtualization
network i/o virtualization
0.212013
Performance Enhancement for Network I/O Virtualization with Efficient Interrupt Coalescing and Virtual Receive-Side Scaling · IEEE Trans. Parallel Distributed Syst. 2013
Cloud and datacenter computing › virtualization
virtual machine networking
0.212013
Performance Enhancement for Network I/O Virtualization with Efficient Interrupt Coalescing and Virtual Receive-Side Scaling · IEEE Trans. Parallel Distributed Syst. 2013
Distributed systems
fault tolerance
0.112012
ReNIC: Architectural extension to SR-IOV I/O virtualization for efficient replication · ACM Trans. Archit. Code Optim. 2012
Distributed systems › replication
virtual machine replication
0.112012
ReNIC: Architectural extension to SR-IOV I/O virtualization for efficient replication · ACM Trans. Archit. Code Optim. 2012
Operating systems › virtualization
hypervisor
0.122013
SR-IOV Based Network Interrupt-Free Virtualization with Event Based Polling · IEEE J. Sel. Areas Commun. 2013
High performance network virtualization with SR-IOV · HPCA 2010
Operating systems
virtualization
0.112015
Boosting GPU Virtualization Performance with Hybrid Shadow Page Tables · USENIX ATC 2015
Cloud and datacenter computing › virtualization › virtual machine management
server consolidation
0.112014
HYVI: A HYbrid VIrtualization Solution Balancing Performance and Manageability · IEEE Trans. Parallel Distributed Syst. 2014
Cloud and datacenter computing › virtualization
virtual machine migration
0.112014
HYVI: A HYbrid VIrtualization Solution Balancing Performance and Manageability · IEEE Trans. Parallel Distributed Syst. 2014
Operating systems › virtualization
i/o virtualization
0.012013
Performance Enhancement for Network I/O Virtualization with Efficient Interrupt Coalescing and Virtual Receive-Side Scaling · IEEE Trans. Parallel Distributed Syst. 2013
Hardware reliability and fault tolerance
memory reliability
0.012013
kMemvisor: flexible system wide memory mirroring in virtual environments · HPDC 2013
Cloud and datacenter computing
cloud infrastructure
0.012012
ReNIC: Architectural extension to SR-IOV I/O virtualization for efficient replication · ACM Trans. Archit. Code Optim. 2012

Methods — techniques the papers use, named apart from their topics

two-level NUMA design · 1.3memory tiering · 1.3mediated pass-through · 0.8adaptive polling · 0.6sampling pre-filtering · 0.4one-shot pre-copy · 0.4hashing-based dirty page detection · 0.4DMA buffer tracking · 0.4ladder mapping · 0.3fence memory space pool · 0.3paravirtualization comparison · 0.2virtual receive-side scaling · 0.2interrupt coalescing · 0.2event-based polling · 0.2SR-IOV · 0.2
YearPublicationVenuePosition
2023 FlexHM: A Practical System for Heterogeneous Memory with Flexible and Efficient Performance Optimizations
abstract
With the rapid development of cloud computing, numerous cloud services, containers, and virtual machines have been bringing tremendous demands on high-performance memory resources to modern data centers. Heterogeneous memory, especially the newly released Optane memory, offer appropriate alternatives against DRAM in clouds with the advantages of larger capacity, lower purchase cost, and promising performance. However, cloud services suffer serious implementation inconvenience and performance degradation when using hybrid DRAM and Optane memory. This article proposes FlexHM, a practical system to manage transparent heterogeneous memory resources and flexibly optimize memory access performance for all VMs, containers, and native applications. We present an open-source prototype of FlexHM in Linux with several main contributions. First, FlexHM raises a novel two-level NUMA design to manage DRAM and Optane memory as transparent main memory resources. Second, FlexHM provides flexible and efficient memory management, helping optimize memory access performance or save purchase costs of memory resources for differential cloud services with customized management strategies. Finally, the evaluations show that cloud workloads using 50% Optane slow memory on FlexHM can achieve up to 93% of the performance when using all-DRAM, and FlexHM provides up to 5.8× improvement over the previous heterogeneous memory system solution when workloads use the same ratio of DRAM and Optane memory.
Bo Peng 0043, Yaozu Dong, Jianguo Yao 0002, Fengguang Wu, Haibing Guan
ACM Trans. Archit. Code Optim.2
2022 MDev-NVMe: Mediated Pass-Through NVMe Virtualization Solution With Adaptive Polling
abstract
The fast access to data and high parallel processing in high-performance computing instigates an urgent demand on the improvement of the NVMe storage within modern data centers. However, the former NVMe virtualization’s unsatisfactory performance demonstrates that NVMe devices are often underutilized within cloud platforms. An NVMe virtualization mechanism with high performance and device sharing has captured researchers and developers’ attention. This article introduces MDev-NVMe, a new virtualization solution for NVMe storage device with (1) full NVMe storage virtualization for VMs running native NVMe driver, (2) a mediated pass-through mechanism for NVMe management, and (3) adaptive configuration of active polling optimization to simultaneously achieve high throughput, low latency performance, and substantial device scalability. We practically implement the MDev-NVMe as a Linux kernel module. This article subsequently evaluates MDev-NVMe with Intel OPTANE and P3600 SSD by comparing several mainstream NVMe virtualization mechanisms using application-level I/O benchmarks. MDev-NVMe with active polling can demonstrate a 142 percent improvement over native (interrupt-driven) throughput and over 2.5 × theVirtiothroughput with only 70 percent native average latency and 31 percentVirtioaverage latency. Finally, the advantages of MDev-NVMe and the importance of adaptive polling are discussed, offering evidence that MDev-NVMe is a superior virtualization choice for cloud storage.
Bo Peng 0043, Jianguo Yao 0002, Yaozu Dong, Haibing Guan
IEEE Trans. Computers3
2020 coIOMMU: A Virtual IOMMU with Cooperative DMA Buffer Tracking for Efficient Memory Management in Direct I/O
Luwei Kang, Yaozu Dong
USENIX ATC5
2020 gMig: Efficient vGPU Live Migration with Overlapped Software-Based Dirty Page Verification
abstract
This paper introduces gMig, an open-source and practical vGPU live migration solution for full virtualization. Taking the advantage of the dirty pattern of GPU workloads, gMig presents the One-Shot Pre-Copy mechanism combined with the hashing based Software Dirty Page technique to achieve efficient vGPU live migration. Particularly, we propose three core techniques for gMig: 1) Dynamic Graphics Address Remapping, which parses and manipulates GPU commands to adjust the address mapping and adapt to a different environment after migration, 2) Software Dirty Page, which utilizes a hashing based approach with sampling pre-filtering to detect page modification, overcomes the commodity GPU's hardware limitation, and speeds up the migration by only sending the dirtied pages, 3) Overlapped Migration Process, which significantly compresses the hanging overhead by overlapping the dirty page verification and transmission concurrently. Our evaluation shows that gMig achieves GPU live migration with an average downtime of 302 ms on Windows and 119 ms on Linux. With the help of Software Dirty Page, the number of GPU pages transferred during the downtime is effectively reduced by up to 80.0 percent . The design of sampling filter and overlapped processing can bring about further 30.0 and 10.0 percent improvements in page processing.
Qiumin Lu, Jiacheng Ma 0001, Yaozu Dong, Zhengwei Qi, Jianguo Yao 0002, Bingsheng He, Haibing Guan
IEEE Trans. Parallel Distributed Syst.4
2019 ACRN: a big little hypervisor for IoT development
abstract
With the rapid growth of Internet of Things (IoT) and the new emerging IoT computing paradigm such as edge computing, it is prevalent to see that today’s real-time and functional safety devices, particularly in industrial IoT and automotive scenarios, are getting multi-functional by combining multiple platforms into single product. The new trend potentially prompts embedded virtualization as a promising solution in terms of workload consolidation, separation, and cost- effective. However, hypervisors, such as KVM and XEN, are designed to run on a server and can not be easily restructured to fulfill the requirements such as real-time constrains from IoT products. Meanwhile, existing embedded virtualization solutions are normally tailored towards specific IoT scenarios, which makes them hard to extend towards various scenarios. In addition, most commercial solutions are mature and appealing but expensive and closed-source. This paper presents ACRN, a flexible, lightweight, scalable, and open source embedded hypervisor for IoT development. By focusing on CPU and memory partitioning, and mean- while optionally offloading embedded I/O virtualization to a tiny user space device model, ACRN presents a consolidated system satisfying real-time and general-purpose needs simultaneously. By adopting customer-friendly permissive BSD license, ACRN provides a practical industry-grade solution with immediate readiness. In this paper we will de- scribe the design and implementation of ACRN, and conduct thorough evaluations to demonstrate its feasibility and effectiveness. The source code of ACRN has been released at https://github.com/projectacrn/acrn-hypervisor.
Jinkui Ren, Yaozu Dong
VEE4
2018 MDev-NVMe: A NVMe Storage Virtualization Solution with Mediated Pass-Through
Bo Peng 0043, Haozhong Zhang, Jianguo Yao 0002, Yaozu Dong, Haibing Guan
USENIX ATC4
2018 gMig: Efficient GPU Live Migration Optimized by Software Dirty Page for Full Virtualization
abstract
This paper introduces gMig, an open-source and practical GPU live migration solution for full virtualization. By taking advantage of the dirty pattern of GPU workloads, gMig presents the One-Shot Pre-Copy combined with the hashing based Software Dirty Page technique to achieve efficient GPU live migration. Particularly, we propose three approaches for gMig: 1) Dynamic Graphics Address Remapping, which parses and manipulates GPU commands to adjust the address mapping to adapt to a different environment after migration, 2) Software Dirty Page, which utilizes a hashing based approach to detect page modification, overcomes the commodity GPU's hardware limitation, and speeds up the migration by only sending the dirtied pages, 3) One-Shot Pre-Copy, which greatly reduces the rounds of pre-copy of graphics memory. Our evaluation shows that gMig achieves GPU live migration with an average downtime of 302 ms on Windows and 119 ms on Linux. With the help of Software Dirty Page, the number of GPU pages transferred during the downtime is effectively reduced by 80.0%.
Jiacheng Ma 0001, Yaozu Dong, Wentai Li, Zhengwei Qi, Bingsheng He, Haibing Guan
VEE3
2018 Demon: An Efficient Solution for on-Device MMU Virtualization in Mediated Pass-Through
abstract
Memory Management Units (MMUs) for on-device address translation are widely used in modern devices. However, conventional solutions for on-device MMU virtualization, such as shadow page table implemented in mediated pass-through, still suffer from high complexity and low performance.
Jianguo Yao 0002, Yaozu Dong, Haibing Guan
VEE3
2018 Scalable GPU Virtualization with Dynamic Sharing of Graphics Memory Space
abstract
With increasing GPU-intensive workloads deployed on cloud, cloud service providers are seeking for practical and efficient GPU virtualization solutions. However, the cutting-edge GPU virtualization techniques such as gVirt still suffer from the restriction of scalability, which constrains the number of guest virtual GPU instances. This paper presents gScale, a scalable and practical open source GPU virtualization solution based on gVirt. gScale presents a sharing mechanism which combines partition and sharing together to break the hardware limitation of global graphics memory space. Particularly, we propose two approaches for gScale: (1) the private shadow graphics translation table (GTT) , which enables global graphics memory space sharing among virtual GPUs, (2) ladder mapping and fence memory space pool, which allows CPU access host physical memory space (serving the graphics memory) to bypass global graphics memory space. Furthermore, to mitigate the performance degradation caused by switching private shadow GTT when the number of vGPUs scales up, four other mechanisms are proposed: (1) slot sharing, which improves the performance of vGPU by dividing the high global graphics memory into multiple slots, (2) fine-grained slotting, which provides a flexible virtual graphics memory configuration, (3) predictive GTT copy mechanism, which reduces the performance loss by switching private shadow GTT before context switch, (4) predictive-copy aware scheduling, which maximizes the improvement of predictive GTT copy mechanism in cloud environment. Evaluation shows that gScale scales up to 15 guest virtual GPU instances in Linux or 12 guest virtual GPU instances in Windows, which is 5x and 4x, respectively, that of gVirt. At the same time, gScale incurs a slight but acceptable runtime overhead when hosting multiple virtual GPU instances.
Mochi Xue, Jiacheng Ma 0001, Wentai Li, Yaozu Dong, Zhengwei Qi, Bingsheng He, Haibing Guan
IEEE Trans. Parallel Distributed Syst.5
2017 MobiXen: Porting Xen on Android devices for mobile virtualization
abstract
The mobile virtualization technology provides a feasible way to improve the manageability and security for embedded systems. This paper presents an architecture named MobiXen to address these challenges. In the MobiXen, both Xen's physical memory space and virtual address space are shrunk as much as possible and thus Android owns more memory resource; optimizations are developed to reduce the virtualization overhead when Android is accessing system resources; new policies are implemented to achieve low suspend/resume latency. With these work adopted, MobiXen is customized as a high efficient mobile hypervisor. Detailed implementations shows that, most of the performance degradation brought by MobiXen is less than 3%, which is imperceptible by end users.
Yaozu Dong, Jianguo Yao 0002, Haibing Guan, R. Ananth Krishna, Yunhong Jiang
DATE1
2016 gScale: Scaling up GPU Virtualization with Dynamic Sharing of Graphics Memory Space
Mochi Xue, Yaozu Dong, Jiacheng Ma 0001, Zhengwei Qi, Bingsheng He, Haibing Guan
USENIX ATC3
2015 Boosting GPU Virtualization Performance with Hybrid Shadow Page Tables
Yaozu Dong, Mochi Xue, Zhengwei Qi, Haibing Guan
USENIX ATC1
2014 A Full GPU Virtualization Solution with Mediated Pass-Through
Yaozu Dong, David Cowperthwaite
USENIX ATC2
2014 Multi-Granularity Memory Mirroring via Binary Translation in Cloud Environments
abstract
As the size of DRAM memory grows in clusters, memory errors are common. Current memory availability strategies mostly focus on memory backup and error recovery. Hardware solutions like mirror memory needs costly peripheral equipments while existing software approaches reduce the expense but are limited by the high overhead in practical usage. Moreover, in cloud environments, containers such as LXC now can be used as process and application-level virtualization to run multiple isolated systems on a single host. In this paper, we present a novel system called Memvisor to provide high availability memory mirroring. It is a software approach achieving flexible multi-granularity memory mirroring based on virtualization and binary translation. We can flexibly set memory areas to be mirrored or not from process level to the whole user mode applications. Then, all memory write instructions are duplicated. Data written to memory are synchronized to backup space in the instruction level. If memory failures happen, Memvisor will recover the data from the backup space. Compared with traditional software approaches, the instruction level synchronization lowers the probability of data loss and reduces the backup overhead. The results show that Memvisor outperforms the state-of-the-art software approaches even in the worst case.
Zhengwei Qi, Haoliang Dong, Yaozu Dong, Haibing Guan
IEEE Trans. Netw. Serv. Manag.4
2014 HYVI: A HYbrid VIrtualization Solution Balancing Performance and Manageability
abstract
Virtualization is a building block technology of cloud computing. Software virtualization solutions, such as paravirtualized Memory Management Unit and paravirtualized I/O, provide flexible manageability, such as virtual machine migration and IP-based network packet filtering. However, it suffers from performance and scalability issues. Hardware virtualization and its advanced accelerations, such as Extended Page Table and Single Root I/O Virtualization, were introduced later to simplify virtualization implementation and improve performance. However, this solution may suffer from a manageability issue. In this paper, we describe our work on the implementation of hardware virtualization and optimizations of the advanced hardware acceleration support in Xen to improve the virtualization performance and scalability. We further propose HYVI, a hybrid virtualization solution combining the advantages of software virtualization (manageability) and hardware virtualization (performance and scalability), to address performance and manageability issues. Finally, we present performance evaluations and characterizations of these hardware accelerations and hybrid virtualization, using both OS benchmarks and a server consolidation benchmark (vConsolidate). The results show that optimized hardware acceleration can achieve an up to 77 percent and 33 percent performance improvement over software full virtualization and paravirtualization in the server consolidation benchmark, respectively. Hybrid virtualization achieves an up to 1.56 × performance, with 0.69 lower CPU core usage of paravirtualization, and maintains flexible manageability.
Yaozu Dong, Jinquan Dai, Haibing Guan
IEEE Trans. Parallel Distributed Syst.1
2013 COLO: COarse-grained LOck-stepping virtual machines for non-stop service
abstract
Virtual machine (VM) replication provides a software solution of for business continuity and disaster recovery through application-agnostic hardware fault tolerance by replicating the state of primary VM (PVM) to secondary VM (SVM) on a different physical node. Unfortunately, current VM replication approaches suffer from excessive overhead, which severely limit their applicability and suitability. In this paper, we leverage the practical effect of networked server-client system that PVM and SVM are considered as in the same state only if they can generate the same response from the clients' point of view, and this is exploited to optimize performance. To this end, we propose a generic and highly efficient non-stop service solution, named as "COLO" (COarse-grained LOck-stepping virtual machine) utilizing on-demand VM replication. COLO monitors the output responses of the PVM and SVM, and rules the SVM as a valid replica of the PVM according to the output similarity between PVM and SVM. If the responses do not match, the commit of network response is withheld until PVM's state has been synchronized to SVM. Hence, we ensure that the system is always capable of failover by SVM. Although non-determinism may mean a different internal state of SVM from that of the PVM, it is equally valid and remains consistent from external observations. Unlike earlier instruction level lock-stepping deterministic execution approaches, COLO can easily support Multi-Processors (MP) involving workloads with the satisfying performance. Results show that COLO significantly outperforms existing approaches, particularly on server-client workloads such as online databases and web server applications.
Yaozu Dong, Yunhong Jiang, Ian Pratt 0001, Shiqing Ma, Jian Li 0021, Haibing Guan
SoCC1
2013 kMemvisor: flexible system wide memory mirroring in virtual environments
Bin Wang 0062, Zhengwei Qi, Haibing Guan, Haoliang Dong, Yaozu Dong
HPDC6
2013 SR-IOV Based Network Interrupt-Free Virtualization with Event Based Polling
abstract
Along with the developments of networking and virtualization technologies, high speed network connections have become one of the key components in cloud computing and data-centers. Single-Root I/O Virtualization (SR-IOV) enhances the network throughput to the extent of becoming close to the line rate and achieving high scalability in the 10Gbps and higher network environments. However, the overhead of SR-IOV interrupt virtualization remains significant due to some additional trap-and-emulation overhead on the virtual interrupt controller. The higher the virtualization network connection is, the higher the interrupt frequency becomes through high bandwidth network. To mitigate this problem, we propose a smart Event-Based Polling model (sEBP), which leverages existing system events to trigger a regular packet polling such that network interrupts are eliminated from the critical I/O paths in the virtual environment. Due to the many varieties of system events, sEBP can deal with the network workload in a configurable and flexible manner. Based on a hierarchical virtualized environment, it can also be implemented either at the guest OS kernel level or at the Virtual Machine Manager (VMM) level. Since polling is much lighter than interrupt processing, sEBP significantly reduces the network processing overhead. The experimental results prove the efficiency of sEBP, which can achieve up to a 59% performance improvement and a 23% improved scalability ratio.
Haibing Guan, Yaozu Dong, Jian Li 0021
IEEE J. Sel. Areas Commun.2
2013 Performance Enhancement for Network I/O Virtualization with Efficient Interrupt Coalescing and Virtual Receive-Side Scaling
abstract
Virtualization is a key technology in cloud computing; it can accommodate numerous guest VMs to provide transparent services, such as live migration, high availability, and rapid checkpointing. Cloud computing using virtualization allows workloads to be deployed and scaled quickly through the rapid provisioning of virtual machines on physical machines. However, I/O virtualization, particularly for networking, suffers from significant performance degradation in the presence of high-speed networking connections. In this paper, we first analyze performance challenges in network I/O virtualization and identify two problems-conventional network I/O virtualization suffers from excessive virtual interrupts to guest VMs, and the back-end driver does not efficiently use the computing resources of underlying multicore processors. To address these challenges, we propose optimization methods for enhancing the networking performance: 1) Efficient interrupt coalescing for network I/O virtualization and 2) virtual receive-side scaling to effectively leverage multicore processors. These methods are implemented and evaluated with extensive performance tests on a Xen virtualization platform. Our experimental results confirm that the proposed optimizations can significantly improve network I/O virtualization performance and effectively solve the performance challenges.
Haibing Guan, Yaozu Dong, Ruhui Ma, Dongxiao Xu, Jian Li 0021
IEEE Trans. Parallel Distributed Syst.2
2012 Memvisor: Application Level Memory Mirroring via Binary Translation
abstract
Memory failures are common in clusters, and their destructive effects (e.g., increasing downtime and losing data) make users suffer great loss. Current memory availability strategies mostly require extra expensive hardware. Software approaches based on check pointing technologies intend to reduce the expense, but their high overhead limits the practical usage. In this paper, we present a novel system called Memvisor to provide software mirrored memory for applications. Specifically, all memory write instructions are duplicated. Data written to memory are synchronized to backup space. If memory failures happen, Memvisor will recover the data from the backup space. Compared with traditional software approaches, the instruction-level synchronization lowers the probability of data loss and reduces backup overhead. The results show that even in the worst case, Memvisor outperforms the state-of-the-art software approaches.
Haoliang Dong, Bin Wang 0062, Haiyang Sun 0003, Zhengwei Qi, Haibing Guan, Yaozu Dong
CLUSTER7
2012 Built-in Device Simulator for OS Performance Evaluation
abstract
I/O devices are evolving rapidly, while OS optimization is always slower because of its dependence on physical devices. This inevitably prevents latest devices from working with their rating performance, which remains a big problem for performance-critical applications. Though I/O device simulators can help carry out performance evaluation before physical devices are ready, the existing simulator implementations are still unsatisfactory, either having too big overhead or requiring too much extra work. In this paper, we propose kernel built-in device simulation to provide accurate real time evaluations with acceptable extra effort. With the work of simulation well isolated, the overhead is reasonable compared to native environment. A bonding Ethernet interface is implemented in this way and experiments on it confirm the close-to-native performance of the idea.
Junjie Mao, Yu Chen 0004, Yaozu Dong
CLUSTER3
2012 sEBP: Event Based Polling for Efficient I/O Virtualization
abstract
Interrupt virtualization remains a key overhead source in high performance network virtualization (Single-root I/O virtualization or SR-IOV). SR-IOV can give close to line rate network bandwidth and good scalability in the 10 Gbps network environment, however the overhead of the interrupt virtualization in SR-IOV remains non-trivial, due to additional trap-and emulation overhead on the virtual interrupt controller, and high interrupt frequency brought by the high bandwidth network. In this paper we propose sEBP, an event-based polling model to eliminate the interrupts from the critical I/O paths in the virtual environment. A variety of system events are collected by sEBP, either at the guest kernel level or at the VMM level. Upon those events the NIC status is polled. The polling is lightweight, and plenty of system events fulfill the role of the interrupts. By removing the overhead of the interrupts, sEBP manages to achieve up to 59% performance improvement and 23% better scalability ratio.
Yaozu Dong, Xiang Mi, Haibing Guan
CLUSTER2
2012 ANOLE: A Profiling-Driven Adaptive Lock Waiter Detection Scheme for Efficient MP-guest Scheduling
abstract
In today's data center, there is a growing virtualization evolving trend to consolidate multiple servers into a single physical system. New architecture design that couples more and more cores into one processor furthers this trend. However, virtualization also poses new challenges such as lock holder preemption. In this work, we first demonstrate that lock holder preemption could bring dramatic performance degradation in virtualization environment. Then we propose ANOLE, a runtime adaptive lock waiter detection approach for lock holder preemption overhead reduction of MP guests. It leverages the modern hardware feature without any modification to spin lock implementation. ANOLE implements a hyper visor framework to preempt virtual CPUs adaptively and a user agent for guest spin lock profiling on KVM. We present in-depth performance evaluation under different scenarios, covering simple OS workloads, SPECvirt, and windows guest workloads. The experiment results demonstrate the solid performance benefit of ANOLE, it brings up to 50% performance improvement under different usage scenarios, with better lock waiter detection and load balance.
Yaozu Dong, Jiangang Duan
CLUSTER2
2012 CAR: Securing PCM Main Memory System with Cache Address Remapping
abstract
Phase Change Memory (PCM) has emerged as a promising alternative of DRAM to provide energy-efficient and high-capacity memory for high performance servers. A new DRAM + PCM hybrid memory architecture has been proposed to leverage PCM's high density and DRAM's robustness and performance. One of the big challenges of PCM is its limited write endurance (107~ 108times per cell). By knowing the association between DRAM and PCM, malicious software can easily force DRAM cache to be flushed continuously, which produces writes to certain PCM cells repeatedly (known as selective attack) and wears out PCM. Although existing wear-leveling approaches could evenly distribute writes under selective attack, the overall endurance of PCM is still severely impacted, and therefore it is suboptimal. In this paper, we propose Cache Address Remapping (CAR), that can adaptively remap DRAM cache address, to hide the association between DRAM and PCM. Moreover, CAR can minimize the write-back traffic to PCM under selective attack by uniformly distributing the writes to a single cache set into different cache sets. We propose a practical and low overhead implementation of CAR, called RanCAR. Experimental results show that CAR could reduce DRAM cache miss rate by ~4600x under selective attack, and prolong PCM lifetime from several minutes to 13.8 years on average.
Gang Wu 0008, Huxing Zhang, Yaozu Dong, Jingtong Hu
ICPADS3
2012 Lock-Visor: An Efficient Transitory Co-scheduling for MP Guest
abstract
Multiprocessor (MP) virtual machines (VMs) are widely used in cloud environments. However, MP VMs suffer from lock holder preemption (LHP) issue. This causes a tremendous waste of CPU cycles, leading to deteriorated synchronization latency and a significant degradation in system performance. Previous works have addressed the problem with software co-scheduling or lock waiter yielding. However, co-scheduling suffers from CPU utility fragmentation, priority inversion and loss of the flexibility of hyper visor scheduler, which causes inefficiency in CPU usage. Lock waiter yielding, another solution, suffers from a large impact on hyper visor scheduler and issues with response latency. In this paper, we propose Lock-visor, an efficient transitory co-scheduling algorithm, to bypass the guest spin lock loop effectively. Our protocol has little to no impact on the flexibility of hyper visor scheduler, and achieves better system performance. Multiple policies are explored on top of transitory co-scheduling to maximize the efficiency of Lock-visor, i.e. instant transitory, selective instant transitory and deferred transitory co-scheduling. Comprehensive experiments are conducted using CPU-intensive, I/O-intensive and lock-intensive workloads. Our experimental results show that Lock-visor can significantly improve system performance (e.g. Lock-visor has up to 341.3% performance advantage over original Linux kernel 2.6.38 in Sys Bench 4-VM case), while at the same time improve system latency with little to no effect on scheduling fairness.
Lei Zhang 0060, Yu Chen 0004, Yaozu Dong
ICPP3
2012 Virtualization challenges: a view from server consolidation perspective
abstract
Server consolidation, by running multiple virtual machines on top of a single platform with virtualization, provides an efficient solu-tion to parallelism and utilization of modern multi-core processors system. However, the performance and scalability of server con-solidation solution on modern massive advanced server is not well addressed. In this paper, we conduct a comprehensive study of Xen per-formance and scalability characterization running SPECvirt_sc2010, and identify that large memory and cache footprint, due to the unnecessary high frequent context switch, introduce additional challenges to the system performance and scalability. We propose two optimizations (dynamically-allocable tasklets and context-switch rate controller) to improve the performance. The results show the improved memory and cache efficiency with a reduction of the overall CPI, resulting in an improvement of server consolidation capability by 15% in SPECvirt_sc2010. In the meantime, our optimization achieves an up to 50% acceleration of service response, which greatly improves the QoS of Xen virtualization solution.
Hui Lu 0001, Yaozu Dong, Jiangang Duan, Kevin Tian
VEE2
2012 CompSC: live migration with pass-through devices
abstract
Live migration is one of the most important features of virtualization technology. With regard to recent virtualization techniques, performance of network I/O is critical. Current network I/O virtualization (e.g. Para-virtualized I/O, VMDq) has a significant performance gap with native network I/O. Pass-through network devices have near native performance, however, they have thus far prevented live migration. No existing methods solve the problem of live migration with pass-through devices perfectly.
Zhenhao Pan, Yaozu Dong, Yu Chen 0004, Lei Zhang 0060, Zhijiao Zhang
VEE2
2012 High performance network virtualization with SR-IOV
Yaozu Dong, Guangdeng Liao, Haibing Guan
J. Parallel Distributed Comput.1
2012 Optimizing virtual machines using hybrid virtualization
Qian Lin 0002, Zhengwei Qi, Jiewei Wu, Yaozu Dong, Haibing Guan
J. Syst. Softw.4
2012 ReNIC: Architectural extension to SR-IOV I/O virtualization for efficient replication
abstract
Virtualization is gaining popularity in cloud computing and has become the key enabling technology in cloud infrastructure. By replicating the virtual server state to multiple independent platforms, virtualization improves the reliability and availability of cloud systems. Unfortunately, existing Virtual Machine (VM) replication solutions were designed only for software virtualized I/O, which suffers from large performance and scalability overheads. Although hardware-assisted I/O virtualization (such as SR-IOV) can achieve close to native performance and very good scalability, they cannot be properly replicated across different physical machines due to architectural limitations (such as lack of efficient device state read/write, buffering outbound packets, etc.). In this paper, we address those architectural limitations, by proposing ReNIC, an architectural extension to SR-IOV I/O virtualization for efficient I/O replications. We have extended Xen hypervisor and the Remus rapid checkpoint solution to support this new architectural extension. We developed a system simulator on multi-core systems to extensively evaluate ReNIC. The experimental results demonstrate that ReNIC achieves up to 54% CPU usage reduction, compared to software based I/O virtualization at runtime, and up to 16.2% performance advantage over software based I/O virtualization in rapid checkpoint. During migration, ReNIC reduces service shutdown time by about 50%, compared to device emulation and paravirtualized I/O, and over 71% compared to teaming driver.
Yaozu Dong, Zhenhao Pan, Jinquan Dai, Yunhong Jiang
ACM Trans. Archit. Code Optim.1
2011 Optimizing Network I/O Virtualization with Efficient Interrupt Coalescing and Virtual Receive Side Scaling
abstract
Virtualization is a fundamental component in cloud computing because it provides numerous guest VM transparent services, such as live migration, high availability, rapid checkpoint, etc. However, I/O virtualization, particularly for network, is still suffering from significant performance degradation. In this paper, we analyze performance challenges in network I/O virtualization and observe that the conventional network I/O virtualization incurs excessive virtual interrupts to guest VMs, and the backend driver in the driver domain is not parallelized and cannot leverage underlying multi-core processors. Motivated by the above observations, we propose optimizations: efficient interrupt coalescing for network I/O virtualization and virtual receive side scaling to effectively leverage multi-core processors. We implemented those optimizations in Xen and did extensive performance evaluation. Our experimental results reveal that the proposed optimizations significantly improve network I/O virtualization performance and effectively tackle the performance challenges.
Yaozu Dong, Dongxiao Xu, Guangdeng Liao
CLUSTER1
2011 Improving PCM Endurance with Randomized Address Remapping in Hybrid Memory System
abstract
Phase-Change-Memory (PCM) has emerged as a promising alternative of DRAM main memory. A new hybrid memory architecture, where DRAM serves as cache of PCM main memory, has been proposed to leverage PCM's high scalability and DRAM's fast access time. One biggest issue of PCM is the limited number of writes to storage cells. We argue that good cache mechanism will decrease PCM writes dramatically in hybrid memory system. In this paper, we demonstrate that traditional set associative cache is susceptible to malicious attacks, which lead certain PCM cells to wear-out by constant cache flushes. A novel approach called Randomized Address Remapping (RAR) is proposed to hide the mapping details between DRAM and PCM. With this approach, the attacks based on set associative cache do not work, while the efficiency of caching still remains. We present Static Randomized Address Remapping (SRAR) and Dynamic Randomized Address Remapping (DRAR) in this paper. SRAR invalidates set associative cache based attacks by distributing their address accesses to different sets. DRAR uses a region-based approach to change the mapping dynamically, in case that the static mapping relationship is discovered by attacker compromising operating system. Experimental results show that RAR approaches can prevent malicious attacks and improve PCM endurance greatly.
Gang Wu 0008, Huxing Zhang, Yaozu Dong
CLUSTER4
2010 High performance network virtualization with SR-IOV
abstract
Virtualization poses new challenges to I/O performance. The single-root I/O virtualization (SR-IOV) standard allows an I/O device to be shared by multiple Virtual Machines (VMs), without losing runtime performance. We propose a generic virtualization architecture for SR-IOV devices, which can be implemented on multiple Virtual Machine Monitors (VMMs). With the support of our architecture, the SR-IOV device driver is highly portable and agnostic of underlying VMM. Based on our first implementation of network device driver, we applied several optimizations to reduce virtualization overhead. Then, we carried out comprehensive experiments to evaluate SR-IOV performance and compare it with paravirtualized network driver. The results show SR-IOV can achieve line rate (9.48 Gbps) and scale network up to 60 VMs at the cost of only 1.76% additional CPU overhead per VM, without sacrificing throughput. It has better throughout, scalability, and lower CPU utilization than paravirtualization.
Yaozu Dong, Haibing Guan
HPCA1
2009 Towards high-quality I/O virtualization
abstract
High-quality I/O virtualization (that is, complete device semantics, full-feature set, close-to-native performance and real-time response) is critical to both server and client virtualizations. Existing solutions for I/O virtualization (e.g., full device emulation, paravirtualization and direct I/O) cannot meet the requirements of high-quality I/O virtualization due to high overheads, lack of complete semantic or full-feature set support.
Yaozu Dong, Jinquan Dai, Zhiteng Huang, Haibing Guan, Kevin Tian, Yunhong Jiang
SYSTOR1