EDBT 2026 Demo / reviewers in the wild / expert
Wei Zhang 0052
dblp:10/4661-52
· DBLP profile ↗
23ranked-venue papers
7as first author
8since 2021 · last 2026
0009-0004-9512-4192ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Computer networks · 3 · 1 first-authorHuman-computer interaction and ubiquitous computing · 3 · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SG-IOV: Socket-Granular I/O Virtualization for SmartNIC-Based Container NetworksabstractI/O Virtualization (IOV) is a cornerstone of cloud computing, with container networking as a critical form of IOV in modern cloud paradigms. While container networks serve as feature-rich infrastructure, they incur a high CPU tax yet leave room for efficiency improvement. A natural idea is to offload container networks onto hardware such as SmartNICs via IOV interfaces. However, existing IOV mechanisms, such as SR-IOV, are misaligned with container requirements: limited device scalability versus high container density, packet-layer abstraction versus application-layer processing demands, and coarse-grained virtualization versus fine-grained container workloads. Chenxingyu Zhao, Jaehong Min, Shengkai Lin, Wei Zhang 0052, Kaiyuan Zhang 0001, Ming Liu 0027, Arvind Krishnamurthy |
ASPLOS (2) | 5 |
| 2026 | TrustWeave: Integrity Measurement and Attestation For Multi-Cloud LLMsabstractMulti-cloud deployed multi-agent systems powered by Large Language Models (LLMs) introduce critical security challenges. While Intel Trust Domain Extensions (TDX)-based Confidential Virtual Machines (CVMs) provide strong isolation and boot-time attestation, they lack dynamic runtime integrity verification, a capability essential for trusted agent systems that frequently load new models and coordinate across distributed services. We present TrustWeave, a runtime integrity measurement and attestation framework that extends the Linux Integrity Measurement Architecture (IMA) with support for Intel TDX's Runtime Measurement Registers (RTMRs). TrustWeave enables userspace attestation of dynamically loaded agent components throughout their life-cycle, providing runtime trust guarantees for secure and scalable LLM agent deployments. Jianchang Su, Kexin Chu, Youyou Lu, Wei Zhang 0052 |
EuroSys | 7 |
| 2025 | MCaM : Efficient LLM Inference with Multi-tier KV Cache ManagementabstractThe KV cache in current LLM serving system is primarily used to accelerate processing within a single request and is aggressively deleted once the response is generated. However, in scenarios like virtual assistants and multi-turn conversations, the KV cache can be reused across requests, which can dramatically reduce computation costs and improve serving latency. Caching historical tokens, however, significantly increases memory requirements. Furthermore, existing serving systems treat the request scheduler and KV cache separately, despite their tight coupling.MCaM is a multi-tier cache system that enables the KV cache reuse and sharing across requests. It leverages DRAM as slow- tier memory for storing the KV cache of historical prompts. To efficiently utilize fast-tier Hign Bandwidth Memory(HBM) on GPU, we co-designed the KV cache manager and scheduler to coordinate request scheduling and token placement across tiers. To hide the reload time, MCaM employs a pipeline prefetcher that overlaps communication and computation. Additionally, MCaM incorporates a quality-aware sparsification algorithm to heterogeneously compress the KV cache in each layer. This approach not only reduces data transfer size but also decreases the overall KV cache size. To remove data offloading from a request’s critical path, we designed an asynchronous offload engine that swaps data from HBM to DRAM in the background. Our experiments show that MCaM can reduce TTFT by up to 69% and improve prompt prefilling throughput by 3.3X. It can also reduce the end-to-end latency of LLM inference by up to 58% when request length increase to 4096 tokens. Kexin Chu, Zixu Shen, Sheng-Ru Cheng, Dawei Xiang, Ziqin Liu, Wei Zhang 0052 |
ICDCS | 6 |
| 2024 | PISeL: Pipelining DNN Inference for Serverless ComputingabstractServerless computing offers resource efficiency, cost efficiency, and a "pay-as-you-go" pricing model, which makes it highly attractive to both users and cloud providers. However, serverless computing faces serious cold start problem, especially for deep neural network (DNN) inference, which requires low latency. Existing cold start optimization focuses only on quick container start and fast runtime and library loading. However, DNN application bootstrap (DNN framework load and start, model initialization, model download, deserialization and copy) is the leading factor during the overall cold start time. As the model size grows, the application-level bootstrap becomes more severe. Masoud Rahimi Jafari, Jianchang Su, Oliver Wang, Wei Zhang 0052 |
CIKM | 5 |
| 2024 | FastMatch: Enhancing Data Pipeline Efficiency for Accelerated Distributed TrainingabstractTraining a deep learning model typically involves two interdependent stages: preprocessing input data on the CPU and training the model on the GPU. This process begins with the CPU performing resource-intensive I/O operations to load raw data, which it then processes and supplies to the GPU. This creates a typical provider-consumer relationship between the CPU and GPU. However, the inefficiency of I/O and CPU operations, in contrast to the high capabilities of GPUs, often turns preprocessing into a significant bottleneck. The imbalance between CPU and GPU resources leads to inefficiency, as the underutilized GPU cannot fully exploit its computational potential. To address this issue, we propose FastMatch, a system designed to selectively and precisely cache key data, thereby reducing the reliance on extensive I/O operations, and adjusting the CPU's performance to better align with the GPU's capabilities swiftly. FastMatch dynamically identifies important data and prompts the caching system to store it in fast DRAM. This cached data significantly enhances I/O performance while maintaining comparable accuracy levels. Additionally, FastMatch monitors and rectifies mismatches in CPU and GPU resources by tuning the number of preprocessing processes. This helps distribute preprocessing workloads more evenly, ensuring optimal utilization of both CPU and GPU resources. We implemented a prototype of FastMatch and successfully integrated it with TorchData. Our evaluations showed up to 2.2 × reduction in training time compared to vanilla PyTorch Dataloader. Jianchang Su, Masoud Rahimi Jafari, Wei Zhang 0052 |
ICCD | 4 |
| 2024 | Data-Enhanced Prediction with Decomposition and Amplitude-Aware Permutation Entropy in Distributed Computing SystemsabstractIn recent years, distributed computing has wit-nessed widespread applications across numerous organizations. Predicting workload and computing resource data can facilitate proactive service operation management, leading to substantial improvements in quality of service and cost efficiency. However, these data often exhibit non-linearity, high volatility, and inter-dependencies across different categories, presenting challenges for accurate forecasting. Consequently, there is a critical need to develop a method that thoroughly and comprehensively analyzes all available data to forecast future trends effectively. This work proposes a novel integrated data-enhanced prediction model named SVI for achieving high-accuracy workload prediction in distributed computing systems. SVI employs the Savitzky-Golay filter and variational mode decomposition for feature processing, whose features are subsequently utilized by Informer for multivariate joint analysis of the enhanced data, achieving high-precision prediction. Ablation and comparative experiments with advanced prediction models are conducted on the Google cluster trace and other typical datasets. Realistic data-driven results indicate that SVI improves the prediction accuracy by 35.4% compared to the original Informer, with each module contributing to the performance enhancement. Furthermore, compared with Autoformer, SVI enhances the prediction accuracy of workload, CPU, and memory by 62.5%, 65.6%, and 69.1 %, respectively. Haitao Yuan 0001, Qinglong Hu, Jing Bi 0001, Wei Zhang 0052, Jia Zhang 0001, MengChu Zhou |
SMC | 4 |
| 2023 | Accel-GCN: High-Performance GPU Accelerator Design for Graph Convolution NetworksabstractGraph Convolutional Networks (GCNs) are pivotal in extracting latent information from graph data across various domains, yet their acceleration on mainstream GPUs is challenged by workload imbalance and memory access irregularity. To address these challenges, we present Accel-GCN, a GPU accelerator architecture for GCNs. The design of Accel-GCN encompasses: (i) a lightweight degree sorting stage to group nodes with similar degree; (ii) a block-level partition strategy that dynamically adjusts warp workload sizes, enhancing shared memory locality and workload balance, and reducing metadata overhead compared to designs like GNNAdvisor; (iii) a combined warp strategy that improves memory coalescing and computational parallelism in the column dimension of dense matrices. Utilizing these principles, we formulate a kernel for SpMM in GCNs that employs block-level partitioning and combined warp strategy. This approach augments performance and multi-level memory efficiency and optimizes memory bandwidth by exploiting memory coalescing and alignment. Evaluation of Accel-GCN across 18 benchmark graphs reveals that it outperforms cuSPARSE, GNNAdvisor, and graph-BLAST by factors of 1.17×, 1.86×, and 2.94× respectively. The results underscore Accel-GCN as an effective solution for enhancing GCN computational efficiency. The implementation can be found on Github*. Hongwu Peng, Amit Hasan 0001, Shaoyi Huang, Haowen Fang, Wei Zhang 0052, Tong Geng, Omer Khan, Caiwen Ding |
ICCAD | 7 |
| 2021 | Binary Complex Neural Network Acceleration on FPGA : (Invited Paper)abstractBeing able to learn from complex data with phase information is imperative for many signal processing applications. Today’s real-valued deep neural networks (DNNs) have shown efficiency in latent information analysis but fall short when applied to the complex domain. Deep complex networks (DCN), in contrast, can learn from complex data, but have high computational costs; therefore, they cannot satisfy the instant decision-making requirements of many deployable systems dealing with short observations or short signal bursts. Recent, Binarized Complex Neural Network (BCNN), which integrates DCNs with binarized neural networks (BNN), shows great potential in classifying complex data in real-time. In this paper, we propose a structural pruning based accelerator of BCNN, which is able to provide more than 5000 frames/s inference throughput on edge devices. The high performance comes from both the algorithm and hardware sides. On the algorithm side, we conduct structural pruning to the original BCNN models and obtain 20 × pruning rates with negligible accuracy loss; on the hardware side, we propose a novel 2D convolution operation accelerator for the binary complex neural network. Experimental results show that the proposed design works with over 90% utilization and is able to achieve the inference throughput of 5882 frames/s and 4938 frames/s for complex NIN-Net and ResNet-18 using CIFAR-10 dataset and Alveo U280 Board. Hongwu Peng, Shanglin Zhou, Scott Weitze, Sahidul Islam, Tong Geng, Ang Li 0006, Wei Zhang 0052, Minghu Song, Mimi Xie, Hang Liu 0001, Caiwen Ding |
ASAP | 8 |
| 2020 | Fine-grained Task Scheduling in Cloud Data Centers Using Simulated-annealing-based Bees AlgorithmabstractCloud computing is increasingly implemented by a growing number of organizations in recent years. Their critical business applications are deployed in distributed cloud data centers (CDCs) for fast response and low cost. The ever-increasing consumption of energy makes it highly important to schedule tasks efficiently in CDCs. In addition, many factors in CDCs, e.g., the wind and solar energy and prices of power grid have spatial differences. It becomes a challenging problem of how to achieve the energy cost minimization for CDCs in such a market. This work applies a G/G/1 queuing system to evaluate the optimization of servers in each CDC. Furthermore, a single-objective constrained optimization problem is given and addressed by a proposed Simulated-annealing-based Bees Algorithm to yield a close-to-optimal solution. Based on it, a Fine-grained Task Scheduling (FTS) algorithm is designed to minimize the energy cost of CDCs by intelligently scheduling heterogeneous tasks among distributed CDCs. In addition, it also determines running speeds of servers and the number of switched-on servers in each CDC while strictly meeting tasks' delay bounds. Realistic data-driven results demonstrate that FTS outperforms its typical benchmark scheduling peers in terms of energy cost and throughput. Haitao Yuan 0001, Jing Bi 0001, MengChu Zhou, Jia Zhang 0001, Wei Zhang 0052 |
SMC | 5 |
| 2020 | Profit-Maximized Task Offloading with Simulated-annealing-based Migrating Birds Optimization in Hybrid Cloud-Edge SystemsabstractAs an emerging framework, edge computing achieves Internet of Things by providing computing, storage and network resources. It moves computation to edge devices located near users. Nevertheless, nodes in the edge often own limited resources and constrained energy capacities. It is impossible to entirely execute tasks in the edge due to their unsatisfied quality of service. Cloud data centers (CDCs) own almost unlimited resources yet they might cause large transmission delay and high resource cost. Consequently, it is highly needed to intelligently offload tasks between CDC and edge. This work proposes a task offloading algorithm for hybrid cloud-edge systems to achieve profit maximization of a system provider with response time bound assurance. It comprehensively investigates CPU, memory and bandwidth limits of nodes in the edge, and constraints of available energy and servers in CDC. These factors are integrated into a single-objective constrained optimization problem, which is solved by a simulated-annealing-based migrating birds optimization algorithm to yield a close-to-optimal offloading policy between CDC and the edge. Real-life data-driven experimental results show that its profit outperforms its four typical peers. Haitao Yuan 0001, Jing Bi 0001, MengChu Zhou, Jia Zhang 0001, Wei Zhang 0052 |
SMC | 5 |
| 2020 | NFVnice: Dynamic Backpressure and Scheduling for NFV Service ChainsabstractManaging Network Function (NF) service chains requires careful system resource management. We propose NFVnice, a user space NF scheduling and service chain management framework to provide fair, efficient and dynamic resource scheduling capabilities on Network Function Virtualization (NFV) platforms. The NFVnice framework monitors load on a service chain at high frequency (1000Hz) and employs backpressure to shed load early in the service chain, thereby preventing wasted work. Borrowing concepts such as rate proportional scheduling from hardware packet schedulers, CPU shares are computed by accounting for heterogeneous packet processing costs of NFs, I/O, and traffic arrival characteristics. By leveraging cgroups, a user space process scheduling abstraction exposed by the operating system, NFVnice is capable of controlling when network functions should be scheduled. NFVnice improves NF performance by complementing the capabilities of the OS scheduler but without requiring changes to the OS's scheduling mechanisms. Our controlled experiments show that NFVnice provides the appropriate rate-cost proportional fair share of CPU to NFs and significantly improves NF performance (throughput and latency) by reducing wasted work across an NF chain, compared to using the default OS scheduler. NFVnice achieves this even for heterogeneous NFs with vastly different computational costs and for heterogeneous workloads. Sameer G. Kulkarni, Wei Zhang 0052, Jinho Hwang, Shriram Rajagopalan, K. K. Ramakrishnan, Timothy Wood 0001, Mayutan Arumaithurai, Xiaoming Fu 0001 |
IEEE/ACM Trans. Netw. | 2 |
| 2018 | Memory-Oriented Distributed Computing at Rack ScaleabstractNo abstract available. Haris Volos 0001, Kimberly Keeton, Milind Chabbi, Se Kwon Lee, Mark Lillibridge, Yuvraj Patel, Wei Zhang 0052 |
SoCC | 8 |
| 2017 | NFVnice: Dynamic Backpressure and Scheduling for NFV Service ChainsabstractManaging Network Function (NF) service chains requires careful system resource management. We propose NFVnice, a user space NF scheduling and service chain management framework to provide fair, efficient and dynamic resource scheduling capabilities on Network Function Virtualization (NFV) platforms. The NFVnice framework monitors load on a service chain at high frequency (1000Hz) and employs backpressure to shed load early in the service chain, thereby preventing wasted work. Borrowing concepts such as rate proportional scheduling from hardware packet schedulers, CPU shares are computed by accounting for heterogeneous packet processing costs of NFs, I/O, and traffic arrival characteristics. By leveraging cgroups, a user space process scheduling abstraction exposed by the operating system, NFVnice is capable of controlling when network functions should be scheduled. NFVnice improves NF performance by complementing the capabilities of the OS scheduler but without requiring changes to the OS's scheduling mechanisms. Our controlled experiments show that NFVnice provides the appropriate rate-cost proportional fair share of CPU to NFs and significantly improves NF performance (throughput and loss) by reducing wasted work across an NF chain, compared to using the default OS scheduler. NFVnice achieves this even for heterogeneous NFs with vastly different computational costs and for heterogeneous workloads. Sameer G. Kulkarni, Wei Zhang 0052, Jinho Hwang, Shriram Rajagopalan, K. K. Ramakrishnan, Timothy Wood 0001, Mayutan Arumaithurai, Xiaoming Fu 0001 |
SIGCOMM | 2 |
| 2016 | Flurries: Countless Fine-Grained NFs for Flexible Per-Flow CustomizationabstractThe combination of Network Function Virtualization (NFV) and Software Defined Networking (SDN) allows flows to be flexibly steered through efficient processing pipelines. As deployment of NFV becomes more prevalent, the need to provide fine-grained customization of service chains and flow-level performance guarantees will increase, even as the diversity of Network Functions (NFs) rises. Existing NFV approaches typically route wide classes of traffic through pre-configured service chains. While this aggregation improves efficiency, it prevents flexibly steering and managing performance of flows at a fine granularity. Wei Zhang 0052, Jinho Hwang, Shriram Rajagopalan, K. K. Ramakrishnan, Timothy Wood 0001 |
CoNEXT | 1 |
| 2016 | Multi-cache: Dynamic, Efficient Partitioning for Multi-tier Caches in Consolidated VM EnvironmentsabstractEvery physical machine in today's typical datacenter is backed by storage devices with hundreds of Gigabytes to Terabytes in size. Data center vendors usually use hard disk drives for their back-end storage as it is cheap and reliable. However, the increase in the I/O accesses to the back-end storage from one or many of the VMs hosted on a physical machine can reduce its overall accesses time significantly due to contention. This may not be suitable for interactive applications requiring low latency that might be co-located with other I/O intensive applications. In this paper we present Multi-Cache, a multi-layer cache management system that uses a combination of cache devices of varied speed and cost such as solid state drives, non-volatile memories, etc to mitigate this problem. Multi-Cache partitions each device dynamically at runtime according to the workload of each VM and its priority. We use a heuristic optimization technique that ensures maximum utilization of the caches resulting in a high hit rate. We use a weighted partitioning policy that improves latency by up to 72% for individual workloads, and a overall hit rate increase of up to 31% for host running several workloads together in comparison to standard LRU caching algorithms. Sundaresan Rajasekaran, Shaohua Duan, Wei Zhang 0052, Timothy Wood 0001 |
IC2E | 3 |
| 2016 | OpenNetVM: Flexible, high performance NFV (Demo)abstractNetwork Function Virtualization promises to enable dynamic management of software-based network functions. We envision a dynamic and flexible network that can support a smarter data plane than just simple switches that forward packets. This network architecture supports complex stateful routing of flows where processing by network functions (NFs) can transform packet data, customized on a per-flow basis, as it moves between end points. This demo will present OpenNetVM, a highly efficient packet processing framework that greatly simplifies the development of network functions, as well as their management and optimization. OpenNetVM runs network functions in lightweight Docker containers that start in less than a second. The OpenNetVM platform manager provides load balancing, flexible flow management, and service name abstractions. OpenNetVM uses DPDK for high performance I/O, and efficiently routes packets through dynamically created service chains. We will demonstrate how the research community can easily build new network functions and rapidly deploy them to see their effectiveness in high performance network environments. Wei Zhang 0052, Guyue Liu, Phil Lopreiato, Grégoire Todeschi, K. K. Ramakrishnan, Timothy Wood 0001 |
LANMAN | 1 |
| 2016 | SDNFV: Flexible and Dynamic Software Defined Control of an Application- and Flow-Aware Data Plane
Wei Zhang 0052, Guyue Liu, Ali Mohammadkhan, Jinho Hwang, K. K. Ramakrishnan, Timothy Wood 0001 |
Middleware | 1 |
| 2015 | Virtual function placement and traffic steering in flexible and dynamic software defined networksabstractThe integration of network function virtualization (NFV) and software defined networks (SDN) seeks to create a more flexible and dynamic software-based network environment. The line between entities involved in forwarding and those involved in more complex middle box functionality in the network is blurred by the use of high-performance virtualized platforms capable of performing these functions. A key problem is how and where network functions should be placed in the network and how traffic is routed through them. An efficient placement and appropriate routing increases system capacity while also minimizing the delay seen by flows. In this paper, we formulate the problem of network function placement and routing as a mixed integer linear programming problem. This formulation not only determines the placement of services and routing of the flows, but also seeks to minimize the resource utilization. We develop heuristics to solve the problem incrementally, allowing us to support a large number of flows and to solve the problem for incoming flows without impacting existing flows. Ali Mohammadkhan, Sheida Ghapani, Guyue Liu, Wei Zhang 0052, K. K. Ramakrishnan, Timothy Wood 0001 |
LANMAN | 4 |
| 2014 | UniCache: Hypervisor Managed Data Storage in RAM and FlashabstractApplication and OS-level caches are crucial for hiding I/O latency and improving application performance. However, caches are designed to greedily consume memory, which can cause memory-hogging problems in a virtualized data centers since the hypervisor cannot tell for what a virtual machine uses its memory. A group of virtual machines may contain a wide range of caches: database query pools, memcached key-value stores, disk caches, etc., each of which would like as much memory as possible. The relative importance of these caches can vary significantly, yet system administrators currently have no easy way to dynamically manage the resources assigned to a range of virtual machine data caches in a unified way. To improve this situation, we have developed UniCache, a system that provides a hypervisor managed volatile data store that can cache data either in hypervisor controlled main memory (hot data) or on Flash based storage (cold data). We propose a two-level cache management system that uses a combination of recency information, object size, and a prediction of the cost to recover an object to guide its eviction algorithm. We have built a prototype of UniCache using Xen, and have evaluated its effectiveness in a shared environment where multiple virtual machines compete for storage resources. Jinho Hwang, Wei Zhang 0052, Ron Chi-Lung Chiang, Timothy Wood 0001, H. Howie Huang |
IEEE CLOUD | 2 |
| 2014 | MIMP: Deadline and Interference Aware Scheduling of Hadoop Virtual MachinesabstractVirtualization promised to dramatically increase server utilization levels, yet many data centers are still only lightly loaded. In some ways, big data applications are an ideal fit for using this residual capacity to perform meaningful work, but the high level of interference between interactive and batch processing workloads currently prevents this from being a practical solution in virtualized environments. Further, the variable nature of spare capacity may make it difficult to meet big data application deadlines. In this work we propose two schedulers: one in the virtualization layer designed to minimize interference on high priority interactive services, and one in the Hadoop framework that helps batch processing jobs meet their own performance deadlines. Our approach uses performance models to match Hadoop tasks to the servers that will benefit them the most, and deadline-aware scheduling to effectively order incoming jobs. The combination of these schedulers allows data center administrators to safely mix resource intensive Hadoop jobs with latency sensitive web applications, and still achieve predictable performance for both. We have implemented our system using Xen and Hadoop, and our evaluation shows that our schedulers allow a mixed cluster to reduce web response times by more than ten fold, while meeting more Hadoop deadlines and lowering total task execution times by 6.5%. Wei Zhang 0052, Sundaresan Rajasekaran, Timothy Wood 0001, Mingfa Zhu |
CCGRID | 1 |
| 2012 | LVMCI: Efficient and Effective VM Live Migration Selection Scheme in Virtualized Data CentersabstractVirtualization can provide significant benefits in virtualized data centers by enabling efficient and effective live migration to ensure service level agreement(SLA). Most of existing studies make decision on which bad virtual machines (VMs) should be migrated to which appropriate physical machines (PMs) in terms of resource utilizations. However, migration actions may degrade migrated application performance due to extra CPU and bandwidth consumptions. Furthermore, negative performance interferences amongst applications scheduled to the same PM may arise given the poor performance isolations of VMs on a PM. We design and implement a VM migration selection system with less migration costs and application performance interferences, called LVMCI (Live Virtual machine Migration with less Costs and application Interference). We propose a migration cost evaluation model to analyze quantitatively the aspects (i.e. throughput and response latency) of application performance degradation. Dirty rate and frequent dirty rate are two key factors that affect iteration time and downtime. We implement a tool that measures these parameters before VMs are migrated. We distinguish the performance degradation of migrated applications caused by memory iteration phase and stop-and-copy phase, which helps to select VM migrated. Besides that, we propose a performance interference model which helps to select the destination PM. The experimental results show that our system can estimate memory iteration time and downtime with high accuracy, and ensures a high level of SLAs by minimizing performance degradation during migration process and performance interference among co-located VMs at the destination PM. Wei Zhang 0052, Mingfa Zhu, Yiduo Mei, Yunwei Gao, Yuzhong Sun |
ICPADS | 1 |
| 2012 | Autonomic Resource Allocation in Virtualized Data CentersabstractVirtualization has been widely adopted in data centers for improving efficiency and flexibility. Multiple applications are co-hosted in virtualized data centers. In order to meet the Service Level Agreements (SLA), how to allocate resources for multiple applications is an important and challenging task, especially when dealing with fluctuating workloads and complex server applications. Virtual Machine Monitor provides fine-grained resource allocation and live migration. In this paper, we develop RTCOIN-Qclouds, a response time-aware, cost-aware and interference-aware control framework that tunes resource allocation, which ensures a high level of meeting the SLAs. Every physical machine's resources are assigned to multiple virtual machines which run on it based on application's response time, rather than traditional methods based on resource utilization. Virtual machine migration allows data centers to rebalance workloads across physical machines. However, migration actions may lead to performance impact during the migration process. Current virtualization techniques do not provide effective performance isolation between virtual machines (VMs). Specially, hidden contention for physical resources impacts performance differently in different virtual machines. As to the problem of selecting which virtual machines to be migrated, we consider migration cost. As to the problem of selecting which physical machine to be placed, we consider performance interference. Furthermore, we experimentally validate the effectiveness of response time-aware resource allocation in our framework using microbenchmarks. Wei Zhang 0052, Mingfa Zhu, Qimeng Wu, Yuzhong Sun |
ISPA | 1 |
| 2012 | Performance Degradation-Aware Virtual Machine Live Migration in Virtualized ServersabstractLive migration of virtual machines(VMs) is widely used for system management in virtualized servers. When the loads increase and SLAs of some applications are violated, dynamic migration of virtual machines across physical machines (PMs) has the potential to ensure a high level of meeting the SLAs. Because of consuming extra CPU and bandwidth, application performance may be degraded during the migration process. However, different applications have different performance degradation. We design and implement a VM migration selection method that decides which VMs should be migrated. It can not only eliminate resouce competition on the PM, but also have less performance degradation during the migration process. We propose a performance degration-aware model to analyze applications' performance degradation which is directly sensitive to users. We analyze migration source code and find that memory size, dirty rate and frequent dirty rate are key factors that affect iteration time and downtime. We implement a tool that measures dirty rate and frequent dirty rate before VMs are migrated. we make a distinction between memory iteration phrase and stop-and-copy phrase owing to different performance degradation. The experimental results show that our method is effective. Wei Zhang 0052, Mingfa Zhu, Yiduo Mei, Yuzhong Sun |
PDCAT | 1 |