EDBT 2026 Demo / reviewers in the wild / expert
Tao Lu 0014
dblp:03/5189-14
· DBLP profile ↗
21ranked-venue papers
7as first author
10since 2021 · last 2026
0000-0003-2362-2446ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 18 · 6 first-author · 9 since 2021Computer networks · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CEMU: Enabling Full-System Emulation of Computational Storage Beyond Hardware LimitsabstractComputational storage drives (CSDs) present a promising approach to improve system performance through near data processing in SSDs. However, current research platforms are fragmented and inadequate to explore the full design space of CSD systems. Existing hardware and emulator platforms are constrained by physical compute resources, while simulators lack full-system fidelity. To address the problems, we introduce CEMU, a new software-based CSD emulation platform that enables full-system research. It consists of a CSD device emulator and a CSD-oriented software stack. Through a novel virtual machine freezing mechanism, CSD emulation achieves high configurability. While the CSD can utilize the host CPU to physically perform computation to preserve full-system behaviors, the computational delay can be modeled separately to emulate CSDs with CPU-unbounded high computing power. The software stack is designed with two principles, adhering to recent industry CSD standards and being compatible with the existing I/O stack, which is achieved via a newly developed file system FDMFS. We verify CEMU's emulation fidelity across a range of applications by benchmarking against actual CSD hardware, demonstrating average end-to-end performance accuracy of 95% or higher. We also use two case studies on large language model training and LevelDB to demonstrate that CEMU is effective in exploring CSD system research and can uncover insights that have not been discovered in previous research platforms. Jiapin Wang, You Zhou 0009, Kai Lu 0002, Jiguang Wan 0001, Fei Wu 0005, Tao Lu 0014 |
ASPLOS (2) | 8 |
| 2026 | ASIC-based Compression Accelerators for Storage Systems: Design, Placement, and Profiling Insights
Tao Lu 0014, Jiapin Wang, Yelin Shan, Xiang Chen 0028 |
EuroSys | 1 |
| 2026 | Computational Burst Buffers: Accelerating HPC I/O via In-Storage Compression OffloadingabstractBurst buffers (BBs) act as an intermediate storage layer between compute nodes and parallel file systems (PFS), effectively alleviating the I/O performance gap in high-performance computing (HPC). As scientific simulations and AI workloads generate larger checkpoints and analysis outputs, BB capacity shortages and PFS bandwidth bottlenecks are emerging, and CPU-based compression is not an effective solution due to its high overhead. We introduceComputational Burst Buffers(CBBs), a storage paradigm that embeds hardware compression engines such as application-specific integrated circuit (ASIC) inside computational storage drives (CSDs) at the BB tier. CBB transparently offloads both lossless and error-bounded lossy compression from CPUs to CSDs, thereby (i) expanding effective SSD-backed BB capacity, (ii) reducing BB–PFS traffic, and (iii) eliminating contention and energy overheads of CPU-based compression. Unlike prior CSD-based compression designs targeting databases or flash caching, CBB co-designs the burst-buffer layer and CSD hardware for HPC and quantitatively evaluates compression offload in BB–PFS hierarchies. We prototype CBB using a PCIe 5.0 CSD with an ASIC Zstd-like compressor and an FPGA prototype of an SZ entropy encoder, and evaluate CBB on a 16-node cluster. Experiments with four representative HPC applications and a large-scale workflow simulator show up to 61% lower application runtime, 8–12× higher cache hit ratios, and substantially reduced compute-node CPU utilization compared to software compression and conventional BBs. These results demonstrate that compression-aware BBs with CSDs provide a practical, scalable path to next-generation HPC storage. Xiang Chen 0028, Bing Lu 0001, Haoquan Long, Huizhang Luo, Yili Ma, Guangming Tan, Dingwen Tao, Fei Wu 0005, Tao Lu 0014 |
IEEE Trans. Parallel Distributed Syst. | 9 |
| 2025 | StreamCSD: SSD-Autonomous Stream Management via In-Storage Content LearningabstractWrite amplification (WA) from migrating valid pages during garbage collection (GC) degrades SSD performance and lifespan. Although stream management based on high-level software semantics reduces WA, existing solutions require host modifications, hindering their adoption. We introduce StreamCSD, an SSD-autonomous stream management approach using in-storage content learning, eliminating host-side changes. Leveraging compression ratios from embedded compressors in computational storage drives (CSDs), StreamCSD employs a streaming Kmeans algorithm to cost-efficiently cluster data into streams. Evaluations show that StreamCSD reduces WA from 1.7 to 1.06 under multimodal generative AI workloads, matching state-of-the-art methods with minimal impact on bandwidth. StreamCSD operates without host modifications, promoting broader adoption of multi-stream SSDs. Xiang Chen 0028, Yelin Shan, Jiapin Wang, Yunxin Huang, Yafei Yang, Tao Lu 0014, You Zhou 0009, Fei Wu 0005 |
DAC | 7 |
| 2025 | Garbage Collection Does Not Only Collect Garbage: Piggybacking-Style Defragmentation for Deduplicated Backup StorageabstractDeduplication is widely used in backup storage and reduces storage overhead by allowing backups to share common data chunks. However, it naturally disrupts the sequential layout of backup images, leading to fragmentation, which slows down backup restoration. Existing solutions to this issue often come with trade-offs, either reducing deduplication effectiveness or introducing significant I/O overhead. Dingbang Liu, Xiangyu Zou, Tao Lu 0014, Philip Shilane, Wen Xia, Yanqi Pan |
EuroSys | 3 |
| 2024 | Balloon-ZNS: Constructing High-Capacity and Low-Cost ZNS SSDs with Built-in CompressionabstractZNS SSDs are emerging storage devices promising low cost, high performance, and software definability. This paper proposes Balloon-ZNS that enables transparent compression in ZNS SSDs to enhance cost efficiency. ZNS SSDs, unlike traditional SSDs, require data pages to be stored in logical zones and flash blocks with aligned offsets, conflicting with the management of compressed, variable-length pages. Motivated by a key observation that compressibility locality widely exists in data streams, Balloon-ZNS employs a compressibility-adaptive, slot-aligned storage management scheme to address the intractable conflict. Evaluation with RocksDB shows Balloon-ZNS can reap more than 80% of the compression gain while achieving 7.3% lower to 14.4% higher throughput than a vanilla ZNS SSD, on average, when data compressibility is not poor. Yu Wang 0168, Zibin Sun, You Zhou 0009, Tao Lu 0014, Changsheng Xie 0001, Fei Wu 0005 |
DAC | 4 |
| 2024 | HA-CSD: Host and SSD Coordinated Compression for Capacity and PerformanceabstractIntegrating data compression capability into SSDs has demonstrated great potential to improve the utilization and lifetime of the storage device and also the performance of the entire system. It is advocated to add a hardware engine into the SSD for low-latency compression and decompression. However, this requires a new and long hardware product development cycle, which would prevent current storage systems from reaping the benefits of in-SSD compression. In this paper, we explore a software-based in-SSD compression solution, which can be delivered to users quickly through a simple SSD firmware update. The most critical challenge is the severe performance bottleneck caused by compression and decompression, as the in-SSD embedded CPU has quite limited computing power. To tackle this challenge, we propose a host-assisted computational storage device, called HA-CSD. It employs an offline, data hotness- and compressibility-aware compression strategy to remove compression from the critical write I/O path. A novel decompression architecture is devised to utilize the powerful host CPU for fast decompression. We implement HA-CSD in a commercial enterprise SSD with a code change of more than 25K lines in the host NVMe driver and SSD firmware. Experimental results show that HA-CSD achieves 2.1GB/s and 5.2GB/s read and write bandwidth. Compared with RocksDB built-in compression, HA-CSD can increase the YCSB benchmark throughput by up to 5.7×, and improve the host CPU efficiency significantly. Xiang Chen 0028, Tao Lu 0014, Jiapin Wang, Guangchun Xie, Xueming Cao, Yuanpeng Ma, Bing Si, Yunxin Huang, Yafei Yang, You Zhou 0009, Fei Wu 0005 |
IPDPS | 2 |
| 2024 | A Unified Computational Storage and Memory Architecture in the CXL EraabstractComputational Storage Drives (CSDs) integrate compute engines directly within SSDs for efficient near-data processing. With the introduction of Compute Express Link (CXL), memory expanders can similarly become Computational Memory (CM) by offloading computations. However, the integration of specific hardware accelerators in CSDs has posed substantial mass production challenges, a pitfall we anticipate will also affect CM. To address this, we propose an innovative, decoupled architecture that uses CXL switches to separate accelerators from storage and memory devices. Our analysis suggests this architecture effectively sidesteps the anti-mass production issues faced by CSD and CM with an affordable (e.g.,$< 20\%$) performance degradation. Xiang Chen 0028, Fei Wu 0005, Tao Lu 0014 |
NAS | 3 |
| 2023 | ADT-FSE: A New Encoder for SZabstractSZ is a lossy floating-point data compressor that excels in compression ratio and throughput for high-performance computing (HPC), time series databases, and deep learning applications. However, SZ performs poorly for small chunks and has slow decompression. We pinpoint the Huffman tree in the quantization factor encoder as the bottleneck of SZ. In this paper, we propose ADT-FSE, a new quantization factor encoder for SZ. Based on the Gaussian distribution of quantization factors, we design an adaptive data transcoding (ADT) scheme to map quantization factors to codes for better compressibility, and then use finite state entropy (FSE) to compress the codes. Experiments show that ADT-FSE improves the quantization factor compression ratio, compression and decompression throughput by up to 5×, 2× and 8×, respectively, over the original SZ Huffman encoder. On average, SZ_ADT is over 2× faster than ZFP in decompression. Case studies of the TDengine time series database and HDF5 file store confirm that SZ_ADT significantly boosts user-perceived application performance. In addition, ADT-FSE makes the compression ratio prediction of SZ_ADT easy and accurate, and has the potential to dramatically reduce the area size of SZ hardware implementation. Tao Lu 0014, Zibin Sun, Xiang Chen 0028, You Zhou 0009, Fei Wu 0005, Yunxin Huang, Yafei Yang |
SC | 1 |
| 2023 | High-Ratio Lossy Compression: Exploring the Autoencoder to Compress Scientific DataabstractScientific simulations on high-performance computing (HPC) systems can generate large amounts of floating-point data per run. To mitigate the data storage bottleneck and lower the data volume, it is common for floating-point compressors to be employed. As compared to lossless compressors, lossy compressors, such as SZ and ZFP, can reduce data volume more aggressively while maintaining the usefulness of the data. However, a reduction ratio of more than two orders of magnitude is almost impossible without seriously distorting the data. In deep learning, the autoencoder technique has shown great potential for data compression, in particular with images. Whether the autoencoder can deliver similar performance on scientific data, however, is unknown. In this article, we for the first time conduct a comprehensive study on the use of autoencoders to compress real-world scientific data and illustrate several key findings on using autoencoders for scientific data reduction. We implement an autoencoder-based compression prototype to reduce floating-point data. Our study shows that the out-of-the-box implementation needs to be further tuned in order to achieve high compression ratios and satisfactory error bounds. Our evaluation results show that, for most of the test datasets, the tuned autoencoder outperforms SZ by up to 4X, and ZFP by up to 50X in compression ratios, respectively. Our practices and lessons learned in this work can direct future optimizations for using autoencoders to compress scientific data. Tong Liu 0030, Jinzhen Wang, Qing Liu 0002, Shakeel Alibhai, Tao Lu 0014, Xubin He |
IEEE Trans. Big Data | 5 |
| 2020 | Performance Optimization for Relative-Error-Bounded Lossy Compression on Scientific DataabstractScientific simulations in high-performance computing (HPC) environments generate vast volume of data, which may cause a severe I/O bottleneck at runtime and a huge burden on storage space for postanalysis. Unlike traditional data reduction schemes such as deduplication or lossless compression, not only can error-controlled lossy compression significantly reduce the data size but it also holds the promise to satisfy user demand on error control. Pointwise relative error bounds (i.e., compression errors depends on the data values) are widely used by many scientific applications with lossy compression since error control can adapt to the error bound in the dataset automatically. Pointwise relative-error-bounded compression is complicated and time consuming. We develop efficient precomputation-based mechanisms based on the SZ lossy compression framework. Our mechanisms can avoid costly logarithmic transformation and identify quantization factor values via a fast table lookup, greatly accelerating the relative-error-bounded compression with excellent compression ratios. In addition, we reduce traversing operations for Huffman decoding, significantly accelerating the decompression process in SZ. Experiments with eight well-known real-world scientific simulation datasets show that our solution can improve the compression and decompression rates (i.e., the speed) by about 40 and 80 p, respectively, in most of cases, making our designed lossy compression strategy the best-in-class solution in most cases. Xiangyu Zou, Tao Lu 0014, Wen Xia, Xuan Wang 0002, Weizhe Zhang, Haijun Zhang 0002, Sheng Di, Dingwen Tao, Franck Cappello |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2019 | Accelerating Relative-error Bounded Lossy Compression for HPC datasets with Precomputation-Based MechanismsabstractScientific simulations in high-performance computing (HPC) environments are producing vast volume of data, which may cause a severe I/O bottleneck at runtime and a huge burden on storage space for post-analysis. Unlike the traditional data reduction schemes (such as deduplication or lossless compression), not only can error-controlled lossy compression significantly reduce the data size but it can also hold the promise to satisfy user demand on error control. Point-wise relative error bounds (i.e., compression errors depends on the data values) are widely used by many scientific applications in the lossy compression, since error control can adapt to the precision in the dataset automatically. Point-wise relative error bounded compression is complicated and time consuming. In this work, we develop efficient precomputation-based mechanisms in the SZ lossy compression framework. Our mechanisms can avoid costly logarithmic transformation and identify quantization factor values via a fast table lookup, greatly accelerating the relative-error bounded compression with excellent compression ratios. In addition, our mechanisms also help reduce traversing operations for Huffman decoding, and thus significantly accelerate the decompression process in SZ. Experiments with four well-known real-world scientific simulation datasets show that our solution can improve the compression rate by about 30% and decompression rate by about 70% in most of cases, making our designed lossy compression strategy the best choice in class in most cases. Xiangyu Zou, Tao Lu 0014, Wen Xia, Xuan Wang 0002, Weizhe Zhang, Sheng Di, Dingwen Tao, Franck Cappello |
MSST | 2 |
| 2018 | Understanding and Modeling Lossy Compression Schemes on HPC Scientific DataabstractScientific simulations generate large amounts of floating-point data, which are often not very compressible using the traditional reduction schemes, such as deduplication or lossless compression. The emergence of lossy floating-point compression holds promise to satisfy the data reduction demand from HPC applications; however, lossy compression has not been widely adopted in science production. We believe a fundamental reason is that there is a lack of understanding of the benefits, pitfalls, and performance of lossy compression on scientific data. In this paper, we conduct a comprehensive study on state-of-the-art lossy compression, including ZFP, SZ, and ISABELA, using real and representative HPC datasets. Our evaluation reveals the complex interplay between compressor design, data features and compression performance. The impact of reduced accuracy on data analytics is also examined through a case study of fusion blob detection, offering domain scientists with the insights of what to expect from fidelity loss. Furthermore, the trial and error approach to understanding compression performance involves substantial compute and storage overhead. To this end, we propose a sampling based estimation method that extrapolates the reduction ratio from data samples, to guide domain scientists to make more informed data reduction decisions. Tao Lu 0014, Qing Liu 0002, Xubin He, Huizhang Luo, Eric Suchyta, Jong Choi 0001, Norbert Podhorszki, Scott Klasky, Matthew Wolf, Tong Liu 0030, Zhenbo Qiao |
IPDPS | 1 |
| 2017 | Canopus: A Paradigm Shift Towards Elastic Extreme-Scale Data Analytics on HPC StorageabstractScientific simulations on high performance computing (HPC) platforms generate large quantities of data. To bridge the widening gap between compute and I/O, and enable data to be more efficiently stored and analyzed, simulation outputs need to be refactored, reduced, and appropriately mapped to storage tiers. However, a systematic solution to support these steps has been lacking on the current HPC software ecosystem. To that end, this paper develops Canopus, a progressive JPEGlike data management scheme for storing and analyzing big scientific data. It co-designs the data decimation, compression and data storage, taking the hardware characteristics of each storage tier into considerations. With reasonably low overhead, our approach refactors simulation data into a much smaller, reduced-accuracy base dataset, and a series of deltas that is used to augment the accuracy if needed. The base dataset and deltas are compressed and written to multiple storage tiers. Data saved on different tiers can then be selectively retrieved to restore the level of accuracy that satisfies data analytics. Thus, Canopus provides a paradigm shift towards elastic data analytics and enables end users to make trade-offs between analysis speed and accuracy on-the-fly. We evaluate the impact of Canopus on unstructured triangular meshes, a pervasive data model used by scientific modeling and simulations. In particular, we demonstrate the progressive data exploration of Canopus using the “blob detection” use case on the fusion simulation data. Tao Lu 0014, Eric Suchyta, David Pugmire, Jong Choi 0001, Scott Klasky, Qing Liu 0002, Norbert Podhorszki, Mark Ainsworth, Matthew Wolf |
CLUSTER | 1 |
| 2017 | Canopus: Enabling Extreme-Scale Data Analytics on Big HPC Storage via Progressive Refactoring
Tao Lu 0014, Eric Suchyta, Jong Choi 0001, Norbert Podhorszki, Scott Klasky, Qing Liu 0002, David Pugmire, Matthew Wolf, Mark Ainsworth |
HotStorage | 1 |
| 2017 | Toward Managing HPC Burst Buffers Effectively: Draining Strategy to Regulate Bursty I/O BehaviorabstractHPC (high-performance computing) applications usually show bursty I/O behaviors. In order to expedite the applications, permanent storage systems are usually provisioned to serve such I/O bursts. Approaching the era of exascale computing, non-volatile RAM is introduced as burst buffers, to absorb the bursty bulk data and relax the I/O provisioning requirement of the permanent storage systems. However, without judiciously draining the burst buffers, I/O bursts are passed down to the underlying storage systems, which causes severe I/O contention issues.In order to minimize the I/O provisioning requirement and resolve the issues caused by I/O bursts, we propose a proactive draining scheme to manage the draining process of distributed node-local burst buffers. In addition, we develop an I/O provisioning model to predict the minimized I/O provisioning requirement for permanent storage systems. Evaluation results show that applying the proactive draining scheme largely relaxes the I/O provisioning requirement while preserving the I/O performance of underlying storage systems. Ping Huang 0001, Xubin He, Tao Lu 0014, Sudharshan S. Vazhkudai, Devesh Tiwari |
MASCOTS | 4 |
| 2016 | Successor: Proactive cache warm-up of destination hosts in virtual machine migration contextsabstractIn virtualization platforms, host-side storage caches can serve virtual machines (VM) disk I/O requests, which originally target network storage servers. When these requests hit host-side caches, network and disk access latencies are obviated, and thus VMs perceive improved storage performance. VM migration is common in cloud environments, however, VM migration does not transfer host-side cache states. As a result, a newly migrated VM suffers performance degradation until the cache is fully rebuilt. The performance degradation period can be hours long if the cache is naturally warmed up. Employing existing cache warm-up solutions such as migrating host-side cache and Bonfire, VMs may either have a prolonged total migration time or undergo a performance degradation period of tens of minutes due to the warm-up caused storage contention. We propose Successor, which proactively warms up caches of destination hosts before migration completes. Specifically, accessibility of destination hosts during migration enables Successor to parallelize cache warm-up and VM migration. Compared with migrating host-side cache and Bonfire, Successor achieves zero VM-perceived cache warm-up time with low resource costs and performance penalties. We have implemented a prototype of Successor on QEMU/KVM based virtualization platform and verified its efficiency. Tao Lu 0014, Ping Huang 0001, Morgan Stuart, Yuhua Guo, Xubin He, Ming Zhang 0026 |
INFOCOM | 1 |
| 2016 | LAMS: A latency-aware memory scheduling policy for modern DRAM systemsabstractThis paper introduces a new memory scheduling policy called LAMS, which is inspired by a recently proposed memory architecture and targets for future high capacity memory systems. As memory capacity increases, the bit-lines connected to memory row buffers become much longer, dramatically lengthening memory access latency, due to increased parasitic capacitance. Recent study has proposed to partition long bit-lines into near and far (relative to the row buffer) segments via inserting isolation transistors such that access to near segment can be accomplished much faster, while access to far segment remains nearly the same. However, how to effectively leverage the new memory architecture still remains unexplored. We suggest to take advantage of this new memory architecture via performing latency-aware memory scheduling for pending requests to explore their performance potentials. In this scheduling policy, each memory request is classified to one of the following three categories, row-buffer hit, near-buffer, and far-buffer. Based on the classification, it issues requests in the order of row-buffer hit → near-buffer → far-buffer. In doing so, it avoids long-latency requests blocking short-latency memory requests, reducing total memory queuing time in the memory controller and improving overall memory performance. Our evaluation results on a simulated memory system show that comparing with the commonly used FR-FCFS scheduler, our LAMS improves performance and energy efficiency by up to 20.6% and 34%, respectively. Even comparing with the four competitive schedulers chosen from memory scheduling champion (MSC), LAMS still improves performance and energy efficiency by up to 6.1% and 23.4%, respectively. Wenjie Liu 0002, Ping Huang 0001, Tang Kun, Tao Lu 0014, Ke Zhou 0001, Chun-hua Li, Xubin He |
IPCCC | 4 |
| 2016 | Improve Restore Speed in Deduplication Systems Using Segregated CacheabstractThe chunk fragmentation problem inherently associated with deduplication systems significantly slows down the restore performance, as it causes the restore process to assemble chunks which are distributed in a large number of containers as a result of storage indirection. Existing solutions attempting to address the fragmentation problem either sacrifice deduplication efficiency or require additional memory resources. In this work, we propose a new restore cache scheme, which accelerates the restore process using the same amount of cache space as that of the traditional LRU restore cache. We leverage the recipe knowledge to recognize the containers which will soon be accessed for restoring a backup version and classify those containers into bursty containers which are differentiated from other regular containers. Bursty and regular containers are then put in two separate caches, respectively. Bursty containers, containing many chunks that will be needed for restore within a short period of time, are put in a smaller cache managed at the container granularity. On the contrary, regular containers are put in the other bigger cache managed at the chunk granularity, with chunks which will not be used dropped off at the time when the containers are brought in. In doing so, bursty containers have better chances to be quickly evicted from the restore cache, avoiding their unnecessarily occupying cache space for too long. Our evaluation results have demonstrated that our proposed cache scheme can improve restore speed factor by up to 3.05X and reduce the number of container reads by 67.3% on average, relative to a conventional LRU restore cache. Wenjie Liu 0002, Ping Huang 0001, Tao Lu 0014, Xubin He, Hua Wang 0008, Ke Zhou 0001 |
MASCOTS | 3 |
| 2015 | Alleviating DRAM Refresh Overhead Via Inter-rank Piggyback CachingabstractDRAM cells leak charge over time, causing stored data to be lost. Therefore, periodic refreshes are required to ensure data integrity. Modern DRAM usually refreshes cells at rank level, resulting in an entire rank being unavailable during a refresh period. As DRAM density keeps increasing, more rows need to be refreshed during a single refresh operation, which causes higher refresh latency and significantly degrades the overall memory system performance. To mitigate DRAM refresh overhead, we propose a caching scheme, called Rank-level Piggyback Caching, or RPC for short, based on the fact that ranks in the same channel are refreshed in a staggered manner. The key idea is to cache the to-be-read data in a rank (e.g. Rank 1) to its adjacent rank (e.g. Rank 2) before Rank 1 is locked for refresh. Each rank reserves or over-provisions a very small area, denoted as a cache region, to store the cached data. The cache regions from all ranks are organized in a rotated fashion. In other words, the cached data for the last rank is stored in the first rank. When a read request arrives at a rank undergoing refresh, the memory controller first checks the cache region in the next rank in the same channel, if the requested data is cached, the memory controller services the request from the cache without waiting for the refresh operation to complete, which reduces memory access latency and improves system performance. Our experimental results show that RPC outperforms existing Fine Granularity Refresh modes. In a single-core and four-rank system, it improves system performance by 8.7% and 10.8% on average for the PARSEC 2.1 and SPLASH-2 benchmark suites, respectively. In a four-core and four-rank system, the improvement of system performance for these two benchmark suites is 8.6% and 12.2%, respectively. Yuhua Guo, Ping Huang 0001, Tao Lu 0014, Xubin He, Qing Gary Liu |
MASCOTS | 4 |
| 2014 | Clique Migration: Affinity Grouping of Virtual Machines for Inter-cloud Live MigrationabstractAffinity is common among Virtual Machines (VMs) in cloud environments. If VMs collaborating on a job are split in geographically distributed clouds, the low bandwidth and high latency inter-cloud communication via a wide area network (WAN) will dramatically degrade the system performance. A potential solution is migrating all of the VMs collaborating on a job in parallel, so as to avoid wide area communication. However, if the job is too large, it becomes impractical to migrate all of the VMs simultaneously due to limited WAN bandwidth and high block dirty rate. We propose a migration optimization mechanism called Clique Migration to partition a large group of VMs into subgroups based on the traffic affinities among VMs. Then, subgroups are migrated one at a time. Based on Clique Migration, we propose and implement two algorithms called R-Min-Cut and Kmeans-SF. Analysis of the traffic trace of 68 VMs in an IBM production cluster shows that our algorithms can reduce inter-cloud traffic by 25% to 60%, when the degree of parallel migration is from 2 to 32. Tests of MPI multi-Ping Ping benchmark running on simulated inter-cloud environments, show that our algorithms can significantly shorten the period during which applications undergo performance degradation. Tests of MPI Reduce scatter benchmark show that R-Min-Cut can keep the performance during migration at 26% to 75% of the non-migration scenario. Tao Lu 0014, Morgan Stuart, Xubin He |
NAS | 1 |