Zhaoning Zhang 0001

dblp:67/11341-1 · DBLP profile ↗
← Back
29ranked-venue papers
4as first author
13since 2021 · last 2026
0000-0001-7518-1385ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 7 since 2021Systems, architecture and hardware · 10 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 DMHM: Density-aware Manifold Learning and Hybrid Mahalanobis Energy for LLMs-generated Text Detection
abstract
Tianle Liu, Zhiliang Tian, Zhen Huang, Tianlun Liu, Jingyuan Huang, Zhaoning Zhang, Chengcheng Shao, Dongsheng Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhiliang Tian, Zhen Huang 0006, Tianlun Liu, Zhaoning Zhang 0001, Chengcheng Shao, Dongsheng Li 0001
ACL (1)6
2026 Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts Inference
abstract
Mixture-of-Experts (MoE) has become a dominant architecture for scaling large language models due to their sparse activation mechanism.However, the substantial number of expert activations creates a critical latency bottleneck during inference, especially in resourceconstrained deployment scenarios.Existing approaches that reduce expert activations potentially lead to severe model performance degradation.In this work, we introduce the concept of activation budget as a constraint on the number of expert activations and propose Alloc-MoE, a unified framework that optimizes budget allocation coordinately at both the layer and token levels to minimize performance degradation.At the layer level, we introduce Alloc-L, which leverages sensitivity profiling and dynamic programming to determine the optimal allocation of expert activations across layers.At the token level, we propose Alloc-T, which dynamically redistributes activations based on routing scores, optimizing budget allocation without increasing latency.Extensive experiments across multiple MoE models demonstrate that Alloc-MoE maintains model performance under a constrained activation budget.Especially, Alloc-MoE achieves 1.15× prefill and 1.34× decode speedups on DeepSeek-V2-Lite at half of the original budget.
Baihui Liu, Kaiyuan Tian, Zhaoning Zhang 0001, Linbo Qiao, Dongsheng Li 0001
ACL (1)4
2026 DOA: Dataflow Optimization for Attention on Multi-core DSPs with Three-Level Memory Hierarchy
Zhiquan Lai, Shun Ouyang, Zhaoning Zhang 0001, Menghan Jia, Huayou Su, Dongsheng Li 0001
APPT6
2026 LayerScope: Predictive Cross-Layer Scheduling for Efficient Multi-Batch MoE Inference on Legacy Servers
abstract
Mixture-of-Experts (MoE) models face memory and PCIe latency bottlenecks when deployed on commodity hardware. Offloading expert weights to CPU memory results in PCIe transfer latency that exceeds GPU computation by several folds. We present PreScope, a prediction-driven expert scheduling system that addresses three key challenges: inaccurate activation prediction, PCIe bandwidth competition, and cross-device scheduling complexity. Our solution includes: 1) Learnable Layer-Aware Predictor (LLaPor) that captures layer-specific expert activation patterns; 2) Prefetch-Aware Cross-Layer Scheduling (PreSched) that generates globally optimal plans balancing prefetching costs and loading overhead; 3) Asynchronous I/O Optimizer (AsyncIO) that decouples I/O from computation, eliminating waiting bubbles. PreScope achieves 141% higher throughput and 74.6% lower latency than state-of-the-art solutions.
Enda Yu, Dezun Dong, Zhaoning Zhang 0001, Zhe Bai, Weiling Yang, Haojie Wang 0004, Dongsheng Li 0001, Yongwei Wu 0001, Xiangke Liao
ICS3
2026 Balance divergence for knowledge distillation
Yafei Qi, Chen Wang 0074, Zhaoning Zhang 0001, Yongmin Zhang
Eng. Appl. Artif. Intell.3
2025 Correlation-Aware Example Selection for In-Context Learning with Nonsymmetric Determinantal Point Processes
abstract
Qiunan Du, Zhiliang Tian, Zhen Huang, Kailun Bian, Tianlun Liu, Zhaoning Zhang, Xinwang Liu, Feng Liu, Dongsheng Li. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Qiunan Du, Zhiliang Tian, Zhen Huang 0006, Kailun Bian, Tianlun Liu, Zhaoning Zhang 0001, Xinwang Liu 0002, Dongsheng Li 0001
EMNLP6
2025 ALM-KD: Adaptive Layer Mapping Knowledge Distillation for LLMs
abstract
Transformer-based models are renowned for their proficiency in computationally demanding NLP tasks. However, we have identified several shortcomings in current distillation methods for transformer-based Large Language Models (LLMs), including the disregard for layer importance, inappropriate mapping granularity, and the omission of dimetsionality compression. We propose Adaptive Layer Mapping Knowledge Dis-tillation(ALM-KD), a method designed to enhance efficiency by adaptively selecting the most relevant hidden layers and dynamically adjusting the distillation weights based on their significance. ALM-KD utilizes an ada-learner composed of RNN and MLP, along with an ada-loss function, to dynamically prioritize teacher layers according to their relevance to downstream tasks, while also integrating horizontal and vertical compression for substantial model knowledge reduction. Our experiments on the GLUE benchmark demonstrate that the student model developed using ALM-KD achieves over 96% of the performance of its teacher model, BERTbase, with an 80% compression rate, nearly tripling inference speed, and reducing memory usage by half. ALM-KD thus renders deep neural networks viable for real-time and high-concurrency applications.
Xu Baizhou, Niu Xin, Zhaoning Zhang 0001, Huang Zheng, Linbo Qiao, Yafei Qi
IJCNN4
2025 Modal Mimicking Knowledge Distillation for monocular three-dimensional object detection
Menghao Yang, Yafei Qi, Bing Xiong 0001, Zhaoning Zhang 0001
Eng. Appl. Artif. Intell.4
2025 SelfDN: Adaptive Self-Denoising for multi-view 3D object detection
Yafei Qi, Menghao Yang, Yongmin Zhang, Bing Xiong 0001, Zhaoning Zhang 0001
Knowl. Based Syst.6
2024 Deep Tiny Network for Recognition-Oriented Face Image Quality Assessment
Baoyun Peng, Min Liu 0019, Zhaoning Zhang 0001, Kai Xu 0004, Dongsheng Li 0001
CVM (2)3
2024 Adaptive Selective Knowledge Distillation: Not Blindly Accepting Teachers as Oracles
Baoyun Peng, Zhaoning Zhang 0001, Yafei Qi, Linbo Qiao
PRCV (3)2
2024 DELTA: Memory-Efficient Training via Dynamic Fine-Grained Recomputation and Swapping
abstract
To accommodate the increasingly large-scale models within limited-capacity GPU memory, various coarse-grained techniques, such as recomputation and swapping, have been proposed to optimize memory usage. However, these methods have encountered limitations, either in terms of inefficient memory reduction or diminished training performance. In response to this, our article introduces dynamic tensor offloading and recomputation (DELTA), an innovative approach for memory-efficient large-scale model training that combines fine-grained memory optimization and prefetching technology to reduce memory usage while maintaining high training throughput concurrently. Initially, we formulate the problem of memory-throughput joint optimization as an easy-solving 0/1 Knapsack problem. Leveraging this formalization, we use an improving polynomial complexity heuristic algorithm to address the problem effectively. Furthermore, we introduce, to the best of our knowledge, a novel bidirectional prefetching technology into dynamic memory management that significantly accelerates the model training when compared to relying solely on recomputation or swapping. Finally, DELTA offers users an automated training execution library, eliminating the need for manual configuration or specialized expertise. Experimental results demonstrate the effectiveness of DELTA in reducing GPU memory consumption. Compared to state-of-the-art methods, DELTA achieves substantial memory savings ranging from 40% to 72%, while maintaining comparable convergence performance for various models, including ResNet-50, ResNet-101, and BERT-Large. Notably, DELTA enables the training of GPT2-Large and GPT2-XL with batch sizes increased by 5.5× and 6×, respectively, showcasing its versatility and practicality in enabling large-scale model training on GPU hardware.
Qiao Li 0001, Lujia Yin, Dongsheng Li 0001, Yiming Zhang 0003, Xingcheng Zhang, Linbo Qiao, Zhaoning Zhang 0001, Kai Lu 0001
ACM Trans. Archit. Code Optim.9
2021 ParaX: Boosting Deep Learning for Big Data Analytics on Many-Core CPUs
abstract
Despite the fact that GPUs and accelerators are more efficient in deep learning (DL), commercial clouds like Facebook and Amazon now heavily use CPUs in DL computation because there are large numbers of CPUs which would otherwise sit idle during off-peak periods. Following the trend, CPU vendors have not only released high-performance many-core CPUs but also developed efficient math kernel libraries. However, current DL platforms cannot scale well to a large number of CPU cores, making many-core CPUs inefficient in DL computation. We analyze the memory access patterns of various layers and identify the root cause of the low scalability, i.e., the per-layer barriers that are implicitly imposed by current platforms which assign one single instance (i.e., one batch of input data) to a CPU. The barriers cause severe memory bandwidth contention and CPU starvation in the access-intensive layers (like activation and BN). This paper presents a novel approach called ParaX, which boosts the performance of DL on many-core CPUs by effectively alleviating bandwidth contention and CPU starvation. Our key idea is to assign one instance to each CPU core instead of to the entire CPU, so as to remove the per-layer barriers on the executions of the many cores. ParaX designs an ultralight scheduling policy which sufficiently overlaps the access-intensive layers with the compute-intensive ones to avoid contention, and proposes a NUMA-aware gradient server mechanism for training which leverages shared memory to substantially reduce the overhead of per-iteration parameter synchronization. We have implemented ParaX on MXNet. Extensive evaluation on a two-NUMA Intel 8280 CPU shows that ParaX significantly improves the training/inference throughput for all tested models (for image recognition and natural language processing) by 1.73X ~ 2.93X.
Lujia Yin, Yiming Zhang 0003, Zhaoning Zhang 0001, Yuxing Peng 0001
Proc. VLDB Endow.3
2019 Rise the Momentum: A Method for Reducing the Training Error on Multiple GPUs
Lujia Yin, Zhaoning Zhang 0001, Dongsheng Li 0001
ICA3PP (2)3
2019 Correlation Congruence for Knowledge Distillation
abstract
Most teacher-student frameworks based on knowledge distillation (KD) depend on a strong congruent constraint on instance level. However, they usually ignore the correlation between multiple instances, which is also valuable for knowledge transfer. In this work, we propose a new framework named correlation congruence for knowledge distillation (CCKD), which transfers not only the instance-level information but also the correlation between instances. Furthermore, a generalized kernel method based on Taylor series expansion is proposed to better capture the correlation between instances. Empirical experiments and ablation studies on image classification tasks (including CIFAR-100, ImageNet-1K) and metric learning tasks (including ReID and Face Recognition) show that the proposed CCKD substantially outperforms the original KD and other SOTA KD-based methods. The CCKD can be easily deployed in the majority of the teacher-student framework such as KD and hint-based learning methods.
Baoyun Peng, Dongsheng Li 0001, Shunfeng Zhou, Yichao Wu, Zhaoning Zhang 0001, Yu Liu 0015
ICCV7
2019 ThunderNet: Towards Real-Time Generic Object Detection on Mobile Devices
abstract
Real-time generic object detection on mobile platforms is a crucial but challenging computer vision task. Prior lightweight CNN-based detectors are inclined to use one-stage pipeline. In this paper, we investigate the effectiveness of two-stage detectors in real-time generic detection and propose a lightweight two-stage detector named ThunderNet. In the backbone part, we analyze the drawbacks in previous lightweight backbones and present a lightweight backbone designed for object detection. In the detection part, we exploit an extremely efficient RPN and detection head design. To generate more discriminative feature representation, we design two efficient architecture blocks, Context Enhancement Module and Spatial Attention Module. At last, we investigate the balance between the input resolution, the backbone, and the detection head. Benefit from the highly efficient backbone and detection part design, ThunderNet surpasses previous lightweight one-stage detectors with only 40% of the computational cost on PASCAL VOC and COCO benchmarks. Without bells and whistles, ThunderNet runs at 24.1 fps on an ARM-based device with 19.2 AP on COCO. To the best of our knowledge, this is the first real-time detector reported on ARM platforms. Code will be released for paper reproduction.
Zheng Qin 0002, Zhaoning Zhang 0001, Yiping Bao, Gang Yu 0002, Yuxing Peng 0001, Jian Sun 0001
ICCV3
2019 HPDL: Towards a General Framework for High-performance Distributed Deep Learning
abstract
With growing scale of the data volume and neural network size, we have come into the era of distributed deep learning. High-performance training and inference on distributed computing systems has been attracting increasing research attention in both academia and industry. Meanwhile, diversity of existing machine learning frameworks (e.g. TensorFlow, Pytorch and MXNet) and the explosion of deep learning hardwares (e.g. CPUs, GPUs, FPGAs and ASICs) bring more challenges for users to leverage new deep learning technologies and accelerating capability of hardware devices. We firstly search around the state-of-the-art work in the area which open our mind to take a vision upon the future deep learning framework. Then, we propose HPDL, a general framework for high-performance distributed deep learning which is compatible with existing frameworks and adaptive to various hardware architectures. At last, we discuss and foresee the key technologies fulfilling high-performance and large-scale deep learning, including optimization algorithm, hybrid communication mechanism, model parallelization, resource scheduling and single-node execution optimization.
Dongsheng Li 0001, Zhiquan Lai, Ke-shi Ge, Yiming Zhang 0003, Zhaoning Zhang 0001, Huaimin Wang 0001
ICDCS5
2018 An Empirical Study of Face Recognition under Variations
abstract
Face recognition (FR) has recently made remark- able progress given the extraordinary capabilities of modern deep learning(DL) models. Though superior performance of DL based methods over human has been reported on benchmark dataset, it remains an open problem how those systems work in real-world condition with variations such as head pose, lighting, occlusion and image noises, in particular, in comparison to conventional FR approaches. It is hard to answer this question in a quantitative manner as current benchmark datasets are in lack of full range of variations. In this paper we propose a flexible approach to simulate face images under different variations with controllable degrees and based on the simulated dataset we quantitatively study how modern DL based and conventional FR methods perform. Based on the observations on a large number of synthesized face images, we draw several conclusions such as how head pose pitch and yaw variations will influence the FR systems and which part on the face is the most significant region, and in which situation conventional methods still show some advantages. The findings will not only be useful to assess current FR system in a quantitative manner but also shed light on future FR system design and data augmentation.
Baoyun Peng, Dongsheng Li 0001, Zhaoning Zhang 0001
FG4
2018 Fd-Mobilenet: Improved Mobilenet with a Fast Downsampling Strategy
abstract
We present Fast-Downsampling MobileNet (FD-MobileNet), an efficient and accurate network for very limited computational budgets (e.g., 10-140 MFLOPs). Our key idea is applying a fast downsampling strategy to Mobile Net framework. In FD-Mobile Net, we perform 32× downsampling within 12 layers, only half the layers in the original MobileNet. This design brings three advantages: (i) It remarkably reduces the computational cost. (ii) It increases the information capacity and achieves significant performance improvements. (iii) It is engineering-friendly and provides fast actual inference speed. Experiments on ILSVRC 2012 and PASCAL VOC datasets demonstrate that FD-Mobile Net consistently outperforms MobileNet and achieves comparable results with ShufflieNet under different computational budgets, for instance, surpassing Mobile-Net by 5.5% on the ILSVRC 2012 top-l accuracy and 8.3% on the VOC 2007 mAP under a complexity of 12 MFLOPs. On an ARM-based device, FD-Mobile Net achieves 1.11× inference speedup over Mobile Net and 1.82× over Shufflie Net under the same complexity.
Zheng Qin 0002, Zhaoning Zhang 0001, Xiaotao Chen, Yuxing Peng 0001
ICIP2
2018 A Quick Survey on Large Scale Distributed Deep Learning Systems
abstract
Deep learning have been widely used in various fields and has worked very well as a major role. While the gradual penetration into various fields, data quantity of each applications is increasing tremendously, and so as the computation complexity and model parameters. As an obvious result, the training and inference is time consuming. For example, a classic Resnet50 classification model will be trained in 14 days on a NVIDIA M40 GPU with ImageNet data set. Thus, distributed acceleration is a very useful way to dispatch the computation of training and even inference to scale of nodes in parallel and accelerate the whole process. Facebook's work and UC Berkeley's acceleration can training the Resnet-50 model within hour and minutes by distributed deep learning algorithm and system, representatively. As other distributed accelerations, it gives a possibility to accelerate large models on large data sets from weeks to minutes, which gives researchers and developers more space to explore and search. However, besides acceleration, what other issues will be confronted of the distributed deep learning system? Where is the upper limit of acceleration? What application will acceleration be used for? What is the price and cost of acceleration? In this paper, we will take a simple and quick survey on the distributed deep learning system from algorithm perspective, distributed system perspective and applications perspective. We will present several recent excellent works, and bring analysis on the restricts and prospects of the distributed methods.
Zhaoning Zhang 0001, Lujia Yin, Yuxing Peng 0001, Dongsheng Li 0001
ICPADS1
2018 Nominal Data Similarity: A Hierarchical Measure
abstract
Similarity of nominal data plays fundamental roles in numerous fields of both machine learning and data mining. Unlike the similarity of numerical data, that of nominal data is much more difficult to describe, and few efforts have been done for it. Although existing nominal similarity measures can reveal a part of data properties, they suffer from low accuracy due to ignoring value relationships or integrating multi-view relationships inappropriately. In this paper, we propose a novel hierarchical measure for nominal data similarity (HNS). The HNS leverages the intrinsic data characteristics by considering low-level information both within and between attributes, and hierarchically seizes the value distributions, attribute interactions and attribute to object contributions. Meanwhile, it aggregates multi-view relationships trough a bottom to top framework, remaining consistency as well as complementary details. We theoretically analyzed this measure, and experiments on six UCI data sets demonstrate that the HNS outperforms the state-of-the-art nominal similarity measures in term of target alignment and clustering accuracy.
Hao Yu 0010, Zhaoning Zhang 0001, Gen Zhang
IJCNN2
2018 Diagonalwise Refactorization: An Efficient Training Method for Depthwise Convolutions
abstract
Depthwise convolutions provide significant performance benefits owing to the reduction in both parameters and mult-adds. However, training depthwise convolution layers with GPUs is slow in current deep learning frameworks because their implementations cannot fully utilize the GPU capacity. To address this problem, in this paper we present an efficient method (called diagonalwise refactorization) for accelerating the training of depthwise convolution layers. Our key idea is to rearrange the weight vectors of a depthwise convolution into a large diagonal weight matrix so as to convert the depthwise convolution into one single standard convolution, which is well supported by the cuDNN library that is highly-optimized for GPU computations. We have implemented our training method in five popular deep learning frameworks. Evaluation results show that our proposed method gains 15.4× training speedup on Darknet, 8.4× on Caffe, 5.4× on PyTorch, 3.5× on MXNet, and 1.4× on TensorFlow, compared to their original implementations of depthwise convolutions.
Zheng Qin 0002, Zhaoning Zhang 0001, Dongsheng Li 0001, Yiming Zhang 0003, Yuxing Peng 0001
IJCNN2
2018 Merging and Evolution: Improving Convolutional Neural Networks for Mobile Applications
abstract
Compact neural networks are inclined to exploit “sparsely-connected” convolutions such as depthwise convolution and group convolution for employment in mobile applications. Compared with standard “fully-connected” convolutions, these convolutions are more computationally economical. However, “sparsely-connected” convolutions block the inter-group informa-tion exchange, which induces severe performance degradation. To address this issue, we present two novel operations named merging and evolution to leverage the inter-group information. Our key idea is encoding the inter-group information with a narrow feature map, then combining the generated features with the original network for better representation. Taking advantage of the proposed operations, we then introduce the Merging-and- Evolution (ME) module, an architectural unit specifically designed for compact networks. Finally, we propose a family of compact neural networks called MENet based on ME modules. Extensive experiments on ILSVRC 2012 dataset and PASCAL VOC 2007 dataset demonstrate that MENet consistently outperforms other state-of -the-art compact networks under different computational budgets. For instance, under the computational budget of 140 MFLOPs, MENet surpasses ShuffleNet by 1% and MobileNet by 1.95% on ILSVRC 2012 top-l accuracy, while by 2.3% and 4.1% on PASCAL VOC 2007 mAP, respectively.
Zheng Qin 0002, Zhaoning Zhang 0001, Shiqing Zhang, Hao Yu 0010, Jincai Li, Yuxing Peng 0001
IJCNN2
2018 Loss Rank Mining: A General Hard Example Mining Method for Real-time Detectors
abstract
Modern object detectors usually suffer from low accuracy issues, as foregrounds always drown in tons of back-grounds and become hard examples during training. Compared with those proposal-based ones, real-time detectors are in far more serious trouble since they renounce the use of region-proposing stage which is used to filter a majority of back-grounds for achieving real-time rates. Though foregrounds as hard examples are in urgent need of being mined from tons of backgrounds, a considerable number of state-of-the-art real-time detectors, like YOLO series, have yet to profit from existing hard example mining methods, as using these methods need detectors fit series of prerequisites. In this paper, we propose a general hard example mining method named Loss Rank Mining (LRM) to fill the gap. LRM is a general method for real-time detectors, as it utilizes the final feature map which exists in all real-time detectors to mine hard examples. By using LRM, some elements representing easy examples in final feature map are filtered and detectors are forced to concentrate on hard examples during training. Extensive experiments validate the effectiveness of our method. With our method, the improvements of YOLOv2 detector on auto-driving related dataset KITTI and more general dataset PASCAL VOC are over 5% and 2% mAP, respectively. In addition, LRM is the first hard example mining strategy which could fit YOLOv2 perfectly and make it better applied in series of real scenarios where both real-time rates and accurate detection are strongly demanded.
Hao Yu 0010, Zhaoning Zhang 0001, Zheng Qin 0002, Hao Wu 0031, Dongsheng Li 0001, Xicheng Lu
IJCNN2
2017 GraphA: Adaptive Partitioning for Natural Graphs
abstract
Large-scale graph computation is central to applications ranging from language processing to social networks. However, natural graphs tend to have skewed power-law distributions where a small subset of the vertices have a large number of neighbors. Existing graph-parallel systems suffer from load imbalance, high communication cost, or suboptimal and complex processing. In this paper we present GraphA, an Adaptive approach to efficient partitioning and computation of large-scale natural graphs. GraphA provides an adaptive and uniform graph partitioning algorithm, which partitions the datasets in a load-balanced manner by using an incremental number of hash functions. We have implemented GraphA both on Spark and on GraphLab. Extensive evaluation shows that GraphA remarkably outperforms state-of-the-art graph-parallel systems (GraphX and PowerLyra) in ingress time, execution time and storage overhead, for both real-worldand synthetic graphs.
Dongsheng Li 0001, Chengfei Zhang, Zhaoning Zhang 0001, Yiming Zhang 0003
ICDCS4
2016 DSS: A Scalable and Efficient Stratified Sampling Algorithm for Large-Scale Datasets
Minne Li, Dongsheng Li 0001, Zhaoning Zhang 0001, Xicheng Lu
NPC4
2016 Large-scale virtual machines provisioning in clouds: challenges and approaches
Zhaoning Zhang 0001, Dongsheng Li 0001, Kui Wu 0001
Frontiers Comput. Sci.1
2014 RAFlow: Read Ahead Accelerated I/O Flow through Multiple Virtual Layers
abstract
Virtualization is the foundation for cloud computing, and the virtualization can not be achieved without software defined, elastic, flexible and scalable virtual layers. Unfortunately, if multiple virtual storage devices are chained together, the system may be subject to severe performance degradation. While the read-ahead (RA) mechanism in storage devices plays a very important role to improve I/O performance, RA may not be effective as expected for multiple virtualization layers, since it is originally designed for one layer only. When I/O requests are passed through a long I/O path, they may trigger a chain reaction and lead to unnecessary data transmission and thus bandwidth waste. In this paper, we study the dynamic behavior of RA through multiple I/O layers and demonstrate that if controlled well, RA can greatly accelerate I/O speed. We present RAFlow, a RA control mechanism, to effectively improve I/O performance by strategically expanding RA window at each layer. Our real-world experiments show that it can achieve 20% to 50% performance improvement in I/O paths with up to 8 virtualized storage devices.
Zhaoning Zhang 0001, Kui Wu 0001, Huiba Li, Jinghua Feng, Yuxing Peng 0001, Xicheng Lu
NAS1
2014 VMThunder: Fast Provisioning of Large-Scale Virtual Machine Clusters
abstract
Infrastructure as a service (IaaS) allows users to rent resources from the Cloud to meet their various computing requirements. The pay-as-you-use model, however, poses a nontrivial technical challenge to the IaaS cloud service providers: how to fast provision a large number of virtual machines (VMs) to meet users' dynamic computing requests? We address this challenge with VMThunder, a new VM provisioning tool, which downloads data blockson demandduring the VM booting process and speeds up VM image streaming by strategically integrating peer-to-peer (P2P) streaming techniques with enhanced optimization schemes such as transfer on demand, cache on read, snapshot on local, and relay on cache. In particular, VMThunder stores the original images in a share storage and in the meantime it adopts a tree-based P2P streaming scheme so that common image blocks are cached and reused across the nodes in the cluster. We implement VMThunder in CentOS Linux and thoroughly test its performance. Comprehensive experimental results show that VMThunder outperforms the state-of-the-art VM provisioning methods, with respect to scalability, latency, and VM runtime I/O performance.
Zhaoning Zhang 0001, Ziyang Li 0003, Kui Wu 0001, Dongsheng Li 0001, Huiba Li, Yuxing Peng 0001, Xicheng Lu
IEEE Trans. Parallel Distributed Syst.1