Fangming Liu

dblp:31/6019 · DBLP profile ↗
← Back
154ranked-venue papers
16as first author
69since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 63 · 7 first-author · 31 since 2021Computer networks · 62 · 7 first-author · 20 since 2021Artificial intelligence and machine learning · 13 · 1 first-author · 13 since 2021Software engineering, systems software and programming languages · 10 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 3 since 2021Security and privacy · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 FP=XINT: Representing Neural Networks via Low-Bit Series Basis Functions
abstract
Deep neural networks are often over-parameterized, resulting in prohibitive storage and computational costs. A fundamental question is whether a complex network can be re-expressed in terms of a compact set of basis functions without sacrificing accuracy. Motivated by this perspective, we aim to approximate a dense model by decomposing it into a small number of lightweight components that capture the essential functional structure of the network. To this end, we propose a series expansion framework that rewrites a neural network as a linear combination of low-bit basis models. Within the post-training quantization setting, the full-precision model is expanded hierarchically at the tensor, layer, and model levels into a structured set of basis functions. We theoretically prove that this expansion converges exponentially to the original model. Furthermore, we design AbelianAdd and AbelianMul operations between isomorphic basis models, endowing the expansion with an Abelian group structure that naturally supports commutative and parallel computation. Experimental results across diverse architectures show that our series expansion method leverages a set of ultra-low-bit basis functions, not only preserving full-precision performance without the need for calibration data or fine-tuning, but also featuring a parallel-friendly design that enables efficient and scalable deployment.
Daning Cheng, Yunquan Zhang, Jiake Tian, Fangming Liu
AAAI6
2026 Hardware-Constrained Online Coordination for Hybrid CXL-RDMA Disaggregated Memory
abstract
Hybrid CXL–RDMA disaggregated memory systems expose heterogeneous access costs and explicit hardware limits, making static tiering and fabric-specific optimizations insufficient under high concurrency. We identify a per-request coordination problem: how to schedule memory requests over heterogeneous access paths and execution contexts under explicit transaction and queue bounds. We formulate this problem as a hardware-constrained online scheduling model. We then propose HAPES, a hardware-aware runtime that jointly coordinates access path, execution context, and request granularity based on workload state and instantaneous hardware pressure.
Fangming Liu
CF4
2026 Optimal Compilation of Syndrome Extraction Circuits for General Quantum LDPC Codes
abstract
Quantum error correcting codes (QECC) are essential for constructing large-scale quantum computers that deliver faithful results. As strong competitors to the conventional surface code, quantum low-density parity-check (qLDPC) codes are emerging rapidly: they offer high encoding rates while maintaining reasonable physical-qubit connectivity requirements. Despite the existence of numerous code constructions, a notable gap persists between these designs—some of which remain purely theoretical—and their circuit-level deployment.In this work, we propose Auto-Stabilizer-Check (ASC), a universal compilation framework that generates depth-optimal syndrome extraction circuits for arbitrary qLDPC codes. ASC leverages the sparsity of parity-check matrices and exploits the commutativity of X and Z stabilizer measurement subroutines to search for optimal compilation schemes. By iteratively invoking an SMT solver, ASC returns a depth-optimal solution if a satisfying assignment is found, and a near-optimal solution in cases of solver timeouts. Notably, ASC provides the first definitive answer to one of IBM’s open problems: for all instances of bivariate bicycle (BB) code reported in their work, our compiler certifies that no depth-6 syndrome extraction circuit exists.Furthermore, by integrating ASC with an end-to-end evaluation framework—one that assesses different compilation settings under a circuit-level noise model—ASC reduces circuit depth by approximately 50% and achieves an average 7x-8x suppression of the logical error rate for general qLDPC codes, compared with as-soon-as-possible (ASAP) and coloration-based scheduling. ASC thus substantially reduces manual design overhead and demonstrates its strong potential to serve as a key component in accelerating hardware deployment of qLDPC codes.
Dingchao Gao, Runshi Zhou, Fangming Liu, Zheng-Feng Ji
DATE5
2026 Pulse: Training Acceleration for Large Diffusion Models with Automatic Pipeline Parallelism
Boran Sun, Guoyong Jiang, Yuechen Tao, Zhishu Che, Jieling Yu, Shan Chang, Huaxi Gu, Fangming Liu
ICDCS10
2026 CxlRatioOpt : A Predictive Optimizer for Performance Trade-offs in CXL-Tiered Memory
Bowen Wang 0013, Haiyuan Wan, Zhirun Yue, Fangming Liu
ISCAS8
2026 QRoute: A QoE-aware route planning system for enhancing in-vehicle video streaming
Jiahai Hu, Qiangyu Pei, Fangming Liu
Comput. Networks5
2026 DroneNav: Unified text-visual representation and structured spatial reasoning for robust UAV vision-and-language navigation
Fangming Liu, Guohua Li, Linfeng Zou
Neurocomputing1
2026 Cooling as You Wish: Component-Level Cooling for Heterogeneous Edge Datacenters
abstract
As computing shifts toward the edge, edge datacenters are becoming essential for supporting diverse real-time applications. Unlike traditional cloud datacenters, edge datacenters face unique cooling challenges due to their requirements forproximity to end users, high density, and hardware heterogeneity. While warm water cooling is a promising technique for this infrastructure, current one-size-fits-all cooling strategies significantly compromise efficiency due to severe inter- and intra-component hotspots. In this work, we present CoolEdge+, a cost-effective component–level water cooling system for enhancing the cooling efficiency of edge datacenters. Specifically, CoolEdge+dynamically adjusts the inlet water temperature for each component through a carefully designed water circulation architecture to mitigate inter-component hotspots. To address intra-component hotspots, it employs vapor chamber–based cold plates that rapidly dissipate heat without manual intervention or additional energy consumption. We further design a fine-grained cooling control framework that leverages a well-managed power capping approach to decide on customized inlet water temperatures and hardware power limits. Based on a hardware prototype and a real-world trace from Alibaba PAI, evaluation results show that CoolEdge+reduces cooling energy consumption by up to 27.19% compared to existing coarse-grained systems, while maintaining performance guarantees. Compared to the state-of-the-art CoolEdge, CoolEdge+saves 35.24% more cooling costs with comparable energy consumption and no latency violations.
Fangming Liu, Qiangyu Pei, Yongjie Yuan, Qixia Zhang, Ziyang Jia, Fei Xu 0009, Bingheng Yan
IEEE Trans. Computers1
2026 Efficient Tensor Offloading Based on CXL Memory Pool for Extreme Scale Deep Learning
abstract
The exponential growth of deep learning models imposes severe memory constraints on GPUs, significantly increasing training costs. The prevailing solution for addressing memory constraints is tensor offloading, exemplified by ZeRO Infinity, which leverages GPUs, CPUs, and NVMe SSDs to enable large-scale model training. However, ZeRO-Infinity faces significant performance bottlenecks due to memory access imbalance, coarse-grained tensor transfers, NVMe bandwidth and latency limits, and software complexity. Compute Express Link (CXL) emerges as a promising technology for building disaggregated memory pools, yet its integration into large-scale training remains underexplored.This paper introduces an efficient CXL memory pool into the system for tensor offloading and leveraging CXL protocol features for hardware acceleration. The proposed design incorporates NUMA-aware memory allocation and a Dynamic Adaptive Pipelining (DAP) strategy to enhance communication–computation overlap, along with a hybrid object-based memory management scheme to mitigate fragmentation and improve allocation efficiency. Experimental evaluation on an 8- GPU system with CXL Type-3 expansion cards demonstrates up to 72.7% throughput improvement, 62.9% latency reduction, and support for training 1.14× larger models compared with SSD based ZeRO-Infinity. This is the first work to integrate CXL with ZeRO-Infinity for large-scale training, offering practical insights for future CXL-based heterogeneous systems.
Dongwei Xu, Fangming Liu, Bowen Wang 0013, Haiyuan Wan, Zhirun Yue
IEEE Trans. Computers7
2026 Detection and Mitigation Data Poisoning Attacks in Multimodal Online Federated Learning
Heqiang Wang, Xiaoxiong Zhong, Hualong Wu, Fangming Liu, Weizhe Zhang
IEEE Trans. Inf. Forensics Secur.5
2026 Multimodal Online Federated Learning With Modality Missing in Internet of Things
abstract
The Internet of Things (IoT) ecosystem generates vast amounts of multimodal data from heterogeneous sources such as sensors, cameras, and microphones. As edge intelligence continues to evolve, IoT devices have progressed from simple data collection units to nodes capable of executing complex computational tasks. This evolution necessitates the adoption of distributed learning strategies to effectively handle multimodal data in an IoT environment. Furthermore, the real-time nature of data collection and limited local storage on edge devices in IoT call for an online learning paradigm. To address these challenges, we introduce the concept of Multimodal Online Federated Learning (MMO-FL), a novel framework designed for dynamic and decentralized multimodal learning in IoT environments. Building on this framework, we further account for the inherent instability of edge devices, which frequently results in missing modalities during the learning process. We conduct a comprehensive theoretical analysis under both complete and missing modality scenarios, providing insights into the performance degradation caused by missing modalities. To mitigate the impact of modality missing, we propose the Prototypical Modality Mitigation (PMM) algorithm, which leverages prototype learning to effectively compensate for missing modalities. Experimental results on two multimodal datasets further demonstrate the superior performance of PMM compared to benchmarks.
Heqiang Wang, Xiang Liu 0004, Xiaoxiong Zhong, Lixing Chen, Fangming Liu, Weizhe Zhang
IEEE Trans. Mob. Comput.5
2026 Denoising and Adaptive Online Vertical Federated Learning for Sequential Multi-Sensor Data in IIoT
abstract
With the advancement of computational capabilities in edge devices such as intelligent sensors in the Industrial Internet of Things (IIoT), these sensors evolving beyond simple data collection to support complex computational tasks. This advancement provides new opportunities for adopting distributed learning approaches in IIoT. In this study, we focus on enhancing learning performance in an industrial assembly line scenario where multiple distributed sensors sequentially collect real-time data with distinct feature spaces. However, existing research lacks an online distributed learning framework tailored for such IIoT settings. To address this gap, we propose the Denoising and Adaptive Online Vertical Federated Learning (DAO-VFL) algorithm, a novel algorithm that leverages the computing potential of edge sensors while addressing key challenges such as communication overhead and data privacy. DAO-VFL effectively manages continuous data streams and adapts to shifting learning objectives. Furthermore, it can address critical challenges prevalent in industrial environment, such as communication noise and heterogeneity of sensor capabilities. To support the proposed algorithm, we provide a comprehensive theoretical analysis, highlighting the effects of noise reduction and adaptive local iteration decisions on the regret bound. Experimental results on two real-world datasets further demonstrate the superior performance of DAO-VFL compared to benchmarks.
Heqiang Wang, Xiaoxiong Zhong, Fangming Liu, Weizhe Zhang
IEEE Trans. Mob. Comput.4
2026 TrimCaching: Parameter-Sharing Edge Caching for AI Model Downloading
abstract
Next-generation mobile networks are expected to facilitate fast AI model downloading to end users. By caching models on edge servers, mobile networks can deliver models to end users with low latency, resulting in a paradigm of edge model caching. In this paper, we develop a novel model placement framework, called parameter-sharing model caching (TrimCaching). TrimCaching exploits the key observation that a wide range of AI models, such as convolutional neural networks or large language models, can share a significant proportion of parameter blocks containing reusable knowledge, thereby improving storage efficiency. To this end, we formulate a parameter-sharing model placement problem to maximize the cache hit ratio in multi-edge wireless networks by balancing the fundamental tradeoff between storage efficiency and service latency. We show that the formulated problem is a submodular maximization problem with submodular constraints, for which no polynomial-time approximation algorithm exists. To tackle this challenge, we study an important special case, where a small fixed number of parameter blocks are shared across models, which often holds in practice. In such a case, a polynomial-time algorithm with a $\left(1-ε\right)/2$-approximation guarantee is developed. Subsequently, we address the original problem for the general case by developing a greedy algorithm. Simulation results demonstrate that the proposed TrimCaching framework significantly improves the cache hit ratio compared with state-of-the-art content caching without exploiting shared parameters in AI models.
Guanqiao Qu, Zheng Lin 0001, Qian Chen 0012, Jian Li 0031, Fangming Liu, Xianhao Chen, Kaibin Huang
IEEE Trans. Netw.5
2025 Knowledge Editing with Dynamic Knowledge Graphs for Multi-Hop Question Answering
abstract
Multi-hop question answering (MHQA) poses a significant challenge for large language models (LLMs) due to the extensive knowledge demands involved. Knowledge editing, which aims to precisely modify the LLMs to incorporate specific knowledge without negatively impacting other unrelated knowledge, offers a potential solution for addressing MHQA challenges with LLMs. However, current solutions struggle to effectively resolve issues of knowledge conflicts. Most parameter-preserving editing methods are hindered by inaccurate retrieval and overlook secondary editing issues, which can introduce noise into the reasoning process of LLMs. In this paper, we introduce KEDKG, a novel knowledge editing method that leverages a dynamic knowledge graph for MHQA, designed to ensure the reliability of answers. KEDKG involves two primary steps: dynamic knowledge graph construction and knowledge graph augmented generation. Initially, KEDKG autonomously constructs a dynamic knowledge graph to store revised information while resolving potential knowledge conflicts. Subsequently, it employs a fine-grained retrieval strategy coupled with an entity and relation detector to enhance the accuracy of graph retrieval for LLM generation. Experimental results on benchmarks show that KEDKG surpasses previous state-of-the-art models, delivering more accurate and reliable answers in environments with dynamic information.
Yigeng Zhou, Jing Li 0034, Yequan Wang, Xuebo Liu 0002, Daojing He, Fangming Liu, Min Zhang 0005
AAAI7
2025 Multi-Modality Expansion and Retention for LLMs through Parameter Merging and Decoupling
abstract
Junlin Li, Guodong Du, Jing Li, Sim Kuan Goh, Wenya Wang, Yequan Wang, Fangming Liu, Ho-Kin Tang, Saleh Alharbi, Daojing He, Min Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Guodong Du 0002, Jing Li 0034, Sim Kuan Goh, Wenya Wang 0001, Yequan Wang, Fangming Liu, Ho-Kin Tang, Saleh Alharbi, Daojing He, Min Zhang 0005
ACL (1)7
2025 Learning to Generate Structured Output with Schema Reinforcement Learning
abstract
Yaxi Lu, Haolun Li, Xin Cong, Zhong Zhang, Yesai Wu, Yankai Lin, Zhiyuan Liu, Fangming Liu, Maosong Sun. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yaxi Lu, Haolun Li 0003, Xin Cong, Zhong Zhang 0004, Yesai Wu, Yankai Lin 0001, Zhiyuan Liu 0001, Fangming Liu, Maosong Sun 0001
ACL (1)8
2025 Demystifying Small Language Models for Edge Deployment
abstract
Small language models (SLMs) have emerged as a promising solution for deploying resource-constrained devices, such as smartphones and Web of Things. This work presents the first comprehensive study of over 60 SLMs such as Microsoft Phi and Google Gemma that are publicly accessible. Our findings show that state-of-the-art SLMs outperform 7B models in general tasks, proving their practical viability. However, SLMs’ in-context learning capabilities remain limited, and their efficiency has significant optimization potential. We identify key SLM optimization opportunities, including dynamic task-specific routing, model-hardware co-design, and vocabulary/KV cache compression. Overall, we expect the work to reveal an all-sided landscape of SLMs, benefiting the research community across algorithm, model, system, and hardware levels.
Zhenyan Lu, Xiang Li 0067, Dongqi Cai 0001, Rongjie Yi, Fangming Liu, Wei Liu 0302, Jian Luan 0001, Nicholas D. Lane, Mengwei Xu 0001
ACL (1)5
2025 Safety Alignment via Constrained Knowledge Unlearning
abstract
Zesheng Shi, Yucheng Zhou, Jing Li, Yuxin Jin, Yu Li, Daojing He, Fangming Liu, Saleh Alharbi, Jun Yu, Min Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Zesheng Shi, Yucheng Zhou 0001, Jing Li 0034, Yu Li 0007, Daojing He, Fangming Liu, Saleh Alharbi, Jun Yu 0002, Min Zhang 0005
ACL (1)7
2025 Impromptu Cybercrime Euphemism Detection
abstract
Detecting euphemisms is essential for content security on various social media platforms, but existing methods designed for detecting euphemisms are ineffective in impromptu euphemisms. In this work, we make a first attempt to an exploration of impromptu euphemism detection and introduce the Impromptu Cybercrime Euphemisms Detection (ICED) dataset. Moreover, we propose a detection framework tailored to this problem, which employs context augmentation modeling and multi-round iterative training. Our detection framework mainly consists of a coarse-grained and a fine-grained classification model. The coarse-grained classification model removes most of the harmless content in the corpus to be detected. The fine-grained model, impromptu euphemisms detector, integrates context augmentation and multi-round iterations training to better predicts the actual meaning of a masked token. In addition, we leverage ChatGPT to evaluate the mode’s capability. Experimental results demonstrate that our approach achieves a remarkable 76-fold improvement compared to the previous state-of-the-art euphemism detector.
Xiang Li 0001, Yucheng Zhou 0001, Laiping Zhao, Jing Li 0034, Fangming Liu
COLING5
2025 RECALL: REpresentation-aligned Catastrophic-forgetting ALLeviation via Hierarchical Model Merging
abstract
Bowen Wang, Haiyuan Wan, Liwen Shi, Chen Yang, Peng He, Yue Ma, Haochen Han, Wenhao Li, Tiao Tan, Yongjian Li, Fangming Liu, Gong Yifan, Sheng Zhang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Haiyuan Wan, Liwen Shi, Haochen Han, Tiao Tan, Fangming Liu, Yifan Gong 0010
EMNLP11
2025 Archer: Adaptive Memory Compression with Page-Association-Rule Awareness for High-Speed Response of Mobile Devices
Changlong Li 0006, Zongwei Zhu, Chao Wang 0003, Fangming Liu, Edwin H.-M. Sha, Xuehai Zhou
FAST4
2025 Unlearning the Noisy Correspondence Makes CLIP More Robust
abstract
The data appetite for Vision-Language Models (VLMs) has continuously scaled up from the early millions to billions today, which faces an untenable trade-off with data quality and inevitably introduces Noisy Correspondence (NC) samples. Undoubtedly, such semantically unrelated data significantly impairs the performance of VLMs. Previous efforts mainly address this challenge by estimating refined alignment for more precise guidance. However, such resource-intensive pipelines that train VLMs from scratch struggle to meet realistic data demands. In this paper, we present a brand new perspective that seeks to directly eliminate the harmful effects of NC in pre-trained VLMs. Specifically, we propose NCU, a Noisy Correspondence Unlearning fine-tuning framework that efficiently enhances VLMs' robustness by forgetting learned noisy knowledge. The key to NCU is learning the hardest negative information, which can provide explicit unlearning direction for both false positives and false negatives. Such twin goals unlearning process can be formalized into one unified optimal transport objective for fast fine-tuning. We validate our approach with the prevailing CLIP model over various downstream tasks. Remarkably, NCU surpasses the robust pre-trained method on zero-shot transfer while with lower computational overhead. The code will be released upon acceptance.
Haochen Han, Alex Jinpeng Wang, Peijun Ye 0002, Fangming Liu
ICCV4
2025 Proactive Agent: Shifting LLM Agents from Reactive Responses to Active Assistance
abstract
Agents powered by large language models have shown remarkable abilities in solving complex tasks. However, most agent systems remain reactive, limiting their effectiveness in scenarios requiring foresight and autonomous decision-making. In this paper, we tackle the challenge of developing proactive agents capable of anticipating and initiating tasks without explicit human instructions. We propose a novel data-driven approach for this problem. Firstly, we collect real-world human activities to generate proactive task predictions. These predictions are then labeled by human annotators as either accepted or rejected. The labeled data is used to train a reward model that simulates human judgment and serves as an automatic evaluator of the proactiveness of LLM agents. Building on this, we develop a comprehensive data generation pipeline to create a diverse dataset, ProactiveBench, containing 6,790 events. Finally, we demonstrate that fine-tuning models with the proposed ProactiveBench can significantly elicit the proactiveness of LLM agents. Experimental results show that our fine-tuned model achieves an F1-Score of 66.47% in proactively offering assistance, outperforming all open-source and close-source models. These results highlight the potential of our method in creating more proactive and effective agent systems, paving the way for future advancements in human-agent collaboration.
Yaxi Lu, Shenzhi Yang, Cheng Qian 0008, Guirong Chen, Qinyu Luo, Yesai Wu, Xin Cong, Zhong Zhang 0004, Yankai Lin 0001, Weiwen Liu, Yasheng Wang, Zhiyuan Liu 0001, Fangming Liu, Maosong Sun 0001
ICLR14
2025 Towards Federated Inference: An Online Model Ensemble Framework for Cooperative Edge AI
Zhi Zhou 0006, Mengke Huang, Tao Ouyang, Fangming Liu, Xu Chen 0004
INFOCOM5
2025 Espresso: Cost-Efficient Large Model Training by Exploiting GPU Heterogeneity in the Cloud
Qiannan Zhou, Fei Xu 0009, Lingxuan Weng, Ruixing Li, Li Chen 0019, Zhi Zhou 0006, Fangming Liu
INFOCOM8
2025 It Takes Two to Tango: Serverless Workflow Serving via Bilaterally Engaged Resource Adaptation
abstract
Serverless platforms typically adopt an earlybinding approach for function sizing, requiring developers to specify an immutable size for each function within a workflow beforehand. Accounting for potential runtime variability, developers must size functions for worst-case scenarios to ensure service-level objectives (SLOs), resulting in significant resource inefficiency. To address this issue, we propose Janus, a novel resource adaptation framework for serverless platforms. Janus employs a late-binding approach, allowing function sizes to be dynamically adapted based on runtime conditions. The main challenge lies in the information barrier between the developer and the provider: developers lack access to runtime information, while providers lack domain knowledge about the workflow. To bridge this gap, Janus allows developers to provide hints containing rules and options for resource adaptation. Providers then follow these hints to dynamically adjust resource allocation at runtime based on real-time function execution information, ensuring compliance with SLOs. We implement Janus and conduct extensive experiments with real-world serverless workflows. Our results demonstrate that Janus enhances resource efficiency by up to 34.7% compared to the state-of-the-art.
Jing Wu 0024, Lin Wang 0015, Quanfeng Deng, Chen Yu 0003, Bingheng Yan, Fangming Liu
IPDPS7
2025 A comparative measurement study of cross-layer 5G performance under different mobility scenarios
Jiahai Hu, Qiangyu Pei, Fangming Liu
Comput. Networks5
2025 mmDigit: A Real-Time Digit Recognition Framework in Air-Writing Using FMCW Radar
abstract
Millimeter-wave (mmWave) radar sensors show significant promise in noncontact human-computer interaction. Using air-writing as a substitute for conventional input devices, such as keyboards and mice has become a pivotal topic in contemporary research. In response to the absence of a dataset in existing studies about air-writing digits and the insufficient exploration of real-time recognition in edge devices, we propose a real-time air-writing digit recognition framework based on mmWave radar, termed mmDigit. Initially, we use mmWave radar equipped with frequency-modulated continuous wave (FMCW) technology to collect digital echo data and design a data processing pipeline to track and reconstruct digital trajectory images. These images are subsequently fed into a lightweight neural network, which is only 6.9K in parameter size, for exploring the images’ quality and the recognition and cross-user capabilities of small-scale air-writing datasets. To enhance mmDigit’s performance, we implement a transfer learning strategy to accommodate a broader range of digit writing styles and habits, achieving a recognition accuracy of 99.14% and a cross-user capability of 94.13%. Additionally, applying a knowledge distillation strategy enables the lightweight network to extract and learn deep-layer features, thereby improving the cross-user recognition accuracy to 96.22%.
Jiake Tian, Yi Zou 0001, Jiale Lai, Fangming Liu
IEEE Internet Things J.4
2025 Toward Transformer-compatible multivariate time series learning via visibility graph-based structural encoding
Ting Chen 0009, Xinyue Ren, Jinzhou Lai, Hongming Tan, Fangming Liu, Wai Kin Chan
Knowl. Based Syst.5
2025 Opara: Exploiting Operator Parallelism for Expediting DNN Inference on GPUs
abstract
GPUs have become thedefactohardware devices for accelerating Deep Neural Network (DNN) inference workloads. However, the conventionalsequential execution mode of DNN operatorsin mainstream deep learning frameworks cannot fully utilize GPU resources, even with the operator fusion enabled, due to the increasing complexity of model structures and a greater diversity of operators. Moreover, theinadequate operator launch orderin parallelized execution scenarios can lead to GPU resource wastage and unexpected performance interference among operators. In this paper, we proposeOpara, a resource- and interference-aware DNNOperatorparallel scheduling framework to accelerate DNN inference on GPUs. Specifically,Oparafirst employsCUDA StreamsandCUDA Graphtoparallelizethe execution of multiple operators automatically. To further expedite DNN inference,Oparaleverages the resource demands of operators to judiciously adjust the operator launch order on GPUs, overlapping the execution of compute-intensive and memory-intensive operators. We implement and open source a prototype ofOparabased on PyTorch in anon-intrusivemanner. Extensive prototype experiments with representative DNN and Transformer-based models demonstrate thatOparaoutperforms the default sequentialCUDA Graphin PyTorch and the state-of-the-art operator parallelism systems by up to$1.68\boldsymbol{\times}$and$1.29\boldsymbol{\times}$, respectively, yet with acceptable runtime overhead.
Aodong Chen, Fei Xu 0009, Li Han 0001, Li Chen 0019, Zhi Zhou 0006, Fangming Liu
IEEE Trans. Computers7
2025 Communication-Efficient Task-Offloading in Mobile Edge Computing System: A Multi-Agent Multi-Armed Bandit Approach
abstract
Mobile devices within the Mobile Edge Computing (MEC) system can offload tasks from the near server to reduce the retrieval latency. Combining collaborative learning for MEC will make progress in both speeding up the convergence of the task-offloading algorithms applied at the end devices side and relieving resources lacking pressure on servers at edge server or cloud server sides. However, stochastic bandit algorithms are incompetent in the presence of complex environments and rarely consider device-to-device communication among users, lacking flexibility under a centralized system. Thus, in this paper, we propose an Upper Confidence Bound (UCB) basedAdaptiveProbabilityUpdating algorithm (APU) and apply it at users side guiding them offloading tasks efficiently. APU could let users manage their own probability-based distributed task offloading strategy and based on it select the most adaptive server. Drawing on a long line of research in network communication and traditional MAB frameworks, we build a multi-agent variant of APU and its corresponding MEC system. Furthermore, we take the regret analysis of the proposed algorithms by rigorous mathematical proof of the sub-linearity convergence, and at the end of this paper, we do some simulation experiments to demonstrate the effectiveness.
Xiaoxiong Zhong, Hangfan Li, Fangming Liu
IEEE Trans. Cloud Comput.5
2025 DeVA: An Edge-Assisted Video Analytics Framework for Depth Estimation
abstract
Edge-assisted video analytics frameworks, which offload vision-based tasks to edge servers, offer a promising approach to enhance accuracy while minimizing network resource overhead. However, these frameworks often overlook depth estimation, a critical task for applications like augmented reality and intelligent surveillance. Depth estimation, which calculates the distance between objects and the camera, generates depth images with unique characteristics, making existing approaches impractical or inefficient for video analytics in this context. In this work, we present DeVA, an edge-assisted video analytics framework for depth estimation that ensures accuracy with minimal network resource overhead. We examine the impact of various video analytics configurations, including resolution and quantization parameter (QP), on accuracy. Additionally, we analyze the region of interest (RoI) for depth estimation and propose methods for tracking RoI areas locally on the device. DeVA features an adaptive video encoding mechanism that dynamically adjusts the resolution for offloaded video and optimizes QPs for RoI and non-RoI areas. We implement DeVA and evaluate its performance using public video datasets. The results show that DeVA reduces 57.12% of the bandwidth overhead while keeping depth estimation errors within acceptable limits, demonstrating a great balance between accuracy and network resource usage.
Jingwen Yin, Ruichao Zhong, Fangming Liu
IEEE Trans. Mob. Comput.4
2025 Working Smarter Not Harder: Hybrid Cooling for Deep Learning in Edge Datacenters
abstract
The proliferation of deep-learning-based mobile and IoT applications has driven the increasing deployment of edge datacenters equipped with domain-specific accelerators. The unprecedented computing power offered by these accelerators puts a heavy burden on the cooling system, motivating more potent cooling techniques like cold water cooling. However, we observe that cold water cooling results in significant energy waste in edge datacenters due to the fluctuating resource utilization both spatially and temporally. To tackle this issue, we propose the concept of “working smarter” by slowing down accelerators deliberately whenever possible and enabling warm water cooling during these times to achieve cooling efficiency. Based on this concept, we develop Hyco—a hybrid water cooling system tailored for edge datacenters running deep learning workloads. First, Hyco features a zone-based cooling architecture enabling dynamic switching between cold water and warm water cooling. Then, based on a lightweight latency estimation method, Hyco incorporates a learning-based scheduling scheme to determine “which” accelerator workers and “when” to slow down through an adaptive and intelligent power-latency trade-off for deep learning models. The simulation with real-world traces shows that Hyco reduces the cooling energy consumption by up to 34.74× while satisfying latency constraints more than 99% of the time for deep-learning-based applications.
Qiangyu Pei, Yongjie Yuan, Haichuan Hu, Lin Wang 0015, Bingheng Yan, Chen Yu 0003, Fangming Liu
IEEE Trans. Sustain. Comput.8
2024 InferCool: Enhancing AI Inference Cooling through Transparent, Non-Intrusive Task Reassignment
abstract
The increasing power consumption of AI inference in modern datacenters has escalated cooling demands significantly, necessitating the adoption of potent cooling approaches like water cooling. Unlike traditional cloud workloads, AI inference has unique characteristics that create substantial gaps in achieving optimal cooling efficiency. In this work, we present the first comprehensive measurement study of AI inference cooling across various models within an industrial-ready scheduling framework, highlighting significant inefficiencies and their causes. To fill the gap while following the fundamental requirements of cooling systems, we explore a new opportunity presented by modern Multi-Instance GPU-enabled inference serving, where the scheduling dimension is naturally orthogonal to the cooling dimension. Building on this insight, we develop InferCool, a cooling middleware designed to enhance cooling efficiency for inference serving through transparent, non-intrusive task reassignment. It includes a streamlined power and temperature prediction approach and a thermal-aware, adaptive application deployment and request scheduling mechanism. Real-world experiments on a water-cooled testbed and a three-node cluster demonstrate that InferCool can reduce the maximum GPU temperature by 5°C across eight A100 GPUs, equivalent to cooling energy savings of about 20%. Importantly, InferCool requires no modifications to existing cooling infrastructures and is compatible with existing scheduling systems.
Qiangyu Pei, Lin Wang 0015, Bingheng Yan, Chen Yu 0003, Fangming Liu
SoCC6
2024 TrimCaching: Parameter-Sharing AI Model Caching in Wireless Edge Networks
abstract
Next-generation mobile networks are expected to facilitate fast AI model downloading to end users. By caching models on edge servers, mobile networks can deliver models to end users with low latency, resulting in a paradigm called edge model caching. In this paper, we develop a novel model placement scheme, called parameter-sharing model caching (TrimCaching). TrimCaching exploits the key observation that a wide range of AI models, such as convolutional neural networks or large language models, can share a significant proportion of parameter blocks containing reusable knowledge, thereby improving storage efficiency. To this end, we formulate a parameter-sharing model placement problem to maximize the cache hit ratio in multi-edge wireless networks by balancing the fundamental tradeoff between storage efficiency and service latency. We show that the formulated problem is a submodular maximization problem with submodular constraints, for which no polynomial-time approximation algorithm exists. To overcome this challenge, we study an important special case, where a small fixed number of parameter blocks are shared across models, which often holds in practice. In such a case, a polynomial-time algorithm with (1 - E) /2-approximation guarantee is developed. Subsequently, we address the original problem for the general case by developing a greedy algorithm. Simulation results demonstrate that the proposed TrimCaching framework significantly improves the cache hit ratio compared with state-of-the-art content caching without exploiting shared parameters in AI models.
Guanqiao Qu, Zheng Lin 0001, Fangming Liu, Xianhao Chen, Kaibin Huang
ICDCS3
2024 X-Stream: A Flexible, Adaptive Video Transformer for Privacy-Preserving Video Stream Analytics
abstract
Video stream analytics (VSA) systems fuel many exciting applications that facilitate people’s lives, but also raise critical concerns about exposing too much individuals’ privacy. To alleviate these concerns, various frameworks have been presented to enhance the privacy of VSA systems. Yet, existing solutions suffer two limitations: (1) being scenario-customized, thus limiting the generality of adapting to multifarious scenarios, (2) requiring complex, imperative programming, and tedious process, thus largely reducing the usability of such systems. In this paper, we present X-Stream, a privacy-preserving video transformer that achieves flexibility and efficiency for a large variety of VSA tasks. X-Stream features three major novel designs: (1) a declarative query interface that provides a simple yet expressive interface for users to describe both their privacy protection and content exposure requirements, (2) an adaptation mechanism that dynamically selects the most suitable privacy-preserving techniques and their parameters based on the current video context, and (3) an efficient execution engine that incorporates optimizations for multi-task deduplication and inter-frame inference. We implement X-Stream and evaluate it with representative VSA tasks and public video datasets. The results show that X-Stream achieves significantly improved privacy protection quality and performance over the state-of-the-art, while being simple to use.
Dou Feng, Lin Wang 0015, Lingching Tung, Fangming Liu
INFOCOM5
2024 HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
abstract
Deep Neural Network (DNN) inference on serverless functions is gaining prominence due to its potential for substantial budget savings. Existing works on serverless DNN inference solely optimize batching requests from one application with a single Service Level Objective (SLO) on CPU functions. However, production serverless DNN inference traces indicate that the request arrival rate of applications is surprisingly low, which inevitably causes a long batching time and SLO violations. Hence, there is an urgent need for batching multiple DNN inference requests with diverse SLOs (i.e., multi-SLO DNN inference) in serverless platforms. Moreover, the potential performance and cost benefits of deploying heterogeneous (i.e., CPU and GPU) functions for DNN inference have received scant attention.In this paper, we present HarmonyBatch, a cost-efficient resource provisioning framework designed to achieve predictable performance for multi-SLO DNN inference with heterogeneous serverless functions. Specifically, we construct an analytical performance and cost model of DNN inference on both CPU and GPU functions, by explicitly considering the GPU time-slicing scheduling mechanism and request arrival rate distribution. Based on such a model, we devise a two-stage merging strategy in HarmonyBatch to judiciously batch the multi-SLO DNN inference requests into application groups. It aims to minimize the budget of function provisioning for each application group while guaranteeing diverse performance SLOs of inference applications. We have implemented a prototype of HarmonyBatch on Alibaba Cloud Function Compute. Extensive prototype experiments with representative DNN inference workloads demonstrate that HarmonyBatch can provide predictable performance to serverless DNN inference workloads while reducing the monetary cost by up to 82.9% compared to the state-of-the-art methods.
Fei Xu 0009, Yikun Gu, Li Chen 0019, Fangming Liu, Zhi Zhou 0006
IWQoS5
2024 An MTD-driven Hybrid Defense Method Against DDoS Based on Markov Game in Multi-controller SDN-enabled IoT Networks
abstract
The widespread deployment of low-cost, vulnerable IoT devices allows attackers to exploit them to generate botnets and launch distributed denial-of-service (DDoS) attacks, which has become a serious security challenge for ensuring quality of service (QoS). For cost-effective defense against DDoS, we propose a novel hybrid defense method that includes proactive moving target defense (MTD) and passive security control to resist DDoS threats at different stages in IoT networks in this paper. We construct a multi-stage Markov game model to portray the game as a competition between the attacker and the defender for the control duration of the attack surface, and design an optimal defense strategy algorithm. In particular, we introduce a new parameter of action execution interval expectation in the game and add node importance evaluation in the reward quantification so that the optimal action execution interval of each defense technique can be output. We also consider the possibility that advanced attackers may launch DDoS on the SDN controller in the game. The experimental results demonstrate that our proposed method can defend against DDoS cost-effectively and ensure the QoS in IoT networks with acceptable overhead.
Yuming Feng 0002, Weizhe Zhang, Zijun Feng, Xiaoxiong Zhong, Fangming Liu
IWQoS5
2024 YuanRong: A Production General-purpose Serverless System for Distributed Applications in the Cloud
abstract
We design, implement, and evaluate YuanRong, the first production general-purpose serverless platform with a unified programming interface, multi-language runtime, and a distributed computing kernel for cloud-based applications. YuanRong addresses many limitations of existing Function-as-a-Service (FaaS) systems, particularly in performance and lack of important features. First, our fast function system supports sub-millisecond function invocation and locality-aware hierarchical scheduling. Second, our multi-semantic built-in data system achieves object exchange latency of 200 microseconds, enabling end-to-end latency of 2 milliseconds for streaming elements at 5Gbps throughput. Third, the extensible, portable Service Bridge bridges stateless and stateful operations, allowing connection reuse and distributed transactions, and offers unified backend abstraction for multi-cloud portability. YuanRong has been deployed for over 3 years at Huawei across nearly 20 datacenter regions, processing up to 30 billion requests per day on more than 100,000 CPU cores, with a daily average CPU usage of 53%. It serves various serverless workloads, including enhanced FaaS services, microservices, data analytics, deep model training and serving, and some HPC workloads. Our experience shows that Spring-based microservices can be migrated to YuanRong within one day, reducing resource costs by 90%, demonstrating its generality and efficiency in supporting a broad spectrum of applications.
Jianmin Qian, Yulin Che, Licheng Song, Jie Wu 0042, Wei Liu 0248, Fangming Liu, Kun Tan 0003
SIGCOMM13
2024 mmHPE: Human Pose Estimation Based on Point Cloud from Millimeter-wave Radar
abstract
Rehabilitation therapy involving repetitive exer-cises targeting specific human joints under the supervision of a doctor is crucial for patients with movement disorders. Yet the cost of commuting and the demand for medical resources are inconvenient for patients. Human-computer interaction can provide remote rehabilitation guidance for patients at home through human pose estimation technology, but privacy concerns with optical sensors and the cost and discomfort of wearable sensors have hindered progress in this field. To address the above challenges, we propose mmHPE, an innovative 3D human pose estimation framework that uses millimeter-wave (mmWave) radar sensors. It initially manipulates the raw data captured from radar sensors to generate a spatiotemporal sequence point cloud dataset. Afterward, we create a Convolutional Neural Network (CNN) that is linked to a Bidirectional Long Short-Term Memory (Bi-LSTM) network. Moreover, a multi-head attention mechanism is employed to boost the network's performance and to accurately estimate the locations of human skeletons. Ultimately, the 21 points with the corre-sponding human pose position are successfully reconstructed. We investigate the mmHPE framework's feasibility and cross-domain stability in different home environments in real-world scenarios. This innovation proffers patients a convenient and privacy-conscious solution for their rehabilitation training req-uisites.
Jiale Lai, Jiake Tian, Yi Zou 0001, Xianfeng Song, Fangming Liu, Dacheng Li
SMC5
2024 λGrapher: A Resource-Efficient Serverless System for GNN Serving through Graph Sharing
abstract
Graph Neural Networks (GNNs) have been increasingly adopted for graph analysis in web applications such as social networks. Yet, efficient GNN serving remains a critical challenge due to high workload fluctuations and intricate GNN operations. Serverless computing, thanks to its flexibility and agility, offers on-demand serving of GNN inference requests. Alas, the request-centric serverless model is still too coarse-grained to avoid resource waste.
Haichuan Hu, Fangming Liu, Qiangyu Pei, Yongjie Yuan, Zichen Xu 0001, Lin Wang 0015
WWW2
2024 Demystifying the Cost of Serverless Computing: Towards a Win-Win Deal
abstract
Serverless is an emerging computing paradigm that greatly simplifies the development, deployment, and maintenance of cloud applications. However, due to potential cost issues brought by the widely adopted pricing, it is difficult to answer how to use and operate serverless computing services from the perspectives of users and providers. To demystify the cost of serverless computing, we present one of the first studies that develops an analytical model for serverless cost from the perspectives of users and providers, by comparing it to Infrastructure-as-a-Service. Based on the model, driven by real-world traces, extensive simulation results verify the following cost issues: 1) For the users, serverless is not always cost-saving, even possibly leading to expense explosion; 2) The serverless providers are in urgent need of widening use scenarios to improve resource utilization and raise revenue; 3) The prevailing pricing fails to neither reduce the risk of expense explosion nor meet the need of attracting more workloads. To remove the cost barrier, we proposefuture function, auction-based pricing for serverless, to offer discounts to the users as well as boost profit for the providers. Experimental results show the duration price of functions can be reduced by 57.5% on average for 13.5% of users yet without harming the revenue of providers.
Fangming Liu, Yipei Niu
IEEE Trans. Parallel Distributed Syst.1
2024 DyLaClass: Dynamic Labeling Based Classification for Optimal Sparse Matrix Format Selection in Accelerating SpMV
abstract
Sparse matrix-vector multiplication (SpMV) is crucial in many scientific and engineering applications, particularly concerning the effectiveness of different sparse matrix storage formats for various architectures, no single format excels across all hardware. Previous research has focused on trying different algorithms to build predictors for the best format, yet it overlooked how to address the issue of the best format changing in the same hardware environment and how to reduce prediction overhead rather than merely considering the overhead in building predictors. This paper proposes a novel classification algorithm for optimizing sparse matrix storage formats, DyLaClass, based on dynamic labeling and flexible feature selection. Particularly, we introduce mixed labels and features with strong correlations, allowing us to achieve ultra-high prediction accuracy with minimal feature inputs, significantly reducing feature extraction overhead. For the first time, we propose the concept of the most suitable storage format rather than the best storage format, which can stably predict changes in the best format for the same matrix across multiple SpMV executions. We further demonstrate the proposed method on the University of Florida’s public sparse matrix collection dataset. Experimental results show that compared to existing work, our method achieves up to 91% classification accuracy. Using two different hardware platforms for verification, the proposed method outperforms existing methods by 1.26 to 1.43 times. Most importantly, the stability of the proposed prediction model is 25.5% higher than previous methods, greatly increasing the feasibility of the model in practical field applications.
Yi Zou 0001, Xianfeng Song, Fangming Liu, Quan Xue
IEEE Trans. Parallel Distributed Syst.5
2024 ComboFunc: Joint Resource Combination and Container Placement for Serverless Function Scaling With Heterogeneous Container
abstract
Serverless computing provides developers with a maintenance-free approach to resource usage, but it also transfers resource management responsibility to the cloud platform. However, the fine granularity of serverless function resources can lead to performance bottlenecks and resource fragmentation on nodes when creating many function containers. This poses challenges in effectively scaling function resources and optimizing node resource allocation, hindering overall agility. To address these challenges, we have introduced ComboFunc, an innovative resource scaling system for serverless platforms. ComboFunc associates function with heterogeneous containers of varying specifications and optimizes their resource combination and placement. This approach not only selects appropriate nodes for container creation, but also leverages the new feature of Kubernetes In-place Pod Vertical Scaling to enhance resource scaling agility and efficiency. By allowing a single function to correspond to heterogeneous containers with varying resource specifications and providing the ability to modify the resource specifications of existing containers in place, ComboFunc effectively utilizes fragmented resources on nodes. This, in turn, enhances the overall resource utilization of the entire cluster and improves scaling agility. We also model the problem of combining and placing heterogeneous containers as an NP-hard problem and design a heuristic solution based on a greedy algorithm that solves it in polynomial time. We implemented a prototype of ComboFunc on the Kubernetes platform and conducted experiments using real traces on a local cluster. The results demonstrate that, compared to existing strategies, ComboFunc achieves up to 3.01 × faster function resource scaling and reduces resource costs by up to 42.6%.
Zhaojie Wen, Quanfeng Deng, Yipei Niu, Fangming Liu
IEEE Trans. Parallel Distributed Syst.6
2024 Joint Optimization of Parallelism and Resource Configuration for Serverless Function Steps
abstract
Function-as-a-Service (FaaS) offers a fine-grained resource provision model, enabling developers to build highly elastic cloud applications. User requests are handled by a series of serverless functions step by step, which forms a multi-step workflow. The developers are required to set proper configurations for functions to meet service level objectives (SLOs) and save costs. However, developing the configuration strategy is challenging. This is mainly because the execution of serverless functions often suffers from cold starts and performance fluctuation, which requires a dynamic configuration strategy to guarantee the SLOs. In this article, we present StepConf, a framework that automates the configuration as the workflow runs. StepConf optimizes memory size for each function step in the workflow and takes inter and intra-function parallelism into consideration, which has been overlooked by existing work. StepConf intelligently predicts the potential configurations for subsequent function steps, and proactively prewarms function instances in a configuration-aware manner to reduce the cold start overheads. We evaluate StepConf on AWS and Knative. Compared to existing work, StepConf improves performance by up to 5.6× under the same cost budget and achieves up to a 40% cost reduction while maintaining the same level of performance.
Zhaojie Wen, Yipei Niu, Quanfeng Deng, Fangming Liu
IEEE Trans. Parallel Distributed Syst.6
2024 Graft: Efficient Inference Serving for Hybrid Deep Learning With SLO Guarantees via DNN Re-Alignment
abstract
Deep neural networks (DNNs) have been widely adopted for various mobile inference tasks, yet their ever-increasing computational demands are hindering their deployment on resource-constrained mobile devices. Hybrid deep learning partitions a DNN into two parts and deploys them across the mobile device and a server, aiming to reduce inference latency or prolong battery life of mobile devices. However, such partitioning produces (non-uniform) DNN fragments which are hard to serve efficiently on the server. This article presents Graft—an efficient inference serving system for hybrid deep learning with latency service-level objective (SLO) guarantees. Our main insight is to mitigate the non-uniformity by a core concept called DNN re-alignment, allowing multiple heterogeneous DNN fragments to be restructured to share layers. To fully exploit the potential of DNN re-alignment, Graft employs fine-grained GPU resource sharing. Based on that, we propose efficient algorithms for merging, grouping, and re-aligning DNN fragments to maximize request batching opportunities, minimizing resource consumption while guaranteeing the inference latency SLO. We implement a Graft prototype and perform extensive experiments with five types of widely used DNNs and real-world network traces. Our results show that Graft improves resource efficiency by up to 70% compared with the state-of-the-art inference serving systems.
Jing Wu 0024, Lin Wang 0015, Qirui Jin, Fangming Liu
IEEE Trans. Parallel Distributed Syst.4
2024 Tetris: Proactive Container Scheduling for Long-Term Load Balancing in Shared Clusters
abstract
Long-running containerized workloads (e.g., machine learning), which typically showtime-varyingpatterns, are increasingly prevailing in shared production clusters. To improve workload performance, current schedulers mainly focus on optimizingshort-termbenefits of cluster load balancing orinitial container placementon servers. However, this would inevitably bring manyinvalid migrations(i.e., containers are migrated back and forth among servers over a short time window), leading to significant service level objective (SLO) violations. This paper introducesTetris, amodel predictive control(MPC)-based container scheduling strategy to proactively migrate long-running workloads for cluster load balancing. Specifically, we first build a discrete-time dynamic model forlong-termoptimization of container scheduling. To solve such an optimization problem,Tetristhen employs two main components: (1) a container resource predictor, which leverages time-series analysis approaches to accurately predict the container resource consumption; (2) an MPC-based container scheduler that jointly optimizes the cluster load balancing and container migration costover a certain sliding time window. We implement and open source a prototype ofTetrisbased on K8s. Extensive prototype experiments and trace-driven simulations demonstrate thatTetriscan improve the cluster load balancing degree by up to 77.8% without incurring any SLO violations, compared to the state-of-the-art container scheduling strategies.
Fei Xu 0009, Xiyue Shen, Shuohao Lin, Li Chen 0019, Zhi Zhou 0006, Fen Xiao, Fangming Liu
IEEE Trans. Serv. Comput.7
2023 AsyFunc: A High-Performance and Resource-Efficient Serverless Inference System via Asymmetric Functions
abstract
Recent advances in deep learning (DL) have spawned various intelligent cloud services with well-trained DL models. Nevertheless, it is nontrivial to maintain the desired end-to-end latency under bursty workloads, raising critical challenges on high-performance while resource-efficient inference services. To handle burstiness, some inference services have migrated to the serverless paradigm for its rapid elasticity. However, they neglect the impact of the time-consuming and resource-hungry model-loading process when scaling out function instances, leading to considerable resource inefficiency for maintaining high performance under burstiness.
Qiangyu Pei, Yongjie Yuan, Haichuan Hu, Fangming Liu
SoCC5
2023 New Problems in Active Sampling for Mobile Robotic Online Learning
abstract
AI models deployed in real-world tasks (e.g., surveillance, implicit mapping, health care) typically need to be online trained for better modelling of the changing real-world environments and various online training methods (e.g., domain adaptation, few shot learning) are proposed for refining the AI models based on training input incrementally sampled from the real world. However, in the whole loop of AI model online training, there is a section rarely discussed: how to sample training input from the real world. In this paper, we show from the perspective of online training of AI models deployed on edge devices (e.g., robots) that several problems in sampling of training input on the device are affecting the time and energy consumption for the online training process to reach high performance. Notably, the online training relies on training input consecutively sampled from the real world and the consecutive samples from nearby states (e.g., position and orientation of a camera) are too similar and would limit the training accuracy gain per training iteration; on the other hand, while we can choose to sample more about the inaccurate samples to better final training accuracy, it is costly to obtain the accuracy statistics of samples via traditional ways such as validating, especially for AI models deployed on edge devices. These findings aim to raise research effort for practical online training of AI models, so that they can achieve resiliently and sustainably high performance in real-world tasks.
Xiuxian Guan, Junming Wang 0001, Zekai Sun, Zongyuan Zhang, Tianyang Duan, Shengliang Deng, Fangming Liu, Heming Cui
COMPSAC7
2023 Fast and Scalable Gate-Level Simulation in Massively Parallel Systems
abstract
The natural bijection between a proposed circuit design and its graph representation shall allow any graph optimization algorithm deploying into many-core systems efficiently. However, this process suffers from the exponentially growing overhead and heavy memory footprint with the signal propagation. To conquer the unique challenge, we systematically study the simulation with millions of gates, and identify that the processing complexity could grow exponentially from the signal inputs, the skewness of the computational graph stays. Thus, we present ZhouBi, a fast and scalable gate-level simulation framework to fully exploit the parallelism from many-core systems. ZhouBi contributes in threefolds, (I) a graph representation that colors gate-level netlists and identifies skew partitions based on the graph skewness; (II) A set of heuristic algorithms that picks opportunistic and conservative algorithms to accelerate the simulation; (III) A system facility that supports selective mapping between simulation and many-core, providing a tradeoff between the risk of concurrent simulation fail and performance gain. We have prototyped ZhouBi and evaluated it with practical baselines. ZhouBi can achieve a 27.6× performance gain, as compared to the state-of-the-practice Veriwell without compromising any correctness. Our framework supports large graphs enabling scale-out gate-level simulations for chip design.
Haichuan Hu, Zichen Xu 0001, Yuhao Wang 0001, Fangming Liu
ICCAD4
2023 Two-Stage Coded Distributed Learning: A Dynamic Partial Gradient Coding Perspective
abstract
Distributed learning has been widely adopted to train a global model from local data. However, its performance can be severely affected by stragglers. Recently, some research has been dedicated to resolving the straggler problem by adopting gradient coding, the essence of gradient coding is to solve the straggler problem by adding data redundancy. However, the large amount of data redundancy as well as computation and communication overhead that it brings is still hard to be resolved. Besides, the complexity of the encoding and decoding will increase linearly with the number of the local workers. To this end, in this paper, we design a lightweight coding method in the computing phase and seek to ensure fair transmission in the communication phase. Specifically, to tolerate stragglers in computing phase, we propose a two-stage dynamic coding scheme, part of the workers start computing the partial gradients from the data partitions assigned in the first stage, and the remaining workers for computation in the second stage is decided based on which workers have finished in the first stage. To further tolerate stragglers in the communication phase, a perturbed Lyapunov function is designed to maximize admission data balancing fairness as well as the throughput. The experimental result verifies the derived properties and demonstrates that our proposed solution can achieve a better performance for practical network parameters and benchmark data in terms of accuracy and resource utilization in the distributed learning system.
Xinghan Wang 0001, Xiaoxiong Zhong, Jiahong Ning, Tingting Yang 0001, Yuanyuan Yang 0001, Guoming Tang, Fangming Liu
ICDCS7
2023 spotDNN: Provisioning Spot Instances for Predictable Distributed DNN Training in the Cloud
abstract
Distributed Deep Neural Network (DDNN) training on cloud spot instances is increasingly compelling as it can significantly save the user budget. To handle unexpected instance revocations, provisioning a heterogeneous cluster using the asynchronous parallel mechanism becomes the dominant method for DDNN training with spot instances. However, blindly provisioning a cluster of spot instances can easily result in unpre-dictable DDNN training performance, mainly because bottlenecks occur on the parameter server network bandwidth and PCIe bandwidth resources, as well as the inadequate cluster heterogeneity. To address the challenges above, we propose spotDNN, a heterogeneity-aware spot instance provisioning framework that provides predictable performance for DDNN training in the cloud. By explicitly considering the contention for bottle-neck resources, we first build an analytical performance model of DDNN training in heterogeneous clusters. It leverages the weighted average batch size and convergence coefficient to quantify the DDNN training loss in heterogeneous clusters. Through a lightweight workload profiling, we further design a cost-efficient instance provisioning strategy which incorporates the bounds calculation and sliding window techniques to effectively guarantee the training performance service level objectives (SLOs). We have implemented a prototype of spotDNN and conducted extensive experiments on Amazon EC2. Experiment results show that spotDNN can deliver predictable DDNN training performance while reducing the monetary cost by up to 68.1% compared to the existing solutions, yet with acceptable runtime overhead.
Ruitao Shang, Fei Xu 0009, Zhuoyan Bai, Li Chen 0019, Zhi Zhou 0006, Fangming Liu
IWQoS6
2023 Online MEC Offloading for V2V Networks
abstract
As an enabling technology for vehicle-to-vehicle (V2V) networks, multi-access edge computing (MEC) provides a feasible platform for sharing power and resources, and offloading some of the computation-intensive tasks between vehicles. This, however, is challenging with the unpredictable variations in road traffic conditions and vehicle mobility in MEC-enabled V2V networks. Consequently, such computation task offloading can be easily disrupted, which may require frequent switching of task offloading between vehicles and degrade the Quality of Service (QoS). In this paper, we focus on the computation offloading problem under unstable connections in MEC-enabled V2V networks. We first model this as a distributed online service optimization problem, which is proved to be NP-hard. In order to minimize the out-of-service time (i.e., the service mismatching, switching and compromise time), we propose a distributed Online Instability-aware Computation Offloading (OICO) heuristic algorithm to improve the service efficiency and quality. Specifically, in order to minimize the service mismatching rate, we design an efficient Service Path Matching (SPM) algorithm for matching pairs of customer vehicles (which require offload computing services) and server vehicles (which provide edge computing services) that share the longest matching path. We evaluate OICO through real-world traces, i.e., GAIA open dataset from DiDi. Extensive simulation results demonstrate that OICO can increase the service matching rate by 25% and reduce the power consumption by about 54% per customer vehicle compared with the existing schemes.
Fangming Liu, Qixia Zhang, Bo Li 0001
IEEE Trans. Mob. Comput.1
2023 iGniter: Interference-Aware GPU Resource Provisioning for Predictable DNN Inference in the Cloud
abstract
GPUs are essential to accelerating the latency-sensitive deep neural network (DNN) inference workloads in cloud datacenters. To fully utilize GPU resources,spatial sharingof GPUs among co-located DNN inference workloads becomes increasingly compelling. However, GPU sharing inevitably bringssevere performance interferenceamong co-located inference workloads, as motivated by an empirical measurement study of DNN inference on EC2 GPU instances. While existing works on guaranteeing inference performance service level objectives (SLOs) focus on eithertemporal sharingof GPUs orreactiveGPU resource scaling and inference migration techniques, how toproactivelymitigate such severe performance interference has received comparatively little attention. In this paper, we proposeiGniter, aninterference-awareGPU resource provisioning framework for cost-efficiently achieving predictable DNN inference in the cloud.iGniteris comprised of two key components: (1) alightweightDNN inference performance model, which leverages the system and workload metrics that are practically accessible to capture the performance interference; (2) Acost-efficientGPU resource provisioning strategy thatjointlyoptimizes the GPU resource allocation and adaptive batching based on our inference performance model, with the aim of achieving predictable performance of DNN inference workloads. We implement a prototype ofiGniterbased on the NVIDIA Triton inference server hosted on EC2 GPU instances. Extensive prototype experiments on four representative DNN models and datasets demonstrate thatiGnitercan guarantee the performance SLOs of DNN inference workloads with practically acceptable runtime overhead, while saving the monetary cost by up to$25\%$in comparison to the state-of-the-art GPU resource provisioning strategies.
Fei Xu 0009, Jianian Xu, Li Chen 0019, Ruitao Shang, Zhi Zhou 0006, Fangming Liu
IEEE Trans. Parallel Distributed Syst.7
2022 CoolEdge: hotspot-relievable warm water cooling for energy-efficient edge datacenters
abstract
As the computing frontier drifts to the edge, edge datacenters play a crucial role in supporting various real-time applications. Different from cloud datacenters, the requirements of proximity to end-users, high density, and heterogeneity, present new challenges to cool the edge datacenters efficiently. Although warm water cooling has become a promising cooling technique for this infrastructure, the one-size-fits-all cooling control would lower the cooling efficiency considerably because of the severe thermal imbalance across servers, hardware, and even inside one hardware component in an edge datacenter. In this work, we propose CoolEdge, a hotspot-relievable warm water cooling system for improving the cooling efficiency and saving costs of edge datacenters. Specifically, through the elaborate design of water circulations, CoolEdge can dynamically adjust the water temperature and flow rate for each heterogeneous hardware component to eliminate the hardware-level hotspots. By redesigning cold plates, CoolEdge can quickly disperse the chip-level hotspots without manual intervention. We further quantify the power saving achieved by the warm water cooling theoretically, and propose a custom-designed cooling solution to decide an appropriate water temperature and flow rate periodically. Based on a hardware prototype and real-world traces from SURFsara, the evaluation results show that CoolEdge reduces the cooling energy by 81.81% and 71.92%, respectively, compared with conventional and state-of-the-art water cooling systems.
Qiangyu Pei, Qixia Zhang, Fangming Liu, Ziyang Jia, Yishuo Wang, Yongjie Yuan
ASPLOS5
2022 Optimal Admission Control Mechanism Design for Time-Sensitive Services in Edge Computing
abstract
Edge computing is a promising solution for reducing service latency by provisioning time-sensitive services directly from the network edge. However, upon workload peaks at the resource-limited edge, an edge service has to queue service requests, incurring high waiting time. Such quality of service (QoS) degradation ruins the reputation and reduces the long-term revenue of the service provider.To address this issue, we propose an admission control mechanism for time-sensitive edge services. Specifically, we allow the service provider to offer admission advice to arriving requests regarding whether to join for service or balk to seek alternatives. Our goal is twofold: maximizing revenue of the service provider and ensuring QoS if the provided admission advice is followed. To this end, we propose a threshold structure that estimates the highest length of the request queue. Leveraging such a threshold structure, we propose O2A, a mechanism to balance the trade-off between increasing revenue from accepting more requests and guaranteeing QoS by advising requests to balk. Rigorous analysis shows that O2A achieves the goal and that the provided admission advice is optimal for end-users to follow. We further validate O2A through trace-driven simulations with both synthetic and real-world service request traces.
Lin Wang 0015, Fangming Liu
INFOCOM3
2022 Retention-Aware Container Caching for Serverless Edge Computing
abstract
Serverless edge computing adopts an event-based model where Internet-of-Things (IoT) services are executed in lightweight containers only when requested, leading to significantly improved edge resource utilization. Unfortunately, the startup latency of containers degrades the responsiveness of IoT services dramatically. Container caching, while masking this latency, requires retaining resources thus compromising resource efficiency. In this paper, we study the retention-aware container caching problem in serverless edge computing. We leverage the distributed and heterogeneous nature of edge platforms and propose to optimize container caching jointly with request distribution. We reveal step by step that this joint optimization problem can be mapped to the classic ski-rental problem. We first present an online competitive algorithm for a special case where request distribution and container caching are based on a set of carefully designed probability distribution functions. Based on this algorithm, we propose an online algorithm called O-RDC for the general case, which incorporates the resource capacity and network latency by opportunistically distributing requests. We conduct extensive experiments to examine the performance of the proposed algorithms with both synthetic and real-world serverless computing traces. Our results show that ORDC outperforms existing caching strategies of current serverless computing platforms by up to 94.5% in terms of the overall system cost.
Lin Wang 0015, Fangming Liu
INFOCOM4
2022 StepConf: SLO-Aware Dynamic Resource Configuration for Serverless Function Workflows
abstract
Function-as-a-Service (FaaS) offers a fine-grained resource provision model, enabling developers to build highly elastic cloud applications. User requests are handled by a series of serverless functions step by step, which forms a function-based workflow. The developers are required to set proper resource configuration for functions, so as to meet service level objectives (SLOs) and save cost. However, developing the resource configuration strategy is challenging. It is mainly because execution of cloud functions often suffers from cold start and performance fluctuation, which requires a dynamic configuration strategy to guarantee the SLOs. In this paper, we present StepConf, a framework that automates the resource configuration for functions as the workflow runs. StepConf optimizes memory size for each function step in the workflow and takes inter and intra-function parallelism into consideration. We evaluate StepConf on AWS Lambda. Compared with baselines, the experimental results show that StepConf can save cost up to 40.9% while ensuring the SLOs.
Zhaojie Wen, Yishuo Wang, Fangming Liu
INFOCOM3
2022 λDNN: Achieving Predictable Distributed DNN Training With Serverless Architectures
abstract
Serverless computing is becoming a promising paradigm for Distributed Deep Neural Network (DDNN) training in the cloud, as it allows users to decompose complex model training into a number offunctionswithout managing virtual machines or servers. Though provided with a simpler resource interface (i.e., function number and memory size), inadequate function resource provisioning (either under-provisioning or over-provisioning) easily leads tounpredictableDDNN training performance in serverless platforms. Our empirical studies on AWS Lambda indicate that, suchunpredictable performanceof serverless DDNN training is mainly caused by the resource bottleneck of Parameter Servers (PS) and small local batch size. In this article, we design and implement$\lambda$λDNN, a cost-efficient function resource provisioning framework to provide predictable performance for serverless DDNN training workloads, while saving the budget of provisioned functions. Leveraging the PS network bandwidth and function CPU utilization, we build alightweightanalytical DDNN training performance model to enable our design of$\lambda$λDNNresource provisioning strategy, so as to guarantee DDNN training performance with serverless functions. Extensive prototype experiments on AWS Lambda and complementary trace-driven simulations demonstrate that,$\lambda$λDNNcan deliver predictable DDNN training performance and save the monetary cost of function resources by up to 66.7 percent, compared with the state-of-the-art resource provisioning strategies, yet with an acceptable runtime overhead.
Fei Xu 0009, Yiling Qin, Li Chen 0019, Zhi Zhou 0006, Fangming Liu
IEEE Trans. Computers5
2022 MarVeLScaler: A Multi-View Learning-Based Auto-Scaling System for MapReduce
abstract
To promote cloud computing from current pay-per-request model to truly pay-per-use, tenants are crying for automatic tools to auto-estimate the amount of resources for MapReduce jobs. Such tools call for accurately quantifying the relationship among workload, resources and completion time. Various prediction models have been proposed. However, none of these models takes virtual machines’ (VMs) performance variance during a job's execution into consideration, leading to underestimate the required resources and exceed the job's deadline. To address this problem, we propose a multi-view deep learning model to capture real-time performance variance and automatically scale out the cloud cluster whenever necessary. We implementMarVeLScaler, a prototype system including two useful modules, namely,Scale EstimatorandScale Controller. Scale Estimator preliminarily estimates the required cluster size for a MapReduce job with given a concrete workload and deadline. During the runtime, Scale Controller adjusts the scale of the cluster according to its real-time running status to guarantee the job finished on time. We evaluate the performance of MarVeLScaler based on Hadoop in Alibaba Cloud. Experiments show that MarVeLScaler can provide 98.4 percent accuracy of prediction in determining initial cluster size, and save 30.8 percent of expense while still guaranteeing similar performance compared with the state-of-the-art methods.
Fangming Liu, Yibing Sheng, Miao Zhao, Jianping Wang 0001
IEEE Trans. Cloud Comput.2
2022 When FPGA Meets Cloud: A First Look at Performance
abstract
Cloud service providers promote their new field programmable gate array (FPGA) infrastructure as a service (IaaS) as the new era of cloud product. This FPGA IaaS wraps virtualized compute resources with FPGA boards, e.g., Amazon AWS F1, and reserves acceleration capability for specific applications. Though this acceleration technique sounds promising, questions like real world performance, best-fit scenarios, portability, etc., still need further clarification. In this article, we present one of the first few empirical studies that take a close look at FPGA clouds from the tenants’ perspective. We have conducted measurement studies on Amazon AWS, Alibaba, and Huawei clouds for over one year. The experimental results show that: (1) Tenants experience severe performance-cost imbalance on FPGA IaaS platforms; (2) The inter-communication performance in FPGA clouds is tightly constrained by hardware drivers, e.g., small optimization of DMA drivers for PCIe can harvest significant performance gain; (3) The virtualized FPGA clouds are far from mature, e.g., small-sized jobs can greatly degrade the performance of FPGA clouds due to underutilized PCIe bandwidth. Our study not only provides useful hints to help tenants with FPGA service selection, but also sheds some lights for cloud providers to improve the performance of FPGA clouds.
Xiuxiu Wang, Yipei Niu, Fangming Liu, Zichen Xu 0001
IEEE Trans. Cloud Comput.3
2022 An Online Framework for Joint Network Selection and Service Placement in Mobile Edge Computing
abstract
With the rapid development and deployment of 5G wireless technology, mobile edge computing (MEC) has emerged as a new computing paradigm to facilitate a large variety of infrastructures at the network edge to reduce user-perceived communication delay. One of the fundamental problems in this new paradigm is to preserve satisfactory quality-of-service (QoS) for mobile users in light of densely dispersed wireless communication environment and often capacity-constrained MEC nodes. Such user-perceived QoS, typically in terms of the end-to-end delay, is highly vulnerable to both access network bottleneck and communication delay. Previous works have primarily focused on optimizing the communication delay through dynamic service placement, while ignoring the critical effect of access network selection on the access delay. In this work, we study the problem of jointly optimizing the access network selection and service placement for MEC, with the objective of improving the QoS in a cost-efficient manner by judiciously balancing the access delay, communication delay, and service switching cost. Specifically, we propose an efficient online framework to decompose a long-term time-varying optimization problem into a series of one-shot subproblems. To address the NP-hardness of the one-shot problem, we design a computationally-efficient two-phase algorithm based on matching and game theory, which achieves a near-optimal solution. Both rigorous theoretical analysis on the optimality gap and extensive trace-driven simulations are conducted to validate the efficacy of our proposed solution.
Bin Gao 0013, Zhi Zhou 0006, Fangming Liu, Fei Xu 0009, Bo Li 0001
IEEE Trans. Mob. Comput.3
2022 EdgeDR: An Online Mechanism Design for Demand Response in Edge Clouds
abstract
The computing frontier is moving from centralized mega datacenters towards distributed cloudlets at the network edge. We argue that cloudlets are well-suited for handling power demand response to help the grid maintain stability due to more flexible workload management attributed to their distributed nature. However, they also require computing demand response to avoid overload and maintain reliability. To this end, we propose a novel online market mechanism, EdgeDR, to achieve cost efficiency in edge demand response programs. At a high level, we observe that the cloudlet operator can dynamically switch on/off entire cloudlets to compensate for the energy reduction required by the power grid or provide enough computing resources to the edge service. We formulate a long-term social cost minimization problem and decompose it into a series of one-round procurement auctions. In each auction instance, we propose to let the cloudlet tenants bid with cost functions of their two-dimension service quality degradation tolerance, and let the cloudlet operator choose the service quality, manage the workload, and schedule the cloudlet activation status. In addition, we present a dynamic payment mechanism for the operator to balance the tradeoff between short-term profit and long-term benefit in more practical scenarios. Via rigorous analysis, we exhibit that our bidding policy is individually rational and truthful; our workload management algorithm has near-optimal performance in each auction; and our overall online algorithm achieves a provable competitive ratio. We further confirm the performance of our mechanism through extensive trace-driven simulations.
Lei Jiao 0002, Fangming Liu, Lin Wang 0015
IEEE Trans. Parallel Distributed Syst.3
2022 PostMan: Rapidly Mitigating Bursty Traffic via On-Demand Offloading of Packet Processing
abstract
Unexpected bursty traffic brought by certain sudden events, such as news in the spotlight on a social network or discounted items on sale, can cause severe load imbalance in backend services. Migrating hot data - the standard approach to achieve load balance - meets a challenge when handling such unexpected load imbalance, because migrating data will slow down the server that is already under heavy pressure. This article proposes PostMan, an alternative approach to rapidly mitigate load imbalance for services processing small requests. Motivated by the observation that processing large packets incurs far less CPU overhead than processing small ones, PostMan deploys a number of middleboxes called helpers to assemble small packets into large ones for the heavily-loaded server. This approach essentially offloads the overhead of packet processing from the heavily-loaded server to helpers. To minimize the overhead, PostMan activates helpers on demand, only when bursty traffic is detected. The heavily-loaded server determines when clients connect/disconnect to/from helpers based on the real-time load statistics. To tolerate helper failures, PostMan can migrate connections across helpers and can ensure packet ordering despite such migration. Driven by real-world workloads, our evaluation shows that, with the help of PostMan, a Memcached server can mitigate bursty traffic within hundreds of milliseconds, while migrating data takes tens of seconds and increases the latency during migration.
Yipei Niu, Panpan Jin, Yikai Xiao, Rong Shi, Fangming Liu, Chen Qian 0001, Yang Wang 0009
IEEE Trans. Parallel Distributed Syst.6
2022 HiTDL: High-Throughput Deep Learning Inference at the Hybrid Mobile Edge
abstract
Deep neural networks (DNNs) have become a critical component for inference in modern mobile applications, but the efficient provisioning of DNNs is non-trivial. Existing mobile- and server-based approaches compromise either the inference accuracy or latency. Instead, a hybrid approach can reap the benefits of the two by splitting the DNN at an appropriate layer and running the two parts separately on the mobile and the server respectively. Nevertheless, the DNN throughput in the hybrid approach has not been carefully examined, which is particularly important for edge servers where limited compute resources are shared among multiple DNNs. This article presents HiTDL, a runtime framework for managing multiple DNNs provisioned following the hybrid approach at the edge. HiTDL's mission is to improve edge resource efficiency by optimizing the combined throughput of all co-located DNNs, while still guaranteeing their SLAs. To this end, HiTDL first builds comprehensive performance models for DNN inference latency and throughout with respect to multiple factors including resource availability, DNN partition plan, and cross-DNN interference. HiTDL then uses these models to generate a set of candidate partition plans with SLA guarantees for each DNN. Finally, HiTDL makes global throughput-optimal resource allocation decisions by selecting partition plans from the candidate set for each DNN via solving a fairness-aware multiple-choice knapsack problem. Experimental results based on a prototype implementation show that HiTDL improves the overall throughput of the edge by$4.3\times$compared with the state-of-the-art.
Jing Wu 0024, Lin Wang 0015, Qiangyu Pei, Xingqi Cui, Fangming Liu, Tingting Yang 0001
IEEE Trans. Parallel Distributed Syst.5
2022 An Edge Based Data-Driven Chiller Sequencing Framework for HVAC Electricity Consumption Reduction in Commercial Buildings
abstract
It is well-known that the HVAC (heating, ventilation, and air conditioning) dominates electricity consumption in commercial buildings. In this paper, we focus on one of the core problems in building operation, namelychiller sequencingto reduce HVAC electricity consumption. Our contributions are threefold. First, we make a case for why it is important to quantify the performance profile of a chiller, namely coefficient of performance (COP), atrun-time, by developing a data-driven COP estimation methodology. Second, we show that predicting COP accurately is a non trivial problem, requiring considerable computation time. To overcome this barrier, we develop a data-driven COP prediction model and an edge-based chiller sequencing framework integrating the COP predictions, and show that they strike a good balance between electricity saving and ease of use for real-world deployment. Finally, we evaluate the performance of our scheme by applying it to real-world data, spanning four years, obtained from multiple chillers across three large commercial buildings in Hong Kong. The results show an electricity saving of over 30 percent compared to baselines. We offer our edge based data-driven chiller sequencing framework as a cost-effective and practical mechanism to reduce electricity consumption associated with HVAC operation in commercial buildings.
Zimu Zheng, Cheng Fan 0002, Nan Guan, Arun Vishwanath, Dan Wang 0002, Fangming Liu
IEEE Trans. Sustain. Comput.7
2021 Deep discriminative image feature learning for cross-modal semantics understanding
Hong Zhang 0022, Fangming Liu, Bo Li 0002, Yihai Zhu
Knowl. Based Syst.2
2021 Online Adaptive Interference-Aware VNF Deployment and Migration for 5G Network Slice
abstract
Based on network function virtualization (NFV) and software defined network (SDN),network slicingis proposed as a new paradigm for building service-customized 5G network. In each network slice, service-required virtual network functions (VNFs) can be flexibly deployed in an on-demand manner, which will support a variety of 5G use cases. However, due to the real-time network variations and diverse performance requirements among different 5G scenarios, online adaptive VNF deployment and migration are needed to dynamically accommodate to service-specific requirements. In this paper, we first propose a time-slot based 5G network slice model, which jointly includes both edge cloud servers and core cloud servers. Since VNF consolidation may cause severe performance degradation, we adopt a demand-supply model to quantify the VNF interference. To achieve our objective—maximizing the total reward of accepted requests (i.e., the total throughput minus the weighted total VNF migration cost), we propose an Online Lazy-migration Adaptive Interference-aware Algorithm (OLAIA) for real-time VNF deployment and cost-efficient VNF migration in a 5G network slice, where an Adaptive Interference-aware Algorithm (AIA) is proposed as OLAIA’s core function for placing a given set of requests’ VNFs with maximized total throughput. Through trace-driven evaluations on two typical 5G network slices, we demonstrate that OLAIA can efficiently handle the real-time network variations and the VNF interference when deploying VNFs for real-time arriving requests. In particular, OLAIA improves the total reward by 22.18% in the autonomous driving scenario and by 51.10% in the 4K/8K HD video scenario, as compared with other state-of-the-art solutions.
Qixia Zhang, Fangming Liu, Chaobing Zeng
IEEE/ACM Trans. Netw.2
2021 Precise Power Capping for Latency-Sensitive Applications in Datacenter
abstract
Power capping is widely used in cloud datacenters to mitigate power over-provisioning problem, thus improve datacenter capacity and cut off their operation cost. However, inappropriate or aggressive power capping may lead to performance degradation of applications (especially latency-sensitive ones), and there are few effective methods that can accurately evaluate and control such negative impact caused by aggressive power capping. In this paper, we proposeFine-Grained Differential Method(FGD) to quantitatively analyze how inappropriate power capping degrades the performance of latency-sensitive applications. By using FGD, we can minimize the provisioned power for each server by setting a precise power budget according to application’sService Level Agreement(SLA). And we further proposePrecise Power Capping(PPCapping) which is designed to increase the datacenter capacity with a fixed power supply by means of FGD. Our research also provides an insight of precise tradeoff between applications’ SLAs and datacenter capacity. We verify FGD and PPCapping by using real world traces from Tencent’s datecenter with 25,328 servers. The experimental results show that FGD can accurately analyze the impact of power capping on the performance of latency-sensitive applications, and PPCapping can effectively increase datacenter capacity compared with the typical power provisioning strategy.
Song Wu 0001, Xinhou Wang, Hai Jin 0001, Fangming Liu, Haibao Chen, Chuxiong Yan
IEEE Trans. Sustain. Comput.5
2020 Latency-aware VNF Chain Deployment with Efficient Resource Reuse at Network Edge
abstract
With the increasing demand of low-latency network services, mobile edge computing (MEC) emerges as a new paradigm, which provides server resources and processing capacities in close proximity to end users. Based on network function virtualization (NFV), network services can be flexibly provisioned as virtual network function (VNF) chains deployed at edge servers. However, due to the resource shortage at the network edge, how to efficiently deploy VNF chains with latency guarantees and resource efficiency remains as a challenging problem. In this work, we focus on jointly optimizing the resource utilization of both edge servers and physical links under the latency limitations. Specifically, we formulate the VNF chain deployment problem as a mixed integer linear programming (MILP) to minimize the total resource consumption. We design a novel two-stage latency-aware VNF deployment scheme: highlighted by a constrained depth-first search algorithm (CDFSA) for selecting paths, and a path-based greedy algorithm (PGA) for assigning VNFs by reusing as many VNFs as possible. We demonstrate that our proposed algorithm can efficiently achieve a near-optimal solution with a theoretically-proved worstcase performance bound. Extensive simulation results show that the proposed algorithm outperforms three previous heuristic algorithms.
Panpan Jin, Xincai Fei, Qixia Zhang, Fangming Liu, Bo Li 0001
INFOCOM4
2020 Heat to Power: Thermal Energy Harvesting and Recycling for Warm Water-Cooled Datacenters
abstract
Warm water cooling has been regarded as a promising method to improve the energy efficiency of water-cooled datacenters. In warm water-cooling systems, hot spots occur as a common problem where the hybrid cooling architecture integrating thermoelectric coolers (TECs) emerges as a new remedy. Equipped with this architecture, the inlet water temperature can be raised higher, which provides more opportunities for heat recycling. However, currently, the heat absorbed from the server components is ejected directly into the water without being recycled, which leads to energy wasting. In order to further improve the energy efficiency, we propose Heat to Power (H2P), an economical and energy-recycling warm water cooling architecture, where thermoelectric generators (TEGs) harvest thermal energy from the “used” warm water and generate electricity for reusing in datacenters. Specifically, we propose some efficient optimization methods, including an economical water circulation design, fine-grained adjustments of the cooling setting and dynamic workload scheduling for increasing the power generated by TEGs. We evaluate H2P based on a real hardware prototype and cluster traces from Google and Alibaba. Experiment results show that TEGs equipped with our optimization methods can averagely generate 4.349 W, 4.203 W, and 3.979 W (4.177 W averagely) electricity on one CPU under the drastic, irregular and common workload traces, respectively. The power reusing efficiency (PRE) can reach 12.8%~16.2% (14.23% averagely) and the total cost of ownership (TCO) of datacenters can be reduced by up to 0.57%.
Weixiang Jiang, Fangming Liu, Qixia Zhang, Ziyang Jia
ISCA3
2020 Finedge: A Dynamic Cost-Efficient Edge Resource Management Platform for NFV Network
abstract
With the evolution of network function virtualization (NFV) and edge computing, software-based network functions (NFs) can be deployed on closer-to-end-user edge servers to support a broad range of new services with high bandwidth and low latency. However, due to the resource limitation, strict QoS requirements and real-time flow fluctuations in edge network, existing cloud-based resource management strategy in NFV platforms is inefficient to be applied to the edge. Thus, we propose Finedge, a dynamic, fine-grained and cost-efficient edge resource management platform for NFV network. First, we conduct empirical experiments to find out the effect of NFs' resource allocation and their flow-level characteristics on performance. Then, by jointly considering these factors and QoS requirements (e.g., latency and packet loss rate), Finedge can automatically assign the most suitable CPU core and tune the most cost-efficient CPU quota to each NF. Finedge is also implemented with some key strategies including real-time flow monitoring, elastic resource scaling up and down, and also flexible NF migration among cores. Through extensive evaluations, we validate that Finedge can efficiently handle heterogeneous flows with the lowest CPU quota and the highest SLA satisfaction rate as compared with the default OS scheduler and other state-of-the-art resource management schemes.
Qixia Zhang, Fangming Liu
IWQoS3
2020 eDirect: Energy-Efficient D2D-Assisted Relaying Framework for Cellular Signaling Reduction
abstract
Mobile Instant Messaging (IM) apps, such as WhatsApp and WeChat, frequently send heartbeat messages to remote servers to maintain their always-online status. Periodic heartbeat messages are small in size, but their transmissions incur heavy signaling traffic due to frequently establishing and releasing communication channels between Base Stations (BSs) and smartphones, known as the signaling storm. Meanwhile, smartphones also need to activate the cellular data communication module frequently for transmitting short heartbeat messages, resulting in substantial energy consumption. To address these issues, we present eDirect, an energy-efficient D2D-assIsted Relaying framEwork for Cellular signaling reducTion. eDirect selects active smartphones as relays to opportunistically collect heartbeat messages from nearby smartphones using energy-efficient D2D communication. The collected heartbeat messages are transmitted to the BS in an aggregated manner to reduce cellular signaling traffic. Based on the beating frequencies and deadlines of the collected heartbeat messages, eDirect schedules transmissions of the collected heartbeat messages to minimize signaling overhead and energy consumption while meeting the deadline constraints. We implement and evaluate our solution on Android smartphones. The results from real-world experiments show that our solution reduces signaling traffic by at least 50% and energy consumption by up to 36%.
Xiaomeng Yi, Yanqi Jin, Fangming Liu, Minghua Chen 0001
IEEE/ACM Trans. Netw.4
2020 On-Edge Multi-Task Transfer Learning: Model and Practice With Data-Driven Task Allocation
abstract
On edge devices, data scarcity occurs as a common problem where transfer learning serves as a widely-suggested remedy. Nevertheless, transfer learning imposes heavy computation burden to the resource-constrained edge devices. Existing task allocation works usually assume all submitted tasks are equally important, leading to inefficient resource allocation at a task level when directly applied in Multi-task Transfer Learning (MTL). To address these issues, we first reveal that it is crucial to measure the impact of tasks on overall decision performance improvement and quantify task importance. We then show that task allocation with task importance for MTL (TATIM) is a variant of NP-complete Knapsack problem, where the complicated computation to solve this problem needs to be conducted repeatedly under varying contexts. To solve TATIM with high computational efficiency, we propose a Data-driven Cooperative Task Allocation (DCTA) approach. Finally, we evaluate the performance of DCTA by not only a trace-driven simulation, but also a new comprehensive real-world AIOps case study which bridges model and practice via a new architecture and main components design within AIOps system. Extensive experiments show that our DCTA reduces 3.24 times of processing time, and saves 48.4 percent energy consumption compared with the state-of-the-art when solving TATIM.
Zimu Zheng, Chuang Hu, Dan Wang 0002, Fangming Liu
IEEE Trans. Parallel Distributed Syst.5
2020 Errata to "On-Edge Multi-Task Transfer Learning: Model and Practice With Data-Driven Task Allocation"
abstract
Presents corrections to author information for the above named paper.
Zimu Zheng, Chuang Hu, Dan Wang 0002, Fangming Liu
IEEE Trans. Parallel Distributed Syst.5
2020 A Truthful and Efficient Incentive Mechanism for Demand Response in Green Datacenters
abstract
Datacenter demand response is envisioned as a promising tool for mitigating operational stability issues faced by smart grids. It enables significant potentials in peak load reduction and facilitates the incorporation of distributed generation. Monetary refund from the smart grid can also alleviate the cloud's burden in escalating electricity cost. However, the current demand response paradigm is inefficient towards incentivizing a cloud service provider (CSP) that operates geo-distributed datacenters. To incentivize CSP participation, this work presents an auction mechanism that enables smart grids to voluntarily submit bids to the CSP to procure diverse amounts of demand response with different payments. To maximize the social welfare of the auction, the CSP that acts as the auctioneer needs to solve the winner determination problem at large-scale. By applying the proximal Jacobian alternating direction method of multipliers, we propose a distributed algorithm for each datacenter to solve a small-scale problem in a parallel fashion. Desirable properties of the proposed auction, such as social welfare maximization and truthfulness are achieved through Vickrey-Clarke-Groves (VCG) payment. Through extensive evaluations based on real datacenter workload traces and IEEE 14-bus test systems, we demonstrate that our incentive mechanism constitutes a win-win mechanism for both the geo-distributed cloud and the smart grid.
Zhi Zhou 0006, Fangming Liu, Zongpeng Li
IEEE Trans. Parallel Distributed Syst.2
2019 Data-driven Task Allocation for Multi-task Transfer Learning on the Edge
abstract
Edge computing for machine learning has become a heated research topic. On edge devices, data scarcity occurs as a common problem where transfer learning serves as a widely-suggested remedy. Nevertheless, one obstacle is that transfer learning imposes heavy computation burden to the resource-constrained edge devices. Motivated by the fact that only a few tasks of Multi-task Transfer Learning (MTL) have a higher potential for overall decision performance improvement, we design a novel task allocation scheme, which assigns more important tasks to more powerful edge devices to maximize the overall decision performance. In this paper, we focus on task allocation under multi-task scenarios by introducing task importance and make the following contributions. First, we reveal that it is important to measure the impact of tasks on overall decision performance improvement and quantify task importance. We also observe the long-tail property of task importance, i.e., only a few tasks are important, which facilitates more efficient task allocation. Second, we show that task allocation with task importance for MTL (TATIM) is in fact a variant of the NP-complete Knapsack problem, where the complicated computation to solve this problem needs to be conducted repeatedly under varying contexts. To solve TATIM with high computational efficiency, we innovatively propose a Data-driven Cooperative Task Allocation (DCTA) approach. Third, we evaluate the performance of our DCTA approach by applying it to a real-world industrial operation (e.g., AIOps) scenario. Experiments show that our DCTA approach can reduce 3.24 times of processing time compared with the state-of-the-art when solving TATIM. We offer our DCTA approach as an effective and practical mechanism for reducing the required resource associated with performing MTL on edge devices.
Zimu Zheng, Chuang Hu, Dan Wang 0002, Fangming Liu
ICDCS5
2019 Stage Delay Scheduling: Speeding up DAG-style Data Analytics Jobs with Resource Interleaving
abstract
To increase the resource utilization of datacenters, big data analytics jobs are commonly running stages in parallel which are organized into and scheduled according to the Directed Acyclic Graph (DAG). Through an in-depth analysis of the latest Alibaba cluster trace and our motivation experiments on Amazon EC2, however, we show that the CPU and network resources are still under-utilized due to the unwise stage scheduling, thereby prolonging the completion time of a DAG-style job (e.g., Spark). While existing works on reducing the job completion time focus on either task scheduling or job scheduling, stage scheduling has received comparably little attention. In this paper, we design and implement DelayStage, a simple yet effective stage delay scheduling strategy to interleave the cluster resources across the parallel stages, so as to increase the cluster resource utilization and speed up the job performance. With the aim of minimizing the makespan of parallel stages, DelayStage judiciously arranges the execution of stages in a pipelined manner to maximize the performance benefits of resource interleaving. Extensive prototype experiments on 30 Amazon EC2 instances and complementary trace-driven simulations show that DelayStage can improve the cluster resource utilization by up to 81.8% and reduce the job completion time by up to 41.3%, in comparison to the stock Spark and the state-of-the-art stage scheduling strategies, yet with acceptable runtime overhead.
Wujie Shao, Fei Xu 0009, Li Chen 0019, Haoyue Zheng, Fangming Liu
ICPP5
2019 Cynthia: Cost-Efficient Cloud Resource Provisioning for Predictable Distributed Deep Neural Network Training
abstract
It becomes an increasingly popular trend for deep neural networks with large-scale datasets to be trained in a distributed manner in the cloud. However, widely known as resource-intensive and time-consuming, distributed deep neural network (DDNN) training suffers from unpredictable performance in the cloud, due to the intricate factors of resource bottleneck, heterogeneity and the imbalance of computation and communication which eventually cause severe resource under-utilization. In this paper, we propose Cynthia, a cost-efficient cloud resource provisioning framework to provide predictable DDNN training performance and reduce the training budget. To explicitly explore the resource bottleneck and heterogeneity, Cynthia predicts the DDNN training time by leveraging a lightweight analytical performance model based on the resource consumption of workers and parameter servers. With an accurate performance prediction, Cynthia is able to optimally provision the cost-efficient cloud instances to jointly guarantee the training performance and minimize the training budget. We implement Cynthia on top of Kubernetes by launching a 56-docker cluster to train four representative DNN models. Extensive prototype experiments on Amazon EC2 demonstrate that Cynthia can provide predictable training performance while reducing the monetary cost for DDNN workloads by up to 50.6%, in comparison to state-of-the-art resource provisioning strategies, yet with acceptable runtime overhead.
Haoyue Zheng, Fei Xu 0009, Li Chen 0019, Zhi Zhou 0006, Fangming Liu
ICPP5
2019 An Online Market Mechanism for Edge Emergency Demand Response via Cloudlet Control
abstract
The computing frontier is moving from centralized mega datacenters towards distributed cloudlets at the network edge. We argue that cloudlets are well-suited for participation in Emergency Demand Response (EDR) programs due to their enormous energy consumption and flexible workload distribution, while existing EDR mechanisms for clouds and colocation datacenters are not suitable for cloudlets. We propose a novel online market mechanism, EdgeEDR, to incentivize cloudlets to participate in EDR, featuring multiple cloudlet-specific designs. At a high level, we observe that cloudlet operators can dynamically switch on/off entire cloudlets to compensate for the energy reduction required by the power grid. We formulate a long-term social cost minimization problem and decompose it into a series of one-round procurement auctions. In each auction instance, we propose to let the cloudlet tenants bid with cost functions of their service quality degradation tolerance, and let the cloudlet operator choose the service quality, allocate the workload, and shut down the cloudlets. Via rigorous analysis, we exhibit that our bidding policy is individually rational and truthful; our workload distribution algorithm has near-optimal performance in each auction; and our overall online algorithm achieves a provable competitive ratio. We further confirm the performance of our mechanism through extensive trace-driven simulations.
Lei Jiao 0002, Lin Wang 0015, Fangming Liu
INFOCOM4
2019 Winning at the Starting Line: Joint Network Selection and Service Placement for Mobile Edge Computing
abstract
Mobile Edge Computing (MEC) is an emerging computing paradigm in which computational capabilities are pushed from the central cloud to the network edges. However, preserving the satisfactory quality-of-service (QoS) for user applications is non-trivial among multiple densely dispersed yet capacity constrained MEC nodes. This is mainly because both the access network and edge nodes are vulnerable to network congestion. Previous works are mostly limited to optimizing the QoS through dynamic service placement, while ignoring the critical effects of access network selection on the network congestion. In this paper, we study the problem of jointly optimizing the access network selection and service placement for MEC, towards the goal of improving the QoS by balancing the access, switching and communication delay. Specifically, we first design an efficient online framework to decompose the long-term optimization problem into a series of one-shot problems. To address the NP-hardness of the one-shot problem, we further propose an iteration-based algorithm to derive a computation efficient solution. Both rigorous theoretical analysis on the optimality gap and extensive trace-driven simulations validate the efficacy of our proposed solution.
Bin Gao 0013, Zhi Zhou 0006, Fangming Liu, Fei Xu 0009
INFOCOM3
2019 Adaptive Interference-Aware VNF Placement for Service-Customized 5G Network Slices
abstract
Based on network function virtualization (NFV) and software defined network (SDN), network slicing is proposed as a new paradigm for building service-customized 5G network. In each network slice, service-required virtual network functions (VNFs) can be flexibly deployed in an on-demand manner, which will support a variety of 5G use cases. However, due to the diverse performance requirements among different 5G scenarios, an adaptive VNF placement approach is needed to automatically accommodate to service-specific requirements. In this paper, we tackle the VNF placement problem by first proposing a general 5G network slice framework, which jointly contains both edge cloud and core cloud servers. Specially, based on the fact that VNF consolidation may cause severe performance degradation, we adopt a demand-supply model to quantity the VNF interference. With an aim to maximize the total throughput of accepted requests, we propose an Adaptive Interference-Aware (AIA) heuristic approach to automatically place VNFs in 5G service-customized network slices. Through simulations on two typical 5G scenarios, we demonstrate that AIA can efficiently handle traffic variation especially caused by VNF interference and improve the total throughput by 20.11% and 24.21% in autonomous driving and 4K/8K HD video network slices as compared with the state-of-the-art methods.
Qixia Zhang, Fangming Liu, Chaobing Zeng
INFOCOM2
2019 Fine-grained warm water cooling for improving datacenter economy
abstract
Driven by the increasing power consumption of datacenters, the industry is focusing more on water cooling for improving the energy efficiency. Using warm water to cool servers has been considered as an efficient method to reduce the cooling energy. However, warm water cooling may lead to the risk of cooling failure and its energy efficiency suffers from the thermal imbalance among servers, due to the lack of fine-grained cooling control. In this paper, we propose a hybrid cooling architecture design that incorporates thermoelectric cooler into the water cooling system, to deal with cooling mismatching in a fine-grained manner. We exploit the warm water cooling strategy and design an adaptive cooling control framework according to workload variations, to make water cooling system more economical for datacenters. We evaluate the hybrid water cooling design based on a real hardware prototype and cluster traces from Google and Alibaba. Compared with conventional water cooling system, our hybrid water cooling system can reduce the energy consumption by 58.72%~78.43% to handle the cooling mismatching.
Weixiang Jiang, Ziyang Jia, Sirui Feng, Fangming Liu, Hai Jin 0001
ISCA4
2019 NFVdeep: adaptive online service function chain deployment with deep reinforcement learning
abstract
With the evolution of network function virtualization (NFV), diverse network services can be flexibly offered as service function chains (SFCs) consisted of different virtual network functions (VNFs). However, network state and traffic typically exhibit unpredictable variations due to stochastically arriving requests with different quality of service (QoS) requirements. Thus, an adaptive online SFC deployment approach is needed to handle the real-time network variations and various service requests. In this paper, we firstly introduce a Markov decision process (MDP) model to capture the dynamic network state transitions. In order to jointly minimize the operation cost of NFV providers and maximize the total throughput of requests, we propose NFVdeep, an adaptive, online, deep reinforcement learning approach to automatically deploy SFCs for requests with different QoS requirements. Specifically, we use a serialization-and-backtracking method to effectively deal with large discrete action space. We also adopt a policy gradient based method to improve the training efficiency and convergence to optimality. Extensive experimental results demonstrate that NFVdeep converges fast in the training process and responds rapidly to arriving requests especially in large, frequently transferred network state space. Consequently, NFVdeep surpasses the state-of-the-art methods by 32.59% higher accepted throughput and 33.29% lower operation cost on average.
Yikai Xiao, Qixia Zhang, Fangming Liu, Jia Wang 0009, Miao Zhao
IWQoS3
2019 PostMan: Rapidly Mitigating Bursty Traffic by Offloading Packet Processing
Panpan Jin, Yikai Xiao, Rong Shi, Yipei Niu, Fangming Liu, Chen Qian 0001, Yang Wang 0009
USENIX ATC6
2018 Non-IT Energy Accounting in Virtualized Datacenter
abstract
Energy accounting plays a crucial role in datacenter energy management, wherein the energy consumption of non-IT units (e.g., UPS and cooling system) makes up a significant portion. However, it is challenging to fairly account for non-IT energy on an individual VM basis, because the non-IT units are shared by multiple VMs in a virtualized datacenter and only the system-level non-IT energy consumption can be measured. Existing policies, e.g., equally or proportionally allocating non-IT energy to VMs based on their IT energy, are not fair, in the sense that they can not satisfy a set of desired axiomatic principles of fair allocation. In this paper, we propose LEAPS, a Lightweight Energy Accounting Policy based on a provably fair methodology called Shapley value. We evaluate it using real-world datacenter trace and demonstrate that, compared to original Shapley value approach that has an exponential complexity, LEAPS yields almost the same energy accounting result within a maximum relative error less than 6.97%, while having a negligible computation time.
Weixiang Jiang, Shaolei Ren, Fangming Liu, Hai Jin 0001
ICDCS3
2018 DHL: Enabling Flexible Software Network Functions with FPGA Acceleration
abstract
Network function virtualization (NFV) aims to run software network functions (NFs) in commodity servers. As CPU is general-purpose hardware, one has to use many CPU cores to handle complex packet processing at line rate. Owing to its performance and programmability, FPGA has emerged as a promising platform for NFV. However, the programmable logic blocks on an FPGA board are limited and expensive. Implementing the entire NFs on FPGA is thus resource-demanding. Further, FPGA needs to be reprogrammed when the NF logic changes which can take hours to synthesize the code. It is thus inflexible to use FPGA to implement the entire NFV service chain. We present dynamic hardware library (DHL), a novel CPU-FPGA co-design framework for NFV with both high performance and flexibility. DHL employs FPGA as accelerators only for complex packet processing. It abstracts accelerator modules in FPGA as a hardware function library, and provides a set of transparent APIs for developers. DHL supports running multiple concurrent software NFs with distinct accelerator functions on the same FPGA and provides data isolation among them. We implement a prototype of DHL with Intel DPDK. Experimental results demonstrate that DHL greatly reduces the programming efforts to access FPGA, brings significantly higher throughput and lower latency over CPU-only implementation, and minimizes the CPU resources.
Xiuxiu Wang, Fangming Liu, Hong Xu 0001
ICDCS3
2018 Adaptive VNF Scaling and Flow Routing with Proactive Demand Prediction
abstract
With the evolution of Network Function Virtual-izaiton (NFV), enterprises are increasingly outsourcing their network functions to the cloud. However, using virtualized network functions (VNFs) to provide flexible services in today's cloud is challenging due to the inherent difficulty in intelligently scaling VNFs to cope with traffic fluctuations. To best utilize cloud resources, NFV providers need to dynamically scale the VNF deployments and reroute traffic demands for their customers. Since most existing work is reactive in nature, we seek a proactive approach to provision new instances for overloaded VNFs ahead of time based on the estimated flow rates. We formulate the VNF provisioning problem in order that the cost incurred by inaccurate prediction and VNF deployment is minimized. In the proposed online algorithm, we first employ an efficient online learning method which aims at minimizing the error in predicting the service chain demands. We then derive the requested instances with adaptive processing capacities and call two other algorithms for new instance assignment and service chain rerouting, respectively, while achieving good competitive ratios. The joint online algorithm is proven to provide good performance guarantees by both theoretical analysis and trace-driven simulation.
Xincai Fei, Fangming Liu, Hong Xu 0001, Hai Jin 0001
INFOCOM2
2018 Load Balancing Across Microservices
abstract
With the advent of cloud container technology, enterprises develop applications through microservices, breaking monolithic software into a suite of small services whose instances run independently in containers. User requests are served by a series of microservices forming a chain, and the chains often share microservices. Existing load balancing strategies either incur significant networking overhead or ignore the competition for shared microservices across chains. Furthermore, typical load balancing solutions leverage a hybrid technique by combining HTTP with message queue to support microservice communications, bringing additional operational complexity. To address these challenges, we propose a chain-oriented load balancing algorithm (COLBA) based solely on message queues, which balances load based on microservice requirements of chains to minimize response time. We model the load balancing problem as a non-cooperative game, and leverage Nash bargaining to coordinate microservice allocation across chains. Employing convex optimization with rounding, we efficiently solve the problem that is proven NP-hard. Extensive trace-driven simulations demonstrate that COLBA reduces the overall average response time at least by 13% compared with existing load balancing strategies.
Yipei Niu, Fangming Liu, Zongpeng Li
INFOCOM2
2018 Demystifying the Performance Interference of Co-Located Virtual Network Functions
abstract
Network function virtualization (NFV) decouples network functions from the dedicated hardware and enables them running on commodity servers, facilitating widespread deployment of virtualized network functions (VNFs). Network operators tend to deploy VNFs in virtual machines (VMs) due to VM's ease of duplication and migration, which enables flexible VNF placement and scheduling. Efforts have been paid to provide efficient VNF placement approaches, aiming at minimizing the resource cost of VNF deployment and reducing the latency of service chain. However, existing placement approaches may result in hardware resource competition of co-located VNFs, leading to performance degradation. In this paper, we present a measurement study on the performance interference among different types of co-located VNFs and analyze how VNFs' competitive hardware resources and the characteristics of packet affect the performance interference. We disclose that the performance interference between co-located VNFs is ubiquitous, which causes the performance degradation, in terms of VNFs' throughput, ranging from 12.36% to 50.3%, and the competition of network I/O bandwidth plays a key role in the performance interference. Based on our measurement results, we give some advices on how to design more efficient VNF placement approaches.
Chaobing Zeng, Fangming Liu, Weixiang Jiang
INFOCOM2
2018 JouleMR: Towards Cost-Effective and Green-Aware Data Processing Frameworks
abstract
Interests have been growing in energy management of the cluster effectively in order to reduce the energy consumption as well as the electricity cost. Renewable energy and dynamic pricing schemes in smart grids are two major emerging trends in energy markets. However, current data processing frameworks are not aware of the efficiency of each joule consumed by the data center workloads in the context of these two major trends. In fact, not all joules are equal in the sense that the amount of work that can be done by a joule can vary significantly in data centers. Ignoring this fact leads to significant energy waste (by 25 percent of the total energy consumption in Hadoop YARN on a Facebook production trace according to our study). In this paper, we propose JouleMR, a cost-effective and green-aware data processing framework. Specifically, we investigate how to exploit such joule efficiency to maximize the benefits of renewable energy as well as dynamic pricing schemes for MapReduce framework. We develop job/task scheduling algorithms with a particular focus on the factors on joule efficiency in the data center, including the energy efficiency of MapReduce workloads, renewable energy supply, dynamic pricing and the battery usage. We further develop a simple yet effective performanceenergy consumption model to guide our scheduling decisions. We have implemented JouleMR on top of Hadoop YARN. The experiments demonstrate the accuracy of our models, and the effectiveness of our cost-effective and green-aware optimizations outperform the state-of-the-art implementations over Hadoop YARN.
Zhaojie Niu, Bingsheng He, Fangming Liu
IEEE Trans. Big Data3
2017 Aemon: Information-agnostic Mix-flow Scheduling in Data Center Networks
abstract
Data center networks carry a mix of flows, some with deadlines and some without. Existing mix-flow transport designs assume prior knowledge of flow sizes, which may not hold in practice. Without such information, mix-flow scheduling becomes particularly challenging due to (1) the lack of precise rate control of deadline flows with minimal impact on non-deadline flows; (2) difficulty in assigning priority to both two types of flows.
Tao Wang 0088, Hong Xu 0001, Fangming Liu
APNet3
2017 Virtual Machine Power Accounting with Shapley Value
abstract
The ever-increasing power consumption of datacenters has eaten up a large portion of their profit. One possible solution is to charge datacenter users for their actual power usage. However, it poses a great technical challenge as the power of VMs co-existing in a physical machine cannot be measured directly. It is thus critical to develop a fair method to disaggregate the power of a physical machine to individual VMs. We tackle the above challenge by modeling the power disaggregation problem as a cooperative game and propose non-deterministic Shapley value to discover the fair power share of VMs (in the sense of satisfying four desired axiomatic principles), while compensating the negative impact of VM power variation. We demonstrate that the results from existing power model-based solution can deviate from the "ground truth" by 25.22% ~46.15%. And compared with the exact Shapley value, our non-deterministic Shapley value can achieve less than 5% error for 90% of the time.
Weixiang Jiang, Fangming Liu, Guoming Tang, Kui Wu 0001, Hai Jin 0001
ICDCS2
2017 Reducing Cellular Signaling Traffic for Heartbeat Messages via Energy-Efficient D2D Forwarding
abstract
Mobile Instant Messaging (IM) apps, such as WhatsApp and WeChat, frequently send heartbeat messages to remote servers to maintain always-online status. Periodic heartbeat messages are small in size, but their transmissions incur heavy signaling traffic to frequently establish and release communication channels between base stations and smartphones, known as signaling storm. Meanwhile, smartphones also need to activate cellular data communication module frequently for transmitting short heartbeat messages, resulting in substantial energy consumption. To address these issues, we propose a Device-to-Device (D2D) based heartbeat relaying framework, in order to reduce signaling traffic and energy consumption in heartbeat transmission. The framework selects the smartphones as relays to opportunistically collect heartbeat messages from nearby smartphones using energy-efficient D2D communication. The collected heartbeat messages are transmitted to the BS in an aggregated manner to reduce cellular signaling traffic. Based on the periods and the expiration time of the collected heartbeat messages, the framework schedules the transmissions of collected heartbeat messages to minimize signaling and energy consumption while satisfying time constrains. We implement and evaluate our solution on Android smartphones. The results from real-world experiments show that our solution achieves more than 50% signaling traffic reduction and up to 36% energy saving.
Yanqi Jin, Fangming Liu, Xiaomeng Yi, Minghua Chen 0001
ICDCS2
2017 Multi-resource Load Balancing for Virtual Network Functions
abstract
Middleboxes are widely deployed to perform various network functions to ensure security and improve performance. The recent trend of Network Function Virtualization (NFV) makes it easy for operators to deploy software implementations of these network functions on commodity servers. However, virtual network functions consume different amounts of resources when processing packets. Thus a multi-resource load balancing (MRLB) mechanism is needed to efficiently utilize server resources. MRLB problem in the context of NFV is fundamentally different from multi-resource allocation problems, as well as traditional single-resource load balancing and multi-resource load balancing problems in task scheduling. In this paper, we tackle the MRLB problem in NFV by first proposing dominant load-the load of the most stressed resource on a server-as the load balancing metric. We then formulate the MRLB problem as an optimization to minimize the maximum dominant load of all NFV servers given the demand. Based on proximal Jacobian ADMM, we propose an efficient algorithm to solve the problem in large scale settings. Through extensive trace-driven simulations and prototype experiments on a testbed, we show that our MRLB algorithm with dominant load performs significantly better and faster than benchmarking algorithms.
Tao Wang 0088, Hong Xu 0001, Fangming Liu
ICDCS3
2017 Joint Optimization of Chain Placement and Request Scheduling for Network Function Virtualization
abstract
Compared with executing Network Functions (NFs) on dedicated hardwares, the recent trend of Network Function Virtualization (NFV) holds the promise for operators to flexibly deploy software-based NFs on commodity servers. However, virtual NFs (VNFs) are normally "chained" together to provide a specific network service. Thus, an efficient scheme is needed to place the VNF chains across the network and effectively schedule requests to service instances, which can maximize the average resource utilization of each node in service and simultaneously minimize the average response latency of each request. To this end, we formulate first VNF chains placement problem as a variant of bin-packing problem, which is NP-hard, and we model request scheduling problem based on the key concepts from open Jackson network. To jointly optimize the performance of NFV, we propose a priority-driven weighted algorithm to improve resource utilization and a heuristic algorithm to reduce response latency. Through extensive trace-driven simulations, we show that our methods can indeed enhance performance in diverse scenarios. In particular, we can improve the average resource utilization by 33.4% and can reduce the average total latency by 19.9% as compared with the state-of-the-art methods.
Qixia Zhang, Yikai Xiao, Fangming Liu, John C. S. Lui
ICDCS3
2017 Handling flash deals with soft guarantee in hybrid cloud
abstract
Flash deal applications, which offer significant benefits (e.g., discount) to subscribers within a short period of time, are becoming increasingly prevalent. Motivated by such transient profit, flash crowds of subscribers request services simultaneously. Considering the unique business logic, a hybrid cloud with soft guarantee, i.e., bounding the response time of delay-tolerant requests, has great potential to handle flash crowds. In this paper, to cost-effectively withstand flash crowds with soft guarantee, we propose a solution that makes smart decisions on scheduling requests in the hybrid cloud and adjusting the capacity of the public cloud. In respect of scheduling requests, we apply Sequential Quadratic Programming (SQP) to achieve soft guarantee. Furthermore, for adjusting capacity, we design an online algorithm to tune the scale of the public cloud towards jointly minimizing cost and response time, yet without a priori knowledge of request arrival rate. We prove that the online algorithm can obtain a competitive ratio of 1-6ε against the optimal solution, where ε can be tuned close to 0. By conducting extensive trace-driven experiments in a website prototype deployed on OpenStack Mitaka and Amazon Web Service, our solution reduces response time by 15% compared with previous work under given budget.
Yipei Niu, Fangming Liu, Xincai Fei, Bo Li 0001
INFOCOM2
2017 Towards load-balanced VNF assignment in geo-distributed NFV Infrastructure
abstract
Network functions virtualization (NFV) is increasingly adopted by telecommunications (telecos) service providers for cost savings and flexible management. However, deploying virtual network functions (VNFs) in geo-distributed central offices (COs) is not straightforward. Unlike most existing centralized schemes in clouds, VNFs of a service chain usually need to be deployed in multiple COs due to limited resource capacity and uneven setup cost at various locations. To ensure the Quality of Service of service chains, a key problem for service providers is to determine where a VNF should go, in order to achieve cost-efficiency and load balancing of both computing and bandwidth resources, across all selected COs. To this end, we present a framework of CO Selection (CS) and VNF Assignment (VA) for distributed deployment of NFV. Specifically, we first select a set of COs that minimizes the communication cost among the selected COs. Then, we employ a shadow-routing based approach, which minimizes the maximum of appropriately defined CO utilizations, to jointly solve the VNF-CO and VNF-server assignment problem. Simulations demonstrate the effectiveness of CS algorithm, and asymptotic optimality, scalability and high adaptivity of the VNF assignment approach.
Xincai Fei, Fangming Liu, Hong Xu 0001, Hai Jin 0001
IWQoS2
2017 Pricing Intra-Datacenter Networks with Over-Committed Bandwidth Guarantee
Fangming Liu, John C. S. Lui
USENIX ATC2
2017 Cost-Effective Service Provisioning for Hybrid Cloud Applications
Fangming Liu, Yipei Niu
Mob. Networks Appl.1
2017 NIPD: Non-Intrusive Power Disaggregation in Legacy Datacenters
abstract
Fine-grained power monitoring, which refers to power monitoring at the server level, is critical to the efficient operation and energy saving of datacenters. Fined-grained power monitoring, however, is extremely challenging in legacy datacenters that host server systems not equipped with power monitoring sensors. Installing power monitoring hardware at the server level not only incurs high costs but also complicates the maintenance of high-density server clusters and enclosures. In this paper, we present a zero-cost, purely software-based solution to this challenging problem. We use a novel technique of non-intrusive power disaggregation (NIPD) that establishes power mapping functions (PMFs) between the states of servers and their power consumption, and infer the power consumption of each server with the aggregated power of the entire datacenter. The PMFs that we have developed can support both linear and nonlinear power models via the state feature transformation. To reduce the training overhead, we further develop adaptive PMFs update strategies and ensure that the training data and state features are appropriately selected. We implement and evaluate NIPD over a real-world datacenter with 326 nodes. The results show that our solution can provide high precision power estimation at both rack level and server level. In specific, with PMFs including only two nonlinear terms, our power estimation i) at rack level has mean relative error of 2.18 percent, and ii) at server level has mean relative errors of 9.61 and 7.53 percent corresponding to the idle and peak power, respectively.
Guoming Tang, Weixiang Jiang, Fangming Liu, Kui Wu 0001
IEEE Trans. Computers4
2017 eBA: Efficient Bandwidth Guarantee Under Traffic Variability in Datacenters
abstract
Datacenter networks suffer unpredictable performance due to a lack of application level bandwidth guarantees. A lot of attention has been drawn to solve this problem such as how to provide bandwidth guarantees for virtualized machines (VMs), proportional bandwidth share among tenants, and high network utilization under peak traffic. However, existing solutions fail to cope with highly dynamic traffic in datacenter networks. In this paper, we propose eBA, an efficient solution to bandwidth allocation that provides end-to-end bandwidth guarantee for VMs under large numbers of short flows and massive bursty traffic in datacenters. eBA leverages a novel distributed VM-to-VM rate control algorithm based on the logistic model under the control-theoretic framework. eBA's implementation requires no changes to hardware or applications and can be deployed in standard protocol stack. The theoretical analysis and the experimental results show that eBA not only guarantees the bandwidth for VMs, but also provides fast convergence to efficiency and fairness, as well as smooth response to bursty traffic.
Fangming Liu, Xiaomeng Huang, John C. S. Lui
IEEE/ACM Trans. Netw.1
2017 An Efficient Online Algorithm for Dynamic SDN Controller Assignment in Data Center Networks
abstract
Software defined networking is increasingly prevalent in data center networks for it enables centralized network configuration and management. However, since switches are statically assigned to controllers and controllers are statically provisioned, traffic dynamics may cause long response time and incur high maintenance cost. To address these issues, we formulate the dynamic controller assignment problem (DCAP) as an online optimization to minimize the total cost caused by response time and maintenance on the cluster of controllers. By applying the randomized fixed horizon control framework, we decompose DCAP into a series of stable matching problems with transfers, guaranteeing a small loss in competitive ratio. Since the matching problem is NP-hard, we propose a hierarchical two-phase algorithm that integrates key concepts from both matching theory and coalitional games to solve it efficiently. Theoretical analysis proves that our algorithm converges to a near-optimal Nash stable solution within tens of iterations. Extensive simulations show that our online approach reduces total cost by about 46%, and achieves better load balancing among controllers compared with static assignment.
Tao Wang 0088, Fangming Liu, Hong Xu 0001
IEEE/ACM Trans. Netw.2
2016 Not All Joules are Equal: Towards Energy-Efficient and Green-Aware Data Processing Frameworks
abstract
Interests have been growing in integrating renewable energy into data centers, which attracts many research efforts in developing green-aware algorithms and systems. However, little attention was paid to the efficiency of each joule consumed by data center workloads. In fact, not all joules are equal in the sense that the amount of work that can be done by a joule can vary significantly in data centers. Ignoring this fact leads to significant energy waste (by 25% of the total energy consumption in Hadoop YARN on a Facebook production trace according to our study). In this paper, we investigate how to exploit such joule efficiency to maximize the benefits of renewable energy for MapReduce framework. We develop job/task scheduling algorithms with a particular focus on the factors on joule efficiency in the data center, including the energy efficiency of MapReduce workloads, renewable energy supply and the battery usage. We further develop a simple yet effective performance-energy consumption model to guide our scheduling decisions. We have implemented GreenMR, an energy-efficient and green-aware MapReduce framework, on top of Hadoop YARN. The experiments demonstrate the accuracy of our models, and the effectiveness of our energy-efficient and green-aware optimizations over Hadoop YARN and a state-ofthe-art green-aware Hadoop YARN implementation.
Zhaojie Niu, Bingsheng He, Fangming Liu
IC2E3
2016 Flexible Instance: Meeting Deadlines of Delay Tolerant Jobs in the Cloud with Dynamic Pricing
abstract
A wide range of cloud computing jobs are delay tolerant up to a predefined deadline. Existing IaaS services offer either high cost and high fulfillment ratio or low cost without fulfillment ratio guarantee, where the fulfillment ratio is the ratio of job execution time to the time between job submission and completion. Neither of the services represents a cost-effective way to exploit job elasticity. This work proposes flexible instance, a cloud service where user-specified service fulfillment ratio, as a new pricing factor, is guaranteed by the provider to meet deadlines. Job elasticity is exploited by the provider to enhance resource utilization, by regulating demand fluctuation through computation arbitrage across the temporal domain. We leverage a two-stage pricing framework to agilely adapt cloud resource price to the demand-supply dynamics. The first stage uses an online strategy to reserve resources for each cloud user to guarantee its specified fulfillment ratio. We set the price of resources dynamically according to resource utilization, with a pricing curve O(ln p)-competitive to the optimal fixed-price offline strategy in provider revenue. The second stage allows cloud users to submit a small budget to compete for extra service fulfillment ratio for execution speedup. A Nash bargaining framework is explored to achieve fairness, resource efficiency, and revenue maximization simultaneously. Extensive simulations driven by real-world traces show that flexible instance can reduce user cost for job execution while increasing provider revenue.
Xiaomeng Yi, Fangming Liu, Zongpeng Li, Hai Jin 0001
ICDCS2
2016 Dynamic SDN controller assignment in data center networks: Stable matching with transfers
abstract
Software defined networking is becoming increasingly prevalent in data center networks for its programmability that enables centralized network configuration and management. However, since switches are statically assigned to controllers, traffic dynamics cause load imbalance among the controllers. As a result, some controllers are not fully utilized, while switches connected to overloaded controllers may experience long response times. In this paper, we consider dynamic controller assignment so as to minimize the average response time of the control plane. We formulate this problem as a stable matching problem with transfers, and propose a hierarchically two-phase algorithm that integrates key concepts from both matching theory and coalitional games to solve it efficiently. Theoretical analysis proves that our algorithm converges to a near-optimal Nash stable solution within tens of iterations. Extensive simulations show that our approach reduces response time by about 86%, and achieves better load balancing among controllers compared to static assignment.
Tao Wang 0088, Fangming Liu, Hong Xu 0001
INFOCOM2
2016 On the performance of cloud storage applications with global measurement
abstract
In recent years, Dropbox, Google, and Microsoft have been competing in the market of consumer cloud storage (CCS) services. While once the key comparative metric, storage capacity per user has outgrown the needs of most users. Today, third-party applications based on CCS's RESTful Web APIs are becoming a primary way for users to utilize their expanded storage resources. Unfortunately, there is very little visibility into the performance of these Web APIs, even though they are primary determinants of the end user experience on these storage applications. In this paper, we report results from a comprehensive measurement study of the Web APIs of five popular CCS providers. Our results reveal significant differences and limitations in API performance, which result in performance bottlenecks visible to the user through the storage application. We analyze the underlying system designs of the five providers' Web APIs, and present the performance implications of their different design choices. Our research provides practical guidance for service providers to optimize their API performance, for developers to improve the experience of third-party applications, and for users to pick appropriate services that best match their requirements.
Guangyuan Wu, Fangming Liu, Haowen Tang, Keke Huang, Qixia Zhang, Zhenhua Li 0001, Ben Y. Zhao, Hai Jin 0001
IWQoS2
2016 Demystifying the energy efficiency of Network Function Virtualization
abstract
Middleboxes are prevalent in today's enterprise and data center networks. Network function virtualization (NFV) is a promising technology to replace dedicated hardware middleboxes with virtualized network functions (VNFs) running on commodity servers. However, no prior study has examined the energy efficiency of different NFV implementations. In this paper, we conduct a measurement study on the power efficiency of software data planes, the virtual I/O and the software middleboxes, which are important parts of the NFV implementations. We run two popular software middleboxes (Snort and Bro) on three common software data planes (i.e., DPDK-OVS, Click Modular Router and Netmap). Our results show significant differences on power among those different NFV implementations. We analyze the underlying design choices and give implications on how to build more power efficient NFV implementations.
Fangming Liu, Tao Wang 0088, Hong Xu 0001
IWQoS2
2016 Bilateral Electricity Trade Between Smart Grids and Green Datacenters: Pricing Models and Performance Evaluation
abstract
Datacenter demand response is a promising approach for mitigating operational instability faced by smart grids. It enables significant potentials in peak load shedding and facilitates the incorporation of distributed generation and intermittent energy sources. This paper considers two key aspects toward real-time electricity pricing for eliciting demand response: 1) two-way electricity flow between smart grids and large datacenters with hybrid green generation capabilities and 2) the geo-distributed nature of large cloud systems, and hence the potential competition among smart grids that serve different datacenters of the cloud. We propose a pricing scheme tailored for geo-distributed green datacenters, from a multi-leader (smart grids) single-follower (cloud) game point of view. At the cloud side, in quest for scalability, robustness, and performance, the energy cost minimization problem is solved in a distributed manner, based on the technique of alternating direction method of multipliers. At the smart grid side, a practical equilibrium of the multi-leader single-follower pricing game is desired. To this end, we employ the technique of equilibrium problem with equilibrium constraints and exact linearization, to accurately transform the multi-leader single-follower pricing game, which is non-convex into a mixed integer linear system that can be readily solved. The effectiveness of the proposed solutions is evaluated based on the real datacenter workload traces and the IEEE 14-bus test systems with real generation and demand data.
Zhi Zhou 0006, Fangming Liu, Zongpeng Li
IEEE J. Sel. Areas Commun.2
2016 Heterogeneity and Interference-Aware Virtual Machine Provisioning for Predictable Performance in the Cloud
abstract
Infrastructure-as-a-service (IaaS) cloud providers offer tenants elastic computing resources in the form of virtual machine (VM) instances to run their jobs. Recently, providing predictable performance (i.e., performance guarantee) for tenant applications is becoming increasingly compelling in IaaS clouds. However, the hardware heterogeneity and performance interference across the same type of cloud VM instances can bring substantial performance variation to tenant applications, which inevitably stops the tenants from moving their performance-sensitive applications to the IaaS cloud. To tackle this issue, this paper proposes Heifer, a Heterogeneity and interference-aware VM provisioning framework for tenant applications, by focusing on MapReduce as a representative cloud application. It predicts the performance of MapReduce applications by designing a lightweight performance model using the online-measured resource utilization and capturing VM interference. Based on such a performance model, Heifer provisions the VM instances of the good-performing hardware type (i.e., the hardware that achieves the best application performance) to achieve predictable performance for tenant applications, by explicitly exploring the hardware heterogeneity and capturing VM interference. With extensive prototype experiments in our local private cloud and a real-world public cloud (i.e., Microsoft Azure) as well as complementary large-scale simulations, we demonstrate that Heifer can guarantee the job performance while saving the job budget for tenants. Moreover, our evaluation results show that Heifer can improve the job throughput of cloud datacenters, such that the revenue of cloud providers can be increased, thereby achieving a win-win situation between providers and tenants.
Fei Xu 0009, Fangming Liu, Hai Jin 0001
IEEE Trans. Computers2
2016 A Framework for Truthful Online Auctions in Cloud Computing with Heterogeneous User Demands
abstract
Auction-style pricing policies can effectively reflect the underlying trends in demand and supply for the cloud resources, and thereby attracted a research interest recently. In particular, a desirable cloud auction design should be (1) online to timely reflect the fluctuation of supply-demand relations, (2) expressive to support the heterogeneous user demands, and (3) truthful to discourage users from cheating behaviors. Meeting these requirements simultaneously is non-trivial, and most existing auction mechanism designs do not directly apply. To meet these goals, this paper conducts the first work on a framework for truthful online cloud auctions where users with heterogeneous demands could come and leave on the fly. Concretely speaking, we first design a novel bidding language, wherein users' heterogeneous requirement on their desired allocation time, application type, and even how they value among different possible allocations can be flexibly and concisely expressed. Besides, building on top of our bidding language we propose COCA, an incentive-Compatible (truthful) Online Cloud Auction mechanism. To ensure truthfulness with heterogenous and online user demand, the design of COCA is driven by a monotonic payment rule and a utility-maximizing allocation rule. Moreover, our theoretical analysis shows that the worst-case performance of COCA can be well-bounded, and our further discussion shows that COCA performs well when some other important factors in online auction design are taken into consideration. Finally, in simulations the performance of COCA is seen to be comparable to the well-known off-line Vickrey-Clarke-Groves (VCG) mechanism [19].
Hong Zhang 0025, Hongbo Jiang 0001, Bo Li 0001, Fangming Liu, Athanasios V. Vasilakos, Jiangchuan Liu
IEEE Trans. Computers4
2016 Fair Network Bandwidth Allocation in IaaS Datacenters via a Cooperative Game Approach
abstract
With wide application of virtualization technology, tenants are able to access isolated cloud services by renting the shared resources in Infrastructure-as-a-Service (IaaS) datacenters. Unlike resources such as CPU and memory, datacenter network, which relies on traditional transport-layer protocols, suffers unfairness due to a lack of virtual machine (VM)-level bandwidth guarantees. In this paper, we model the datacenter bandwidth allocation as a cooperative game, toward VM-based fairness across the datacenter with two main objectives: 1) guarantee bandwidth for VMs based on their base bandwidth requirements, and 2) share residual bandwidth in proportion to the weights of VMs. Through a bargaining game approach, we propose a bandwidth allocation algorithm, Falloc, to achieve the asymmetric Nash bargaining solution (NBS) in datacenter networks, which exactly meets our objectives. The cooperative structure of the algorithm is exploited to develop an online algorithm for practical real-world implementation. We validate Falloc with experiments under diverse scenarios and show that by adapting to different network requirements of VMs, Falloc can achieve fairness among VMs and balance the tradeoff between bandwidth guarantee and proportional bandwidth sharing. Our large-scale trace-driven simulations verify that Falloc achieves high utilization while maintaining fairness among VMs in datacenters.
Fangming Liu, John C. S. Lui, Hai Jin 0001
IEEE/ACM Trans. Netw.2
2016 Carbon-Aware Online Control of Geo-Distributed Cloud Services
abstract
Recently, datacenter carbon emission has become an emerging concern for the cloud service providers. Previous works are limited on cutting down the power consumption of datacenters to defuse such a concern. In this paper, we show how the spatial and temporal variabilities of the electricity carbon footprint can be fully exploited to further green the cloud running on top of geographically distributed datacenters. Specifically, we first verify that electricity cost minimization conflicts with carbon emission minimization, based on an empirical study of several representative geo-distributed cloud services. We then jointly consider the electricity cost, service level agreement (SLA) requirement, and emission reduction budget. To navigate such a three-way tradeoff, we take advantage of Lyapunov optimization techniques to design and analyze a carbon-aware control framework, which makes online decisions on geographical load balancing, capacity right-sizing, and server speed scaling. Results from rigorous mathematical analysis and real-world trace-driven evaluation demonstrate the effectiveness of our framework in reducing both electricity cost and carbon emission.
Zhi Zhou 0006, Fangming Liu, Ruolan Zou, Jiangchuan Liu, Hong Xu 0001, Hai Jin 0001
IEEE Trans. Parallel Distributed Syst.2
2015 Cost-Effective Service Provisioning for Hybrid Cloud Applications
Yipei Niu, Fangming Liu
CollaborateCom3
2015 Harnessing Dynamic Interests of Crowd in Chinese Online Shopping Festivals
abstract
Online shopping is popular in promotion festivals like Singles' Day in China. With the presence of a flash crowd of customers during the limited time in E-commerce websites, the traditional system provides multi-dimensional product ranking lists including popularity based bought lists and viewed lists due to the customers' conformity to the popular products. However, the overloaded information may cause a long path to purchase for customers, who can not efficiently buy satisfied products within the time-limited festivals. To shorten customers' purchasing path, this paper designs an interests-based ranking method to automatically and efficiently aggregate an unified ranking list, by analyzing thousands of customers' shopping behavior to harness the dynamic and wide interests of crowd. Initially, we transform the local interaction behavior of flash crowd into the interests of crowd, which then is further personalized to an unified ranking for each customer. We combine user's temporal interests and current crowd wisdom to generate a new personalized unified ranking. The combined ranking is continuously adjusted with the recent fixed past time window data iteratively according to the evolution growth of crowd. An implementation test is carried out on a real world trace in large-scale interaction logs, which is from Alibaba Tmall with 12 million users and 29 thousand brands. It shows that combining the user's temporal interests and the aggregated crowd wisdom ranking improve the quality of item recommendations distinctly.
Linhai He, Hai Jin 0001, Fangming Liu, Xijiang Ke
COMPSAC4
2015 eTrain: Making Wasted Energy Useful by Utilizing Heartbeats for Mobile Data Transmissions
abstract
With the rapid proliferation of smartphones, hundreds of millions of mobile users are attracted to Instant Messaging (IM) apps. While such apps have brought convenience to our life, it comes with the price of great energy consumption, as these apps keep sending heartbeat messages to the server periodically in order to maintain an always-online connection. These frequent and fragmented transmissions result in a considerable amount of energy waste. In this paper, we investigate the "cost and potential of heartbeats". We quantify power consumption of heartbeats of real-world IM apps through extensive measurements. The measurement results confirm that huge power consumption is induced by heartbeats. The goal of this paper is to save energy by turning the energy wastage of heartbeats into transmitting useful data. Thus, we develop eTrain, a transmission management system running on Android phones, which takes advantage of IM heartbeats (as trains) to piggyback aggregated delay-tolerant apps' data such as e-mail and Weibo (as cargoes) via an online transmission strategy, so as to minimize the cumulative tail energy without sacrificing user-specified deadlines. Compared to other existing works, eTrain can reduce more energy consumption under the same settings. Experiments conducted on smartphones show that eTrain can achieve 12%-33% energy saving in various application scenarios.
Tan Zhang, Fangming Liu, Hongkun Leng, Guanfeng Liang
ICDCS3
2015 When hybrid cloud meets flash crowd: Towards cost-effective service provisioning
abstract
With rapid development in online shopping, e-commerce websites are facing intensive user requests from an increasing number of customers. Especially in promotion seasons, these websites may encounter flash crowds which pull heavy pressure o private infrastructure and even make he website unavailable. Such severe flash crowds can be addressed by leveraging hybrid cloud solution, which relieves workloads of the private cloud by offloading the excessive user requests to the IaaS public cloud. However, the bursty and fluctuation of flash crowds bring challenges to distributing user requests with targest of delay-minimizing and cost-saving. In his paper, we apply the queueing theory to evaluate the average response time and explore the tradeoff between performance and cost in the hybrid cloud. By taking advantage of Lyapunov optimization techniques, we design an online decision algorithm for request distribution which achieves the average response time arbitrarily close to the theoretically optimum and controls he outsourcing cost based on a given budge. The simulation results demonstrate ha in a hybrid cloud, our solution can reduce he cost of e-commerce services as well as guarantee performance when encountering flash crowds.
Yipei Niu, Fangming Liu, Jiangchuan Liu, Bo Li 0001
INFOCOM3
2015 When smart grid meets geo-distributed cloud: An auction approach to datacenter demand response
abstract
Datacenter demand response is envisioned as a promising tool for mitigating operational stability issues faced by smart grids. It enables significant potentials in peak load reduction and facilitates the incorporation of distributed generation. Monetary refund from the smart grid can also alleviate the cloud's burden in escalating electricity cost. However, the current demand response paradigm is inefficient towards incentivizing a cloud that runs over geo-distributed datacenters. Leveraging auction theory, this work presents an efficient incentive mechanism to elicit demand response from geo-distributed clouds. To determine the winning bids and their corresponding payments, the cloud that acts as the auctioneer needs to solve a set of winner determination problems that are highly challenging. By integrating techniques from the Gibbs sampling method and the alternating direction method of multipliers, we propose a decentralized algorithm for each datacenter to make autonomous decisions on winning bid selection and workload management, striking a balance among the economic efficiency, truthfulness and the computational efficiency. Through extensive trace-driven evaluations, we demonstrate that our incentive mechanism constitutes a win-win mechanism for both the geo-distributed cloud and the smart grid.
Zhi Zhou 0009, Fangming Liu, Zongpeng Li, Hai Jin 0001
INFOCOM2
2015 Zero-Cost, Fine-Grained Power Monitoring of Datacenters Using Non-Intrusive Power Disaggregation
abstract
Fine-grained power monitoring, which refers to power monitoring at the server level, is critical to the efficient operation and energy saving of datacenters. Fined-grained power monitoring, however, is extremely challenging in legacy datacenters that host server systems not equipped with power monitoring sensors. Installing power monitoring hardware at the server level not only incurs high costs but also complicates the maintenance of high-density server clusters and enclosures. In this paper, we present a zero-cost, purely software-based solution to this challenging problem. We use a novel technique of non-intrusive power disaggregation (NIPD) that establishes power mapping functions (PMFs) between the states of servers and their power consumption, and infer the power consumption of each server with the aggregated power of the entire datacenter. We implement and evaluate NIPD over a real-world datacenter with 326 nodes. The results show that our solution can provide high precision power estimation at the rack level, with mean relative error of 2.63%, and the server level, with mean relative error of 10.27% and 8.17% for the estimation of idle power and peak power, respectively.
Guoming Tang, Weixiang Jiang, Fangming Liu, Kui Wu 0001
Middleware4
2015 UniDrive: Synergize Multiple Consumer Cloud Storage Services
abstract
Consumer cloud storage (CCS) services have become popular among users for storing and synchronizing files via apps installed on their devices. A single CCS, however, has intrinsic limitations on networking performance, service reliability, and data security. To overcome these limitations, we present UniDrive, a CCS app that synergizes multiple CCSs (multi-cloud) by using only few simple public RESTful Web APIs. UniDrive follows a server-less, client-centric design, in which synchronization logic is purely implemented at client devices and all communication is conveyed through file upload and download operations. Strong consistency of the metadata is guaranteed via a quorum-based distributed mutual-exclusive lock mechanism. UniDrive improves reliability and security by judiciously distributing erasure coded files across multiple CCSs. To boost networking performance, UniDrive leverages all available clouds to maximize parallel transfer opportunities, but the key insight behind is the concept of data block over-provisioning and dynamic scheduling. This suite of techniques masks the diversified and varying network conditions of the underlying clouds, and exploits more the faster clouds via a simple yet effective in-channel probing scheme. Extensive experimental results on the global Amazon EC2 platform and a real-world trial by 272 users confirmed significantly superior and consistent sync performance of UniDrive over any single CCS.
Haowen Tang, Fangming Liu, Guobin Shen, Chuanxiong Guo
Middleware2
2015 Pricing Bilateral Electricity Trade between Smart Grids and Hybrid Green Datacenters
abstract
Datacenter demand response is envisioned as a promising approach for mitigating operational instability faced by smart grids. It enables significant potentials in peak load shedding and facilitates the incorporation of distributed generation and intermittent energy sources. This work considers two key aspects towards realtime electricity pricing for eliciting demand response: (i) Two-way electricity flow between smart grids and large datacenters with hybrid green generation capabilities. (ii) The geo-distributed nature of large cloud systems, and hence the potential competition among smart grids that serve different datacenters of the cloud. We propose a pricing scheme tailored for geo-distributed green datacenters, from a multi-leader single-follower game point of view. At the cloud side, in quest for performance, scalability and robustness, the energy cost is minimized in a distributed manner, based on the technique of alternating direction of multipliers (ADMM). At the smart grid side, a practical equilibrium of the pricing game is desired. To this end, we employ mathematical programming with equilibrium constraints (MPEC), equilibrium problem with equilibrium constraints (EPEC) and exact linearization, to transform the multi-leader single-follower pricing game into a mixed integer linear program (MILP) that can be readily solved. The effectiveness of the proposed solutions is evaluated based on trace-driven simulations.
Zhi Zhou 0009, Fangming Liu, Zongpeng Li
SIGMETRICS2
2015 AppATP: An Energy Conserving Adaptive Mobile-Cloud Transmission Protocol
abstract
Many mobile applications require frequent wireless transmissions between the content provider and mobile devices, consuming much energy in mobile devices. Motivated by the popularity of prefetch-friendly or delay-tolerant apps (e.g., social networking, app updates, cloud storage), we design and implement an application-layer transmission protocol, AppATP, which leverages cloud computing to manage data transmissions for mobile apps, transferring data to and from mobile devices in an energy efficient manner. Measurements show that significantly amount of energy is consumed by mobile devices during poor connectivity. Based on this observation, AppATP adaptively seizes periods of good bandwidth condition to prefetch frequently used data with minimum energy consumption, while deferring delay-tolerant data during poor network connectivity. Using the stochastic control framework, AppATP only relies on the current network information and data queue sizes to make an online decision on transmission scheduling, and performs well under unpredictable wireless network conditions. We implement AppATP on Samsung Note 2 smartphones and Amazon EC2. Results from both trace-driven simulations and extensive real-world experiments show that AppATP can be applied to a variety of application scenarios while achieving 30-50 percent energy savings for mobile devices.
Fangming Liu, Peng Shu, John C. S. Lui
IEEE Trans. Computers1
2014 Fuel Cell Generation in Geo-Distributed Cloud Services: A Quantitative Study
abstract
The demand for capping carbon emission has promoted the use of fuel cell energy in cloud computing, yet it is unclear what and how much benefit it may bring. This paper, for the first time, attempts to quantitatively examine the benefits brought by fuel cell generation, and to illustrate how such benefits can be realized with an intelligent coordination between grid power and fuel cell generation. Specifically, we propose UFC, a quantitative index called the utility of the cloud using fuel cells, which captures the level of the data enters operator's overall satisfaction from energy cost, carbon emission, and workload performance. We formulate the UFC maximization problem to jointly optimize both fuel cell generation and geographical request routing. In order to avoid centralized solutions with high complexity and low scalability, we develop a distributed algorithm blending the advantages of Alternating Direction Method of Multipliers (ADMM) and the auxiliary variable method, whose performance is evaluated and verified through our extensive simulations based on real-world data enter workload traces, electricity prices and generation data sets.
Zhi Zhou 0009, Fangming Liu, Bo Li 0001, Baochun Li, Hai Jin 0001, Ruolan Zou
ICDCS2
2014 On efficient bandwidth allocation for traffic variability in datacenters
abstract
Datacenter networks suffer unpredictable performance due to a lack of application level bandwidth guarantees. A lot of attentions have been drawn to solve this problem such as how to provide bandwidth guarantees for Virtualized Machines (VMs), proportional bandwidth share among tenants, and high network utilization under peak traffic. However, existing solutions fail to cope with highly dynamic traffic in datacenter networks. In this paper, we consider the effects of large numbers of short flows and massive bursty traffic in the datacenter, and design a novel distributed rate allocation algorithm based on the Logistic model under the control-theoretic framework. The theoretical analysis and experimental results using OpenFlow show that our algorithm not only guarantees the bandwidth for VMs, but also provides fast convergence to efficiency and fairness, and smooth response to bursty traffic.
Fangming Liu, Xiaomeng Huang, John C. S. Lui, Mi Hu, Qiao Gao, Hai Jin 0001
INFOCOM2
2014 Server Selection and Topology Control for Multi-Party Video Conferences
abstract
This paper proposes new methods for multi-server placement and topology control in multi-party video conferences. Given a large server pool available from CDN infrastructures and datacenter networks, our lightweight methods can rapidly determine the network topology and select the best physical servers to deploy virtualized server instances on, with the objective of minimizing the mean end-to-end delay between clients. We propose D-Grouping, a ping-based clustering algorithm, which is used in combination with convex optimization to determine the network topology and fine-tune server selection. To verify the proposed methods, we present extensive simulation studies based on the ping traces collected from 518 PlanetLab nodes as well as real-world experiment results based on a prototype implementation.
Shuopeng Zhang, Di Niu 0002, Yaochen Hu 0001, Fangming Liu
NOSSDAV4
2014 Managing Performance Overhead of Virtual Machines in Cloud Computing: A Survey, State of the Art, and Future Directions
abstract
Infrastructure-as-a-Service (IaaS) cloud computing offers customers (tenants) a scalable and economical way to provision virtual machines (VMs) on demand while charging them only for the leased computing resources by time. However, due to the VM contention on shared computing resources in datacenters, this new computing paradigm inevitably brings noticeable performance overhead (i.e., unpredictable performance) of VMs to tenants, which has become one of the primary issues of the IaaS cloud. Consequently, increasing efforts have recently been devoted to guaranteeing VM performance for tenants. In this survey, we review the state-of-the-art research on managing the performance overhead of VMs, and summarize them under diverse scenarios of the IaaS cloud, ranging from the single-server virtualization, a single mega datacenter, to multiple geodistributed datacenters. Specifically, we unveil the causes of VM performance overhead by illustrating representative scenarios, discuss the performance modeling methods with a particular focus on their accuracy and cost, and compare the overhead mitigation techniques by identifying their effectiveness and implementation complexity. With the obtained insights into the pros and cons of each existing solution, we further bring forth future research challenges pertinent to the modeling methods and mitigation techniques of VM performance overhead in the IaaS cloud.
Fei Xu 0009, Fangming Liu, Hai Jin 0001, Athanasios V. Vasilakos
Proc. IEEE2
2014 iAware: Making Live Migration of Virtual Machines Interference-Aware in the Cloud
abstract
Large-scale datacenters have been widely used to host cloud services, which are typically allocated to different virtual machines (VMs) through resource multiplexing across shared physical servers. Although recent studies have primarily focused on harnessing live migration of VMs to achieve load balancing and power saving among different servers, there has been little attention on the incurred performance interference and cost on both source and destination servers during and after such VM migration. To avoid potential violations of service-level-agreement (SLA) demanded by cloud applications, this paper proposes iAware, a lightweight interference-aware VM live migration strategy. It empirically captures the essential relationships between VM performance interference and key factors that are practically accessible through realistic experiments of benchmark workloads on a Xen virtualized cluster platform. iAware jointly estimates and minimizes both migration and co-location interference among VMs, by designing a simple multi-resource demand-supply model. Extensive experiments and complementary large-scale simulations are conducted to validate the performance gain and runtime overhead of iAware in terms of I/O and network throughput, CPU consumption, and scalability, compared to the traditional interference-unaware VM migration approaches. Moreover, we demonstrate that iAware is flexible enough to cooperate with existing VM scheduling or consolidation policies in a complementary manner, such that the load balancing or power saving can still be achieved without sacrificing performance.
Fei Xu 0009, Fangming Liu, Linghui Liu, Hai Jin 0001, Bo Li 0001, Baochun Li
IEEE Trans. Computers2
2014 On Arbitrating the Power-Performance Tradeoff in SaaS Clouds
abstract
In this paper, we present an analytical framework for characterizing and optimizing the power-performance tradeoff in Software-as-a-Service (SaaS) cloud platforms. Our objectives are two-folded: 1) We maximize the operating revenue when serving heterogeneous SaaS applications with unpredictable user requests. 2) We minimize the power consumption when processing the user requests. To achieve these objectives, we construct a unified profit-maximizing objective to jointly consider revenue and cost in an economic view. An offline solution to maximize the supreme bound of the objective is first developed, to 1) justify the validity of our theoretical model, and 2) establish a benchmark to examine the effectiveness of other control solutions. As a highlight of our contributions, we take advantage of the Lyapunov optimization techniques to design and analyze an optimal yet practical control framework, which makes online decisions on request admission control, routing, and virtual machine (VMs) scheduling. Our control framework can accommodate a variety of design choices and operational requirements in a datacenter. Specifically, buffering facilities can be introduced to alleviate the bursty admitted requests and to improve the robustness of the system, and a power budget can be enforced to improve the datacenter performance (dollar) per watt. Our mathematical analyses and simulations have demonstrated both the optimality (in terms of the cost-effective power-performance tradeoff) and stability (in terms of robustness and adaptivity to time-varying and bursty user requests) achieved by our proposed control framework.
Fangming Liu, Zhi Zhou 0009, Hai Jin 0001, Bo Li 0001, Baochun Li, Hongbo Jiang 0001
IEEE Trans. Parallel Distributed Syst.1
2013 Online control of datacenter power supply under uncertain demand and renewable energy
abstract
Modern Cloud Service Providers (CSPs) equip their Datacenter Power Supply System (DPSS) with multisources to mitigate power cost, carbon emission, and power outage: (1) multi-markets grid with time-varying energy prices, (2) finite capacity of Uninterrupted Power Supply (UPS), and (3) certain volumes of intermittent renewable energy. With the presence of uncertain renewable sources and datacenter power demand, CSPs have a critical challenge: how to design systematical online control policies that best utilize different characteristics of multisources in a complementary manner to deliver reliable energy to datacenters while minimizing DPSS operation cost. Based on a stochastic optimization model that captures characteristics of DPSS, we apply two-stage Lyapunov optimization to design and analyze an online DPSS control algorithm (OnDPSS). OnDPSS makes decisions on fully utilizing renewable energy, two-timescales power purchasing, and UPS charging/discharging without requiring substantial statistics of system dynamics. Our mathematical analyses and one-month trace-driven simulations have demonstrated both the optimality (in terms of tradeoff between minimization of DPSS operational cost and constraint satisfaction on datacenter availability and UPS lifetime) and system stability (in terms of robustness to time-varying power demand and supply) achieved by OnDPSS algorithm.
Fangming Liu, Hai Jin 0001, Xiaofei Liao
ICC2
2013 SmartDPSS: Cost-Minimizing Multi-source Power Supply for Datacenters with Arbitrary Demand
abstract
To tackle soaring power costs, significant carbon emission and unexpected power outage, Cloud Service Providers (CSPs) typically equip their Datacenters with a Power Supply System (DPSS) nurtured by multiple sources: (1) smart grid with time-varying electricity prices, (2) uninterrupted power supply (UPS), and (3) renewable energy with intermittent and uncertain supply. It remains a significant challenge how to operate among multiple power supply sources in a complementary manner, to deliver reliable energy to datacenter users with arbitrary demand over time, while minimizing a CSP's operation cost over the long run. This paper proposes an efficient, online control algorithm for DPSS, SmartDPSS, based on the two-timescale Lyapunov optimization techniques. Without requiring a priori knowledge of system statistics, SmartDPSS allows CSPs to make online decisions on how much power demand, including delay-sensitive demand and delay-tolerant demand, to serve at each time, the amount of power to purchase from the long-term-ahead and realtime grid markets, and charging and discharging of UPS over time, in order to fully leverage the available renewable energy and time-varying prices from the grid markets, for minimum operational cost. We thoroughly analyze the performance of our online control algorithm with rigorous theoretical analysis. We also demonstrate its optimality in terms of operational cost, demand service delay, datacenter availability, system robustness and scalability, using extensive simulations based on one-month worth of traces from live power systems.
Fangming Liu, Hai Jin 0001, Chuan Wu 0001
ICDCS2
2013 Falloc: Fair network bandwidth allocation in IaaS datacenters via a bargaining game approach
abstract
With wide application of virtualization technology, tenants are able to access isolated cloud services by renting the shared resources in datacenters. Unlike resources such as CPU and memory, datacenter network, which relies on traditional transport-layer protocols, suffers unfairness due to a lack of VM-level network isolation. In this paper, we propose Falloc, a new bandwidth allocation protocol, towards VM-based fairness across the datacenter with two main objectives: (i) guarantee bandwidth for VMs based on their base bandwidth requirements, and (ii) share residual bandwidth in proportion to weights of VMs. To design Falloc, we model the datacenter bandwidth allocation as a bargaining game and propose a distributed algorithm to achieve the asymmetric Nash bargaining solution (NBS). We apply the theory to practice by implementing Falloc with OpenFlow in experiments under diversed scenarios, which shows that Falloc can achieve fairness by adapting to different network requirements of VMs, and balance the tradeoff between bandwidth guarantee and proportional bandwidth share. By carrying out large scale trace-driven simulations using real-world Mapreduce workload, we show that Falloc achieves high utilization and maintains fairness among VMs in datacenters.
Fangming Liu, Haowen Tang, Yingnan Lian, Hai Jin 0001, John C. S. Lui
ICNP2
2013 A cooperative game based allocation for sharing data center networks
abstract
In current IaaS datacenters, tenants are suffering unfairness since the network bandwidth is shared in a besteffort manner. To achieve predictable network performance for rented virtual machines (VMs), cloud providers should guarantee minimum bandwidth for VMs or allocate the network bandwidth in a fairness fashion at VM-level. At the same time, the network should be efficiently utilized in order to maximize cloud providers' revenue. In this paper, we model the bandwidth sharing problem as a Nash bargaining game, and propose the allocation principles by defining a tunable base bandwidth for each VM. Specifically, we guarantee bandwidth for those VMs with lower network rates than their base bandwidth, while maintaining fairness among other VMs with higher network rates than their base bandwidth. Based on rigorous cooperative game-theoretic approaches, we design a distributed algorithm to achieve efficient and fair bandwidth allocation corresponding to the Nash bargaining solution (NBS). With simulations under typical scenarios, we show that our strategy can meet the two desirable requirements towards predictable performance for tenants as well as high utilization for providers. And by tuning the base bandwidth, our solution can enable cloud providers to flexibly balance the tradeoff between minimum guarantees and fair sharing of datacenter networks.
Fangming Liu, John C. S. Lui, Hai Jin 0001
INFOCOM2
2013 eTime: Energy-efficient transmission between cloud and mobile devices
abstract
Mobile cloud computing, promising to extend the capabilities of resource-constrained mobile devices, is emerging as a new computing paradigm which has fostered a wide range of exciting applications. In this new paradigm, efficient data transmission between the cloud and mobile devices becomes essential. This, however, is highly unreliable and unpredictable due to several uncontrollable factors, particularly the instability and intermittency of wireless connections, fluctuation of communication bandwidth, and user mobility. Consequently, this puts a heavy burden on the energy consumption of mobile devices. Confirmed by our experiments, significantly more energy is consumed during “bad” connectivity. Inspired by the feasibility to schedule data transmissions for prefetching-friendly or delay-tolerant applications, in this paper, we present eTime, a novel Energy-efficient data Transmission strategy between cloud and Mobile dEvices, based on Lyapunov optimization. It aggressively and adaptively seizes the timing of good connectivity to prefetch frequently used data while deferring delay-tolerant data in bad connectivity. To cope with the randomness and unpredictability of wireless connectivity, eTime only relies on the current status information to make a global energy-delay tradeoff decision. Our evaluations from both trace-driven simulation and realworld implementation show that eTime can be applied to various popular applications while achieving 20%-35% energy saving.
Peng Shu, Fangming Liu, Hai Jin 0001, Min Chen 0003, Yupeng Qu, Bo Li 0001
INFOCOM2
2013 A framework for truthful online auctions in cloud computing with heterogeneous user demands
abstract
The paradigm of cloud computing has spontaneously prompted a wide interest in market-based resource allocation mechanisms by which a cloud provider aims at efficiently allocating cloud resources among potential users. Among these mechanisms, auction-style pricing policies, as they can effectively reflect the underlying trends in demand and supply for the computing resources, have attracted a research interest recently. This paper conducts the first work on a framework for truthful online cloud auctions where users with heterogeneous demands could come and leave on the fly. Our framework desirably supports a variety of design requirements, including (1) dynamic design for timely reflecting fluctuation of supply-demand relations, (2) joint design for supporting the heterogeneous user demands, and (3) truthful design for discouraging bidders from cheating behaviors. Concretely speaking, we first design a novel bidding language, wherein users' heterogeneous demands are generalized to regulated and consistent forms. Besides, building on top of our bidding language we propose COCA, an incentive-Compatible (truthful) Online Cloud Auction mechanism based on two proposed guidelines. Our theoretical analysis shows that the worst-case performance of COCA can be well-bounded. Further, in simulations the performance of COCA is seen to be comparable to the well-known off-line Vickrey-Clarke-Groves (VCG) mechanism [11].
Hong Zhang 0025, Bo Li 0001, Hongbo Jiang 0001, Fangming Liu, Athanasios V. Vasilakos, Jiangchuan Liu
INFOCOM4
2013 On arbitrating the power-performance tradeoff in SaaS clouds
abstract
In this paper, we present an analytical framework for characterizing and optimizing the power-performance tradeoff in Software-as-a-Service (SaaS) cloud platforms. Our objectives are two-fold: (1) We maximize the operating profit when serving heterogeneous SaaS applications with unpredictable user requests, and (2) we minimize the power consumption when processing user requests. To achieve these objectives, we take advantage of Lyapunov Optimization techniques to design and analyze an optimal control framework to make online decisions on request admission control, routing, and virtual machine (VMs) scheduling. In particular, our control framework can be flexibly extended to incorporate various design choices and practical requirements of a data-center in the cloud, such as enforcing a certain power budget for improving the performance (dollar) per watt. Our mathematical analyses and simulations have demonstrated both the optimality (in terms of a cost-effective power-performance tradeoff) and system stability (in terms of robustness and adaptivity to time-varying and bursty user requests) achieved by our proposed control framework.
Zhi Zhou 0009, Fangming Liu, Hai Jin 0001, Bo Li 0001, Baochun Li, Hongbo Jiang 0001
INFOCOM2
2013 Carbon-Aware Load Balancing for Geo-distributed Cloud Services
abstract
Recently, data center carbon emission has become an emerging concern for the cloud service providers. Previous works are limited on cutting down the power consumption of the data centers to defuse such a concern. In this paper, we show how the spatial and temporal variabilities of the electricity carbon footprint can be fully exploited to further green the cloud running on top of geographically distributed data centers. We jointly consider the electricity cost, service level agreement (SLA) requirement, and emission reduction budget. To navigate such a three-way tradeoff, we take advantage of Lyapunov optimization techniques to design and analyze a carbon-aware control framework, which makes online decisions on geographical load balancing, capacity right-sizing, and server speed scaling. Results from rigorous mathematical analyses and real-world trace-driven empirical evaluation demonstrate its effectiveness in both minimizing electricity cost and reducing carbon emission.
Zhi Zhou 0009, Fangming Liu, Ruolan Zou, Hong Xu 0001, John C. S. Lui, Hai Jin 0001
MASCOTS2
2013 Cinematic-Quality VoD in a P2P Storage Cloud: Design, Implementation and Measurements
abstract
In this paper, we explore the design space and practice of a new peer-to-peer (P2P) storage cloud, which is capable of replicating, refreshing and on-demand streaming of cinematic-quality video streams, in a decentralized fashion using local storage spaces of end users. We identify key design challenges and tradeoffs in such a P2P storage cloud, and how these are addressed by making informed design choices in a step-by-step fashion. Following our design choices, we have implemented a real-world Video-on-Demand (VoD) system with over 100,000 lines of code, called Novasky, which features new coding-aware peer storage replacement and server push-to-peer strategies, in order to maintain media availability and to balance the system-wide supply-demand relationship in the P2P storage cloud. Since September 2009, it has been deployed in the Tsinghua University campus network, attracting 10,000 users during our measurement studies from February to July 2010, and providing over 1,000 cinematic-quality video streams with bit rates of 1 - 2 Mbps. Based on real-world traces collected over 6 months, we show that Novasky can achieve rapid startups within 4 - 9 seconds and extremely short seek latencies within 3 seconds, while maintaining reasonable operational overhead and server bandwidth costs. Our general understanding on the design tradeoffs of P2P storage cloud and practical experiences with Novasky may bring valuable guidelines to future designs of production-quality P2P storage cloud systems.
Fangming Liu, Shijun Shen, Bo Li 0001, Baochun Li, Hai Jin 0001
IEEE J. Sel. Areas Commun.1
2013 Peer-Assisted On-Demand Streaming: Characterizing Demands and Optimizing Supplies
abstract
Nowadays, there has been significant deployment of peer-assisted on-demand streaming services over the Internet. Two of the most unique and salient features in a peer-assisted on-demand streaming system are the differentiation in the demand (or request) and the prefetching capability with caching. In this paper, we develop a theoretical framework based on queuing models, in order to 1) justify the superiority of service prioritization based on a taxonomy of requests, and 2) understand the fundamental principles behind optimal prefetching and caching designs in peer-assisted on-demand streaming systems. The focus is to instruct how limited uploading bandwidth resources and peer caching capacities can be utilized most efficiently to achieve better system performance. To achieve these objectives, we first use priority queuing analysis to prove how service quality and user experience can be statistically guaranteed, by prioritizing requests in the order of significance, including urgent playback (e.g., random seeks or initial startup), normal playback, and prefetching. We then proceed to construct a fine-grained stochastic supply-demand model to investigate peer caching and prefetching as a global optimization problem. This not only provides insights in understanding the fundamental characterization of demand, but also offers guidelines toward optimal prefetching and caching strategies in peer-assisted on-demand streaming systems.
Fangming Liu, Bo Li 0001, Baochun Li, Hai Jin 0001
IEEE Trans. Computers1
2012 Lifetime or energy: Consolidating servers with reliability control in virtualized cloud datacenters
abstract
Server consolidation using virtualization technologies allow cloud-scale datacenters to improve resource utilization and energy efficiency. However, most existing consolidation strategies solely focused on balancing the tradeoff between service-level-agreements (SLAs) desired by cloud applications and energy costs consumed by hosting servers. With the presence of fluctuating workloads in datacenters, the lifetime and reliability of servers under dynamic power-aware consolidation could be adversely impacted by repeated on-off thermal cycles, wear-and-tear and temperature rise. In this paper, we propose a Reliability-Aware server Consolidation stratEgy, named RACE, to address when and how to perform energy-efficient server consolidation in a reliability-friendly and profitable way. The focus is on the characterization and analysis of this problem as a multi-objective optimization, by developing an utility model that unifies multiple constraints on performance SLAs, reliability factors, and energy costs in a holistic manner. An improved grouping genetic algorithm is proposed to search the global optimal solution, which takes advantage of a collection of reliability-aware resource buffering, and virtual machines-to-servers re-mapping heuristics for generating good initial solutions and improving the convergence rate. Extensive simulations are conducted to validate the effectiveness, scalability and overhead of RACE in improving the overall utility of datacenters while avoiding unprofitable consolidation in the long term - compared with pMapper and PADD strategies for server consolidation.
Fangming Liu, Hai Jin 0001, Xiaofei Liao, Haikun Liu, Li Chen 0019
CloudCom2
2012 Discovering a Large Scale Internet Topology: Complementary and Contrast View
abstract
The Chinese Internet, hosting the worlds largest population of Internet users, has largely remained as a black box due to the lack of large-scale measurement infrastructure and various other factors. In this paper, we characterize the Chinese Internet AS topology using measurements from both data and control planes. Our analysis emphasizes on the distinct characteristics of the regional network in comparison with the global Internet. To obtain a complete and accurate view, we combine multiple data sources in a complementary manner, including the traceroute conducted on both domestic servers and international Planetlab hosts at the data plane, as well as the RouteViews-based BGP routing information at the control plane. Our AS graph successfully captures 97% of the observable ASes in Chinese Internet. The topology is validated to be consistent with the well-known PFP model. Based on the measured topology, we further investigate the topological properties of the Chinese Internet by analyzing various graph-theoretic metrics. We found that the Chinese Internet preserves basic topological properties of global internet, such as power-law in node degree distribution and Pareto principle in international link distribution. This offers new insights for the Chinese Internet characterization and practical guidelines for future engineering optimization.
Heungsun Chang, Fangming Liu, Tongyu Zhan, Bo Li 0001
TrustCom3
2012 Collaborative Caching in Wireless Video Streaming Through Resource Auctions
abstract
Recent advances in wireless communications and mobile networking have dramatically increased the popularity of multimedia services for mobile users, with wireless video streaming at their fingertips. To facilitate efficient acquisition of video content, proxy caching has been widely used by wireless service providers (WSPs), which typically deploy cache servers at mobile switching centers (MSCs). However, capacity provisioning of cache servers is challenging, given the dynamic user demands and the limited cache server resources. With increased densities of wireless service deployment, it is increasingly common that mobile users are covered by more than one WSP within an area. This brings opportunities of a collaborative caching paradigm among the cache servers deployed at different MSCs. In this paper, we explore the benefits of collaborative caching in wireless streaming services, addressing both challenges of incentives and truthfulness of selfish WSPs. We propose a collaborative mechanism that maximizes the social welfare in the context of Vickrey-Clarke-Groves (VCG) auctions, in which cache servers cooperate in the trading of their resources in a self-enforcing manner. Experimental results demonstrate that superior performance can be achieved with respect to the quality of video streaming.
Fangming Liu, Bo Li 0001, Baochun Li, Jiangchuan Liu
IEEE J. Sel. Areas Commun.2
2012 Flash Crowd in P2P Live Streaming Systems: Fundamental Characteristics and Design Implications
abstract
Peer-to-peer (P2P) live video streaming systems have recently received substantial attention, with commercial deployment gaining increased popularity in the internet. It is evident from our practical experiences with real-world systems that, it is not uncommon for hundreds of thousands of users to choose to join a program in the first few minutes of a live broadcast. Such a severe flash crowd phenomenon in live streaming poses significant challenges in the system design. In this paper, for the first time, we develop a mathematical model to: 1) capture the fundamental relationship between time and scale in P2P live streaming systems under a flash crowd, and 2) explore the design principle of population control to alleviate the impact of the flash crowd. We carry out rigorous analysis that brings forth an in-depth understanding on effects of the gossip protocol and peer dynamics. In particular, we demonstrate that there exists an upper bound on the system scale with respect to a time constraint. By trading peer startup delays in the initial stage of a flash crowd for system scale, we design a simple and flexible population control framework that can alleviate the flash crowd without the requirement of otherwise costly server deployment.
Fangming Liu, Bo Li 0001, Lili Zhong, Baochun Li, Hai Jin 0001, Xiaofei Liao
IEEE Trans. Parallel Distributed Syst.1
2011 Collaborative Caching for Video Streaming among Selfish Wireless Service Providers
abstract
Video streaming is now at the fingertips of mobile users with recent advances in wireless communications and mobile networking. Caching has been widely deployed by wireless service providers (WSPs) to facilitate video content dissemination. Yet, capacity provisioning of cache servers is challenging given dynamic user demands and limited wireless bandwidth resources available.With increased densities of wireless service deployment, it is common that mobile users are now covered by more than one WSP within a geographical region. This brings both challenges and opportunities towards a collaborative caching paradigm among cache servers that are deployed by different WSPs. This paper explores the benefits of collaborative caching for wireless video streaming services, addressing challenges related to both incentives and truthfulness of selfish WSPs. We propose a collaborative mechanism that aims to maximize the social welfare in the context of Vickrey-Clarke- Groves (VCG) auctions, which encourages cache servers to spontaneously cooperate for trading their resources in a self-enforcing manner. Results from simulations demonstrate significant performance improvements with respect to video streaming quality.
Bo Li 0001, Fangming Liu, Baochun Li, Jiangchuan Liu
GLOBECOM3
2011 On the efficiency of collaborative caching in ISP-aware P2P networks
abstract
Abstract—Collaborative ISP caching has been advocated to reduce the otherwise significant amount of costly inter-ISP traffic generated by peer-to-peer (P2P) applications. The fundamental design criteria employed by ISP cache servers are, however, not well understood, with respect to dynamic P2P traffic patterns, ISP peering policies and cache server capacity constraints. In particular, there is a lack of investigations on the design and analysis of resource allocation mechanisms with awareness of inter-ISP traffic and ISP policies in the context of collaborative ISP caching — which is our focus in this study. In this paper, by characterizing practical inter-ISP traffic patterns, we have developed a theoretical framework to analyze representative cache resource allocation schemes within the design space of collaborative caching, with a particular focus on minimizing costly inter-ISP traffic. The optimization framework incorporates both locality-aware and locality-unaware peer selection strategies and ISP peering agreements, in order to examine their respective effects on the design of ISP collaborative caching mechanisms. Our analyses not only help us understand the traffic characteristics of existing P2P systems in light of realistic elements, but also offer fundamental insights into designing collaborative ISP caching mechanisms. I.
Bo Li 0001, Fangming Liu, Baochun Li, Hai Jin 0001
INFOCOM3
2011 Novasky: Cinematic-quality VoD in a P2P storage cloud
abstract
In this paper, we present Novasky, a real-world Video-on-Demand (VoD) system capable of delivering cinematic-quality video streams to end users. The foundation of the Novasky design is a peer-to-peer (P2P) storage cloud, storing and refreshing media streams in a decentralized fashion using local storage spaces of end users. We present our design objectives in Novasky, and how these objectives are achieved using a collection of unique mechanisms, with respect to caching strategies, coding mechanisms, and the maintenance of the supply-demand relationship when it comes to media availability in the P2P storage cloud. The production Novasky system has been implemented with over 100,000 lines of code. It has been deployed in the Tsinghua University campus network, operational since September 2009, attracting 10,000 users to date, and providing over 1,000 cinematic-quality video streams with bit rates of 1 - 2 Mbps. Based on real-world traces collected over 6 months, we show that Novasky can achieve rapid startups within 4 - 9 seconds, and extremely short seek latencies within 3 seconds. Our empirical experiences with Novasky may bring valuable insights to future designs of production-quality P2P storage cloud systems.
Fangming Liu, Shijun Shen, Bo Li 0001, Baochun Li, Sanli Li
INFOCOM1
2010 Volumetric-based detection scheme for multi-antenna FH/MFSK systems in the presence of multi-follower jamming
Fangming Liu
Signal Process.1
2010 FS2You: Peer-Assisted Semipersistent Online Hosting at a Large Scale
abstract
It has been widely acknowledged that online file hosting systems within the “cloud” of the Internet have provided valuable services to end users who wish to share files of any size. Such online hosting services are typically provided by dedicated servers, either in content distribution networks (CDNs) or large data centers. Server bandwidth costs, however, are prohibitive in these cases, especially when serving large volumes of files to a large number of users. Though it seems intuitive to take advantage of peer upload bandwidth to mitigate such server bandwidth costs in a complementary fashion, it is not trivial to design and fine-tune important aspects of such peer-assisted online hosting in a real-world large-scale deployment. This paper presents FS2You, a large-scale and real-world online file hosting system with peer assistance and semipersistent file availability. FS2You is designed to dramatically mitigate server bandwidth costs. In this paper, we show a number of key challenges involved in such a design objective, our architectural and protocol design in response to these challenges, as well as an extensive measurement study at a large scale to demonstrate the effectiveness of our design, using real-world traces that we have collected. To our knowledge, this paper represents the first attempt to design, implement, and evaluate a new peer-assisted semipersistent online file hosting system at a realistic scale. Since the launch of FS2You, it has quickly become one of the most popular online file hosting systems in mainland China, and a favorite in many online forums across the country.
Fangming Liu, Bo Li 0001, Baochun Li
IEEE Trans. Parallel Distributed Syst.1
2009 Understanding the Roles of Servers in Large-Scale Peer-Assisted Online Storage Systems
abstract
Online storage systems that provide versatile and convenient platforms for content distribution have attracted significant attention over the Internet. To guarantee adequate levels of service quality and to minimize server cost, such systems typically deploy dedicated servers while effectively utilizing peer bandwidth in a complementary fashion. It is essential to understand the role of servers and critical factors that influence the server contributions. In this paper, with full knowledge of internal mechanisms of a large-scale peer-assisted online storage system, namely FS2You, and large set of real-world traces, we examine the role of servers in such a system. Specifically, through analyzing server traffic volumes versus various critical factors including file popularity, time period (both "cold" and "hot" periods), and peer types, we not only reveal empirical observations that are contrary to general belief with in-depth rationales, but also exploit potential flaws of current design and strategy, which further draw practical implications on future design.
Fangming Liu, Bo Li 0001
ICC1
2009 Quota: Rationing Server Resources in Peer-Assisted Online Hosting Systems
abstract
The increasingly popular online hosting systems are designed to provide versatile and convenient platforms for content hosting and sharing. To guarantee adequate levels of service quality while conserving prohibitive server costs, such systems are often designed to integrate peer bandwidth contributions with strategic server resource provisioning in a complementary and transparent manner. This paper seeks to explore the design space of new protocols to allocate scarce server resources-including both storage space and bandwidth-in peer-assisted online hosting systems. The objective is to maximize the use of limited server storage and bandwidth resources to guarantee adequate levels of service quality, with respect to file availability and downloading performance, while taking full advantage of peer assistance. We identify a number of unique challenges involved in such systems, and propose our design of resource allocation protocols to address these challenges, based on both mathematical analysis and practical implementations. Using real world data sets that we have collected, we evaluate our protocol design through extensive experimental studies from different perspectives, which demonstrate the effectiveness of our design and offer a number of practical guidelines.
Fangming Liu, Bo Li 0001, Baochun Li
ICNP1
2009 FS2You: Peer-Assisted Semi-Persistent Online Storage at a Large Scale
abstract
It has been widely acknowledged that online storage systems within the "cloud" of the Internet provide services of a substantial value to end users who wish to share files of any sizes within a group. Such online storage services are typically provided by dedicated servers, either in content distribution networks (CDNs) or large data centers. Server bandwidth costs, however, are prohibitive in these cases, especially when serving large volumes of files to a large number of users. Though it seems intuitive to take advantage of peer upload bandwidth to mitigate such server bandwidth costs in a complementary fashion, it is not trivial to design and fine-tune important aspects of such peer-assisted online storage in a real-world large-scale deployment. This paper presents FS2You, a large-scale and real-world online storage system with peer assistance and semi-persistent file availability, in order to dramatically mitigate server bandwidth costs. In this paper, we show a number of challenges involved in such a design objective, our architectural and protocol design in response to these challenges, as well as an extensive measurement study at a large scale to demonstrate the effectiveness of our design, using real-world traces that we have collected. To our knowledge, this paper represents the first attempt to design, implement, and evaluate a new peer-assisted semi-persistent online storage system at a realistic scale. Since the launch of FS2You, it has quickly become one of the most popular online storage systems in mainland China, and a favorite in many online forums across the country.
Fangming Liu, Bo Li 0001, Baochun Li
INFOCOM2
2009 Peer-assisted online storage and distribution: modeling and server strategies
abstract
Peer-assisted online storage and distribution systems have recently enjoyed large-scale deployment gaining increased popularity for multimedia content sharing in the Internet. Such systems typically deploy dedicated servers while effectively leveraging peer bandwidth in a complementary fashion, in order to guarantee adequate levels of service quality and minimize server cost. In this paper, motivated by our recent empirical study on a real-world system, FS2You, we develop a mathematical model to characterize and understand peer-assisted online storage systems serving multiple files of different popularity. Specifically, we examine and compare representative server bandwidth allocation strategies, and investigate the critical performance metrics and factors. We demonstrate that different server strategies may lead to remarkably different service qualities in terms of average downloading times, peer satisfaction levels and service quality differentiation. In particular, the current server strategy in FS2You is able to offer system-wide average downloading times comparable to the theoretical bound derived from our model.
Fangming Liu, Bo Li 0001, Baochun Li
NOSSDAV2
2009 The Sink Node Placement and Performance Implication in Mobile Sensor Networks
Yueming Hu 0001, Yueju Xue, Fangming Liu, Gabriel Yik Keung, Bo Li 0001
Mob. Networks Appl.4
2008 An Empirical Study of Flash Crowd Dynamics in a P2P-Based Live Video Streaming System
abstract
Peer-to-peer (P2P) based live video streaming system has emerged as a promising solution for the Internet video streaming applications, partly evident from commercial deployment of several large-scale P2P streaming systems. The key is to leverage the resources available at end users, which offers great potential to scale in the Internet. Flash crowd poses a unique challenge for live streaming systems, in which there could be potentially hundred of thousands of users joining the system during the initial few minutes of a live program. This adds considerable difficulty in particular for a P2P based system in quickly ramping to a scale that can provide reasonable streaming services for newly incoming peers. In this paper, we examine the system dynamics under flash crowd based on measurements obtained from the Coolstreaming system. We are particularly concerned with the impact and user behaviors during flash crowd. The results reveal a number of interesting observations: (1) the system can scale up to a limit during the flash crowd; (2) there is a strong correlation between the number of short sessions and joining rate due to the resource competition among newly joined peers; (3) the user behavior during flash crowd can be best captured by the number of retries and the impatience time.
Bo Li 0001, Gabriel Yik Keung, Susu Xie, Fangming Liu
GLOBECOM4
2008 A preliminary study of information collection in a mobile sensor network
abstract
Mobile sensor networks are desirable in a variety of application scenarios, in which information collection is no doubt of great importance. In this paper, we present a mobile sensor network architecture consisting of a potentially large number of mobile sensors and a single or multiple stationary s
Yueming Hu 0001, Fangming Liu, Gabriel Yik Keung, Bo Li 0001
QSHINE3