EDBT 2026 Demo / reviewers in the wild / expert
Xiaobo Zhou 0002
dblp:13/6395-2
· DBLP profile ↗
170ranked-venue papers
21as first author
63since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 100 · 8 first-author · 33 since 2021Computer networks · 42 · 9 first-author · 13 since 2021Software engineering, systems software and programming languages · 11 · 2 first-author · 4 since 2021Security and privacy · 9 · 7 since 2021Human-computer interaction and ubiquitous computing · 5 · 5 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | User Perceptions of Responsible Gambling Messages as Nudges for Gambling SafetyabstractNudges are subtle interventions designed to influence user behavior without restricting choice. Responsible gambling messages (RGMs) exemplify such nudges by encouraging safer decision-making in gambling environments. Prior research has examined how pop-up messages influence gambling behavior in experimental settings and has explored the design of effective slogan messages. However, little is known about how different types of RGMs shape users’ real-world gambling behavior and safety. To address this gap, we apply a nudging perspective to examine how RGMs support gambling safety throughout gamblers’ decision-making journey. We conducted semi-structured interviews with 22 gamblers and found that participants were generally aware of RGMs, yet some misunderstood their intended purpose. Participants perceived the safety impact of RGMs as reflected in both attitudinal and behavioral dimensions. We further discuss users’ message reception practices and the effectiveness of RGMs as nudges, and conclude with design implications for promoting gambling safety. Maggie Yongqi Guan, Yaxing Yao, Sio Hong Teng, Xiaobo Zhou 0002, Kanye Ye Wang |
CHI | 4 |
| 2026 | Perceived Impacts and Challenges of Agricultural Information on Short-Form Video Platforms as Rural InfrastructureabstractShort-form video platforms (SVSPs) have rapidly evolved from entertainment applications to essential digital infrastructures in rural areas. As such, they are reshaping how farmers organize production and manage everyday activities. However, little is known about how this transformation impacts agricultural practices directly. To explore this, we conducted semi-structured interviews with 20 farmers engaged in crop, livestock, and aquaculture production. Our findings reveal that farmers perceive SVSPs as infrastructural supports across various farming stages—planning, establishment, protection, and sales—by providing timely access to market opportunities, practical knowledge, and peer networks. However, reliance on SVSPs also introduces challenges, including fragmented and unreliable content, issues of contextual relevance, and tensions between platform dynamics and farming practices. This study contributes to understanding how emerging media infrastructures, like SVSPs, reshape rural production. We also offer design recommendations for building more context-aware, resilient, and inclusive digital support systems for agriculture. Nora Sinong Lu, Kanye Ye Wang, Xiaobo Zhou 0002 |
CHI | 3 |
| 2026 | Exploring the Impacts and Challenges of Vibe Coding Paradigm to Children's Programming Learning and PracticesabstractRecent advances in generative AI have introduced a new programming paradigm—vibe coding, a natural language–driven mode of AI collaboration. While promising for adults, little is known about how children engage with this approach, especially in block-based environments. To explore this gap, we conducted workshops with children of varying Scratch experience (n=41) and interviewed five Scratch teachers. Our study investigates how vibe coding impacts children’s programming learning and practice, and what challenges arise. Findings show that vibe coding has both positive and negative impacts across three key contexts of children’s programming experience: acquisition, application, and creation. Across the stages of vibe coding—goal articulation, information interpretation, and outcome evaluation—children encounter distinct challenges. By examining the mismatches between core assumptions of vibe coding and children’s needs, and analyzing its applicability across different contexts, we offer child-centered design implications for future vibe coding systems and GenAI tools. Janice Jianing Si, Qiuning Wang, Alicia Wanyi Liu, Xin Lin 0006, Yujun Zhu, Xiaobo Zhou 0002, April Yi Wang, Kanye Ye Wang |
CHI | 7 |
| 2026 | Nexus: Communication-Aware Role Differentiation for Adaptive Multi-Robot Exploration
Rui Ge 0010, Huanghuang Liang, Jianqi Ma, Chuang Hu, Xiaobo Zhou 0002, Dazhao Cheng |
INFOCOM | 7 |
| 2026 | JanusQuant: Accurate and Efficient 2-bit KV Cache Quantization for Long-Context InferenceabstractLong-context large language models (LLMs) have seen widespread adoption in recent years. However, during inference, the key-value (KV) cache—which stores intermediate activations—consumes significant memory, particularly as sequence lengths grow. Quantization offers a promising path to compress KV cache, but existing 2-bit approaches fall short of achieving optimal inference efficiency due to hardware-unfriendly algorithms and system implementations. Chengyu Sun 0001, Yaqi Xia, Hulin Wang, Donglin Yang, Xiaobo Zhou 0002, Dazhao Cheng |
PPoPP | 5 |
| 2026 | FlePo: GPU Multitask Scheduling Optimization Framework for Dynamic ScenesabstractDeep Neural Networks (DNNs) are widely used in intelligent applications, driving increasing computational demands on GPUs. However, modern GPU multitasking scheduling algorithms fail to effectively balance real-time task performance and resource utilization, especially under dynamic workloads with highly variable DNN computational demands. The complex and workload-dependent execution times of DNN kernels often lead to inefficient resource allocation, degraded system throughput, and missed real-time constraints. To address these challenges, we propose Flexible Parallel Orchestrator (FlePo), a GPU multitasking scheduling framework designed to optimize resource utilization and maintain real-time task performance within acceptable limits for soft real-time systems. FlePo integrates two key techniques: Adaptive Padding Dispatch (APD), which dynamically schedules best-effort tasks while leveraging the predictable execution characteristics of DNN kernels to maintain real-time predictability; and Dynamic Parallel Fusion (DPF), which employs kernel fusion to create computational isolation, reducing interference in parallel job execution. By combining offline profiling with online adaptation, FlePo efficiently responds to workload variations. We evaluate FlePo on two heterogeneous GPU platforms, NVIDIA Tesla V100 and AMD MI50, achieving up to a 50% increase in throughput while keeping real-time overhead below 2%. Our work enhances GPU multitasking in dynamic environments, with potential applications in autonomous driving, smart homes, and intelligent healthcare. Huanghuang Liang, Rui Ge 0010, Yaqi Xia, Chuang Hu, Xiaobo Zhou 0002, Dazhao Cheng |
ACM Trans. Auton. Adapt. Syst. | 6 |
| 2026 | Malope: Memory-Aware and Locality-Preserved Graph Neural Network Training
Junkun Shen, Yuezhi Che, Xiaobo Zhou 0002, Dazhao Cheng |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2025 | Digital Safety for Children with Intellectual Disabilities When Using Mobile Devices from Parents' and Teachers' PerspectivesabstractAs mobile devices become increasingly integrated into children's daily lives, digital safety has emerged as a pressing concern, particularly for children with intellectual disabilities (ID), who are more vulnerable due to their cognitive and behavioral challenges. Despite their heightened risk, little research has addressed the unique digital safety issues these children face. To bridge this gap, we conducted semi-structured interviews with parents and special education teachers who are key figures for overseeing the digital access and safety of children with ID. Our findings highlight four primary concerns: imitation of harmful behaviors, accidental misoperation of devices, risks from fraud, and exposure to cyberbullying. To address these, parents and teachers largely rely on proactive educational strategies supported by technical controls and device restrictions. We conclude by emphasizing the need to adapt special education practices to the evolving digital landscape and propose inclusive safety strategies applicable to other at-risk user groups. Janice Jianing Si, Xin Lin 0006, Haorui Cui, Xiaobo Zhou 0002, Kanye Ye Wang |
CCS | 4 |
| 2025 | Understanding the Challenges Students Face in Non-English Programming Environments Due to the Programming Language Transition: A Case Study of Keywords in the Chinese Version of Scratch
Janice Jianing Si, Huanghuang Liang, Chuang Hu, Yujun Zhu, Xiaobo Zhou 0002, Kanye Ye Wang, Dazhao Cheng |
CHI | 6 |
| 2025 | Defending against Attribute Inference Attacks in Post-Training of Recommendation Systems via UnlearningabstractAttribute Inference Attacks (AIAs) pose a significant threat to recommendation systems (RS) by enabling adversaries to use threat models to infer sensitive user attributes like gender or race from user embeddings, resulting in privacy breaches such as unauthorized profiling and discriminatory policies against specific groups. Existing attribute protection methods are primarily applied during training, suffering from significant limitations, such as architectural inflexibility, dependence on interaction data, and potential catastrophic degradation in recommendation performance. To overcome these challenges, we propose AttrCloak, an efficient and effective post-training attribute unlearning (AU) framework that removes sensitive information from user embeddings without altering RS training architectures. AttrCloak employs dual-objective optimization with parameter self-sharing to minimize mutual information between user embeddings and sensitive attributes while preserving recommendation quality. Furthermore, it accommodates data-free scenarios by leveraging regularization loss when interaction data is unavailable. Comprehensive evaluations on four real-world datasets demonstrate AttrCloak's good performance in privacy protection and recommendation performance. Yili Gong, Jiawei Jiang 0001, Chuang Hu, Xiaobo Zhou 0002, Dazhao Cheng |
ICDE | 5 |
| 2025 | LightTrace: A Versatile Ebpf-Enabled Toolkit for Lightweight Distributed TracingabstractDistributed tracing is widely employed for troubleshooting distributed systems such as microservices. Existing tracing systems typically improve one or more of data completeness, non-intrusiveness, or lightweight operation through various data generation strategies. However, no current solution optimizes all three aspects simultaneously, which can introduce significant overhead in I/O-intensive environments that demand both high completeness and minimal intrusion. In this paper, we introduce LightTrace, a novel, eBPF-enabled toolkit that optimizes the data transmission mechanism in distributed tracing. LightTrace integrates seamlessly with mainstream tracing systems and leverages eBPF to reduce end-to-end latency and lower overall system overhead. LightTrace is implemented using a combination of kernel-level eBPF and user-space Golang components. Our evaluation demonstrates that LightTrace decreases the average latency overhead by up to 22.8 % and improves the peak throughput of microservice systems by between$\mathbf{1 2. 3 \%}$and$\mathbf{1 8. 6 \%}$. Furthermore, LightTrace's adaptability across diverse tracing platforms underscores its versatility in various microservice environments. Yanze Zhang, Kanye Ye Wang, Shufan Gong, Huanghuang Liang, Chuang Hu, Xiaobo Zhou 0002 |
ICPADS | 6 |
| 2025 | Zero-shot Federated Unlearning via Transforming from Data-Dependent to Personalized Model-CentricabstractFederated Unlearning (FU) addresses the "right to be forgotten" in federated learning by removing specific client data's contribution without retraining from scratch. Existing FUs are data-dependent, which make the assumption that systems can access original training data or stored historical parameter updates during unlearning. However, the assumption cannot always hold in practice, as users usually request the deletion of client data and historical parameter updates due to privacy concerns or storage limitations. Therefore, it is crucial to develop a zero-shot FU method without such data access. The key challenge is how to distinguish and remove the impact of target clients without data-level information. Motivated by the idea that if we can learn client-specific personalized information from the model instead of data, FU can be model-centric and data-free, we present the first zero-shot FU framework ZeroFU. By embedding client contributions into the model during learning via condition computation, ZeroFU enables the model to possess personalized features for unlearning. The unlearning is achieved using a proposed GAN-based distillation framework that obfuscates the personalized feature of the target client. Evaluations demonstrate its effectiveness in unlearning under non-IID settings. Huanghuang Liang, Jingling Yuan, Jiawei Jiang 0001, Kanye Ye Wang, Chuang Hu, Xiaobo Zhou 0002, Dazhao Cheng |
IJCAI | 7 |
| 2025 | Harnessing Inter-GPU Shared Memory for Seamless MoE Communication-Computation FusionabstractThe Mixture of Experts (MoE) architecture enhances model quality by scaling up model parameters. However, its development is hindered in distributed training scenarios due to significant communication overhead and expert load imbalance. Existing methods, which only allow for coarse-grained overlapping of communication and computation, slightly alleviate communication costs but at the same time, they introduce a notable impairment of computational efficiency. Furthermore, current approaches to addressing load imbalance often compromise model quality. Hulin Wang, Yaqi Xia, Donglin Yang, Xiaobo Zhou 0002, Dazhao Cheng |
PPoPP | 4 |
| 2025 | MXBLAS: Accelerating 8-bit Deep Learning with a Unified Micro-Scaled GEMM LibraryabstractMicro-scaling General Matrix Multiplication (MX-GEMM), which leverages 8-bit micro-scaling format (MX-format) inputs, represents a significant step forward in accelerating deep learning workloads. The MX-format space is diverse, encompassing various scaling patterns and granularities. However, current MX-GEMM implementations typically adopt a model-oriented approach, where format customization is tailored to individual models. This results in three key limitations: rigid problem-kernel coupling, inefficient promotion operations, and overlooked quantization overhead. Weihu Wang, Yaqi Xia, Donglin Yang, Xiaobo Zhou 0002, Dazhao Cheng |
SC | 4 |
| 2025 | HyTiS: Hybrid Tile Scheduling for GPU GEMM with Enhanced Wave Utilization and Cache LocalityabstractGeneral matrix-matrix multiplication (GEMM) is a fundamental operation in both deep learning and scientific computing. To accelerate these workloads, GPUs with a large number of streaming multiprocessors (SMs) are widely used. However, as modern GPUs scale in core count and adopt larger tile sizes, the wave quantization problem induced by partially filled waves results in growing hardware underutilization and substantially degraded performance. Existing solutions to this problem often suffer from low execution efficiency or introduce additional synchronization overhead. Zheng Zhang 0036, Hulin Wang, Donglin Yang, Xiaobo Zhou 0002, Dazhao Cheng |
SC | 5 |
| 2025 | Voltrix: Sparse Matrix-Matrix Multiplication on Tensor Cores with Asynchronous and Balanced Kernel Optimization
Yaqi Xia, Weihu Wang, Donglin Yang, Xiaobo Zhou 0002, Dazhao Cheng |
USENIX ATC | 4 |
| 2025 | Investigating the Impact of Online Community Involvement on Safety Practices and Perceived Risks Among People Who Use Drugs
Nora Sinong Lu, Isaak Hanimann, Janice Jianing Si, Dazhao Cheng, Xiaobo Zhou 0002, Kanye Ye Wang |
USENIX Security Symposium | 6 |
| 2025 | Streamlining Data Transfer in Collaborative SLAM Through Bandwidth-Aware Map DistillationabstractEdge intelligence offers a promising solution for Simultaneous Localization and Mapping (SLAM) in large-scale scenarios, where multiple robots collaboratively perceive the environment and upload their local maps to an edge server. However, maintaining mapping accuracy under constrained and dynamic communication resources remains a significant challenge for the practical deployment of robot swarms. Concurrent data uploads from multiple agents can exacerbate network congestion, leading to the loss of critical information, delayed updates, and, ultimately, the inconsistency of the generated maps. This paper presents Hermes, an edge-assisted collaborative mapping system designed for communication-constrained environments. Hermes streamlines data transfer through bandwidth-aware map distillation, ensuring only the most crucial messages are transmitted to the edge server. We quantify the importance of keyframes and landmarks based on their information entropy gain in pose estimation. By selectively sharing essential submaps, Hermes adaptively balances communication bandwidth and information richness during the mapping process. We implemented Hermes on heterogeneous platforms and conducted experiments using public datasets and self-collected campus data. Hermes exceeds SwarmMap by 50% in bandwidth utilization with similar accuracy and surpasses COVINS-G by 65% in trajectory error under highly constrained network resources. Rui Ge 0010, Huanghuang Liang, Chuang Hu, Xiaobo Zhou 0002, Dazhao Cheng |
IEEE Trans. Mob. Comput. | 5 |
| 2025 | Emotions in Fandom Crowdfunding: Investigating How Online Interactions Affect Collaborative Monetary ActivitiesabstractFandom crowdfunding, where fans collectively raise funds for idols, fosters dynamic interactions within fandom communities, evoking a range of emotions. Despite the prevalence of such activities, the specific emotions involved and their effects on participant behavior remain underexplored. Addressing this, our mixed-methods study—encompassing observations, interviews, and analysis of crowdfunding data—investigated emotions during fandom crowdfunding and their influence on behavior across crowdfunding stages: planning, support, encouragement, realization, and auditing. We identified 10 key emotions related to idols and the community, finding these emotions crucial in shaping participant actions. Our findings highlight the dual impact of fandom crowdfunding on the community’s internal dynamics and its relationships with idols and broader society. We propose design recommendations for enhancing fandom crowdfunding and suggest how general crowdfunding can benefit from insights gained from the fandom context, offering a novel understanding of emotions in collaborative monetary activities. Molly Zhuangtong Huang, Zhicong Lu, Caishi Huang, Zhenning Li 0001, Hantao Zhao, Xiaobo Zhou 0002, Dazhao Cheng, Kanye Ye Wang |
ACM Trans. Comput. Hum. Interact. | 6 |
| 2025 | Spread+: Scalable Model Aggregation in Federated Learning With Non-IID DataabstractFederated learning (FL) addresses privacy concerns by training models without sharing raw data, overcoming the limitations of traditional machine learning paradigms. However, the rise of smart applications has accentuated the heterogeneity in data and devices, which presents significant challenges for FL. In particular, data skewness among participants can compromise model accuracy, while diverse device capabilities lead to aggregation bottlenecks, causing severe model congestion. In this article, we introduce Spread+, a hierarchical system that enhances FL by organizing clients into clusters and delegating model aggregation to edge devices, thus mitigating these challenges. Spread+ leverages hedonic coalition formation game to optimize customer organization and adaptive algorithms to regulate aggregation intervals within and across clusters. Moreover, it refines the aggregation algorithm to boost model accuracy. Our experiments demonstrate that Spread+ significantly alleviates the central aggregation bottleneck and surpasses mainstream benchmarks, achieving performance improvements of 49.58% over FAVG and 22.78% over Ring-allreduce. Huanghuang Liang, Boan Liu, Chuang Hu, Dan Wang 0002, Xiaobo Zhou 0002, Dazhao Cheng |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2025 | m$^{2}$2LLM: A Multi-Dimensional Optimization Framework for LLM Inference on Mobile DevicesabstractLarge Language Models (LLMs) are reshaping mobile AI. Directly deploying LLMs on mobile devices is an emerging paradigm that can widely support different mobile applications while preserving data privacy. However, intensive memory footprint, long inference latency and high energy consumption severely bottlenecks on-device inference of LLM in real-world scenarios. In response to these challenges, this work introduces m$^{2}$LLM, an innovative framework that performs joint optimization from multiple dimensions for on-device LLM inference in order to strike a balance among performance, realtimeliness and energy efficiency. Specifically, m$^{2}$LLM features the following four core components including : 1) Hardware-aware Model Customization, 2) Elastic Chunk-wise Pipeline, 3) Latency-guided Prompt Compression and 4) Layer-wise Resource Scheduling. These four components interact with each other in order to guide the inference process from the following three dimensions. At the model level, m$^{2}$LLM designs an elastic chunk-wise pipeline to expand device memory and customize the model according to the hardware configuration, maximizing performance within the memory budget. At the prompt level, facing the stochastic input, m$^{2}$LLM judiciously compresses the prompts in order to guarantee the first token can be generated in time while maintaining the semantic information. Additionally, at the system level, the layer-wise resource scheduler is employed in order to complete the token generation process with minimized energy consumption while guaranteeing the realtimeness in the highly dynamic mobile environment. m$^{2}$LLM is evaluated on off-the-shelf smartphone with represented models and datasets. Compared to baseline methods, m$^{2}$LLM delivers 2.99–13.5× TTFT acceleration and 2.28–24.3× energy savings, with only a minimal model performance loss of 2% –7% . Xiaobo Zhou 0002, Li Li 0064 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2025 | Featherlight Stateful WebAssembly for Serverless Inference WorkflowsabstractIn serverless inference, complex prediction tasks are executed as workflows, relying on efficient state transfer across multiple functions. Serverless platforms typically deploy each function in a separate stateless container, depending on external processes for state management, which often results in suboptimal system utilization and increased latency. We introduce WasmFlow, a novel framework designed for serverless inference that ensures low latency and high throughput. This is achieved through process-level virtualization using WebAssembly. WasmFlow operates functions on a per-thread basis within compact WebAssembly modules, significantly reducing startup times and memory usage. The framework has two key features. (1) Efficient Memory Sharing: WasmFlow facilitates direct and rapid state transfer between functions using threads within the WebAssembly runtime. This is enabled through lightweight, lock-free, zero-copy intra-process communication, complemented by effective inter-process RPC. (2) System Optimizations: We further optimize WasmFlow with an advanced synchronization technique between functions, an affinity-aware workflow scheduler, and adaptive request batching. Implemented and integrated within the Kubernetes ecosystem, WasmFlow's performance was evaluated using synthetic workloads and realworld Azure traces, including typical serverless workflows and ML models. Our results demonstrate that WasmFlow dramatically outperforms existing serverless frameworks. It reduces P90 end-to-end latency by 74x and 78x, increases function density by n1.7x and 223x compared to Faasm and SPRIGHT, and improves system throughput by 12.3x and 8.8x over Knative and WasmEdge, respectively. Xingguo Pang, Yanze Zhang, Zhuofu Chen, Zhijun Ding, Dazhao Cheng, Xiaobo Zhou 0002 |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2024 | High-throughput Real-time Edge Stream Processing with Topology-Aware Resource MatchingabstractWith the proliferation of Internet of Things (IoT) devices, real-time stream processing at the edge of the network has gained significant attention. However, edge stream processing systems face substantial challenges due to the heterogeneity and constraints of computational and network resources and the intricacies of multi-tenant application hosting. An optimized placement strategy for edge application topology becomes crucial to leverage the advantages offered by Edge computing and enhance the throughput and end-to-end latency of data streams. This paper presents Beaver, a resource scheduling framework designed to deploy stream processing topologies across distributed edge nodes efficiently. Its core is a novel scheduler that employs a synergistic integration of graph partitioning within application topologies and a two-sided matching technique to optimize the strategic placement of stream operators. Beaver aims to achieve optimal performance by minimizing bottlenecks in the network, memory, and CPU resources at the edge. We implemented a prototype of Beaver using Apache Storm and Kubernetes orchestration engine and evaluated its performance using an open-source real-time IoT benchmark (RIoTBench). Compared to state-of-the-art techniques, experimental evaluations demonstrate at least 1.6× improvement in the number of tuples processed within a one-second deadline under varying network delay and bandwidth scenarios. Samee Ullah Khan, Xiaobo Zhou 0002, Palden Lama |
CCGrid | 3 |
| 2024 | From Slow Propagation to Partition: Analyzing Bitcoin Over Anonymous RoutingabstractCryptocurrency is designed for anonymous financial transactions to avoid centralized control, censorship, and regulations. To protect anonymity in the underlying P2P networking, Bitcoin adopts and supports anonymous routing of Tor, I2P, and CJDNS. We analyze the networking performances of these anonymous routing with the focus on their impacts on the blockchain consensus protocol. Compared to non-anonymous routing, anonymous routing adds inherent-by-design latency performance costs due to the additions of the artificial P2P relays. However, we discover that the lack of ecosystem plays an even bigger factor in the performances of the anonymous routing for cryptocurrency blockchain. I2P and CJDNS, both advancing the anonymous routing beyond Tor, in particular lack the ecosystem of sizable networking-peer participation. I2P and CJDNS thus result in the Bitcoin experiencing networking partitioning, which has traditionally been researched and studied in cryptocurrency/blockchain security. We focus on I2P and Tor and compare them with the non-anonymous routing because CJDNS has no active public peers resulting in no connectivity. Tor results in slow propagation while I2P yields soft partition, which is a partition effect long enough to have a substantial impact in the PoW mining. To better study and identify the latency and the ecosystem factors of the cryptocurrency networking and consensus costs, we study the behaviors both in the connection manager (directly involved in the P2P networking) and the address manager (informing the connection manager of the peer selections on the backend). This paper presents our analyses results to inform the state of cryptocurrency blockchain with anonymous routing and discusses future work directions and recommendations to resolve the performance and partition issues. Simeon Wuthier, Kelei Zhang, Xiaobo Zhou 0002, Sang-Yoon Chang |
ICBC | 4 |
| 2024 | Accelerating Distributed DLRM Training with Optimized TT Decomposition and Micro-BatchingabstractDeep Learning Recommendation Models (DLRMs) are pivotal in various sectors, yet they are hindered by the high memory demands of embedding tables and the significant communication overhead in distributed training environments. Traditional approaches, like Tensor-Train (TT) decomposition, although effective for compressing these tables, introduce substantial computational burdens. Furthermore, existing frameworks for distributed training are inadequate due to the excessive data exchange requirements.This paper proposes EcoRec, an advanced library designed to expedite the training of DLRMs through a synergistic integration of TT decomposition technology and distributed training. EcoRec introduces a novel computation pattern that eliminates redundancy in TT operations, alongside an efficient multiplication pathway, significantly reducing computational time. Additionally, it provides a unique micro-batching technique with sorted indices to decrease memory demands without additional computational costs. EcoRec also features a novel pipeline training system for embedding layers, ensuring balanced data distribution and enhanced communication efficiency. EcoRec, built on PyTorch and CUDA, has been evaluated on a 32 GPU cluster. The results show EcoRec significantly outperforms the existing ELRec system, achieving up to a $3.1 \times$ speedup and a 38.5% reduction in memory requirements. EcoRec marks a notable advancement in high-performance DLRM training. Weihu Wang, Yaqi Xia, Donglin Yang, Xiaobo Zhou 0002, Dazhao Cheng |
SC | 4 |
| 2024 | Scaling New Heights: Transformative Cross-GPU Sampling for Training Billion-Edge GraphsabstractEfficient training of Graph Neural Networks (GNNs) on billion-edge graphs poses significant challenges due to memory constraints and data transfer bottlenecks, particularly affecting GPU-based sampling. Traditional methods either face severe CPU-GPU data transfer bottlenecks or encounter excessive data shuffling and synchronization overheads in multi-GPU setups. To overcome these challenges in GNN training on large-scale graphs, we introduce HyDRA, a pioneering framework that elevates mini-batch, sampling-based training. HyDRA innovates in multi-GPU memory sharing and multi-node feature retrieval, transforming cross-GPU sampling by seamlessly integrating sampling and data transfer into a single kernel operation. It develops a hybrid pointer-driven data placement technique to enhance neighbor retrieval efficiency, designs a targeted replication strategy for high-degree vertices to reduce communication overhead, and leverages dynamic cross-batch data orchestration with pipelining to minimize redundant data transfers. Evaluated on systems equipped with up to 64 A100 GPUs, HyDRA significantly outperforms current leading methods, achieving $1.4 x$ to 5.3x faster training speeds compared to DSP and DGL-UVA and demonstrating up to a 42x improvement in multi-GPU scalability. HyDRA sets a new benchmark for high-performance GNN training at large scales. Yaqi Xia, Donglin Yang, Xiaobo Zhou 0002, Dazhao Cheng |
SC | 3 |
| 2024 | MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive OperatorsabstractOperator fusion, a key technique to improve data locality and alleviate GPU memory bandwidth pressure, often fails to extend to the fusion of multiple compute-intensive operators due to saturated computation throughput. However, the dynamicity of tensor dimension sizes could potentially lead to these operators becoming memory-bound, necessitating the generation of fused kernels — a task hindered by limited search spaces for fusion strategies, redundant memory access, and prolonged tuning time, leading to sub-optimal performance and inefficient deployment. We introduce MCFuser, a pioneering framework designed to overcome these obstacles by generating high-performance fused kernels for what we define as memory-bound compute-intensive (MBCI) operator chains. Leveraging high-level tiling expressions to delineate a comprehensive search space, coupled with Directed Acyclic Graph (DAG) analysis to eliminate redundant memory accesses, MCFuser streamlines kernel optimization. By implementing guidelines to prune the search space and incorporating an analytical performance model with a heuristic search, MCFuser not only significantly accelerates the tuning process but also demonstrates superior performance. Benchmarked against leading compilers like Ansor on NVIDIA A100 and RTX3080 GPUs, MCFuser achieves up to a 5.9x speedup in kernel performance and outpaces other baselines while reducing tuning time by over $\mathbf{7 0}$-fold, showcasing its agility. Zheng Zhang 0036, Donglin Yang, Xiaobo Zhou 0002, Dazhao Cheng |
SC | 3 |
| 2024 | Expeditious High-Concurrency MicroVM SnapStart in Persistent Memory with an Augmented Hypervisor
Xingguo Pang, Yanze Zhang, Dazhao Cheng, Cheng-Zhong Xu 0001, Xiaobo Zhou 0002 |
USENIX ATC | 6 |
| 2024 | Base station gateway to secure user channel access at the first hop edge
Sang-Yoon Chang, Arijet Sarker, Simeon Wuthier, Jinoh Kim, Jonghyun Kim 0005, Xiaobo Zhou 0002 |
Comput. Networks | 6 |
| 2024 | Secure and Efficient Authentication using Linkage for permissionless Bitcoin network
Hsiang-Jen Hong, Sang-Yoon Chang, Wenjun Fan, Simeon Wuthier, Xiaobo Zhou 0002 |
Comput. Networks | 5 |
| 2024 | A unified hybrid memory system for scalable deep learning and big data applications
Wei Rang, Huanghuang Liang, Kanye Ye Wang, Xiaobo Zhou 0002, Dazhao Cheng |
J. Parallel Distributed Comput. | 4 |
| 2024 | Incendio: Priority-Based Scheduling for Alleviating Cold Start in Serverless ComputingabstractIn serverless computing, cold start results in long response latency. Existing approaches strive to alleviate the issue by reducing the number of cold starts. However, our measurement based on real-world production traces shows that the minimum number of cold starts does not equate to the minimum response latency, and solely focusing on optimizing the number of cold starts will lead to sub-optimal performance. The root cause is that functions have different priorities in terms of latency benefits by transferring a cold start to a warm start. In this paper, we proposeIncendio, a serverless computing framework exploiting priority-based scheduling to minimize the overall response latency from the perspective of cloud providers. We reveal the priority of a function is correlated to multiple factors and design a priority model based on Spearman’s rank correlation coefficient. We integrate a hybrid Prophet-LightGBM prediction model to dynamically manage runtime pools, which enables the system to prewarm containers in advance and terminate containers at the appropriate time. Furthermore, to satisfy the low-cost and high-accuracy requirements in serverless computing, we propose a Clustered Reinforcement Learning-based function scheduling strategy. The evaluations show that Incendio speeds up the native system by 1.4×, and achieves 23% and 14.8% latency reductions compared to two state-of-the-art approaches. Xinquan Cai, Qianlong Sang, Chuang Hu, Yili Gong, Kun Suo, Xiaobo Zhou 0002, Dazhao Cheng |
IEEE Trans. Computers | 6 |
| 2024 | Raptor-T: A Fused and Memory-Efficient Sparse Transformer for Long and Variable-Length SequencesabstractTransformer-based models have made significant advancements across various domains, largely due to the self-attention mechanism’s ability to capture contextual relationships in input sequences. However, processing long sequences remains computationally expensive for Transformer models, primarily due to theO(n2) complexity associated with self-attention. To address this, sparse attention has been proposed to reduce the quadratic dependency to linear. Nevertheless, deploying the sparse transformer efficiently encounters two major obstacles: 1) Existing system optimizations are less effective for the sparse transformer due to the algorithm’s approximation properties leading to fragmented attention, and 2) the variability of input sequences results in computation and memory access inefficiencies. We present Raptor-T, a cutting-edge transformer framework designed for handling long and variable-length sequences. Raptor-T harnesses the power of the sparse transformer to reduce resource requirements for processing long sequences while also implementing system-level optimizations to accelerate inference performance. To address the fragmented attention issue, Raptor-T employs fused and memory-efficient Multi-Head Attention. Additionally, we introduce an asynchronous data processing method to mitigate GPU-blocking operations caused by sparse attention. Furthermore, Raptor-T minimizes padding for variable-length inputs, effectively reducing the overhead associated with padding and achieving balanced computation on GPUs. In evaluation, we compare Raptor-T’s performance against state-of-the-art frameworks on an NVIDIA A100 GPU. The experimental results demonstrate that Raptor-T outperforms FlashAttention-2 and FasterTransformer, achieving an impressive average end-to-end performance improvement of 3.41X and 3.71X, respectively. Hulin Wang, Donglin Yang, Yaqi Xia, Zheng Zhang 0036, Qigang Wang, Jianping Fan 0007, Xiaobo Zhou 0002, Dazhao Cheng |
IEEE Trans. Computers | 7 |
| 2024 | Corrections to "DNN Surgery: Accelerating DNN Inference on the Edge through Layer Partitioning"abstractIn this paper, we reference the previous conference version and complete the grant number mentioned in the acknowledgments of the conference version. Huanghuang Liang, Qianlong Sang, Chuang Hu, Dazhao Cheng, Xiaobo Zhou 0002, Dan Wang 0002, Wei Bao 0001, Yu Wang 0003 |
IEEE Trans. Cloud Comput. | 5 |
| 2024 | Locality-Aware and Fault-Tolerant Batching for Machine Learning on Distributed DatasetsabstractThe performance of distributed ML training is largely determined by workers that generate gradients in the slowest pace, i.e., stragglers. The state-of-the-art load balancing approaches consider that each worker stores a complete dataset locally and the data fetching time can be ignored. They only consider the computation capacity of workers in equalizing the gradient computation time. However, we find that in scenarios of ML on distributed datasets, whether in edge computing or distributed data cache systems, the data fetching time is non-negligible and often becomes the primary cause of stragglers. In this paper, we present LOFT, an adaptive load balancing approach for ML upon distributed datasets at the edge. It aims to balance the time to generate gradients at each worker while ensuring the model accuracy. Specifically, LOFT features a locality-aware batching. It builds performance and optimization models upon data fetching and gradient computation time. Leveraging the models, it develops an adaptive scheme based on grid search. Furthermore, it offers Byzantine gradient aggregation upon Ring All-Reduce, which makes itself fault-tolerant under Byzantine gradients brought by a small batch size. Experiments with twelve public DNN models and four open datasets show that LOFT reduces the training time by up to 46%, while reducing the training loss by up to 67% compared to LB-BSP. Zhijun Ding, Dazhao Cheng, Xiaobo Zhou 0002 |
IEEE Trans. Cloud Comput. | 4 |
| 2024 | Redundancy-Free and Load-Balanced TGNN Training With Hierarchical Pipeline ParallelismabstractRecently, Temporal Graph Neural Networks (TGNNs), as an extension of Graph Neural Networks, have demonstrated remarkable effectiveness in handling dynamic graph data. Distributed TGNN training requires efficiently tackling temporal dependency, which often leads to excessive cross-device communication that generates significant redundant data. However, existing systems are unable to remove the redundancy in data reuse and transfer, and suffer from severe communication overhead in a distributed setting. This work introduces Sven, a co-designed algorithm-system library aimed at accelerating TGNN training on a multi-GPU platform. Exploiting dependency patterns of TGNN models, we develop a redundancy-free graph organization to mitigate redundant data transfer. Additionally, we investigate communication imbalance issues among devices and formulate the graph partitioning problem as minimizing the maximum communication balance cost, which is proved to be an NP-hard problem. We propose an approximation algorithm called Re-FlexBiCut to tackle this problem. Furthermore, we incorporate prefetching, adaptive micro-batch pipelining, and asynchronous pipelining to present a hierarchical pipelining mechanism that mitigates the communication overhead. Sven represents the first comprehensive optimization solution for scaling memory-based TGNN training. Through extensive experiments conducted on a 64-GPU cluster, Sven demonstrates impressive speedup, ranging from 1.9x to 3.5x, compared to State-of-the-Art approaches. Additionally, Sven achieves up to 5.26x higher communication efficiency and reduces communication imbalance by up to 59.2%. Yaqi Xia, Zheng Zhang 0036, Donglin Yang, Chuang Hu, Xiaobo Zhou 0002, Hongyang Chen 0001, Qianlong Sang, Dazhao Cheng |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2024 | MPMoE: Memory Efficient MoE for Pre-Trained Models With Adaptive Pipeline ParallelismabstractIn recent years, the Mixture-of-Experts (MoE) technique has gained widespread popularity as a means to scale pretrained models to exceptionally large sizes. Dynamic activation of experts allows for conditional computation, increasing the number of parameters of neural networks, which is critical for absorbing the vast amounts of knowledge available in many deep learning areas. However, despite the existing system and algorithm optimizations, there are significant challenges to be tackled when it comes to the inefficiencies of communication and memory consumption. In this paper, we present the design and implementation of MPMoE, a high-performance library that accelerates MoE training with adaptive and memory-efficient pipeline parallelism. Inspired by that the MoE training procedure can be divided into multiple independent sub-stages. We design a pipeline parallelism method for reducing communication latency by overlapping with computation operations. Further, we analyze the memory footprint breakdown of MoE training and identify that activations and temporary buffers are the primary contributors to the overall memory footprint. Toward memory efficiency, we propose memory reuse strategies to reduce memory requirements by eliminating memory redundancies. Finally, to optimize pipeline granularity and memory reuse strategies jointly, we propose a profile-based algorithm and a performance model to determine the configurations of MPMoE at runtime. We implement MPMoE upon PyTorch and evaluate it with common MoE models in two physical clusters, including 64 NVIDIA A100 GPU cards and 16 NVIDIA V100 GPU cards. Compared with the state-of-art approach, MPMoE achieves up to 2.3× speedup while reducing more than 30% memory footprint for training large models. Zheng Zhang 0036, Yaqi Xia, Hulin Wang, Donglin Yang, Chuang Hu, Xiaobo Zhou 0002, Dazhao Cheng |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2024 | DeepTM: Efficient Tensor Management in Heterogeneous Memory for DNN TrainingabstractDeep Neural Networks (DNNs) have gained widespread adoption in diverse fields, including image classification, object detection, and natural language processing. However, training large-scale DNN models often encounters significant memory bottlenecks, which ask for efficient management of extensive tensors. Heterogeneous memory system, which combines persistent memory (PM) modules with traditional DRAM, offers an economically viable solution to address tensor management challenges during DNN training. However, existing memory management methods on heterogeneous memory systems often lead to low PM access efficiency, low bandwidth utilization, and incomplete analysis of model characteristics. To overcome these hurdles, we introduce an efficient tensor management approach, DeepTM, tailored for heterogeneous memory to alleviate memory bottlenecks during DNN training. DeepTM employs page-level tensor aggregation to enhance PM read and write performance and executes contiguous page migration to increase memory bandwidth. Through an analysis of tensor access patterns and model characteristics, we quantify the overall performance and transform the performance optimization problem into the framework of Integer Linear Programming. Additionally, we achieve tensor heat recognition by dynamically adjusting the weights of four key tensor characteristics and develop a global optimization strategy using Deep Reinforcement Learning. To validate the efficacy of our approach, we implement and evaluate DeepTM, utilizing the TensorFlow framework running on a PM-based heterogeneous memory system. The experimental results demonstrate that DeepTM achieves performance improvements of up to 36% and 49% compared to the current state-of-the-art memory management strategies AutoTM and Sentinel, respectively. Furthermore, our solution reduces the overhead by 18 times and achieves up to 29% cost reduction compared to AutoTM. Wei Rang, Hongyang Chen 0001, Xiaobo Zhou 0002, Dazhao Cheng |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2023 | Redundancy-Free High-Performance Dynamic GNN Training with Hierarchical Pipeline ParallelismabstractTemporal Graph Neural Networks(TGNNs) extend the success of Graph Neural Networks to dynamic graphs. Distributed TGNN training requires efficiently tackling temporal dependency, which often leads to excessive cross-device communication that generates significant redundant data. However, existing systems are unable to remove the redundancy in data reuse and transfer, and suffer from severe communication overhead in a distributed setting. This paper presents Sven, an algorithm and system co-designed TGNN training library for the end-to-end performance optimization on multi-node multi-GPU systems. Exploiting dependency patterns of TGNN models and characteristics of dynamic graph datasets, we design redundancy-free data organization and load-balancing partitioning strategies that mitigate the redundant data communication and evenly partition dynamic graphs at the vertex level. Furthermore, we develop a hierarchical pipeline mechanism integrating data prefetching, micro-batch pipelining, and asynchronous pipelining to mitigate the communication overhead. As the first scaling study on the memory-based TGNNs training, experiments conducted on an HPC cluster of 64 GPUs show that Sven can achieve up to 1.7x-3.3x speedup over the state-of-art approaches and a factor of up to 5.26x communication efficiency improvement. Yaqi Xia, Zheng Zhang 0036, Hulin Wang, Donglin Yang, Xiaobo Zhou 0002, Dazhao Cheng |
HPDC | 5 |
| 2023 | Let It Go: Relieving Garbage Collection Pain for Latency Critical Applications in GolangabstractGarbage Collection (GC) is a representative automatic memory manager widely deployed in popular programming languages, such as Java, C\#, and Golang (Go). Through GC, these languages provide programmers with flexibility and safety. However, GC leads to non-trivial overhead in compute and memory resources during application runtime. GC threads compete with non-GC threads (mutators) of an application, which particularly impacts latency-critical (LC) applications and causes long tail latency. Existing GC approaches do not efficiently address the interference, as GC is triggered passively without a global insight of the application; or they employ incremental GC to reduce the interference, while the incremental progress is not dynamically tailored during GC process according to runtime characteristics, which leads to significant performance degradation upon bursty requests. Junxian Zhao, Xiaobo Zhou 0002, Sang-Yoon Chang, Cheng-Zhong Xu 0001 |
HPDC | 2 |
| 2023 | MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline ParallelismabstractRecently, Mixture-of-Experts (MoE) has become one of the most popular techniques to scale pre-trained models to extraordinarily large sizes. Dynamic activation of experts allows for conditional computation, increasing the number of parameters of neural networks, which is critical for absorbing the vast amounts of knowledge available in many deep learning areas. However, despite the existing system and algorithm optimizations, there are significant challenges to be tackled when it comes to the inefficiencies of communication and memory consumption.In this paper, we present the design and implementation of MPipeMoE, a high-performance library that accelerates MoE training with adaptive and memory-efficient pipeline parallelism. Inspired by that the MoE training procedure can be divided into multiple independent sub-stages, we design adaptive pipeline parallelism with an online algorithm to configure the granularity of the pipelining. Further, we analyze the memory footprint breakdown of MoE training and identify that activations and temporary buffers are the primary contributors to the overall memory footprint. Toward memory efficiency, we propose memory reusing strategies to reduce memory requirements by eliminating memory redundancies, and develop an adaptive selection component to determine the optimal strategy that considers both hardware capacities and model characteristics at runtime. We implement MPipeMoE upon PyTorch and evaluate it with common MoE models in a physical cluster consisting of 8 NVIDIA DGX A100 servers. Compared with the state-of-art approach, MPipeMoE achieves up to 2.8× speedup and reduces memory footprint by up to 47% in training large models. Zheng Zhang 0036, Donglin Yang, Yaqi Xia, Liang Ding 0006, Dacheng Tao, Xiaobo Zhou 0002, Dazhao Cheng |
IPDPS | 6 |
| 2023 | Auto-tune: An efficient autonomous multi-path payment routing algorithm for Payment Channel Networks
Hsiang-Jen Hong, Sang-Yoon Chang, Xiaobo Zhou 0002 |
Comput. Networks | 3 |
| 2023 | TAPU: A Transmission-Analytics Processing Unit for Accelerating Multifunctions in IoT GatewaysabstractInternet of Things (IoT) gateways integrate various sensors and compute initial decisions before transmitting data to the cloud for further processing. As the functions they need to support become increasingly complex, gateways must upgrade their hardware. Network functions (NF) and video analytics (VAs) are two typical examples of hardware requirements: NFs need specialized hardware accelerators, while VAs need parallel processing power. However, gateways are typically constrained by factors, such as power, size, and cost, leading to a need to multiplex functions and minimize hardware overprovisioning. This article proposes a novel accelerator, the transmission-analytic processing unit (TAPU), which uses multi-image FPGA to accelerate VAs and NFs for IoT gateways. We preconfigure one image for VAs and one image for NFs, then multiplex the FPGA resources in the time dimension. The TAPU system design requires both hardware and software revisions. In the hardware design, we discuss our considerations on hardware choice and present a new abstraction of hardware functions to overcome the challenge of application development on different multi-image FPGAs. For the software, we develop a fully functional TAPU system to adapt to dynamic network and VAs workloads. Our evaluation shows that TAPU utilization can reach 92%, considerably increasing VAs and network processing throughput over the current approach. We further evaluate TAPU through two case studies that support a campus traffic monitoring system and an office surveillance system, demonstrating excellent performance improvement and low overhead. Huanghuang Liang, Qianlong Sang, Chuang Hu, Yili Gong, Dazhao Cheng, Xiaobo Zhou 0002, Yu Wang 0003 |
IEEE Internet Things J. | 6 |
| 2023 | DNN Surgery: Accelerating DNN Inference on the Edge Through Layer PartitioningabstractRecent advances in deep neural networks have substantially improved the accuracy and speed of various intelligent applications. Nevertheless, one obstacle is that DNN inference imposes a heavy computation burden on end devices, but offloading inference tasks to the cloud causes a large volume of data transmission. Motivated by the fact that the data size of some intermediate DNN layers is significantly smaller than that of raw input data, we designed the DNN surgery, which allows partitioned DNN to be processed at both the edge and cloud while limiting the data transmission. The challenge is twofold: (1) Network dynamics substantially influence the performance of DNN partition, and (2) State-of-the-art DNNs are characterized by a directed acyclic graph rather than a chain, so that partition is incredibly complicated. To solve the issues, We design a Dynamic Adaptive DNN Surgery(DADS) scheme, which optimally partitions the DNN under different network conditions. We also study the partition problem under the cost-constrained system, where the resource of the cloud for inference is limited. Then, a real-world prototype based on the selif-driving car video dataset is implemented, showing that compared with current approaches, DNN surgery can improve latency up to 6.45 times and improve throughput up to 8.31 times. We further evaluate DNN surgery through two case studies where we use DNN surgery to support an indoor intrusion detection application and a campus traffic monitor application, and DNN surgery shows consistently high throughput and low latency. Huanghuang Liang, Qianlong Sang, Chuang Hu, Dazhao Cheng, Xiaobo Zhou 0002, Dan Wang 0002, Wei Bao 0001, Yu Wang 0003 |
IEEE Trans. Cloud Comput. | 5 |
| 2023 | Lightweight and Identifier-Oblivious Engine for Cryptocurrency Networking Anomaly DetectionabstractThe distributed cryptocurrency networking is critical because the information delivered through it drives the mining consensus protocol and the rest of the operations. However, the cryptocurrency peer-to-peer (P2P) network remains vulnerable, and the existing security approaches are either ineffective or inefficient because of the permissionless requirement and the broadcasting overhead. We design and build a Lightweight and Identifier-Oblivious eNgine (LION) for the anomaly detection of the cryptocurrency networking. LION is not only effective in permissionless networking but is also lightweight and practical for the computation-intensive miners. We build LION for anomaly detection and use traffic analyses so that it minimally affects the mining rate and is substantially superior in its computational efficiency than the previous approaches based on machine learning. We implement a LION prototype on an active Bitcoin node to show that LION yields less than 1% of mining rate reduction subject to our prototype, in contrast to the state-of-the-art machine-learning approaches costing 12% or more depending on the algorithms subject to our prototype as well, while having detection accuracy of greater than 97% F1-score against the attack prototypes and real-world anomalies. LION therefore can be deployed on the existing miners without the need to introduce new entities in the cryptocurrency ecosystem. Wenjun Fan, Hsiang-Jen Hong, Jinoh Kim, Simeon Wuthier, Makiya Nakashima, Xiaobo Zhou 0002, C. Edward Chow, Sang-Yoon Chang |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2022 | Holmes: SMT Interference Diagnosis and CPU Scheduling for Job Co-locationabstractCo-location of latency-critical services with best-effort batch jobs is commonly adopted in production systems to increase resource utilization. Although memory and CPU isolation have been extensively studied, we find Simultaneous Multi-Threading (SMT) technology imposes non-trivial interference on memory access which jeopardizes efficient co-location and performance assurance of latency-critical services. However, there is not an existing metric to quantitatively measure and lacks a deterministic approach to tackle SMT interference on memory access. Aidi Pi, Xiaobo Zhou 0002, Cheng-Zhong Xu 0001 |
HPDC | 2 |
| 2022 | The Security Investigation of Ban Score and Misbehavior Tracking in Bitcoin NetworkabstractBitcoin P2P networking is especially vulnerable to networking threats because it is permissionless and does not have the security protections based on the trust in identities, which enables the attackers to manipulate the identities for Sybil and spoofing attacks. The Bitcoin node keeps track of its peer’s networking misbehaviors through ban scores. In this paper, we investigate the security problems of the ban-score mechanism and discover that the ban score is not only ineffective against the Bitcoin Message-based DoS (BM-DoS) attacks but also vulnerable to the Defamation attack as the network adversary can exploit the ban score to defame innocent peers. To defend against these threats, we design an anomaly detection approach that is effective, lightweight, and tailored to the networking threats exploiting Bitcoin’s ban-score mechanism. We prototype our threat discoveries against a real-world Bitcoin node connected to the Bitcoin Mainnet and conduct experiments based on the prototype implementation. The experimental results show that the attacks have devastating impacts on the targeted victim while being cost-effective on the attacker side. For example, an attacker can ban a peer in two milliseconds and reduce the victim’s mining rate by hundreds of thousands of hash computations per second. Furthermore, to counter the threats, we empirically validate our detection countermeasure’s effectiveness and performances against the BM-DoS and Defamation attacks. Wenjun Fan, Simeon Wuthier, Hsiang-Jen Hong, Xiaobo Zhou 0002, Sang-Yoon Chang |
ICDCS | 4 |
| 2022 | QUIC Protocol with Post-quantum Authentication
Manohar Raavi, Simeon Wuthier, Pranav Chandramouli, Xiaobo Zhou 0002, Sang-Yoon Chang |
ISC | 4 |
| 2022 | Auto-Tune: Efficient Autonomous Routing for Payment Channel NetworksabstractPayment Channel Network (PCN) is a scaling solution for Cryptocurrency networks. We advance the practicality of the PCN multi-path routing by better modeling the system to incorporate the cost of routing fee and the privacy requirement of the channel balance. We design our Auto-Tune algorithm to optimize the routing concerning both the success rate and the routing fee and utilizing the limited channel capacity information (due to the privacy of the PCN user, the channel balance information is withheld). The simulation result shows Auto-Tune outperforms the current PCN implementation based on single-path routing in the success rate. We compare Auto-Tune against the state-of-the-art Flash algorithm, utilizing the channel-balance information, violating the PCN user privacy, and diverging from current implementation practices. Auto-Tune achieves the routing fee close to the optimal fee obtained by Flash, and its success rate is also close to the success rate achieved by Flash. Hsiang-Jen Hong, Sang-Yoon Chang, Xiaobo Zhou 0002 |
LCN | 3 |
| 2022 | Improving Concurrent GC for Latency Critical Services in Multi-tenant SystemsabstractFor resource utilization efficiency, latency critical (LC) services are commonly co-located with best-effort batch jobs in datacenter servers. Many LC services, such as Cassandra and HBase, run in Java Virtual Machine (JVM). We find that LC services often experience heavy-tailed latency due to performance interference of the concurrent garbage collection (GC) as well as multi-tenancy. The root cause is a semantic gap of resource allocation between JVM and the underlying Linux OS in multi-tenant systems. That is, the OS is unaware of the characteristics of different kinds of threads in JVM (i.e., GC threads and LC worker threads), which may lead to GC threads competing for CPUs; JVM is unaware of the resource utilization in the OS, which may trigger CPU-intensive GC operations when CPUs are busy. Furthermore, we find that co-located batch jobs can interfere with LC services due to Simultaneous Multi-Threading (SMT). Junxian Zhao, Aidi Pi, Xiaobo Zhou 0002, Sang-Yoon Chang, Cheng-Zhong Xu 0001 |
Middleware | 3 |
| 2022 | Greedy Networking in Cryptocurrency Blockchain
Simeon Wuthier, Pranav Chandramouli, Xiaobo Zhou 0002, Sang-Yoon Chang |
SEC | 3 |
| 2022 | Robust P2P networking connectivity estimation engine for permissionless Bitcoin cryptocurrency
Hsiang-Jen Hong, Wenjun Fan, Simeon Wuthier, Jinoh Kim, C. Edward Chow, Xiaobo Zhou 0002, Sang-Yoon Chang |
Comput. Networks | 6 |
| 2022 | A Machine Learning Approach to Anomaly Detection Based on Traffic Monitoring for Secure Blockchain NetworkingabstractWhile blockchain technology provides strong cryptographic protection on the ledger and the system operations, the underlying blockchain networking remains vulnerable due to potential threats such as denial of service (DoS), Eclipse, spoofing, and Sybil attacks. Effectively detecting such malicious events should thus be an essential task for securing blockchain networks and services. Due to its importance, several studies investigated anomaly detection in Bitcoin and blockchain networks, but their analyses mainly focused on the blockchain ledger in the application context (e.g., transactions) and targets specific types of attacks (e.g., double-spending, deanonymization, etc). In this study, we present a security mechanism based on the analysis of blockchain network traffic statistics (rather than ledger data) to detect malicious events, through the functions of data collection and anomaly detection. The data collection engine senses the underlying blockchain traffic and generates multi-dimensional data streams in a periodic, real-time manner. The anomaly detection engine then detects anomalies from the created data instances based on semi-supervised learning, which is capable of detecting previously unseen patterns, and we introduce our profiling-based detection engine implemented on top of AutoEncoder (AE). Our experimental results evaluated with real and simulated traffic data support the effectiveness of our security mechanism and design choices based on the AE structure, with the approximate detection performance to the supervised learning methods only through the profiling of normal instances. The measured time complexity is sufficiently cheap to perform real-time analysis, with less than 1.4 msec for per-instance testing on a single core setting. Jinoh Kim, Makiya Nakashima, Wenjun Fan, Simeon Wuthier, Xiaobo Zhou 0002, Ikkyun Kim, Sang-Yoon Chang |
IEEE Trans. Netw. Serv. Manag. | 5 |
| 2022 | Elastic Parameter Server: Accelerating ML Training With Scalable Resource SchedulingabstractParameter server (PS) based on worker-server communication is designed for distributed machine learning (ML) training in clusters. In feedback-driven exploration of ML model training, users exploit early feedback from each job to decide whether to kill the job or keep it running so as to find the optimal model configuration. However, PS does not support adjusting the number of workers and servers of a job at runtime. It becomes the bottleneck of scalable distributed ML training because the cluster resources cannot be dynamically allocated or deallocated to jobs, resulting in significant early feedback latency and resource under-utilization. This article rethinks the principle of PS architecture. We present Elastic Parameter Server (EPS), a lightweight and user-transparent PS that accelerates feedback-driven exploration for distributed ML training. EPS allows to remove a subset of workers and servers from running jobs and allocate the released resources to an incoming job at runtime so as to reduce its early feedback latency. It can also use the released resources from a killed job to add workers and servers to running jobs to improve resource utilization and the training speed. We develop a heuristic scheduler that leverages EPS and offers scalable resource scheduling for multiple ML jobs. We implement EPS in Tencent Angel and the scheduler in Apache Yarn, and conduct evaluations with various ML models. Experimental results show that EPS achieves up to 1.5x improvement on the ML training speed compared to PS. Aidi Pi, Xiaobo Zhou 0002 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2021 | Security Comparisons and Performance Analyses of Post-quantum Signature Algorithms
Manohar Raavi, Simeon Wuthier, Pranav Chandramouli, Yaroslav Balytskyi, Xiaobo Zhou 0002, Sang-Yoon Chang |
ACNS (2) | 5 |
| 2021 | FlashByte: Improving Memory Efficiency with Lightweight Native StorageabstractIn-memory caching of intermediate data is effective in reducing re-computation and I/O cost in distributed data-analytics frameworks, but it also generates a large amount of data in Java heap which increases the overhead of garbage collection (GC). An alternative off-heap approach caches data in native storage by transmitting the data from heap to native storage so as to reduce GC overhead. However, it incurs severe serialization and de-serialization overheads. Serialization also generates non-trivial metadata of the cached data in native storage. We propose and develop FlashByte, a lightweight native storage that efficiently caches intermediate data. FlashByte improves memory efficiency by achieving low GC overhead, low data transmission overhead, and low memory consumption. Specifically, the cached data are divided into two parts: metadata stored in Java heap and raw data stored in native storage. The metadata is generated based on the profile of workloads. Its size is trivial because it only contains a concise format of raw data, which achieves low memory consumption in the heap as well as low GC overhead. Native storage only stores the raw data to reduce its memory consumption. According to the metadata, the raw data are efficiently transmitted between the heap and native storage without serialization and de-serialization. We implement FlashByte in Spark and conduct evaluation with benchmark workloads. Experimental results show that, compared with the in-heap approach of Vanilla Spark, FlashByte achieves up to 4x speedup of the job execution time, reduces GC time by up to 96%, and reduces the memory consumption in heap by up to 36%. Compared with the alternative off-heap approach, FlashByte achieves up to 2.3x speedup of the job execution time, reduces the data transmission time by up to 84%, and reduces the cache size in native storage by up to 34%. Junxian Zhao, Aidi Pi, Xiaobo Zhou 0002 |
CCGRID | 4 |
| 2021 | A Generic Blockchain Framework to Secure Decentralized ApplicationsabstractBlockchain technology is gaining popularity in industries and governments for information monitoring, distribution, and tracking. Thanks to the built-in security properties, blockchain provides security integrity to various decentralized applications (dApps) involving distributed operations including supply chain, healthcare, banking, internet of things (IoT), and networking. In this paper, we propose a generic blockchain framework (GBF) for applying two blockchains to the dApp systems, one to establish trust and the other to use the trust for securing applications. GBF provides a generally applicable framework and addresses the foundational questions of the blockchain objectives, participants, and the underlying distributed consensus protocol in use. We apply GBF to various case studies from the recent blockchain research to show its effectiveness and generality. We also prototype GBF using smart contract and experiment on CloudLab for preliminary evaluations focusing on the application-general metrics. We propose GBF to facilitate blockchain/dApp research and development by providing the initial framework and enable the preliminary analyses so that the decentralized applications with specific aims can build on GBF. Wenjun Fan, Hsiang-Jen Hong, Xiaobo Zhou 0002, Sang-Yoon Chang |
ICC | 3 |
| 2021 | Performance Characterization of Post-Quantum Digital CertificatesabstractPublic Key Infrastructure (PKI) generates and distributes digital certificates to provide the root of trust for securing digital networking systems. To continue securing digital networking in the quantum era, PKI should transition to use quantum-resistant cryptographic algorithms. The cryptography community is developing quantum-resistant primitives/algorithms, studying, and analyzing them for cryptanalysis and improvements. National Institute of Standards and Technology (NIST) selected finalist algorithms for the post-quantum digital signature cipher standardization, which are Dilithium, Falcon, and Rainbow. We study and analyze the feasibility and the processing performance of these algorithms in memory/size and time/speed when used for PKI, including the key generation from the PKI end entities (e.g., a HTTPS/TLS server), the signing, and the certificate generation by the certificate authority within the PKI. The transition to post-quantum from the classical ciphers incur changes in the parameters in the PKI, for example, Rainbow I significantly increases the certificate size by 163 times when compared with RSA 3072. Nevertheless, we learn that the current X.509 supports the NIST post-quantum digital signature ciphers and that the ciphers can be modularly adapted for PKI. According to our empirical implementations-based study, the post-quantum ciphers can increase the certificate verification time cost compared to the current classical cipher and therefore the verification overheads require careful considerations when using the post-quantum-cipher-based certificates. Manohar Raavi, Pranav Chandramouli, Simeon Wuthier, Xiaobo Zhou 0002, Sang-Yoon Chang |
ICCCN | 4 |
| 2021 | Robust P2P Connectivity Estimation for Permissionless Bitcoin NetworkabstractBlockchain relies on the underlying peer-to-peer (p2p) networking to broadcast and get up-to-date on the blocks and transactions. It is therefore imperative to have high p2p connectivity for the quality of the blockchain system operations. High p2p networking connectivity ensures that a peer node is connected to multiple other peers providing a diverse set of observers of the current state of the blockchain and transactions. However, in a permissionless blockchain network, using the peer identifiers—including the current approach of counting the number of distinct IP addresses and port numbers—can be ineffective in measuring the number of peer connections and estimating the networking connectivity. Such current approach is further challenged by the networking threats manipulating the identifiers. We build a robust estimation engine for the p2p networking connectivity by sensing and processing the p2p networking traffic. We implement a working Bitcoin prototype connected to the Bitcoin Mainnet to validate and improve our engine’s performances and evaluate the estimation accuracy and cost efficiency of our estimation engine. Hsiang-Jen Hong, Wenjun Fan, Simeon Wuthier, Jinoh Kim, Xiaobo Zhou 0002, C. Edward Chow, Sang-Yoon Chang |
IWQoS | 5 |
| 2021 | A Machine Learning Approach to Peer Connectivity Estimation for Reliable Blockchain NetworkingabstractPeer connectivity plays a significant role in a blockchain network since any poor connectivity may result in the nodes operating on outdated data (e.g., cryptocurrency transactions). Although connectivity information is maintained by individual nodes, such identifier-based information might be unreliable due to the possibility of bogus identifiers. This paper tackles the problem of peer connectivity estimation through data-driven analytics of blockchain traffic for reliable blockchain networking. We define a set of variables to represent traffic characteristics and estimate peer connectivity from the collected data using a machine learning methodology. We also investigate the feasibility of feature prioritization to minimize estimation complexities. Our experimental results show that the presented estimation mechanism makes accurate predictions, with less than 0.1 difference between the measurement and estimation for over 99.7% of predictions. The time complexity measured on a commodity machine shows a microsecond scale for completing a single prediction task, enabling real-time operations. Jinoh Kim, Makiya Nakashima, Wenjun Fan, Simeon Wuthier, Xiaobo Zhou 0002, Ikkyun Kim, Sang-Yoon Chang |
LCN | 5 |
| 2021 | Memory at your service: fast memory allocation for latency-critical servicesabstractCo-location and memory sharing between latency-critical services, such as key-value store and web search, and best-effort batch jobs is an appealing approach to improving memory utilization in multi-tenant datacenter systems. However, we find that the very diverse goals of job co-location and the GNU/Linux system stack can lead to severe performance degradation of latency-critical services under memory pressure in a multi-tenant system. Aidi Pi, Junxian Zhao, Xiaobo Zhou 0002 |
Middleware | 4 |
| 2021 | Blockchain-based Secure Coordination for Distributed SDN Control PlaneabstractSoftware-defined wide-area network (SD-WAN) is an emerging and advanced networking platform extending software-defined networking (SDN) across multiple networking domains. Because SD-WAN manages the data plane in the networking domains separated by the public Internet, SDWAN provides a distinct environment and challenges from SDN, including greater risks for the security threats injecting control plane communications from attackers residing outside of the SDN domain. We design and build blockchain-coordinating controllers (BCC) to secure control communications of the SD-WAN controller network formed by the distributed controllers spread across multiple domains. BCC provides resiliency against the security threats in the control plane where an attacker compromises controller communications to manipulate the coordination and the operations of the other controllers. More specifically, BCC provides secure control communications even when up to n controllers’ networking credentials are compromised. BCC is also designed for modularity so that it applies generally across the controller implementations. We prototype BCC using Ethereum and smart contract on CloudLab to validate its effectiveness and efficiency. We experiment on geographically separate nodes on CloudLab and show that BCC achieves the distributed consensus at sub-second level for certificate/key distribution and for network-wide control communication synchronization. Wenjun Fan, Sang-Yoon Chang, Xiaobo Zhou 0002, Younghee Park |
NetSoft | 4 |
| 2021 | Overlapping Communication With Computation in Parameter Server for Scalable DL TrainingabstractScalability of distributed deep learning (DL) training with parameter server (PS) architecture is often communication constrained in large clusters. There are recent efforts that use a layer by layer strategy to overlap gradient communication with backward computation so as to reduce the impact of communication constraint on the scalability. However, the approaches could bring significant overhead in gradient communication. Meanwhile, they cannot be effectively applied to the overlap between parameter communication and forward computation. In this article, we propose and develop iPart, a novel approach that partitions communication and computation in various partition sizes to overlap gradient communication with backward computation and parameter communication with forward computation. iPart formulates the partitioning decision as an optimization problem and solves it based on a greedy algorithm to derive communication and computation partitions. We implement iPart in the open-source DL framework BigDL and perform evaluations with various DL workloads. Experimental results show that iPart improves the scalability of a cluster of 72 nodes by up to 94 percent over the default PS and 52 percent over the layer by layer strategy. Aidi Pi, Xiaobo Zhou 0002, Jun Wang 0001, Cheng-Zhong Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2020 | Blockchain-based Distributed Banking for Permissioned and Accountable Financial Transaction ProcessingabstractDistributed banking platforms and services forgo centralized banks to process financial transactions. For example, M-Pesa provides distributed banking service in the developing regions so that the people without a bank account can deposit, withdraw, or transfer money. The current distributed banking systems lack the transparency in monitoring and tracking of distributed banking transactions and thus do not support auditing of distributed banking transactions for accountability. To address this issue, this paper proposes a blockchain-based distributed banking (BDB) scheme, which uses blockchain technology to leverage its built-in properties to record and track immutable transactions. BDB supports distributed financial transaction processing but is significantly different from cryptocurrencies in its design properties, simplicity, and computational efficiency. We implement a prototype of BDB using smart contract and conduct experiments to show BDB’s effectiveness and performance. We further compare our prototype with the Ethereum cryptocurrency to highlight the fundamental differences and demonstrate the BDB’s superior computational efficiency. Wenjun Fan, Sang-Yoon Chang, Shawn Emery, Xiaobo Zhou 0002 |
ICCCN | 4 |
| 2020 | Optimizing Social Welfare for Task Offloading in Mobile Edge Computing
Hsiang-Jen Hong, Wenjun Fan, C. Edward Chow, Xiaobo Zhou 0002, Sang-Yoon Chang |
Networking | 4 |
| 2020 | An efficient and non-intrusive GPU scheduling framework for deep learning training systemsabstractEfficient GPU scheduling is the key to minimizing the execution time of the Deep Learning (DL) training workloads. DL training system schedulers typically allocate a fixed number of GPUs to each job, which inhibits high resource utilization and often extends the overall training time. The recent introduction of schedulers that can dynamically reallocate GPUs has achieved better cluster efficiency. This dynamic nature, however, introduces additional overhead by terminating and restarting jobs or requires modification to the DL training frameworks.We propose and develop an efficient, non-intrusive GPU scheduling framework that employs a combination of an adaptive GPU scheduler and an elastic GPU allocation mechanism to reduce the completion time of DL training workloads and improve resource utilization. Specifically, the adaptive GPU scheduler includes a scheduling algorithm that uses training job progress information to determine the most efficient allocation and reallocation of GPUs for incoming and running jobs at any given time. The elastic GPU allocation mechanism works in concert with the scheduler. It offers a lightweight and nonintrusive method to reallocate GPUs based on a “SideCar” process that temporarily stops and restarts the job's DL training process with a different number of GPUs. We implemented the scheduling framework as plugins in Kubernetes and conducted evaluations on two 16-GPU clusters with multiple training jobs based on TensorFlow. Results show that our proposed scheduling framework reduces the overall execution time and the average job completion time by up to 45% and 63%, respectively, compared to the Kubernetes default scheduler. Compared to a termination based scheduler, our framework reduces the overall execution time and the average job completion time by up to 20% and 37%, respectively. Oscar J. Gonzalez, Xiaobo Zhou 0002, Thomas Williams, Brian D. Friedman, Martin Havemann, Thomas Y. C. Woo |
SC | 3 |
| 2020 | Blockchain-enabled Collaborative Intrusion Detection in Software Defined NetworksabstractCollaborative intrusion detection system (CIDS) shares the critical detection-control information across the nodes for improved and coordinated defense. Software-defined network (SDN) introduces the controllers for the networking control, including for the networks spanning across multiple autonomous systems, and therefore provides a prime platform for CIDS application. Although previous research studies have focused on CIDS in SDN, the real-time secure exchange of the detection-relevant information (e.g., the detection signature) remains a critical challenge. In particular, the CIDS research still lacks robust trust management of the SDN controllers and the integrity protection of the collaborative defense information to resist against the insider attacks transmitting untruthful and malicious detection signatures to other participating controllers. In this paper, we propose a blockchain-enabled collaborative intrusion detection in SDN, taking advantage of the blockchain's security properties. Our scheme achieves three important security goals: to establish the trust of the participating controllers by using the permissioned blockchain to register the controller and manage digital certificates, to protect the integrity of the detection signatures against malicious detection signature injection, and to attest the delivery/update of the detection signature to other controllers. Our experiments in CloudLab based on a prototype built on Ethereum, Smart Contract, and IPFS demonstrates that our approach efficiently shares and distributes detection signatures in real-time through the trustworthy distributed platform. Wenjun Fan, Younghee Park, Priyatham Ganta, Xiaobo Zhou 0002, Sang-Yoon Chang |
TrustCom | 5 |
| 2020 | ODDS: Optimizing Data-Locality Access for Scientific Data AnalysisabstractWhereas traditional scientific applications are computationally intensive, recent applications require more data-intensive analysis and visualization to extract knowledge from the explosive growth of scientific information and simulation data. As the computational power and size of compute clusters continue to increase, the I/O read rates and associated network for these data-intensive applications have been unable to keep pace. These applications suffer from long I/O latency due to the movement of “big data” from the network/parallel file system, which results in a serious performance bottleneck. To address this problem, we proposed a novel approach called “ODDS” to optimize data-locality access in scientific data analysis and visualization. ODDS leverages a distributed file system (DFS) to provide scalable data access for scientific analysis. Through exploiting the information of underlying data distribution in DFS, ODDS employs a novel data-locality scheduler to transform a compute-centric mapping into a data-centric one and enables each computational process to access the needed data from a local or nearby storage node. ODDS is suitable for parallel applications with dynamic process-to-data scheduling and for applications with static process-to-data assignment. To demonstrate the efficacy of our methods, we present and evaluate ODDS in the context of two state-of-the-art, scientific-analysis applications-mpiBLAST and ParaView-along with the Hadoop distributed file system (HDFS) across a wide variety of computing platform settings. In comparison to existing deployments using NFS, PVFS, or Lustre as the underlying storage systems, ODDS can greatly reduce the I/O cost and double overall performance. Jun Wang 0001, Dezhi Han, Jiangling Yin, Xiaobo Zhou 0002, Changjun Jiang 0002 |
IEEE Trans. Cloud Comput. | 4 |
| 2020 | Preemptive and Low Latency Datacenter Scheduling via Lightweight ContainersabstractDatacenters are evolving to host heterogeneous workloads on shared clusters to reduce the operational cost and achieve higher resource utilization. However, it is challenging to schedule heterogeneous workloads with diverse resource requirements and QoS constraints. On one hand, latency-critical jobs need to be scheduled as soon as they are submitted to avoid any queuing delays. On the other hand, best-effort long jobs should be allowed to occupy the cluster when there are idle resources to improve cluster utilization. The challenge lies in how to minimize the queuing delays of short jobs while maximizing cluster utilization. In this article, we propose and develop BIG-C, a container-based resource management framework for data-intensive cluster computing. The key design is to leverage lightweight virtualization, a.k.a, containers, to make tasks preemptable in cluster scheduling. We devise two types of preemption strategies: immediate and graceful preemptions and show their effectiveness and tradeoffs with loosely-coupled MapReduce workloads as well as iterative, in-memory Spark workloads. Based on the mechanisms for task preemption, we further develop job-level and task-level preemptive policies as well as a preemptive fair share cluster scheduler. Our implementation on Yarn and evaluation with synthetic and production workloads show that low job latency and high resource utilization can be both attained when scheduling heterogeneous workloads on a contended cluster. Wei Chen 0038, Xiaobo Zhou 0002, Jia Rao |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2019 | Scalable Distributed DL Training: Batching Communication and ComputationabstractScalability of distributed deep learning (DL) training with parameter server architecture is often communication constrained in large clusters. There are recent efforts that use a layer by layer strategy to overlap gradient communication with backward computation so as to reduce the impact of communication constraint on the scalability. However, the approaches cannot be effectively applied to the overlap between parameter communication and forward computation. In this paper, we propose and design iBatch, a novel communication approach that batches parameter communication and forward computation to overlap them with each other. We formulate the batching decision as an optimization problem and solve it based on greedy algorithm to derive communication and computation batches. We implement iBatch in the open-source DL framework BigDL and perform evaluations with various DL workloads. Experimental results show that iBatch improves the scalability of a cluster of 72 nodes by up to 73% over the default PS and 41% over the layer by layer strategy. Aidi Pi, Xiaobo Zhou 0002 |
AAAI | 3 |
| 2019 | Proactive Load Shifting for Distributed SDN Control Plane ArchitectureabstractBalancing the workload among distributed SDN controllers plays a critical role for both the network performance and the control plane scalability. Several distributed SDN controller architectures have been proposed to mitigate the risk of controller overload and failures. However, many of these architectures fall short for maintaining the same level of complexity in the control plane. A core implication of complex control plane can translate to a limitation in functional improvements of existing implementation. To address this issue, we propose a novel Proactive Load Shift (PLS) technique that augments the traditional SDN architecture with a shim layer to diminish the complexities in existing distributed SDN controller architectures. While our primary focus is on efficient workload distribution among SDN controllers in a distributed architecture, our shim layer also serves as a programmable abstraction for supporting new network functionalities in the SDN control plane and the data plane without infringing the SDN principles. To achieve optimal network performance in our proposed technique, we eliminate the need for inter-controller synchronization by delegating the synchronization sequence to the shim layer at a per-need based only. Our experimental results show that the PLS technique provides efficient responses to load balancing triggers with less overhead on the control plane. Oluwatobi Akanbi, Amer Aljaedi, Xiaobo Zhou 0002 |
CCNC | 3 |
| 2019 | Pufferfish: Container-driven Elastic Memory Management for Data-intensive ApplicationsabstractData-intensive applications often suffer from significant memory pressure, resulting in excessive garbage collection (GC) and out-of-memory (OOM) errors, harming system performance and reliability. In this paper, we demonstrate how lightweight virtualization via OS containers opens up opportunities to address memory pressure and realize memory elasticity: 1) tasks running in a container can be set to a large heap size to avoid OutOfMemory (OOM) errors, and 2) tasks that are under memory pressure and incur significant swapping activities can be temporarily "suspended" by depriving resources from the hosting containers, and be "resumed" when resources are available. We propose and develop Pufferfish, an elastic memory manager, that leverages containers to flexibly allocate memory for tasks. Memory elasticity achieved by Pufferfish can be exploited by a cluster scheduler to improve cluster utilization and task parallelism. We implement Pufferfish on the cluster scheduler Apache Yarn. Experiments with Spark and MapReduce on real-world traces show Pufferfish is able to avoid OOM errors, improve cluster memory utilization by 2.7x and the median job runtime by 5.5x compared to a memory over-provisioning solution. Wei Chen 0038, Aidi Pi, Xiaobo Zhou 0002 |
SoCC | 4 |
| 2019 | Semantic-aware Workflow Construction and Analysis for Distributed Data Analytics SystemsabstractLogging is a universal approach to recording important events in system workflows of distributed systems. Current log analysis tools ignore the semantic knowledge that is key to workflow construction and analysis. In addition, they focus on infrastructure-level distributed systems. Because of fundamental differences in log features, they are ineffective in distributed data analytics systems. This paper proposes IntelLog, a semantic-aware non-intrusive workflow reconstruction tool for distributed data analytics systems. It is capable of building hierarchical relationships between components and events from logs generated by the targeted systems with little or even no domain knowledge. Leveraging natural language processing, IntelLog automatically extracts and formats semantic information in each log message, including system events, identifiers, locality information, and metrics values. It builds a graph to represent the hierarchical relationship of components in the targeted system via nomenclature conventions. We implement IntelLog for Hadoop MapReduce, Spark and Tez. Evaluation results show that IntelLog provides a fine-grained view of the system workflows with semantics. It outperforms existing tools in automatically detecting anomalies caused by real-world problems, misconfigurations and system bugs. Users can query the formatted semantic knowledge to understand and further troubleshoot the systems. Aidi Pi, Wei Chen 0038, Xiaobo Zhou 0002 |
HPDC | 4 |
| 2019 | Addressing Skewness in Iterative ML Jobs with Parameter PartitionabstractComputational skewness is a significant challenge in multi-tenant data-parallel clusters that introduce dynamic heterogeneity of machine capacity in distributed data processing. Previous efforts to addressing skewness mostly focus on batch jobs based on the assumption that processing time is linearly dependent on the size of partitioned data. However, they are illsuited for iterative machine learning (ML) jobs, which (1) exhibit a non-linear relationship between the size of partitioned parameters and processing time within each iteration, and (2) show an explicit binding relationship between input data and parameters for parameter update. In this paper, we present FlexPara, a parameter partition approach that leverages the non-linear relationship and provisions adaptive tasks to match the distinct machine capacity so as to address the skewness in iterative ML jobs on data-parallel clusters. FlexPara first predicts task processing time based on a capacity model designed for iterative ML jobs without the linear assumption. It then partitions parameters to parallel tasks through proactive parameter reassignment. Such reassignment can significantly reduce network transmission cost incurred by input data movement due to the binding relationship. We implement FlexPara in Spark and evaluate it with various ML jobs. Experimental results show that compared to hash partition, FlexPara speeds up the execution by up to 54% and 43% in private and NSF Chameleon clusters, respectively. Wei Chen 0038, Xiaobo Zhou 0002, Sang-Yoon Chang, Mike Ji |
INFOCOM | 3 |
| 2019 | OS-Augmented Oversubscription of Opportunistic Memory with a User-Assisted OOM KillerabstractExploiting opportunistic memory by oversubscription is an appealing approach to improving cluster utilization and throughput. In this paper, we find the efficacy of memory oversubscription depends on whether or not the oversubscribed tasks can be killed by an OutOf Memory (OOM) killer in a timely manner to avoid significant memory thrashing upon memory pressure. However, current approaches in modern cluster schedulers are actually unable to unleash the power of opportunistic memory because their user space OOM killers are unable to timely deliver a task killing signal to terminate the oversubscribed tasks. Our experiments observe that a user space OOM killer fails to do that because of lacking the memory pressure knowledge from OS while the kernel space Linux OOM killer is too conservative to relieve memory pressure. Wei Chen 0038, Aidi Pi, Xiaobo Zhou 0002 |
Middleware | 4 |
| 2019 | Heterogeneity Aware Workload Management in Distributed Sustainable DatacentersabstractThe tremendous growth of cloud computing and large-scale data analytics highlight the importance of reducing datacenter power consumption and environmental impact of brown energy. While many Internet service operators have at least partially powered their datacenters by green energy, it is challenging to effectively utilize green energy due to the intermittency of renewable sources, such as solar or wind. We find that the geographical diversity of internet-scale services can be carefully scheduled to improve the efficiency of applying green energy in datacenters. In this paper, we propose a holistic heterogeneity-aware cloud workload management approach, sCloud, that aims to maximize the system goodput in distributed self-sustainable datacenters. sCloud adaptively places the transactional workload to distributed datacenters, allocates the available resource to heterogeneous workloads in each datacenter, and migrates batch jobs across datacenters, while taking into account the green power availability and QoS requirements. We formulate the transactional workload placement as a constrained optimization problem that can be solved by nonlinear programming. Then, we propose a batch job migration algorithm to further improve the system goodput when the green power supply varies widely at different locations. Finally, we extend sCloud by integrating a flexible batch job manager to dynamically control the job execution progress without violating the deadlines. We have implemented sCloud in a university cloud testbed with real-world weather conditions and workload traces. Experimental results demonstrate sCloud can achieve near-to-optimal system performance while being resilient to dynamic power availability. sCloud with the flexible batch job management approach outperforms a heterogeneity-oblivious approach by 37 percent in improving system goodput and 33 percent in reducing QoS violations. Dazhao Cheng, Xiaobo Zhou 0002, Zhijun Ding, Yu Wang 0003, Mike Ji |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2019 | Deadline-Aware MapReduce Job Scheduling with Dynamic Resource AvailabilityabstractAs MapReduce is becoming ubiquitous in large-scale data analysis, many recent studies have shown that the performance of MapReduce could be improved by different job scheduling approaches, e.g., Fair Scheduler and Capacity Scheduler. However, most exiting MapReduce job schedulers focus on the scenario that MapReduce cluster is stable and pay little attention to the MapReduce cluster with dynamic resource availability. In fact, MapReduce cluster resources may fluctuate as there is a growing number of Hadoop clusters deployed on hybrid systems, e.g., infrastructure powered by mix of traditional and renewable energy, and cloud platforms hosting heterogeneous workloads. Thus, there is a growing need for providing predictable services to users who have strict requirements on job completion times in such dynamic environments. In this paper, we propose, RDS, a Resource and Deadline-aware Hadoop job Scheduler that takes future resource availability into consideration when minimizing job deadline misses. We formulate the job scheduling problem as an online optimization problem and solve it using an efficient receding horizon control algorithm. To aid the control, we design a self-learning model to estimate job completion times. We further extend the design of RDS scheduler to support flexible performance goals in various dynamic clusters. In particular, we use flexible deadline time bounds instead of the single fixed job completion deadline. We have implemented RDS in the open-source Hadoop implementation and performed evaluations with various benchmark workloads. Experimental results show that RDS substantially reduces the penalty of deadline misses by at least 36 and 10 percent compared with Fair Scheduler and Earliest Deadline First (EDF) scheduler, respectively. In a Hadoop cluster running partially on renewable energy, the experimental result shows the green power based resource prediction approach can further reduce the penalty of deadline misses by 16 percent compared to Auto-Regressive Integrated Moving Average (ARIMA) prediction approach. Dazhao Cheng, Xiaobo Zhou 0002, Yinggen Xu, Changjun Jiang 0002 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2019 | Dependency-Aware Network Adaptive Scheduling of Data-Intensive Parallel JobsabstractDatacenter clusters often run data-intensive jobs in parallel for improving resource utilization and cost efficiency. The performance of parallel jobs is often constrained by the cluster's hard-to-scale network bisection bandwidth. Various solutions have been proposed to address the issue, however, most of them do not consider inter-job data dependencies and schedule jobs independently from one another. In this work, we find that aggregating and co-locating the data and tasks of dependent jobs offer an extra opportunity for data locality improvement that can help to greatly enhance the performance of jobs. We propose and design Dawn, a dependency-aware network-adaptive scheduler that includes an online plan and an adaptive task scheduler. The online plan, taking job dependencies into consideration, determines where (i.e., preferred racks) to place tasks in order to proactively aggregate dependent data. The task scheduler, based on the output of online plan and dynamic network status, adaptively schedules tasks to co-locate with the dependent data in order to take advantage of data locality. We implement Dawn on Apache Yarn and evaluate it on physical and virtual clusters using various machine learning and query workloads. Results show that Dawn effectively improves cluster throughput by up to 73 and 38 percent compared to Fair Scheduler and ShuffleWatcher, respectively. Dawn not only significantly enhances the performance of jobs with dependency, but also works well for jobs without dependency. Wei Chen 0038, Xiaobo Zhou 0002, Liqiang Zhang 0002 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2019 | An I/O Efficient Distributed Approximation Framework Using Cluster SamplingabstractIn this paper, we present an I/O efficient distributed approximation framework to support approximations on arbitrary sub-datasets of a large dataset. Due to the prohibitive storage overhead of caching offline samples for each sub-dataset, existing offline sample-based systems provide high accuracy results for only a limited number of sub-datasets, such as the popular ones. On the other hand, current online sample-based approximation systems, which generate samples at runtime, do not take into account the uneven storage distribution of a sub-dataset. They work well for uniform distribution of a sub-dataset while suffer low I/O efficiency and poor estimation accuracy on unevenly distributed sub-datasets. To address the problem, we develop a distribution aware method called CLAP (cluster sampling based approximation). Our idea is to collect the occurrences of a sub-dataset at each logical partition of a dataset (storage distribution) in the distributed system, and make good use of such information to enable I/O efficient online sampling. There are three thrusts in CLAP. First, we develop a probabilistic map to reduce the exponential number of recorded sub-datasets to a linear one. Second, we apply the cluster sampling with unequal probability theory to implement a distribution-aware method for efficient online sampling for a single or multiple sub-datasets. Third, we enrich CLAP support with more complex approximations such as ratio and regression using bootstrap based estimation beyond the simple aggragation approxiamtions. Forth, we add an option in CLAP to allow users specifying a target error bound when submitting an approximation job. Fifth, we quantitatively derive the optimal sampling unit size in a distributed file system by associating it with approximation costs and accuracy. We have implemented CLAP into Hadoop as an example system and open sourced it on GitHub. Our comprehensive experimental results show that CLAP can achieve a speedup by up to 20× over the precise execution. Xuhong Zhang 0002, Jun Wang 0001, Shouling Ji, Jiangling Yin, Rui Wang 0030, Xiaobo Zhou 0002, Changjun Jiang 0002 |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2019 | A Correlation-Aware Page-Level FTL to Exploit Semantic Links in WorkloadsabstractNAND Flash based Solid State Disks (SSDs) are gaining tremendous popularity in today's storage market due to their unique erase-before-write feature. The Flash Translation Layer (FTL) in the SSDs redirects the incoming writes to a free physical address and manages a logical to physical address mapping table. However, this induces significant performance degradation to the SSDs. One of the main reasons is that current cache management in FTLs is mainly optimized for the temporal or spatial locality. However, because of multiple levels of data buffers in the whole storage architecture, the locality of internal disk I/O is relatively low. What's more, the increasing capacity of SSD not only generates large mapping tables, but also imposes high pressure on the efficiency of page-level address mapping. To overcome this limitation, we propose Correlation-Aware Page-level FTL, a.k.a CPFTL, which exploits I/O correlations in the workloads. In CPFTL, we develop a correlation-aware mapping table based on the correlation in read operations. We then build a correlation prediction table to support fast mapping entry lookup in the correlation-aware mapping table. Finally, we split read and write caches and build a skew-aware dirty entry index to improve the cache hit ratio and reduce the garbage collection overhead. Our emulator and prototype are open-sourced at: https://github.com/janzhou/SSD-Emulator. The experimental results show that CPFTL can reduce the average response time by 63.4 percent for read dominant workloads and 32.9 percent for transaction workloads. Jian Zhou 0004, Dezhi Han, Jun Wang 0001, Xiaobo Zhou 0002, Changjun Jiang 0002 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2018 | Profiling distributed systems in lightweight virtualized environments with logs and resource metricsabstractUnderstanding and troubleshooting distributed systems in the cloud is considered a very difficult problem because the execution of a single user request is distributed to multiple machines. Further, the multi-tenancy nature of cloud environments further introduces interference that causes performance issues. Most existing troubleshooting tools either focus on log analysis or intrusive tracing methods, leaving resource usage monitoring unexplored. Aidi Pi, Wei Chen 0038, Xiaobo Zhou 0002, Mike Ji |
HPDC | 3 |
| 2018 | PETS: Bottleneck-Aware Spark Tuning with Parameter EnsemblesabstractSpark tuning with its dozens of parameters for performance improvement is both a challenge and time consuming effort. Current techniques rely on trial-and-error or best guess utilizing expert knowledge that very few posses. Previous tuning works are not compatible with Spark and also ignore the underlying problem of resource bottlenecks that is both the cause of performance issues, and a potential ally, if its awareness is leveraged in directing tuning to be more effective. We propose and develop PETS, a new method that allows the tuning of associated parameters at the same time, using resource bottleneck awareness to adjust parameter ensemble values in few iterations. Performance evaluation based on testbed implementation shows that with the use of PETS, representative workloads achieve: (1) Significant speedups; (2) Fast convergence speed; (3) Performance gains that are stable with varying workload data sizes, homogenous and heterogenous clusters, and initial parameter settings. The results show that PETS outperforms a machine learning based method, and achieves speedups of up to x4.78 and convergence speed as low as 2 iterations. Tiago B. G. Perez, Wei Chen 0038, Raymond Ji, Xiaobo Zhou 0002 |
ICCCN | 5 |
| 2018 | Reference-distance Eviction and Prefetching for Cache Management in SparkabstractOptimizing memory cache usage is vital for performance of in-memory data-parallel frameworks such as Spark. Current data-analytic frameworks utilize the popular Least Recently Used (LRU) policy, which does not take advantage of data dependency information available in the application's directed acyclic graph (DAG). Recent research in dependency-aware caching, notably MemTune and Least Reference Count (LRC), have made important improvements to close this gap. But they do not fully leverage the DAG structure, which imparts information such as the time-spatial distribution of data references across the workflow, to further improve cache hit ratio and application runtime. Tiago B. G. Perez, Xiaobo Zhou 0002, Dazhao Cheng |
ICPP | 2 |
| 2018 | Improving Utilization and Parallelism of Hadoop Cluster by Elastic ContainersabstractModern datacenter schedulers apply a static policy to partition resources among different tasks. The amount of allocated resource won't get changed during a task's lifetime. However, we found that resource usage during a task's runtime demonstrates high dynamics and it only reaches full usage at few moments. Therefore, the static allocation policy doesn't exploit the dynamic nature of resource usage, leading to low system resource utilization. To address this hard problem, a recently proposed task-consolidation approach packs as many tasks as possible on the same node based on real-time resource demands. However, this approach may cause resource over-allocation and harm application performance. In this paper, we propose and develop ECS, an elastic container based scheduler that leverages resource usage variation within the task lifetime to exploit the potential utilization and parallelism. The key idea is to proactively select and shift tasks backward so that the inherent paralleled tasks can be identified without over-allocation. We formulate the scheduling scheme as an online optimization problem and solves it using a resource leveling algorithm. We have implemented ECS in Apache Yarn and performed evaluations with various MapReduce benchmarks in a cluster. Experimental results show that ECS can efficiently utilize resource and achieves up to 29% reduction on average job completion time while increasing CPU utilization by 25%, compared to stock Yarn. Yinggen Xu, Wei Chen 0038, Xiaobo Zhou 0002, Changjun Jiang 0002 |
INFOCOM | 4 |
| 2018 | Characterizing Scheduling Delay for Low-Latency Data Analytics WorkloadsabstractData analytics workloads are shifting to shorter task execution time, higher degree of parallelism, and execution on faster hardware. As a result, job scheduling is becoming a bottleneck, which needs to offer extreme low-latency, massive throughput, and high scalability. However, few efforts have been focused on systematically understanding the scheduling delay. In this paper, we propose a method and develop a tool, SD-checker, that decomposes the job scheduling delay into multiple components and characterizes each by extensive experiments. SDchecker extracts event messages through mining both cluster scheduler logs and application logs, and constructs a scheduling order graph for the ease of analysis. We use SDchecker to evaluate Spark-SQL on a popular cluster scheduler Yarn. Results show that the scheduling delay may account for 60% of the job runtime of small data analytics workloads. After decomposing the total scheduling delay, we find Spark itself contributes 70% of the delay. Through the evaluation and analysis, we conclude that (1) The causes of scheduling delay are determined by many factors, and (2) The job scheduling is not well optimized yet, and far from ideal for low-latency data analytics workloads. Wei Chen 0038, Aidi Pi, Xiaobo Zhou 0002 |
IPDPS | 4 |
| 2018 | Performance Isolation of Data-Intensive Scale-out Applications in a Multi-tenant CloudabstractData-intensive applications often suffer from performance variability and degradation in the cloud due to intrinsically complex problem of performance interference that arises from multi-tenancy. Although application-level approach of straggler mitigation for scale-out data processing frameworks such as MapReduce and Spark, address the issue to some extent, they incur extra resource and often react after tasks have already slowed down. In this paper, we present PerfCloud, a novel system software that utilizes system level performance metrics for early detection of performance interference in a multi-tenant cloud, and provides non-invasive performance isolation through fine-grained resource control. Unlike existing works, PerfCloud does not require time-consuming workload profiling, or intrusive modification of the application framework and the operating system. We implemented PerfCloud on NSF Cloud's Chameleon testbed using KVM for virtualization, and OpenStack for cloud management. Experimental results with Hadoop MapReduce and Spark benchmarks show that PerfCloud effectively reduces their job completion time, decreases performance variability, and improves resource utilization efficiency while minimizing the performance degradation of other colocated VMs. Palden Lama, Xiaobo Zhou 0002, Dazhao Cheng |
IPDPS | 3 |
| 2018 | PLS: Proactive Load Shifting for Distributed SDN ControllersabstractBalancing the workload among distributed SDN controllers plays a critical role for both the network performance and the control plane scalability. Therefore, various load balancing techniques were proposed for SDN to efficiently utilize the control plane's resources. However, such techniques suffer increased latency and packet loss resulting from load migration and intensive communication among the SDN controllers. The existing solutions adopt load migration based on CPU utilization, which are susceptible to inconsistent load spikes. In this paper, we formally define the problem and present an alternate approach called PLS that constitutes the cornerstone for addressing this problem. We then show through experimental results that our approach provides accurate responses to load migration event triggers. Oluwatobi Akanbi, Amer Aljaedi, Xiaobo Zhou 0002 |
LCN | 3 |
| 2018 | Aggressive Synchronization with Partial Processing for Iterative ML Jobs on ClustersabstractExecuting distributed machine learning (ML) jobs on Spark follows Bulk Synchronous Parallel (BSP) model, where parallel tasks execute the same iteration at the same time and the generated updates must be synchronized on parameters when all tasks are finished. However, the parallel tasks rarely have the same execution time due to sparse data so that the synchronization has to wait for tasks finished late. Moreover, running Spark on heterogeneous clusters makes it even worse because of stragglers, where the synchronization is significantly delayed by the slowest task. Wei Chen 0038, Aidi Pi, Xiaobo Zhou 0002 |
Middleware | 4 |
| 2018 | Energy Efficiency Aware Task Assignment with DVFS in Heterogeneous Hadoop ClustersabstractWhile Hadoop ecosystems become increasingly important for practitioners of large-scale data analysis, they also incur tremendous energy cost. This trend is driving up the need for designing energy-efficient Hadoop clusters in order to reduce the operational costs and the carbon emission associated with its energy consumption. However, despite extensive studies of the problem, existing approaches for energy efficiency have not fully considered the heterogeneity of both workload and machine hardware found in production environments. In this paper, we find that heterogeneity-oblivious task assignment approaches are detrimental to both performance and energy efficiency of Hadoop clusters. Our observation shows that even heterogeneity-aware techniques that aim to reduce the job completion time do not guarantee a reduction in energy consumption of heterogeneous machines. We propose a heterogeneity-aware task assignment approach, E-Ant, that aims to improve the overall energy consumption in a heterogeneous Hadoop cluster without sacrificing job performance. It adaptively schedules heterogeneous workloads on energy-efficient machines, without a priori knowledge of the workload properties. E-Ant employs an ant colony optimization approach that generates task assignment solutions based on the feedback of each task's energy consumption reported by Hadoop TaskTrackers in an agile way. Furthermore, we integrate DVFS technique with E-Ant to further improve the energy efficiency of heterogeneous Hadoop clusters. It relies on a DVFS controller to dynamically scale the CPU frequency of each slave machine in response to time-varying resource demands. Experimental results on a heterogeneous cluster with varying hardware capabilities show that E-Ant with DVFS improves the overall energy savings for a synthetic workload from Microsoft by 23 and 17 percent compared to Fair Scheduler and Tarazu, respectively. Dazhao Cheng, Xiaobo Zhou 0002, Palden Lama, Mike Ji, Changjun Jiang 0002 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2018 | Adaptive Scheduling Parallel Jobs with Dynamic Batching in Spark StreamingabstractToday enterprises have massive stream data that require to be processed in real time due to data explosion in recent years. Spark Streaming as an emerging system is developed to process real time stream data analytics by using micro-batch approach. The unified programming model of Spark Steaming leads to some unique benefits over other traditional streaming systems, such as fast recovery from failures, better load balancing and resource usage. It treats the continuous stream as a series of micro-batches of data and continuously process these micro-batch jobs. However, efficient scheduling of micro-batch jobs to achieve high throughput and low latency is very challenging due to the complex data dependency and dynamism inherent in streaming workloads. In this paper, we propose A-scheduler, an adaptive scheduling approach that dynamically schedules parallel micro-batch jobs in Spark Streaming and automatically adjusts scheduling parameters to improve performance and resource efficiency. Specifically, A-scheduler dynamically schedules multiple jobs concurrently using different policies based on their data dependencies and automatically adjusts the level of job parallelism and resource shares among jobs based on workload properties. Furthermore, we integrate dynamic batching technique with A-Scheduler to further improve the overall performance of the customized Spark Streaming system. It relies on an expert fuzzy control mechanism to dynamically adjust the length of each batch interval in response to time-varying streaming workload and system processing rate. We implemented A-scheduler and evaluated it with a real-time security event processing workload. Our experimental results show that A-scheduler with dynamic batching can reduce end-to-end latency by 38 percent and meanwhile improve workload throughput and energy efficiency by 23 and 15 percent, respectively, compared to the default Spark Streaming scheduler. Dazhao Cheng, Xiaobo Zhou 0002, Yu Wang 0003, Changjun Jiang 0002 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2017 | FLEP: Enabling Flexible and Efficient Preemption on GPUsabstractGPUs are widely adopted in HPC and cloud computing platforms to accelerate general-purpose workloads. However, modern GPUs do not support flexible preemption, leading to performance and priority inversion problems in multi-tasking environments. Bo Wu 0002, Xu Liu 0001, Xiaobo Zhou 0002, Changjun Jiang 0002 |
ASPLOS | 3 |
| 2017 | mBalloon: enabling elastic memory management for big data processingabstractBig Data processing often suffers from significant memory pressure, resulting in excessive garbage collection (GC) and out-of-memory (OOM) errors, harming system performance and reliability. Therefore, users tend to give an excessive heap size to applications to avoid job failure, causing low cluster utilization. Wei Chen 0038, Aidi Pi, Jia Rao, Xiaobo Zhou 0002 |
SoCC | 4 |
| 2017 | Preserving I/O prioritization in virtualized OSesabstractWhile virtualization helps to enable multi-tenancy in data centers, it introduces new challenges to the resource management in traditional OSes. We find that one important design in an OS, prioritizing interactive and I/O-bound workloads, can become ineffective in a virtualized OS. Resource multiplexing between multiple tenants breaks the assumption of continuous CPU availability in physical systems and causes two types of priority inversions in virtualized OSes. In this paper, we present xBalloon, a lightweight approach to preserving I/O prioritization. It uses a balloon process in the virtualized OS to avoid priority inversion in both short-term and long-term scheduling. Experiments in a local Xen environment and Amazon EC2 show that xBalloon improves I/O performance in a recent Linux kernel by as much as 136% on network throughput, 95% on disk throughput, and 125x on network tail latency. Kun Suo, Jia Rao, Luwei Cheng, Xiaobo Zhou 0002, Francis C. M. Lau 0001 |
SoCC | 5 |
| 2017 | Adaptive scheduling of parallel jobs in spark streamingabstractStreaming data analytics has become increasingly vital in many applications such as dynamic content delivery (e.g., advertisements), Twitter sentiment analysis, and security event processing (e.g., intrusion detection systems, and spam filters). Emerging stream processing systems, such as Spark Streaming, treat the continuous stream as a series of micro-batches of data and continuously process these micro-batch jobs. Such micro-batch based stream processing provides several advantages over traditional stream processing systems, which process streaming data one record at a time, including fast recovery from failures, better load balancing and scalability. However, efficient scheduling of micro-batch jobs to achieve high throughput and low latency is very challenging due to the complex data dependency and dynamism inherent in streaming workloads. In this paper, we propose A-scheduler, an adaptive scheduling approach that dynamically schedules parallel micro-batch jobs in Spark Streaming and automatically adjusts scheduling parameters to improve performance and resource efficiency. Specifically, A-scheduler dynamically schedules multiple jobs concurrently using different policies based on their data dependencies and automatically adjusts the level of job parallelism and resource shares among jobs based on workload properties. We implemented A-scheduler and evaluated it with a real-time security event processing workload. Our experimental results show that A-scheduler can reduce end-to-end latency by 42% and improve workload throughput and energy efficiency by 21% and 13%, respectively, compared to the default Spark Streaming scheduler. Dazhao Cheng, Yuan Chen 0001, Xiaobo Zhou 0002, Daniel Gmach, Dejan S. Milojicic |
INFOCOM | 3 |
| 2017 | Addressing Performance Heterogeneity in MapReduce Clusters with Elastic TasksabstractMapReduce applications, which require access to a large number of computing nodes, are commonly deployed in heterogeneous environments. The performance discrepancy between individual nodes in a heterogeneous cluster present significant challenges to attain good performance in MapReduce jobs. MapReduce implementations designed and optimized for homogeneous environments perform poorly on heterogeneous clusters. We attribute suboptimal performance in heterogeneous clusters to significant load imbalance between map tasks. We identify two MapReduce designs that hinder load balancing: (1) static binding between mappers and their data makes it difficult to exploit data redundancy for load balancing; (2) uniform map sizes is not optimal for nodes with heterogeneous performance. To address these issues, we propose FlexMap, a user-transparent approach that dynamically provisions map tasks to match distinct machine capacity in heterogeneous environments. We implemented FlexMap in Hadoop-2.6.0. Experimental results show that it reduces job completion time by as much as 40% compared to stock Hadoop and 30% to SkewTune. Wei Chen 0038, Jia Rao, Xiaobo Zhou 0002 |
IPDPS | 3 |
| 2017 | Preemptive, Low Latency Datacenter Scheduling via Lightweight Virtualization
Wei Chen 0038, Jia Rao, Xiaobo Zhou 0002 |
USENIX ATC | 3 |
| 2017 | Cross-Platform Resource Scheduling for Spark and MapReduce on YARNabstractWhile MapReduce is inherently designed for batch and high throughput processing workloads, there is an increasing demand for non-batch processes on big data, e.g., interactive jobs, real-time queries, and stream computations. Emerging Apache Spark fills in this gap, which can run on an established Hadoop cluster and take advantages of existing HDFS. As a result, the deployment model of Spark-on-YARN is widely applied by many industry leaders. However, we identify three key challenges to deploy Spark on YARN, inflexible reservation-based resource management, inter-task dependency blind scheduling, and the locality interference between Spark and MapReduce applications. The three challenges cause inefficient resource utilization and significant performance deterioration. We propose and develop a cross-platform resource scheduling middleware, iKayak, which aims to improve the resource utilization and application performance in multi-tenant Spark-on-YARN clusters. iKayak relies on three key mechanisms: reservation-aware executor placement to avoid long waiting for resource reservation, dependency-aware resource adjustment to exploit under-utilized resource occupied by reduce tasks, and cross-platform locality-aware task assignment to coordinate locality competition between Spark and MapReduce applications. We implement iKayak in YARN. Experimental results on a testbed show that iKayak can achieve 50 percent performance improvement for Spark applications and 19 percent performance improvement for MapReduce applications, compared to two popular Spark-on-YARN deployment models, i.e., YARN-client model and YARN-cluster model. Dazhao Cheng, Xiaobo Zhou 0002, Palden Lama, Jun Wu 0006, Changjun Jiang 0002 |
IEEE Trans. Computers | 2 |
| 2017 | Improving Performance of Heterogeneous MapReduce Clusters with Adaptive Task TuningabstractDatacenter-scale clusters are evolving toward heterogeneous hardware architectures due to continuous server replacement. Meanwhile, datacenters are commonly shared by many users for quite different uses. It often exhibits significant performance heterogeneity due to multi-tenant interferences. The deployment of MapReduce on such heterogeneous clusters presents significant challenges in achieving good application performance compared to in-house dedicated clusters. As most MapReduce implementations are originally designed for homogeneous environments, heterogeneity can cause significant performance deterioration in job execution despite existing optimizations on task scheduling and load balancing. In this paper, we observe that the homogeneous configuration of tasks on heterogeneous nodes can be an important source of load imbalance and thus cause poor performance. Tasks should be customized with different configurations to match the capabilities of heterogeneous nodes. To this end, we propose a self-adaptive task tuning approach, Ant, that automatically searches the optimal configurations for individual tasks running on different nodes. In a heterogeneous cluster, Ant first divides nodes into a number of homogeneous subclusters based on their hardware configurations. It then treats each subcluster as a homogeneous cluster and independently applies the self-tuning algorithm to them. Ant finally configures tasks with randomly selected configurations and gradually improves tasks configurations by reproducing the configurations from best performing tasks and discarding poor performing configurations. To accelerate task tuning and avoid trapping in local optimum, Ant uses genetic algorithm during adaptive task configuration. Experimental results on a heterogeneous physical cluster with varying hardware capabilities show that Ant improves the average job completion time by 31, 20, and 14 percent compared to stock Hadoop (Stock), customized Hadoop with industry recommendations (Heuristic), and a profilingbased configuration approach (Starfish), respectively. Furthermore, we extend Ant to virtual MapReduce clusters in a multi-tenant private cloud. Specifically, Ant characterizes a virtual node based on two measured performance statistics: I/O rate and CPU steal time. It uses k-means clustering algorithm to classify virtual nodes into configuration groups based on the measured dynamic interference. Experimental results on virtual clusters with varying interferences show that Ant improves the average job completion time by 20, 15, and 11 percent compared to Stock, Heuristic and Starfish, respectively. Dazhao Cheng, Jia Rao, Yanfei Guo, Changjun Jiang 0002, Xiaobo Zhou 0002 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2017 | iShuffle: Improving Hadoop Performance with Shuffle-on-WriteabstractHadoop is a popular implementation of the MapReduce framework for running data-intensive jobs on clusters of commodity servers.Shuffle, the all-to-all input data fetching phase between the map and reduce phase can significantly affect job performance. However, the shuffle phase and reduce phase are coupled together in Hadoop and the shuffle can only be performed by running the reduce tasks. This leaves the potential parallelism between multiple waves of map and reduce unexploited and resource wastage in multi-tenant Hadoop clusters, which significantly delays the completion of jobs in a multi-tenant Hadoop cluster. More importantly, Hadoop lacks the ability to schedule task efficiently and mitigate the data distribution skew among reduce tasks, which leads to further degradation of job performance. In this work, we propose to decouple shuffle from reduce tasks and convert it into a platform service provided by Hadoop. We presentiShuffle, a user-transparent shuffle service that pro-actively pushes map output data to nodes via a novelshuffle-on-writeoperation and flexibly schedules reduce tasks considering workload balance. Experimental results with representative workloads and Facebook workload trace show that iShuffle reduces job completion time by as much as 29.6 and 34 percent in single-user and multi-user clusters, respectively. Yanfei Guo, Jia Rao, Dazhao Cheng, Xiaobo Zhou 0002 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2017 | Moving Hadoop into the Cloud with Flexible Slot Management and Speculative ExecutionabstractLoad imbalance is a major source of overhead in parallel programs such as MapReduce. Due to the uneven distribution of input data, tasks with more data become stragglers and delay the overall job completion. Running Hadoop in a private cloud opens up opportunities for expediting stragglers with more resources but also introduces problems that often outweigh the performance gain: (1) performance interference from co-running jobs may create new stragglers; (2) there exists a semantic gap between the Hadoop task management and resource pool-based virtual cluster management preventing tasks from using resources efficiently. In this paper, we strive to make Hadoop more resilient to data skew and more efficient in cloud environments. We presentFlexSlot, a user-transparent task slot management scheme that automatically identifies map stragglers and resizes their slots accordingly to accelerate task execution. FlexSlot adaptively changes the number of slots on each virtual node to balance the resource usage so that the pool of resources can be efficiently utilized. FlexSlot further improves mitigation of data skew with an adaptive speculative execution strategy. Experimental results show that FlexSlot effectively reduces job completion time up to$47.2$percent compared to stock Hadoop and two recently proposed skew mitigation and speculative execution approaches. Yanfei Guo, Jia Rao, Changjun Jiang 0002, Xiaobo Zhou 0002 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2016 | Online Adaptive Anomaly Detection for Augmented Network FlowsabstractTraditional network anomaly detection involves developing models that rely on packet inspection. However, increasing network speeds and use of encrypted protocols make per-packet inspection unsuited for today’s networks. One method of overcoming this obstacle is aggregating packet header information and performing flow-based analysis where data flow patterns are examined rather than deep packet inspection. Many existing approaches are special purpose limited to detecting specific behavior. Also, the data reduction inherent in identifying anomalous flows hinders alert correlation. In this article, we propose and develop a dynamic anomaly detection approach for augmented network flows. We sketch network state during flow creation, enabling general-purpose threat detection. We describe an efficient flow augmentation approach based on the count-min sketch that provides per-flow-, per-node-, and per-network-level statistics parallel to flow record generation. We design and develop a support vector machine-based adaptive anomaly detection and correlation mechanism, which is capable of aggregating alerts without a priori alert classification and evolving models online. We further develop a lightweight evolving alert aggregation method and combine it with a confidence forwarding mechanism identifying a small percentage predictions for additional processing. We show effectiveness of our methods on both enterprise and backbone traces. Experimental results demonstrate its ability to maintain high accuracy without the need for offline training. Dennis Ippoliti, Changjun Jiang 0002, Zhijun Ding, Xiaobo Zhou 0002 |
ACM Trans. Auton. Adapt. Syst. | 4 |
| 2016 | Elastic Power-Aware Resource Provisioning of Heterogeneous Workloads in Self-Sustainable DatacentersabstractWhile major Cloud service operators have taken various initiatives to operate their datacenters with renewable energy partially or completely, it is challenging to effectively utilize the renewable energy since its generation depends on dynamic natural conditions. In this paper, we propose and develop an elastic power-aware resource provisioning approach (ePower) for heterogeneous workloads in self-sustainable datacenters that completely rely on renewable energy. We aim to maximize the system goodput and control the system power consumption with respect to green power supply. ePower takes challenges and advantages of dynamic power supply, heterogeneous workload characteristics and QoS requirements, and automatically optimizes elastic resource allocations to workloads. The core of ePower design is a novel power-aware simulated annealing algorithm with fuzzy performance modeling for the efficient search of an optimal resource allocation. We have implemented ePower in a university cloud testbed hosting Gridmix2 and RUBiS benchmark applications. We utilize real weather data traces to simulate the green power generation and supply in the experiments. Experimental results demonstrate ePower can achieve near-to-optimal system performance while being resilient to dynamic power availability. It outperforms a representative resource provisioning approach for heterogeneous workloads by at least 24% in improving system goodput and 35 percent in reducing QoS violations. Dazhao Cheng, Jia Rao, Changjun Jiang 0002, Xiaobo Zhou 0002 |
IEEE Trans. Computers | 4 |
| 2016 | Autonomic Performance and Power Control for Co-Located Web Applications in Virtualized DatacentersabstractIn a datacenter, complex and time-varying interactions between various tiers and services of web applications, and the contention of shared resources among co-located virtual machines have significant impact on the user perceived performance and power consumption of the underlying system. We propose and develop APPLEware, an autonomic middleware for joint performance and power control of co-located web applications in virtualized datacenters. It features a distributed control structure that provides predictable performance and energy efficiency for large complex systems. It applies machine learning based self-adaptive modeling to capture the complex and time-varying relationship between the application performance and allocation of resources to various application components, in the face of highly dynamic and bursty workloads. The distributed controllers coordinate with each other and allocate resources to meet the service level agreements of applications in an agile and energy-efficient manner. Experimental results based on a testbed implementation with benchmark applications and large scale simulations demonstrate APPLEware's effectiveness, energy efficiency and scalability. Palden Lama, Yanfei Guo, Changjun Jiang 0002, Xiaobo Zhou 0002 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2015 | Towards Energy Efficiency in Heterogeneous Hadoop Clusters by Adaptive Task AssignmentabstractThe cost of powering servers, storage platforms and related cooling systems has become a major component of the operational costs in big data deployments. Hence, the design of energy-efficient Hadoop clusters has attracted significant research attentions in recent years. However, existing studies do not consider the impact of the complex interplay between workload and hardware heterogeneity on energy efficiency. In this paper, we find that heterogeneity-oblivious task assignment approaches are detrimental to both performance and energy efficiency of Hadoop clusters. Importantly, we make a counterintuitive observation that even heterogeneity-aware techniques that focus on reducing job completion time do not necessarily guarantee energy efficiency. We propose a heterogeneity-aware task assignment approach, E-Ant, that aims to minimize the overall energy consumption in a heterogeneous Hadoop cluster without sacrificing job performance. It adaptively schedules heterogeneous workloads on energy-efficient machines, without a priori knowledge of the workload properties. Furthermore, it provides the flexibility to trade off energy efficiency and job fairness in a Hadoop cluster. E-Ant employs an ant colony optimization approach that generates task assignment solutions based on the feedback of each task's energy consumption reported by Hadoop Task Trackers in an agile way. Experimental results on a heterogeneous cluster with varying hardware capabilities show that E-Ant improves the overall energy savings for a synthetic workload from Microsoft by 17% and 12% compared to Fair Scheduler and Tarazu, respectively. Dazhao Cheng, Palden Lama, Changjun Jiang 0002, Xiaobo Zhou 0002 |
ICDCS | 4 |
| 2015 | StoreApp: A shared storage appliance for efficient and scalable virtualized Hadoop clustersabstractVirtualizing Hadoop clusters provides many benefits, including rapid deployment, on-demand elasticity and secure multi-tenancy. However, a simple migration of Hadoop to a virtualized environment does not fully exploit these benefits. The dual role of a Hadoop worker, acting as both a compute node and a data node, makes it difficult to achieve efficient IO processing, maintain data locality, and exploit resource elasticity in the cloud. We find that decoupling per-node storage from its computation opens up opportunities for IO acceleration, locality improvement, and on-the-fly cluster resizing. To fully exploit these opportunities, we propose StoreApp, a shared storage appliance for virtual Hadoop worker nodes co-located on the same physical host. To completely separate storage from computation and prioritize IO processing, StoreApp pro-actively pushes intermediate data generated by map tasks to the storage node. StoreApp also implements late-binding task creation to take the advantage of prefetched data due to mis-aligned records. Experimental results show that StoreApp achieves up to 61% performance improvement compared to stock Hadoop and resizes the cluster to the (near) optimal degree of parallelism. Yanfei Guo, Jia Rao, Dazhao Cheng, Changjun Jiang 0002, Cheng-Zhong Xu 0001, Xiaobo Zhou 0002 |
INFOCOM | 6 |
| 2015 | Resource and Deadline-Aware Job Scheduling in Dynamic Hadoop ClustersabstractAs Hadoop is becoming increasingly popular in large-scale data analysis, there is a growing need for providing predictable services to users who have strict requirements on job completion times. While earliest deadline first scheduling (EDF) like algorithms are popular in guaranteeing job deadlines in real-time systems, they are not effective in a dynamic Hadoop environment, i.e., a Hadoop cluster with dynamically available resources. As there is a growing number of Hadoop clusters deployed on hybrid systems, e.g., infrastructure powered by mix of traditional and renewable energy, and cloud platforms hosting heterogeneous workloads, variable resource availability becomes common when running Hadoop jobs. In this paper, we propose, RDS, a Resource and Deadline-aware Hadoop job Scheduler that takes future resource availability into consideration when minimizing job deadline misses. We formulate the job scheduling problem as an online optimization problem and solve it using an efficient receding horizon control algorithm. To aid the control, we design a self-learning model to estimate job completion times and use a simple but effective model to predict future resource availability. We have implemented RDS in the open source Hadoop implementation and performed evaluations with various benchmark workloads. Experimental results show that RDS substantially reduces the penalty of deadline misses by at least 36% and 10% compared with Fair Scheduler and EDF scheduler, respectively. Dazhao Cheng, Jia Rao, Changjun Jiang 0002, Xiaobo Zhou 0002 |
IPDPS | 4 |
| 2015 | Fault tolerant MapReduce-MPI for HPC clustersabstractBuilding MapReduce applications using the Message-Passing Interface (MPI) enables us to exploit the performance of large HPC clusters for big data analytics. However, due to the lacking of native fault tolerance support in MPI and the incompatibility between the MapReduce fault tolerance model and HPC schedulers, it is very hard to provide a fault tolerant MapReduce runtime for HPC clusters. We propose and develop FT-MRMPI, the first fault tolerant MapReduce framework on MPI for HPC clusters. We discover a unique way to perform failure detection and recovery by exploiting the current MPI semantics and the new proposal of user-level failure mitigation. We design and develop the checkpoint/restart model for fault tolerant MapReduce in MPI. We further tailor the detect/resume model to conserve work for more efficient fault tolerance. The experimental results on a 256-node HPC cluster show that FT-MRMPI effectively masks failures and reduces the job completion time by 39%. Yanfei Guo, Wesley Bland, Pavan Balaji, Xiaobo Zhou 0002 |
SC | 4 |
| 2015 | Understanding Parallel Performance Under Interferences in Multi-tenant CloudsabstractThe performance of parallel programs is notoriously difficult to reason in virtualized environments. Although performance degradations caused by virtualization and interferences have been well studied, there is little understanding why different parallel programs have unpredictable slow- downs. We find that unpredictable performance is the result of complex interplays between the design of the program, the memory hierarchy of the hosting system, and the CPU scheduling in the hypervisor. We develop a profiling tool, vProfile, to decompose parallel runtime into three parts: compute, steal and synchronization. With the help of time breakdown, we devise two optimizations at the hypervisor to reduce slowdowns. Jia Rao, Xiaobo Zhou 0002, Qing Yi |
SIGMETRICS | 3 |
| 2015 | Introduction Special Section of ICCCN 2014 Conference
Pavan Balaji, Lisong Xu, Changjun Jiang 0002, Xiaobo Zhou 0002 |
Comput. Commun. | 4 |
| 2015 | Self-Tuning Batching with DVFS for Performance Improvement and Energy Efficiency in Internet ServersabstractPerformance improvement and energy efficiency are two important goals in provisioning Internet services in datacenter servers. In this article, we propose and develop a self-tuning request batching mechanism to simultaneously achieve the two correlated goals. The batching mechanism increases the cache hit rate at the front-tier Web server, which provides the opportunity to improve an application’s performance and the energy efficiency of the server system. The core of the batching mechanism is a novel and practical two-layer control system that adaptively adjusts the batching interval and frequency states of CPUs according to the service level agreement and the workload characteristics. The batching control adopts a self-tuning fuzzy model predictive control approach for application performance improvement. The power control dynamically adjusts the frequency of Central Processing Units (CPUs) with Dynamic Voltage and Frequency Scaling (DVFS) in response to workload fluctuations for energy efficiency. A coordinator between the two control loops achieves the desired performance and energy efficiency. We further extend the self-tuning batching with DVFS approach from a single-server system to a multiserver system. It relies on a MIMO expert fuzzy control to adjust the CPU frequencies of multiple servers and coordinate the frequency states of CPUs at different tiers. We implement the mechanism in a test bed. Experimental results demonstrate that the new approach significantly improves the application performance in terms of the system throughput and average response time. At the same time, the results also illustrate the mechanism can reduce the energy consumption of a single-server system by 13% and a multiserver system by 11%, respectively. Dazhao Cheng, Yanfei Guo, Changjun Jiang 0002, Xiaobo Zhou 0002 |
ACM Trans. Auton. Adapt. Syst. | 4 |
| 2015 | Coordinated Power and Performance Guarantee with Fuzzy MIMO Control in Virtualized Server ClustersabstractIt is important but challenging to assure the performance of multi-tier Internet applications with the power consumption cap of virtualized server clusters mainly due to system complexity of shared infrastructure and dynamic and bursty nature of workloads. This paper presents PERFUME, a system that simultaneously guarantees power and performance targets with flexible tradeoffs and service differentiation among co-hosted applications while assuring control accuracy and system stability. Based on the proposed fuzzy MIMO control technique, it effectively controls both the throughput and percentile-based response time of multi-tier applications due to its novel self-adaptive fuzzy modeling that integrates the strengths of fuzzy logic, MIMO control and artificial neural network. Furthermore, we address an important challenge of pro-actively avoiding violations of power and performance targets in anticipation of future workload changes. We implement PERFUME in a testbed of virtualized blade servers hosting multi-tier RUBiS applications. Performance evaluation based on synthetic and real-world Web workloads demonstrates its control accuracy, flexibility in selecting tradeoffs between conflicting targets, service differentiation capability and robustness against highly dynamic and bursty workloads. It outperforms a representative utility based approach in providing guarantee of the system throughput, percentile-based response time and power budget. Palden Lama, Xiaobo Zhou 0002 |
IEEE Trans. Computers | 2 |
| 2014 | Heterogeneity-Aware Workload Placement and Migration in Distributed Sustainable DatacentersabstractWhile major cloud service operators have taken various initiatives to operate their sustainable data enters with green energy, it is challenging to effectively utilize the green energy since its generation depends on dynamic natural conditions. Fortunately, the geographical distribution of data enters provides an opportunity for optimizing the system performance by distributing cloud workloads. In this paper, we propose a holistic heterogeneity-aware cloud workload placement and migration approach, sCloud, that aims to maximize the system good put in distributed self-sustainable data enters. sCloud adaptively places the transactional workload to distributed data enters, allocates the available resource to heterogeneous workloads in each data enter, and migrates batch jobs across data enters, while taking into account the green power availability and QoS requirements. We formulate the transactional workload placement as a constrained optimization problem that can be solved by nonlinear programming. Then, we propose a batch job migration algorithm to further improve the system good put when the green power supply varies widely at different locations. We have implemented sCloud in a university cloud test bed with real-world weather conditions and workload traces. Experimental results demonstrate sCloud can achieve near-to-optimal system performance while being resilient to dynamic power availability. It outperforms a heterogeneity-oblivious approach by 26% in improving system good put and 29% in reducing QoS violations. Dazhao Cheng, Changjun Jiang 0002, Xiaobo Zhou 0002 |
IPDPS | 3 |
| 2014 | Online Adaptive Anomaly Detection for Augmented Network FlowsabstractTraditional network anomaly detection involves developing models that rely on packet inspection. Increasing network speeds and use of encrypted protocols make per-packet inspection unsuited for today's networks. One method of overcoming this obstacle is flow based analysis. Many existing approaches are special purpose, i.e., limited to detecting specific behavior. Also, the data reduction inherent in identifying anomalous flows hinders alert correlation. In this paper we propose a dynamic anomaly detection approach for augmented flows. We sketch network state during flow creation enabling general purpose threat detection. We design and develop a support vector machine based adaptive anomaly detection and correlation mechanism capable of aggregating alerts without a-priori alert classification and evolving models online. We develop a confidence forwarding mechanism identifying a small percentage predictions for additional processing. We show effectiveness of our methods on both enterprise and backbone traces. Experimental results demonstrate the ability to maintain high accuracy without the need for offline training. Dennis Ippoliti, Xiaobo Zhou 0002 |
MASCOTS | 2 |
| 2014 | Improving MapReduce performance in heterogeneous environments with adaptive task tuningabstractThe deployment of MapReduce in datacenters and clouds present several challenges in achieving good job performance. Compared to in-house dedicated clusters, datacenters and clouds often exhibit significant hardware and performance heterogeneity due to continuous server replacement and multi-tenant interferences. As most Mapreduce implementations assume homogeneous clusters, heterogeneity can cause significant load imbalance in task execution, leading to poor performance and low cluster utilizations. Despite existing optimizations on task scheduling and load balancing, MapReduce still performs poorly on heterogeneous clusters. Dazhao Cheng, Jia Rao, Yanfei Guo, Xiaobo Zhou 0002 |
Middleware | 4 |
| 2014 | Towards fair and efficient SMP virtual machine schedulingabstractAs multicore processors become prevalent in modern computer systems, there is a growing need for increasing hardware utilization and exploiting the parallelism of such platforms. With virtualization technology, hardware utilization is improved by encapsulating independent workloads into virtual machines (VMs) and consolidating them onto the same machine. SMP virtual machines have been widely adopted to exploit parallelism. For virtualized systems, such as a public cloud, fairness between tenants and the efficiency of running their applications are keys to success. However, we find that existing virtualization platforms fail to enforce fairness between VMs with different number of virtual CPUs (vCPU) that run on multiple CPUs. We attribute the unfairness to the use of per-CPU schedulers and the load imbalance on these CPUs that incur inaccurate CPU allocations. Unfortunately, existing approaches to reduce unfairness, e.g., dynamic load balancing and CPU capping, introduce significant inefficiencies to parallel workloads. Jia Rao, Xiaobo Zhou 0002 |
PPoPP | 2 |
| 2014 | FlexSlot: Moving Hadoop Into the Cloud with Flexible Slot ManagementabstractLoad imbalance is a major source of overhead in Hadoop where the uneven distribution of input data among tasks can significantly delays the job completion. Running Hadoop in a private cloud opens up opportunities for mitigating data skew with elastic resource allocation, where stragglers are expedited with more resources, yet introduces problems that often cancel out the performance gain: (1) performance interference from co running jobs may create new stragglers, (2) there exist a semantic gap between Hadoop task management and resource pool-based virtual cluster management preventing efficient resource usage. We present Flex Slot, a user-transparent task slot management scheme that automatically identifies map stragglers and resizes their slots accordingly to accelerate task execution. Flex Slot adaptively changes the number of slots on each virtual node to promote efficient usage of resource pool. Experimental results with representative benchmarks show that Flex Slot effectively reduces job completion time by 46% and achieves better resource utilization. Yanfei Guo, Jia Rao, Changjun Jiang 0002, Xiaobo Zhou 0002 |
SC | 4 |
| 2014 | Special section of ICCCN 2013 Conference
Christian Poellabauer, Fan Zhai, Changjun Jiang 0002, Xiaobo Zhou 0002 |
Comput. Commun. | 4 |
| 2014 | Autonomic Performance and Power Control on Virtualized Servers: Survey, Practices, and Trends
Xiaobo Zhou 0002, Changjun Jiang 0002 |
J. Comput. Sci. Technol. | 1 |
| 2014 | Multi-tier service differentiation by coordinated learning-based resource provisioning and admission control
Sireesha Muppala, Guihai Chen, Xiaobo Zhou 0002 |
J. Parallel Distributed Comput. | 3 |
| 2014 | Automated and Agile Server ParameterTuning by Coordinated Learning and ControlabstractAutomated server parameter tuning is crucial to performance and availability of Internet applications hosted in cloud environments. It is challenging due to high dynamics and burstiness of workloads, multi-tier service architecture, and virtualized server infrastructure. In this paper, we investigate automated and agile server parameter tuning for maximizing effective throughput of multi-tier Internet applications. A recent study proposed a reinforcement learning based server parameter tuning approach for minimizing average response time of multi-tier applications. Reinforcement learning is a decision making process determining the parameter tuning direction based on trial-and-error, instead of quantitative values for agile parameter tuning. It relies on a predefined adjustment value for each tuning action. However it is nontrivial or even infeasible to find an optimal value under highly dynamic and bursty workloads. We design a neural fuzzy control based approach that combines the strengths of fast online learning and self-adaptiveness of neural networks and fuzzy control. Due to the model independence, it is robust to highly dynamic and bursty workloads. It is agile in server parameter tuning due to its quantitative control outputs. We implemented the new approach on a testbed of virtualized data center hosting RUBiS and WikiBench benchmark applications. Experimental results demonstrate that the new approach significantly outperforms the reinforcement learning based approach for both improving effective system throughput and minimizing average response time. Yanfei Guo, Palden Lama, Changjun Jiang 0002, Xiaobo Zhou 0002 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2013 | Optimizing virtual machine scheduling in NUMA multicore systemsabstractAn increasing number of new multicore systems use the Non-Uniform Memory Access architecture due to its scalable memory performance. However, the complex interplay among data locality, contention on shared on-chip memory resources, and cross-node data sharing overhead, makes the delivery of an optimal and predictable program performance difficult. Virtualization further complicates the scheduling problem. Due to abstract and inaccurate mappings from virtual hardware to machine hardware, program and system-level optimizations are often not effective within virtual machines. We find that the penalty to access the “uncore” memory subsystem is an effective metric to predict program performance in NUMA multicore systems. Based on this metric, we add NUMA awareness to the virtual machine scheduling. We propose a Bias Random vCPU Migration (BRM) algorithm that dynamically migrates vCPUs to minimize the system-wide uncore penalty. We have implemented the scheme in the Xen virtual machine monitor. Experiment results on a two-way Intel NUMA multicore system with various workloads show that BRM is able to improve application performance by up to 31.7% compared with the default Xen credit scheduler. Moreover, BRM achieves predictable performance with, on average, no more than 2% runtime variations. Jia Rao, Xiaobo Zhou 0002, Cheng-Zhong Xu 0001 |
HPCA | 3 |
| 2013 | pVOCL: Power-Aware Dynamic Placement and Migration in Virtualized GPU EnvironmentsabstractPower-hungry Graphics processing unit (GPU) accelerators are ubiquitous in high performance computing data centers today. GPU virtualization frameworks introduce new opportunities for effective management of GPU resources by decoupling them from application execution. However, power management of GPU-enabled server clusters faces significant challenges. The underlying system infrastructure shows complex power consumption characteristics depending on the placement of GPU workloads across various compute nodes, power-phases and cabinets in a datacenter. GPU resources need to be scheduled dynamically in the face of time-varying resource demand and peak power constraints. We propose and develop a power-aware virtual OpenCL (pVOCL) framework that controls the peak power consumption and improves the energy efficiency of the underlying server system through dynamic consolidation and power-phase topology aware placement of GPU workloads. Experimental results show that pVOCL achieves significant energy savings compared to existing power management techniques for GPU-enabled server clusters, while incurring negligible impact on performance. It drives the system towards energy-efficient configurations by taking an optimal sequence of adaptation actions in a virtualized GPU environment and meanwhile keeps the power consumption below the peak power budget. Palden Lama, Yan Li 0005, Ashwin M. Aji, Pavan Balaji, James Dinan, Shucai Xiao, Yunquan Zhang, Wu-chun Feng, Rajeev Thakur, Xiaobo Zhou 0002 |
ICDCS | 10 |
| 2013 | V-Cache: Towards Flexible Resource Provisioning for Multi-tier Applications in IaaS CloudsabstractAlthough the resource elasticity offered by Infrastructure-as-a-Service (IaaS) clouds opens up opportunities for elastic application performance, it also poses challenges to application management. Cluster applications, such as multi-tier websites, further complicates the management requiring not only accurate capacity planning but also proper partitioning of the resources into a number of virtual machines. Instead of burdening cloud users with complex management, we move the task of determining the optimal resource configuration for cluster applications to cloud providers. We find that a structural reorganization of multi-tier websites, by adding a caching tier which runs on resources debited from the original resource budget, significantly boosts application performance and reduces resource usage. We propose V-Cache, a machine learning based approach to flexible provisioning of resources for multi-tier applications in clouds. V-Cache transparently places a caching proxy in front of the application. It uses a genetic algorithm to identify the incoming requests that benefit most from caching and dynamically resizes the cache space to accommodate these requests. We develop a reinforcement learning algorithm to optimally allocate the remaining capacity to other tiers. We have implemented V-Cache on a VMware-based cloud testbed. Experiment results with the RUBiS and WikiBench benchmarks show that V-Cache outperforms a representative capacity management scheme and a cloud-cache based resource provisioning approach by at least 15% in performance, and achieves at least 11% and 21% savings on CPU and memory resources, respectively. Yanfei Guo, Palden Lama, Jia Rao, Xiaobo Zhou 0002 |
IPDPS | 4 |
| 2013 | Autonomic performance and power control for co-located Web applications on virtualized serversabstractIn a data center, various components of Web applications co-located on virtualized servers exhibit complex time-varying interactions and interference. It has a significant impact on the user perceived performance and power consumption of the underlying system. We propose and develop APPLEware, an autonomic middleware for joint performance and power control of co-located Web applications. It features a distributed control structure that provides performance assurance and energy efficiency for large complex systems. It applies machine learning based self-adaptive modeling to capture the complex and time-varying relationship between the application performance and allocation of resources to various application components, in the presence of highly dynamic and bursty workloads and inter-application performance interference. The distributed controllers perform coordinated resource allocation to meet the service level agreements of applications in an agile and energy-efficient manner. Experimental results based on a testbed implementation with benchmark applications demonstrate APPLEware's effectiveness and energy efficiency. Palden Lama, Yanfei Guo, Xiaobo Zhou 0002 |
IWQoS | 3 |
| 2013 | Self-Tuning Batching with DVFS for Improving Performance and Energy Efficiency in ServersabstractPerformance improvement and energy efficiency are two important goals in provisioning Internet services in data center servers. In this paper, we propose and develop a self-tuning request batching mechanism to simultaneously achieve the two correlated goals. The batching mechanism increases the cache hit rate at the front-tier Web server, which provides the opportunity to improve application's performance and energy efficiency of the server system. The core of the batching mechanism is a novel and practical two-layer control system that adaptively adjusts the batching interval and frequency states of CPUs according to the service level agreement and the workload characteristics. The batching control adopts a self-tuning fuzzy model predictive control approach for application performance improvement. The power control dynamically adjusts the frequency of CPUs with DVFS in response to workload fluctuations for energy efficiency. A coordinator between the two control loops achieves the desired performance and energy efficiency. We implement the mechanism in a test bed and experimental results demonstrate that the new approach significantly improves the application's performance in terms of the system throughput and average response time. The results also illustrate it can reduce the energy consumption of the server system by 13% at the same time. Dazhao Cheng, Yanfei Guo, Xiaobo Zhou 0002 |
MASCOTS | 3 |
| 2013 | Autonomic Provisioning with Self-Adaptive Neural Fuzzy Control for Percentile-Based Delay GuaranteeabstractAutonomic server provisioning for performance assurance is a critical issue in Internet services. It is challenging to guarantee that requests flowing through a multi-tier system will experience an acceptable distribution of delays. The difficulty is mainly due to highly dynamic workloads, the complexity of underlying computer systems, and the lack of accurate performance models. We propose a novel autonomic server provisioning approach based on a model-independent self-adaptive Neural Fuzzy Control (NFC). Existing model-independent fuzzy controllers are designed manually on a trial-and-error basis, and are often ineffective in the face of highly dynamic workloads. NFC is a hybrid of control-theoretical and machine learning techniques. It is capable of self-constructing its structure and adapting its parameters through fast online learning. We further enhance NFC to compensate for the effect of server switching delays. Extensive simulations demonstrate that, compared to a rule-based fuzzy controller and a Proportional-Integral controller, the NFC-based approach delivers superior performance assurance in the face of highly dynamic workloads. It is robust to variation in workload intensity, characteristics, delay target, and server switching delays. We demonstrate the feasibility and performance of the NFC-based approach with a testbed implementation in virtualized blade servers hosting a multi-tier online auction benchmark. Palden Lama, Xiaobo Zhou 0002 |
ACM Trans. Auton. Adapt. Syst. | 2 |
| 2012 | NINEPIN: Non-invasive and energy efficient performance isolation in virtualized serversabstractA virtualized data center faces important but challenging issue of performance isolation among heterogeneous customer applications. Performance interference resulting from the contention of shared resources among co-located virtual servers has significant impact on the dependability of application QoS. We propose and develop NINEPIN, a non-invasive and energy efficient performance isolation mechanism that mitigates performance interference among heterogeneous applications hosted in virtualized servers. It is capable of increasing data center utility. Its novel hierarchical control framework aligns performance isolation goals with the incentive to regulate the system towards optimal operating conditions. The framework combines machine learning based self-adaptive modeling of performance interference and energy consumption, utility optimization based performance targeting and a robust model predictive control based target tracking. We implement NINEPIN on a virtualized HP ProLiant blade server hosting SPEC CPU2006 and RUBiS benchmark applications. Experimental results demonstrate that NINEPIN outperforms a representative performance isolation approach, Q-Clouds, improving the overall system utility and reducing energy consumption. Palden Lama, Xiaobo Zhou 0002 |
DSN | 2 |
| 2012 | Multi-tier Service Differentiation: Coordinated Resource Provisioning and Admission ControlabstractMultiple Internet applications are often hosted in one datacenter and share underlying virtualized server resources. It is important but challenging to provide differentiated treatment to co-hosted applications and improve overall performance with efficient use of limited resources. We propose a coordinated self-adaptive resource management and admission control for multi-tier Internet service differentiation and performance improvement in a shared virtualized platform. We develop reinforcement learning based approaches for virtual machine (VM) auto-configuration and session based admission control. VM auto-configuration simultaneously provisions proportional service differentiation between co-located applications and improves application response time. Admission control improves session throughput of the applications, minimizing resource wastage due to aborted sessions. A shared reward actualizes coordination between the two learning modules. For system agility and scalability, we integrate reinforcement learning with cascade neural networks. We implement the integrated approach in a virtualized blade server system hosting multi-tier RUBiS applications. Experimental results demonstrate that the approach accurately meets differentiation targets and achieves performance improvement of applications. Our approach reacts to dynamic bursty workloads in agile and scalable manner. Sireesha Muppala, Xiaobo Zhou 0002, Guihai Chen |
ICPADS | 2 |
| 2012 | Automated and Agile Server Parameter Tuning with Learning and ControlabstractServer parameter tuning in virtualized data centers is crucial to performance and availability of hosted Internet applications. It is challenging due to high dynamics and burstiness of workloads, multi-tier service architecture, and virtualized server infrastructure. In this paper, we investigate automated and agile server parameter tuning for maximizing effective throughput of multi-tier Internet applications. A recent study proposed a reinforcement learning based server parameter tuning approach for minimizing average response time of multi-tier applications. Reinforcement learning is a decision making process determining the parameter tuning direction based on trial-and-error, instead of quantitative values for agile parameter tuning. It relies on a predefined adjustment value for each tuning action. However it is nontrivial or even infeasible to find an optimal value under highly dynamic and bursty workloads. We design a neural fuzzy control based approach that combines the strengths of fast online learning and self-adaptive ness of neural networks and fuzzy control. Due to the model independence, it is robust to highly dynamic and bursty workloads. It is agile in server parameter tuning due to its quantitative control outputs. We implement the new approach on a test bed of virtualized HP Pro Liant blade servers hosting RUBiS benchmark applications. Experimental results demonstrate that the new approach significantly outperforms the reinforcement learning based approach for both improving effective system throughput and minimizing average response time. Yanfei Guo, Palden Lama, Xiaobo Zhou 0002 |
IPDPS | 3 |
| 2012 | Coordinated VM Resizing and Server Tuning: Throughput, Power Efficiency and ScalabilityabstractPerformance control and power management in virtualized machines (VM) are two major research issues in modern data centers. They are challenging due to complexities of hosted Internet applications, high dynamics in workloads and the shared virtualized infrastructure. Obtaining a model among VM capacity, server configuration, performance and power consumption is a very hard problem even for just one application. In this paper, we propose and develop GARL, a genetic algorithm with multi-agent reinforcement learning approach for coordinated VM resizing and server tuning. In GARL, model-independent reinforcement learning agents generate VM capacity and server configuration options and the genetic algorithm evaluates different combinations of those options for maximizing a global utilization function of system throughput and power efficiency. The multi-agent design makes GARL a scalable approach, which is important as more and more applications are hosted in data centers using cloud services. We build a testbed in a prototype data center and deploy multiple RUBiS benchmark applications. We apply a power budget in the testbed and observe superior system throughput and power efficiency of GARL. Experimental results also find that GARL significantly outperforms a representative reinforcement learning based approach in performance control. GARL shows better scalability when compared to a centralized approach. Yanfei Guo, Xiaobo Zhou 0002 |
MASCOTS | 2 |
| 2012 | A-GHSOM: An adaptive growing hierarchical self organizing map for network anomaly detection
Dennis Ippoliti, Xiaobo Zhou 0002 |
J. Parallel Distributed Comput. | 2 |
| 2012 | Regression-based resource provisioning for session slowdown guarantee in multi-tier Internet servers
Sireesha Muppala, Xiaobo Zhou 0002, Liqiang Zhang 0002, Guihai Chen |
J. Parallel Distributed Comput. | 2 |
| 2012 | Efficient Server Provisioning with Control for End-to-End Response Time Guarantee on Multitier ClustersabstractDynamic virtual server provisioning is critical to quality-of-service assurance for multitier Internet applications. In this paper, we address three important challenging problems. First, we propose an efficient server provisioning approach on multitier clusters based on an end-to-end resource allocation optimization model. It is to minimize the number of virtual servers allocated to the system while the average end-to-end response time guarantee is satisfied. Second, we design a model-independent fuzzy controller for bounding an important performance metric, the 90th-percentile response time of requests flowing through the multitier architecture. Third, to compensate for the latency due to the dynamic addition of virtual servers, we design a self-tuning component that adaptively adjusts the output scaling factor of the fuzzy controller according to the transient behavior of the end-to-end response time. Extensive simulation results, using two representative customer behavior models in a typical three-tier web cluster, demonstrate that the provisioning approach is able to significantly reduce the number of virtual servers allocated for the performance guarantee compared to an existing representative approach. The approach integrated with the model-independent self-tuning fuzzy controller can efficiently assure the average and the 90th-percentile end-to-end response time guarantees on multitier clusters. Palden Lama, Xiaobo Zhou 0002 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2011 | aMOSS: Automated Multi-objective Server Provisioning with Stress-Strain CurvingabstractA modern data center built upon virtualized server clusters for hosting Internet applications has multiple correlated and conflicting objectives. Utility-based approaches are often used for optimizing multiple objectives. However, it is difficult to define a local utility function to suitably represent one objective and to apply different weights on multiple local utility functions. Furthermore, choosing weights statically may not be effective in the face of highly dynamic workloads. In this paper, we propose an automated multi-objective server provisioning with stress-strain curving approach (aMOSS). First, we formulate a multi-objective optimization problem that is to minimize the number of physical machines used, the average response time and the total number of virtual servers allocated for multi-tier applications. Second, we propose a novel stress-strain curving method to automatically select the most efficient solution from a Pareto-optimal set that is obtained as the result of a nondominated sorting based optimization technique. Third, we enhance the method to reduce server switching cost and improve the utilization of physical machines. Simulation results demonstrate that compared to utility-based approaches, aMOSS automatically achieves the most efficient tradeoff between performance and resource allocation efficiency. We implement aMOSS in a test bed of virtualized blade servers and demonstrate that it outperforms a representative dynamic server provisioning approach in achieving the average response time guarantee and in resource allocation efficiency for a multi-tier Internet service. aMOSS provides a unique perspective to tackle the challenging autonomic server provisioning problem. Palden Lama, Xiaobo Zhou 0002 |
ICPP | 2 |
| 2011 | PERFUME: Power and performance guarantee with fuzzy MIMO control in virtualized serversabstractIt is important but challenging to assure the performance of multi-tier Internet applications with the power consumption cap of virtualized server clusters mainly due to system complexity of shared infrastructure and dynamic and bursty nature of workloads. This paper presents PERFUME, a system that simultaneously guarantees power and performance targets with flexible tradeoffs while assuring control accuracy and system stability. Based on the proposed fuzzy MIMO control technique, it accurately controls both the throughput and percentile-based response time of multi-tier applications due to its novel fuzzy modeling that integrates strengths of fuzzy logic, MIMO control and artificial neural network. It is self-adaptive to highly dynamic and bursty workloads due to online learning of control model parameters using a computationally efficient weighted recursive least-squares method. We implement PERFUME in a testbed of virtualized blade servers hosting two multi-tier RUBiS applications. Experimental results demonstrate its control accuracy, system stability, flexibility in selecting tradeoffs between conflicting targets and robustness against highly dynamic variation and burstiness in workloads. It outperforms a representative utility based approach in providing guarantee of the system throughput, percentile-based response time and power budget in the face of highly dynamic and bursty workloads. Palden Lama, Xiaobo Zhou 0002 |
IWQoS | 2 |
| 2011 | Enhanced statistics-based rate adaptation for 802.11 wireless networks
Liqiang Zhang 0002, Yu-Jen Cheng, Xiaobo Zhou 0002 |
J. Netw. Comput. Appl. | 3 |
| 2010 | An Adaptive Growing Hierarchical Self Organizing Map for Network Intrusion DetectionabstractThe growing hierarchical self organizing map (GHSOM) has been shown to be an effective technique to facilitate anomaly detection. However, existing approaches based on GHSOM are not able to adapt online to the ever-changing problem domain of network intrusion. This results in low accuracy in identifying network intrusions, particularly "unknown" attacks. In this paper, we propose an adaptive GHSOM based approach (A-GHSOM) to network intrusion detection. It consists of four significant enhancements: enhanced threshold-based training, dynamic input normalization, feedback-based quantization error threshold adaptation, and prediction confidence filtering and forwarding. We test the capability of the A-GHSOM approach for intrusion detection using the KDD'99 dataset. Extensive experimental results demonstrate that compared with eight representative intrusion detection approaches, A-GHSOM achieves significant overall accuracy improvement and significant improvement in identifying "unknown" attacks while maintaining low false-positive rates. It achieves an overall accuracy rate of 99.63%, and 94.04% accuracy rate in identifying "unknown" attacks while the false positive rate is 1.8%. Dennis Ippoliti, Xiaobo Zhou 0002 |
ICCCN | 2 |
| 2010 | Regression based multi-tier resource provisioning for session slowdown guaranteesabstractAutonomous management of a multi-tier Internet service involves two critical and challenging tasks, one understanding its dynamic behavior when subjected to dynamic workload and second adaptive management of its resources to achieve performance guarantees. In this paper, we propose a statistical machine learning based approach to achieve session slowdown guarantees of a multi-tier Internet service. Session slowdown is the ratio of a session's total queueing delay to its total processing time. It is a compelling performance metric of session-based Internet services because it directly measures user-perceived relative performance. However, there is no analytical model for session slowdown on multi-tier servers. We first conduct training to learn the statistical regression models that quantitatively capture an Internet service's dynamic behavior as relationships between various service parameters. Then, we propose a dynamic resource provisioning approach that utilizes the learned regression models to efficiently achieve session slowdown guarantees under varying workloads. The approach is based on the combination of extensive offline training and online monitoring of the Internet service behavior. Experiments using the industry standard TPC-W benchmark demonstrate the effectiveness and efficiency of the regression based dynamic resource provisioning approach in meeting the session slowdown guarantees of a multi-tier e-commerce application. Sireesha Muppala, Xiaobo Zhou 0002, Liqiang Zhang 0002 |
IPCCC | 2 |
| 2010 | Autonomic Provisioning with Self-Adaptive Neural Fuzzy Control for End-to-end Delay GuaranteeabstractAutonomic server provisioning for performance assurance is a critical issue in data centers. It is important but challenging to guarantee an important performance metric, percentile-based end-to-end delay of requests flowing through a virtualized multi-tier server cluster. It is mainly due to dynamically varying workload and the lack of an accurate system performance model. In this paper, we propose a novel autonomic server allocation approach based on a model-independent and self-adaptive neural fuzzy control. There are model-independent fuzzy controllers that utilize heuristic knowledge in the form of rule base for performance assurance. Those controllers are designed manually on trial and error basis, often not effective in the face of highly dynamic workloads. We design the neural fuzzy controller as a hybrid of control theoretical and machine learning techniques. It is capable of self-constructing its structure and adapting its parameters through fast online learning. Unlike other supervised machine learning techniques, it does not require off-line training. We further enhance the neural fuzzy controller to compensate for the effect of server switching delays. Extensive simulations demonstrate the effectiveness of our new approach in achieving the percentile-based end-to-end delay guarantees. Compared to a rule-based fuzzy controller enabled server allocation approach, the new approach delivers superior performance in the face of highly dynamic workloads. It is robust to workload variation, change in delay target and server switching delays. Palden Lama, Xiaobo Zhou 0002 |
MASCOTS | 2 |
| 2009 | Efficient server provisioning with end-to-end delay guarantee on multi-tier clustersabstractDynamic server provisioning is critical to quality-of-service assurance for multi-tier Internet applications. In this paper, we address three important and challenging problems. First, we propose an efficient server provisioning approach on multi-tier clusters based on an end-to-end resource allocation optimization model. It is to minimize the number of servers allocated to the system while the average end-to-end delay guarantee is satisfied. Second, we design a model-independent fuzzy controller for bounding an important performance metric, the 90th-percentile delay of requests flowing through the multi-tier architecture. Third, to compensate for the latency due to the dynamic addition of servers, we design a self-tuning component that adaptively adjusts the output scaling factor of the fuzzy controller according to the transient behavior of the end-to-end delay. Extensive simulation results, using one representative customer behavior model in a typical three-tier Web cluster, demonstrate that the provisioning approach is able to significantly reduce the server utilization compared to an existing representative approach. The approach integrated with the model-independent self-tuning fuzzy controller can efficiently assure the average and the 90th-percentile end-to-end delay guarantees on multi-tier server clusters. Palden Lama, Xiaobo Zhou 0002 |
IWQoS | 2 |
| 2009 | Introduction to the special issue on self-adaptive and self-organizing wireless networking systemsabstractintroduction Share on Introduction to the special issue on self-adaptive and self-organizing wireless networking systems Authors: Michael Lemmon Department of Electrical Engineering, University of Notre Dame, USA Department of Electrical Engineering, University of Notre Dame, USAView Profile , Christian Poellabauer Department of Computer Science and Engineering, University of Notre Dame, USA Department of Computer Science and Engineering, University of Notre Dame, USAView Profile , Liqiang Zhang Department of Computer and Information Sciences, Indiana University South Bend, USA Department of Computer and Information Sciences, Indiana University South Bend, USAView Profile , Xiaobo Zhou Department of Computer Science, University of Colorado at Colorado Springs, USA Department of Computer Science, University of Colorado at Colorado Springs, USAView Profile Authors Info & Claims ACM Transactions on Autonomous and Adaptive SystemsVolume 4Issue 3Article No.: 15pp 1–4https://doi.org/10.1145/1552297.1552298Published:24 July 2009Publication History 1citation361DownloadsMetricsTotal Citations1Total Downloads361Last 12 Months1Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Michael Lemmon 0001, Christian Poellabauer, Liqiang Zhang 0002, Xiaobo Zhou 0002 |
ACM Trans. Auton. Adapt. Syst. | 4 |
| 2008 | RADAR: Rate-Alert Dynamic RTS/CTS Exchange for Performance Enhancement in Multi-Rate Wireless NetworksabstractRate adaptation is a common technique to exploit channel diversity in wireless networks. Despite the many rate adaptation algorithms proposed for 802.11 networks, the ARF (auto rate fallback) remains the most widely adopted scheme in commercial 802.11 products due to its simplicity. However, ARF suffers from some disadvantages. Our research effort revealed the rate avalanche effect that could significantly degrade the network performance of heavily loaded 802.11 networks. In this work, we propose RADAR (Rate-Alert DynAmic Rts/cts exchange) to judiciously exploit dynamic RTS/CTS exchange in multi-rate 802.11 networks. RADAR could effectively suppress the rate avalanche effect while at the same time minimizes the transmission overhead of RTS/CTS exchanges. Being fully compatible with current 802.11 standards, RADAR can be readily implemented in the NIC driver. Through extensive simulations using realistic channel propagation and reception models, we demonstrate that RADAR is a practical and efficient performance enhancement approach for multi-rate 802.11 networks. Liqiang Zhang 0002, Yu-Jen Cheng, Xiaobo Zhou 0002 |
EUC (1) | 3 |
| 2008 | DIBS: Dual interval bandwidth scheduling for short-term differentiationabstractPacket delay and bandwidth are two important metrics for measuring quality of service (QoS) of Internet services. While proportional delay differentiation (PDD) has been studied intensively in the context of differentiated services, few studies were conducted for per-class bandwidth differentiation. In this paper, we design and evaluate an efficient bandwidth differentiation approach. The DIBS (dual interval bandwidth scheduling) approach focuses on the short-term bandwidth differentiation of multiple classes because many Internet transactions take place in a small time frame. It does so based on the normalized instantaneous bandwidth, measured by the use of packet size and packet delay. It also proposes to use a look-back interval and a look-ahead interval to trade off differentiation accuracy and scheduling overhead. We implemented DIBS in the click modular software router. Extensive experiments have demonstrated its feasibility and effectiveness in achieving short-term bandwidth differentiation. Compared with the representative PDD algorithm WTP, DIBS can achieve better bandwidth differentiation when the inter-class packet size distributions are different. Compared with the representative weighted fair queueing algorithm PGPS, DIBS can achieve more accurate or comparable bandwidth differentiation at various workload situations, with better delay differentiation and lower cost. Humzah Jaffar, Xiaobo Zhou 0002, Liqiang Zhang 0002 |
IPDPS | 2 |
| 2008 | Rate avalanche: The performance degradation in multi-rate 802.11 WLANsabstractThe request-to-send/clear-to-send (RTS/CTS) exchange was defined as an optional mechanism in DCF (distributed coordination function) access method in IEEE 802.11 standard to deal with the hidden node problem. However, in most infrastructure-based WLANs, it is turned off with the belief that the benefit it brings might not even be able to pay off the transmission overhead it introduces. While this is often true for networks using fixed transmission rate, our investigation leads to the opposite conclusion when multiple transmission rates are exploited in WLANs. In particular, through extensive simulations using realistic channel propagation and reception models, we found out that in a heavily loaded multi-rate WLAN, a situation that we call rate avalanche often happens if RTS/CTS is turned off. The rate avalanche effect could significantly degrade the network performance even if no hidden node presents. Our investigation also reveals that, in the absence of effective and practical loss-differentiation mechanisms, simply turning on the RTS/CTS could dramatically enhance the network performance in most cases. Various scenarios/conditions are extensively examined to study their impact on the network performance for RTS/CTS on and off respectively. Our study provides some important insights about using the RTS/CTS exchange in mutlirate 802.11 WLANs. Liqiang Zhang 0002, Yu-Jen Cheng, Xiaobo Zhou 0002 |
IPDPS | 3 |
| 2008 | Fair bandwidth sharing and delay differentiation: Joint packet scheduling with buffer management
Xiaobo Zhou 0002, Dennis Ippoliti, Liqiang Zhang 0002 |
Comput. Commun. | 1 |
| 2008 | Special issue: Resource management and routing in wireless mesh networks
Xiaobo Zhou 0002, Liqiang Zhang 0002 |
Comput. Commun. | 1 |
| 2008 | Special issue: Modeling, testbeds, and applications in wireless mesh networks
Xiaobo Zhou 0002, Liqiang Zhang 0002 |
Comput. Commun. | 1 |
| 2008 | Resource allocation optimization for quantitative service differentiation on server clusters
Xiaobo Zhou 0002, Dennis Ippoliti |
J. Parallel Distributed Comput. | 1 |
| 2007 | Packet Scheduling with Buffer Management for Fair Bandwidth Sharing and Delay DifferentiationabstractPacket delay and bandwidth are two important metrics for measuring quality of service (QoS) of Internet services. Traditionally, packet delay differentiation and fair bandwidth sharing are studied separately. In this paper, we first propose a generalized model for providing fair bandwidth sharing with delay differentiation, namely FBS-DD. It essentially aims to provide multi-dimensional proportional differentiation with respect to both QoS metrics at the same time. We design packet scheduling schemes that take both packet delay and packet size into considerations, without assuming admission control. Furthermore, we propose a control-theoretic buffer management scheme. The packet scheduling with buffer management approach provides delay and bandwidth differentiation in an integrated way, while existing approaches consider delay and loss rate differentiation as orthogonal issues. It enhances the flexibility of network resource management and multi-dimensional QoS provisioning. It is capable of self-adapting to varying workloads from different classes, which automatically builds a firewall around aggressive clients and hence protects network resources from saturation. Simulation results by the use of trace files demonstrate that the approach can provide predictable fair bandwidth sharing with delay differentiation at various situations. The control-theoretic buffer management scheme improves the controllability. Dennis Ippoliti, Xiaobo Zhou 0002, Liqiang Zhang 0002 |
ICCCN | 2 |
| 2007 | Hop-count based probabilistic packet dropping: Congestion mitigation with loss rate differentiation
Xiaobo Zhou 0002, Dennis Ippoliti, Terrance E. Boult |
Comput. Commun. | 1 |
| 2007 | Quality-of-service differentiation on the Internet: A taxonomy
Xiaobo Zhou 0002, Jianbin Wei, Cheng-Zhong Xu 0001 |
J. Netw. Comput. Appl. | 1 |
| 2007 | Efficient algorithms of video replication and placement on a cluster of streaming servers
Xiaobo Zhou 0002, Cheng-Zhong Xu 0001 |
J. Netw. Comput. Appl. | 1 |
| 2007 | Distributed denial-of-service and intrusion detection
Xiaobo Zhou 0002, Cheng-Zhong Xu 0001 |
J. Netw. Comput. Appl. | 1 |
| 2006 | Quantitative Service Differentiation: A Square-Root Proportional Model
Xiaobo Zhou 0002, Cheng-Zhong Xu 0001 |
EUC | 1 |
| 2006 | HPPD: A Hop-count Probabilistic Packet DropperabstractNetwork applications and users have very diverse service expectations and requirements, demanding for provisioning of different levels of quality of service on the Internet. Packet loss rate differentiation has been an active research topic. However, the existing packet dropping schemes for loss rate differentiation have not considered an important issue, that is, the retransmission overhead of dropped packets. In this paper, we design a hop-count based probabilistic packet dropper (HPPD) for congestion mitigation and loss rate differentiation. HPPD aims to meet a two-fold objective by two-dimensional loss rate differentiation: one is the congestion mitigation that aims to reduce congestion in the first place by dropping intra-class packets differently based on their maturity levels to reduce retransmission cost; the other is inter-class proportional loss rate differentiation. The maturity level of a packet, the number of hops travelled, is inferred from its time-to-live value in the IP header. We propose an intra-class nth-root proportional dropping scheme, where n is a controllable parameter trading off dropping fairness for congestion mitigation. Simulation results show that HPPD can significantly mitigate the congestion by reducing the retransmission overhead of dropped packets and achieve the proportional loss rate differentiation at the same time. Xiaobo Zhou 0002, Dennis Ippoliti, Terrance E. Boult |
ICC | 1 |
| 2006 | Landscape-3D; A Robust Localization Scheme for Sensor Networks over Complex 3D TerrainsabstractDespite the fact that sensor networks could often be deployed over three-dimensional (3D) terrains, most approaches on sensor localizations are designed and evaluated considering only two-dimensional (2D) applications. On the other hand, being the foundation of the most previous localization solutions, reliable and sufficient neighborhood-measurements are often hard to achieve for sensor nodes deployed in complex 3D terrains, which makes it difficult to extend those solutions into 3D applications. In the paper, we introduce a robust 3D localization solution called Landscape-3D, in which we treat the localization problem from a novel perspective by taking it as a functional dual of target tracking. Besides several nice features, such as high scalability, high accuracy, zero sensor-to-sensor communication overhead, low computation overhead, etc., one of the most important advantages of Landscape-3D is that it works totally independent of node densities and network topologies, which makes it robust to complex 3D environments. Our simulation model involves various 3D scenarios. Experimental results demonstrate that Landscape-3D is a robust localization approach for sensor networks deployed in complex 3D terrains Liqiang Zhang 0002, Xiaobo Zhou 0002, Qiang Shawn Cheng |
LCN | 2 |
| 2006 | A robust packet scheduling algorithm for proportional delay differentiation services
Jianbin Wei, Cheng-Zhong Xu 0001, Xiaobo Zhou 0002, Qing Li 0007 |
Comput. Commun. | 3 |
| 2006 | An integrated approach with feedback control for robust Web QoS design
Xiaobo Zhou 0002, Yu Cai 0002, C. Edward Chow |
Comput. Commun. | 1 |
| 2006 | Special issue: Security in grid and distributed systems
Weisong Shi, Cheng-Zhong Xu 0001, Xiaobo Zhou 0002 |
J. Parallel Distributed Comput. | 3 |
| 2006 | Resource Allocation for Session-Based Two-Dimensional Service Differentiation on e-Commerce ServersabstractA scalable e-commerce server should be able to provide different levels of quality of service (QoS) to different types of requests based on clients' navigation patterns and the server capacity. E-commerce workloads are composed of sessions. In this paper, we propose a session-based two-dimensional (2D) service differentiation model for online transactions: intersession and intrasession. The intersession model aims to provide different levels of QoS to sessions from different customer classes, and the intrasession model aims to provide different levels of QoS to requests in different states of a session. A primary performance metric of online transactions is slowdown. It measures the waiting time of a request relative to its service time. We present a processing rate allocation scheme for 2D proportional slowdown differentiation. We then introduce service slowdown as a systemwide QoS metric of an e-commerce server. It is defined as the weighted sum of request slowdown in different sessions and in different session states. We formulate the problem of 2D service differentiation as an optimization of processing rate allocation with the objective of minimizing the service slowdown of the server. We prove that the derived rate allocation scheme based on the optimization guarantees client requests' slowdown to be square-root proportional to their prespecified differentiation weights in both intersession and intrasession dimensions. We evaluate this square-root proportional rate allocation scheme and a proportional rate allocation scheme via extensive simulations. Results validate that both schemes can achieve predictable, controllable, and fair 2D service differentiation on e-commerce servers. The square-root proportional rate allocation scheme provides 2D service differentiation at a minimum cost of service slowdown Xiaobo Zhou 0002, Jianbin Wei, Cheng-Zhong Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2005 | Two-Tier Resource Allocation for Slowdown Differentiation on Server ClustersabstractSlowdown, defined as the ratio of a request's queueing delay to its service time, is accepted as an important quality of service metric of Internet servers. In this paper, we investigate the problem of providing proportional slowdown differentiation (PSD) services to various applications and clients on cluster-based Internet servers. We extend a closed-form expression of the expected slowdown of a popular Internet workload model with a typical heavy-tailed service time distribution from a single server mode to a server cluster mode. Based on the closed-form expression, we design a two-tier resource allocation approach, which integrates a dispatcher-based node partitioning scheme and a server-based dynamic process allocation scheme. We evaluate the two-tier resource allocation approach via extensive simulations and compare it with an one-tier node partitioning approach. Simulation results show that the two-tier approach can provide fine-grained PSD services on cluster-based Internet servers. We implement the two-tier approach on a cluster testbed. Experimental results further demonstrate the feasibility of the approach in practice. Xiaobo Zhou 0002, Yu Cai 0002, C. Edward Chow, Marijke F. Augusteijn |
ICPP | 1 |
| 2005 | A Robust Application-Level Approach for Responsiveness DifferentiationabstractThere is a growing demand for provisioning of proportional responsiveness differentiation to various clients on scalable Web servers to meet changing resource availability, and to satisfy different client requirements. Theoretically, a queueing-based processing rate allocation scheme is able to achieve the objective by providing different processing rates to requests of different client classes. However, we find that an implementation of the queueing-theoretical scheme shows weak proportionality with large variance because it does not have fine-grained control over the resources that the kernel consumes and hence the processing rate is not strictly proportional to the number of processes allocated. We design a feedback controller and integrate it with the queueing-theoretical scheme. The integrated application-level approach allocates a certain number of processes to handle requests of different client classes according to the queueing-theoretical scheme. The process allocations are then adjusted according to the difference between target response time and the achieved response time by using proportional integral derivative control. Results demonstrate that this integrated approach can enable Web servers to provide fine-grained response time differentiation. The approach is robust and can be practically deployed on Apache Web servers. Xiaobo Zhou 0002, Yu Cai 0002, Jianbin Wei, Cheng-Zhong Xu 0001 |
ICWS | 1 |
| 2005 | Robust Processing Rate Allocation for Proportional Slowdown Differentiation on Internet ServersabstractA desirable behavior of an Internet server is that a request's queuing delay depends on its service time in a linear fashion. Measuring the quality of service in terms of slowdown, the ratio of a request's queuing delay to its service time, provides a simple way to attain the objective. Moreover, it treats client requests equally regardless of their service time, whereas response time favors requests that need more processing resources. In this paper, we propose a proportional slowdown differentiation (PSD) service model on Internet servers. It aims to maintain prespecified slowdown ratios between different classes of client requests. To provide PSD services, we first derive a closed-form expression of the expected slowdown in an M/G/1 FCFS queuing system with a typical heavy-tailed service time distribution, the bounded Pareto distribution. Based on the closed-form expression, we design a queuing-theoretic strategy of processing-rate allocation. The rate allocation is realized by deploying a virtual server for each class. Simulation results show that the strategy can provide controllable PSD services on Internet servers. It, however, comes along with large variance and weak predictability due to the dynamics of Internet traffic. To address these issues, we design an integral feedback controller and integrate it into the queuing-theoretic strategy. Simulation results demonstrate that the integrated strategy is robust and can deliver predictable PSD services at a superior fine-grained level. We modified the Apache Web server with an implementation of the integrated processing-rate allocation strategy. Experimental results further demonstrate its effectiveness and feasibility in practice. Jianbin Wei, Xiaobo Zhou 0002, Cheng-Zhong Xu 0001 |
IEEE Trans. Computers | 2 |
| 2004 | A robust packet scheduling algorithm for proportional delay differentiation servicesabstractThe proportional delay differentiation (PDD) model is an important approach for relative differentiated services provisioning on the Internet. It aims to maintain pre-specified packet queueing-delay ratios between different classes of traffic at each hop. Existing PDD packet scheduling algorithms are able to achieve the goal in long time-scales when the system is highly utilized. The paper presents a new PDD scheduling algorithm, called Little's average delay (LAD), based on a proof of Little's law. It monitors the arrival rate and the cumulative delays of the packets from each traffic class, and schedules the packets according to their transient queueing properties so as to achieve the desired class delay ratios in both short and long time-scales. Simulation results show that, in comparison with other PDD scheduling algorithms, LAD can provide no worse level of service quality in long time-scales and more accurate and robust control over the delay ratio in short time-scales. In particular, LAD outperforms its main competitors significantly when the desired delay ratio is large. Jianbin Wei, Cheng-Zhong Xu 0001, Xiaobo Zhou 0002 |
GLOBECOM | 3 |
| 2004 | Modeling and Analysis of 2D Service Differentiation on e-Commerce ServersabstractA scalable e-commerce server should be able to provide different levels of quality of service (QoS) to different types of requests according to clients' navigation patterns and the server capacity. In this paper, we propose a two-dimensional (2D) service differentiation (DiffServ) model for online transactions: inter-session and intra-session. The inter-session model aims to provide different levels of QoS to sessions from different customer classes, and the intra-session model aims to provide different levels of QoS to requests in different states of a session. We introduce service slowdown as a QoS metric of e-commerce servers. It is defined as the weighted sum of request slowdown in different sessions and in different session states. We formulate the problem of 2D DiffServ provisioning as an optimization of processing rate allocation with the objective of minimizing service slowdown. We derive the optimal allocations for an M/G/1 server under various server load conditions and prove that the optimal allocations guarantees requests' slowdown to be square-root proportional to their pre-specified differentiation weights in both dimensions. We evaluate the optimal allocation scheme via extensive simulations and compare it with a tailored proportional DiffServ scheme. Simulation results validate that both allocation schemes can achieve predictable, controllable, and fair 2D slowdown differentiation on e-commerce servers. The optimal allocation scheme guarantees 2D DiffServ at a minimum cost of service slowdown. Xiaobo Zhou 0002, Jianbin Wei, Cheng-Zhong Xu 0001 |
ICDCS | 1 |
| 2004 | An Adaptive Process Allocation Strategy for Proportional Responsiveness Differentiation on Web ServersabstractThere is a growing demand for provisioning of different levels of quality of service (QoS) on scalable Web servers to meet changing resource availability and satisfy different client requirements. The proportional differentiation model is getting momentum because of its fairness and differentiation predictability. It states that QoS of different traffic classes should be kept proportional to their pre-specified differentiation parameters, independent of the class loads. In this paper, we present a processing rate allocation scheme for providing proportional response time differentiation on Web servers. A challenging issue is how to achieve processing rates for different request classes in the implementation. We propose a process allocation strategy, which dynamically and adaptively changes the number of processes allocated for handling different request classes while ensuring the ratios of process allocation specified by the processing rate allocation scheme. We implement the process allocation strategy at application level on Apache Web servers. Experimental results show that the processing rate can be achieved by the adaptive process allocation strategy and the Web servers can provide predictable and controllable proportional response time differentiation. Xiaobo Zhou 0002, Yu Cai 0002, Ganesh Godavari, C. Edward Chow |
ICWS | 1 |
| 2004 | Processing Rate Allocation for Proportional Slowdown Differentiation on Internet ServersabstractSummary form only given. A proportional differentiation model states that quality of service of different classes of Internet traffic should be kept proportional to their prespecified differentiation parameters, independent of the class loads. The model has been applied in the proportional queueing delay differentiation (FDD) in both network core and network edges. However, in the server side, an important and interesting performance metric is slowdown, the ratio of a request's queueing delay to its service time. Slowdown is important because it is desirable that a request's delay be proportional to its processing requirement. We investigate the problem of processing rate allocation for proportional slowdown differentiation (PSD) on Internet servers. Existing algorithms for FDD provisioning in the network side are not applicable to PSD provisioning in the server side because slowdown is not only dependent on a job's queueing delay but also on its service time, which varies significantly depending on the requested services. We first derive a closed form expression of the expected slowdown in an M/Gp/1 FCFS queue, which is an M/G/l FCFS queue with a typical heavy-tailed service time distribution (bounded Pareto distribution). PSD provisioning is realized by deploying a task server for handling each request class in a FCFS way. We then develop a strategy of processing rate allocation for the task servers for PSD provisioning. Simulation results have showed that the proposed rate allocation strategy can provide predictable and controllable PSD services on the servers. Xiaobo Zhou 0002, Jianbin Wei, Cheng-Zhong Xu 0001 |
IPDPS | 1 |
| 2004 | Harmonic Proportional Bandwidth Allocation and Scheduling for Service Differentiation on Streaming ServersabstractTo provide ubiquitous access to the proliferating rich media on the Internet, scalable streaming servers must be able to provide differentiated services to various client requests. Recent advances of transcoding technology make network-I/O bandwidth usages at the server communication ports controllable by request schedulers on the fly. In this article, we propose a transcoding-enabled bandwidth allocation scheme for service differentiation on streaming servers. It aims to deliver high bit rate streams to high priority request classes without overcompromising low priority request classes. We investigate the problem of providing differentiated streaming services at application level in two aspects: stream bandwidth allocation and request scheduling. We formulate the bandwidth allocation problem as an optimization of a harmonic utility function of the stream quality factors and derive the optimal streaming bit rates for requests of different classes under various server load conditions. We prove that the optimal allocation, referred to as harmonic proportional allocation, not only maximizes the system utility function, but also guarantees proportional fair sharing between classes with different prespecified differentiation weights. We evaluate the allocation scheme, in combination with two popular request scheduling approaches, via extensive simulations and compare it with an absolute differentiation strategy and a proportional-share strategy tailored from relative differentiation in networking. Simulation results show that the harmonic proportional allocation scheme can meet the objective of relative differentiation in both short and long timescales and greatly enhance the service availability and maintain low queueing delay when the streaming system is highly loaded. Xiaobo Zhou 0002, Cheng-Zhong Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2002 | Optimal Video Replication and Placement on a Cluster of Video-on-Demand ServersabstractA cost-effective approach to building up scalable video-on-demand (VoD) servers is to couple a number of VoD servers together in a cluster. In this article, we study a crucial video replication and placement problem in a distributed storage VoD cluster for high quality and high availability services. We formulate it as a combinatorial optimization problem with objectives of maximizing the encoding bit rate and the number of replicas of each video and balancing the workload of the servers. It is subject to the constraints of the storage capacity and the outgoing network bandwidth of the servers. Under the assumption of single fixed encoding bit rate for all videos, we give an optimal replication algorithm and a bounded-placement algorithm for videos with different popularities. To reduce the complexity of the replication algorithm, we present an efficient algorithm that utilizes the Zipf-like video popularity distributions to approximate the optimal solution. For videos with scalable encoding bit rates, we propose a heuristic algorithm based on simulated annealing. We conduct a comprehensive performance evaluation of the algorithms and demonstrate their effectiveness via simulations over a synthetic workload set. Xiaobo Zhou 0002, Cheng-Zhong Xu 0001 |
ICPP | 1 |
| 2001 | A Video Replacement Policy based on Revenue to Cost Ratio in a Multicast TY-Anytime SystemabstractThis paper analyzes a tree hierarchical network architecture employing video caching and multicasting capacity to support a large scale Video-on-Demand service called TV-Anytime. The host servers are connected to other host servers which collectively store all the videos in the system. The proxy servers are located close to customers and are able to store the most popular videos in an adaptive way. Considering many uncertainties in the future demands for videos, we propose an on-line video replacement policy based on a revenue to cost ratio, with an objective of maximizing the overall revenues generated by the system during the runtime. Simulation results show that this policy leads to an efficient TV-Anytime system and makes the system more adaptive to changes in video popularity which is typically the case in the TV industry. The simulation results also show that multicasting significantly improves the system throughput during the high-load periods and makes the system more scalable. 1 Xiaobo Zhou 0002, Cheng-Zhong Xu 0001, Lars-Olof Burchard, Reinhard Lüling |
IPDPS | 1 |