Tianyu Wo

dblp:58/6260 · DBLP profile ↗
← Back
102ranked-venue papers
1as first author
51since 2021 · last 2026
0000-0002-5331-3364ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 42 · 1 first-author · 19 since 2021Databases, data management, data science and information retrieval · 23 · 11 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 5 since 2021Artificial intelligence and machine learning · 11 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Computer networks · 4 · 3 since 2021Human-computer interaction and ubiquitous computing · 3Software engineering, systems software and programming languages · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MathSE: Improving Multimodal Mathematical Reasoning via Self-Evolving Iterative Reflection and Reward-Guided Fine-Tuning
abstract
Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in vision-language answering tasks. Despite their strengths, these models often encounter challenges in achieving complex reasoning tasks such as mathematical problem-solving. Previous works have focused on fine-tuning on specialized mathematical datasets. However, these datasets are typically distilled directly from teacher models, which capture only static reasoning patterns and leaving substantial gaps compared to student models. This reliance on fixed teacher-derived datasets not only restricts the model's ability to adapt to novel or more intricate questions that extend beyond the confines of the training data, but also lacks the iterative depth needed for robust generalization. To overcome these limitations, we propose MathSE, a Mathematical Self-Evolving framework for MLLMs. In contrast to traditional one-shot fine-tuning paradigms, MathSE iteratively refines the model through cycles of inference, reflection, and reward-based feedback. Specifically, we leverage iterative fine-tuning by incorporating correct reasoning paths derived from previous-stage inference and integrating reflections from a specialized Outcome Reward Model (ORM). To verify the effectiveness of MathSE, we evaluate it on a suite of challenging benchmarks, demonstrating significant performance gains over backbone models. Notably, our experimental results on MathVL-test surpass the leading open-source multimodal mathematical reasoning model QVQ.
Zhen Yang 0034, Jianxin Shi 0004, Tianyu Wo, Jie Tang 0001
AAAI4
2026 Maestro: Workload-Aware Cross-Cluster Scheduling for LLM-based Multi-Agent Systems
Tianyu Wo, Xu Wang 0007, Chunming Hu, Renyu Yang
ICDCS6
2026 Mitigating Privacy Risks in Graph Condensation from a Hyperbolic Geometry Perspective
abstract
Graph condensation reduces large graphs into smaller synthetic ones for efficient training and potential privacy protection. While existing studies demonstrate graph condensation's resilience against membership inference attacks (MIAs), key questions remain unanswered: Can the common MIAs' accuracy truly represent the privacy-preserving capabilities of graph condensation? Does it remain robust against more powerful adversaries? And what are the underlying reasons for its performance? This paper investigates the privacy risks of gradient-matching-based condensation via tailored MIAs. We reveal that existing methods often face a trade-off between performance and generalization, where increasing node diversity can unintentionally amplify privacy leakage. Moreover, existing methods either homogenize nodes of the same class to maximize task-specific performance at the cost of generalization or enhance node diversity by efficiently incorporating additional information to improve model generalization, but such diversity inevitably expands the attack reasoning due to increased data disparity. To better balance performance and privacy, we propose a novel graph condensation framework (HDGC) that investigates privacy issues in graph condensation from a hyperbolic geometric perspective. Specifically, we first leverage hyperbolic geometric properties to constrain gradient-matching directions ( HGGM ), thereby obtaining latent hierarchical semantic guidance when learning the synthetic graph's topology. This mechanism measures node importance in hyperbolic space to enhance model generalization. Subsequently, we introduce hyperbolic adaptive differentially private noise during gradient matching ( HADP ). This perturbation intelligently adjusts noise influence based on local gradient importance and global geometric radius, ensuring diversity among same-class nodes while preserving differential privacy. Finally, relying on the post-processing principle of differential privacy, we incorporate distributionally robust optimization to mitigate excessive utility degradation caused by noise injection without compromising privacy guarantees. Experiments and analyses demonstrate that HDGC effectively captures geometric space characteristics, achieves superior performance, and provides a great foundation for defending inference attacks.
Yuecen Wei, Beining Yang, Qingyun Sun, Hao Peng 0001, Tianyu Wo, Chunming Hu, Xingcheng Fu
KDD (1)6
2026 IR-KG: An Industrial Robot Knowledge Graph Integrating Multi-source Heterogeneous Data
Ningbo Gu, Huanchun Peng, Tianyu Wo, Qingqian Zhou
KSEM (3)5
2026 RobotDiffuse: Diffusion-Based Motion Planning for Redundant Manipulators with the ROP Obstacle Avoidance Dataset
Xudong Mou, Tiejun Wang 0002, Tianyu Wo, Cangbai Xu, Rui Wang 0118, Xudong Liu 0001
KSEM (1)4
2026 Towards Geometry-Consistent Federated Graph Learning
Yuecen Wei, Zhiyu Zhuang, Yisen Gao, Xingcheng Fu, Qingyun Sun, Ziwei Zhang 0001, Tianyu Wo, Chunming Hu
WWW7
2025 IOP: An Idempotent-Like Optimization Method on the Pareto Front of Hypernetwork
abstract
Pareto Front Learning (PFL) has been one of the effective means to resolve multi-objective optimization problems through exploring all optimal solutions to learn the entire Pareto front. Pareto Hypernetwork (PHN) is a new promising way to generate the sequence of Pareto-optimal solutions that can be further used as potential solutions to constitute the Pareto front. However, the existing PHN-based approaches suffer from two performance issues: They take as inputs human-crafted preference vector or chunk embedding, rather than the input data samples, and thus vulnerable to data distribution shifts. Such approaches cannot optimize all potential solutions when forming the Pareto front, as they merely optimize the loss pertaining to one single input at a time of optimization round. To improve the quality of the Pareto front, we propose IOP, a novel Idempotent-like Optimization method to learn the entire Pareto front accurately and enhance Hypernetwork's adaptability to distribution shifts. In particular, IOP performs idempotent-like optimization by exploiting manifold space mapping, so that the target networks generated by the optimized Hypernetwork can effectively handle samples with similar distributions of the input samples, without the pre-defined human-crafted inputs. IOP maximizes the Hypervolume indicator that is composed of all potential solutions at a higher level. Experimental results demonstrate that IOP outperforms the state-of-the-art methods by 4.7% on average in producing the Pareto front and has a 10.5% improvement in adaptability.
Hui Wang 0157, Renyu Yang, Jie Sun 0035, Hao Peng 0001, Xudong Mou, Tianyu Wo, Xudong Liu 0001
AAAI6
2025 Cauchy: A Cost-Efficient LLM Serving System through Adaptive Heterogeneous Deployment
abstract
Recent advances in large language models (LLMs) have intensified the need for serving LLMs that are cost-efficient and QoS-guaranteed. Existing frameworks often co-locate computationally distinct prefill and decode instances on homogeneous GPUs, overlooking their unique resource demands and under-utilizing heterogeneous GPUs. This leads to suboptimal resource utilization and increased capital expenditure. We present Cauchy, a LLM serving framework that adaptively deploys prefill and decode computation to the most suitable heterogeneous GPUs and dynamically schedules user requests. At the core of Cauchy is choosing proper GPU Combo, a conceptual GPU combination encompassing diverse GPU configurations, for their cost efficiency in running prefill-decode pairs. Cauchy deploys a set of combos to satisfy QoS requirements (e.g., goodput) of LLM inference. Cauchy further employs hierarchical scheduling to handle user requests, using opportunistic scheduling within the allocated GPU Combos and a goodput-weighted round-robin policy across GPU Combos. Dynamic autoscaling is used to stabilize the cost-efficiency in the face of surging requests. Experiments show that Cauchy achieves up to a 38.3% improvement in Tokens/USD efficiency over the state-of-the-art baselines, while maintaining strict Service Level Objectives (SLOs). Our work highlights the importance of leveraging workload and GPU heterogeneity to achieve superior cost-efficient LLM serving.
Renyu Yang, Yuxi Luo, Menghao Zhang 0001, Li Li 0029, Chunming Hu, Tianyu Wo, Chengru Song, Jin Ouyang
SoCC9
2025 Cuckoo: Deadline-Aware Job Packing on Heterogeneous GPUs for DL Model Training
abstract
The growing scale and heterogeneity of GPU clusters pose new challenges to deep learning (DL) job scheduling. While existing schedulers primarily focus on GPU utilization, they often ignore multi-dimensional resource demands of DL workloads and lack precise execution time estimation for co-located jobs. While Muri pioneered the use of interleaving execution to improve resource efficiency, it simplified interference when jobs using one resource simultaneously and is agnostic to the deadline constraints. The job grouping also comes to suboptimal when heterogeneous GPU devices are taken into account. In this paper, we propose Cuckoo, a scheduling system that packs deep learning jobs with stringent deadline requirements over a set of heterogeneous GPU devices where multi-dimensional resources are interleaved and shared by a group of jobs. Specifically, the interleaving execution of simultaneous jobs is characterized and modeled through stage-grained execution time estimation considering the runtime performance interference and the impact of GPU heterogeneity on the job performance. The job packing is formulated as a multi-objective optimization problem which is then solved by the maximum weight matching algorithm. Cuckoo then allocates heterogeneous resources to the packed job groups through a graph-based maximum flow and minimum cut algorithm. Experiments show that Cuckoo improves deadline satisfaction rate by 2.38x and reduces average job completion time (JCT) by 1.81x compared with the state-of-the-art approaches. Cuckoo is implemented based on Kubernetes and has been deployed in Kuaishou to serve thousands of model training jobs that can be interleaved on shared heterogeneous GPU clusters
Yuzheng Zhang, Renyu Yang, Weihan Jiang, Tianyu Ye, Yiqiao Liao, Penghao Zhang, Tiezi Zhang, Tianyu Wo, Chunming Hu, Chengru Song, Jin Ouyang
SoCC10
2025 SuperFE: A Scalable and Flexible Feature Extractor for ML-based Traffic Analysis Applications
abstract
The feature extractor component in today's ML-based traffic analysis applications is becoming a key bottleneck. While mainstream software-based approaches can support flexible feature extraction, they fail to scale to multi-100Gbps network speed easily. Meanwhile, hardware-accelerated solutions can scale to high throughput, but cannot flexibly support generic traffic analysis applications. In this paper, we propose SuperFE, a feature extraction framework that allows users to extract traffic features efficiently and flexibly. SuperFE leverages the capabilities of both new-generation programmable switches and SmartNICs, with three key designs. First, SuperFE presents a user-friendly and extensible interface to support customized feature extraction policies, shielding underlying hardware implementation details and complexities. Second, SuperFE introduces a high-performance multi-granularity key-vector cache system in the programmable switches to batch necessary feature metadata for massive amounts of packets. Third, SuperFE exploits the multi-core parallel and hierarchical memory of SoC-based SmartNICs to achieve efficient feature computation with diverse streaming algorithms. Evaluations using our prototype demonstrate that SuperFE enables various state-of-the-art traffic analysis applications to efficiently extract features from multi-100Gbps raw traffic without compromising detection accuracy, and achieves nearly two orders of magnitude higher throughput than the software-based counterparts.
Menghao Zhang 0001, Cheng Guo 0007, Renyu Yang, Han Bao 0011, Xiao Li 0044, Mingwei Xu 0001, Tianyu Wo, Chunming Hu
EuroSys9
2025 Adaptive Parallel SFC Joint Scheduling for Time-Varying Satellite-Terrestrial Integration Networks
abstract
The increasing demand for ultra-low-latency services in satellite-terrestrial integrated networks (STINs) necessitates adaptive resource orchestration to meet stringent QoS requirements under dynamic network conditions. While Network Function Virtualization (NFV) and Service Function Chaining (SFC) offer flexibility, existing approaches primarily focus on static scenarios and fail to adapt to time-varying topology and inter-SFC resource competition in STIN. To address these limitations, we propose an adaptive parallel SFC scheduling framework that optimizes real-time SFC offloading and virtual network function aggregation to minimize end-to-end latency. We formulate the problem as a Mixed-Integer Nonlinear Programming model that jointly optimizes computation offloading, node selection, and link mapping under dynamic network constraints. To avoid the local optimal solution, the proposed framework efficiently estimates and mitigates the global costs caused by offloading and transmission link conflicts. Extensive experimental results demonstrate that our approach reduces the average end-to-end latency by 28.37% compared to state-of-the-art methods under high task density.
Tianyu Wo, Penglin Yang, Fortunatus Kawasa, Xingchen Fu
GLOBECOM2
2025 MF-BERT: A Siamese Pre-training Framework for Motion Forecasting
abstract
Accurately predicting the future motions of traffic agents is essential for autonomous systems. Despite the significant success of existing motion forecasting methods based on supervised learning, they still exhibit two main limitations. First, when annotated data for a scene is limited, these methods often fail to achieve the expected accuracy. Second, they typically rely on complex architectures and extensive prior knowledge to improve performance. To overcome these challenges, we propose MF-BERT, a novel framework that adapts the concept of BERT to motion forecasting, inspired by advancements in the self-supervised pre-training paradigm. During pre-training, we design a siamese sequence modeling task with an asymmetric mask strategy to capture complex behavior patterns of agents. During fine-tuning, the pre-trained representation module initializes the feature encoder of the motion forecasting model, and a multimodal trajectory decoder generates all possible predictions. Experimental results demonstrate the superiority of MF-BERT over state-of-the-art methods.
Jianxin Shi 0004, Jun Ma 0008, Tianyu Wo
ICASSP5
2025 ScNet: Scene-Consistency Network Learning for Multi-Agent Motion Forecasting
abstract
Predicting the motion of traffic agents is a fundamental challenge in autonomous driving, essential for safe and efficient ego-vehicle planning. Traditional methods typically focus on marginal forecasting, where the trajectory of each agent is predicted separately, leading to inconsistencies in scene-level predictions. To address this issue, we propose a scene-consistency network, named ScNet, which jointly predicts the trajectories of multiple agents in a single feedforward pass, ensuring consistency across all predictions. Our method leverages dual independently initialized student models that interact through cross-network contrastive learning at the global feature level, enhancing robustness and scene consistency in the learned representations. To further improve scene coherence, we incorporate a scene-guided strategy that refines these representations. Additionally, we employ a lightweight, anchor-free decoder that generates predictions for all agents, aligning the forecasts with real-world dynamics. Experiments show significant improvements in multi-world prediction metrics across complex environments. Code and models will be publicly available.
Jianxin Shi 0004, Yusen Xie, Fali Wang, Jun Ma 0008, Tianyu Wo
ICME7
2025 FOCA: Foundation-model-based One-Class Anomaly Detection for Time Series
abstract
Time series anomaly detection is challenging due to the rarity of anomalies and the complexity and diversity of normal patterns. Most existing methods rely on a single hypothesis and learn feature patterns from a limited dataset, which restricts their generalization capabilities. At the same time, time series foundation models have shown promising results across multiple tasks due to their strong generalization capabilities. However, time series foundation models are less likely to achieve better performance in complex anomaly detection tasks. To address this issue, this paper introduces FOCA, a novel foundation-model-based one-class anomaly detection approach. This method preserves the generalization ability of the foundation model to capture normal variation patterns and provides a comprehensive feature space for one-class classification. It constrains normal features within a sufficiently small hypersphere to construct a decision boundary for detecting abnormal data. Furthermore, it is observed in practice that the introduction of fine-tuning techniques can further improve the performance of the method. Extensive experiments on two standard benchmark datasets demonstrate that our method outperforms the state-of-the-art approaches.
Tiejun Wang 0002, Rui Wang 0118, Xudong Mou, Tianyu Wo, Xudong Liu 0001
IJCNN5
2025 STCC-Sim: A Satellite-Terrestrial Collaborative Computing Modeling and Simulating Toolkit for Resource Provisioning
abstract
The Satellite-Terrestrial Integrated Network (STIN) technology based on low-earth orbit (LEO) satellites provides a seamless network with low latency and high reliability to global users. The Space-Terrestrial Collaborative Computing (STCC) paradigm based on STIN has become a promising solution for ubiquitous computing. Due to the mobility of satellites and the vulnerability of inter-satellite and satellite-terrestrial network connections, the research on resource scheduling strategy, configuration, and deployment of STCC services is more complex than that of ground cloud services. The lack of simulators further restricts the development of STCC research. In order to solve this problem, we propose an STCC simulation tool for the satellite-terrestrial hybrid cloud in this paper. The tool can simulate the hybrid network models and support the simulation of the flow computing model. By evaluating the performance of the task offloading policy, we demonstrate the effectiveness of the simulation tool.
Tianyu Wo, Xudong Liu 0001
JCC4
2025 LogAD: A Multi-Feature Fusion Approach for Log Anomaly Detection
abstract
With the increasing complexity of software systems, log-based anomaly detection has become critical for ensuring system reliability. However, existing methods often suffer from limited feature integration and insufficient semantic representation, leading to unstable detection performance. To address these challenges, this paper proposes a multi-feature fusion framework for log anomaly detection, leveraging heterogeneous graph neural networks (HGNNs) to capture rich semantic relationships. First, we design a hybrid preprocessing pipeline that combines log parsing (via Drain), session-fixed window grouping, and hybrid label estimation using HDBSCAN clustering and HNSW-based similarity search. This step mitigates label scarcity while enhancing feature representation robustness. Second, we construct a heterogeneous graph with three node types-log sequences, templates, and parameters-to model interdependencies between log events through meta-paths, enabling comprehensive feature fusion. Third, a heterogeneous graph attention network (HGAT) with multi-head attention is developed to prioritize critical patterns across meta-paths, improving anomaly discrimination. Experimental results on benchmark datasets demonstrate that our model outperforms state-of-the-art baselines in accuracy and F1-score. Furthermore, we implement LogAD, an automated detection tool integrating ELK-stack-based log management, multi-feature anomaly detection, and security-focused operational support. The system's visualization interface and efficient processing pipeline provide a practical solution for real-world deployment. This work advances log analysis by bridging feature isolation and semantic sparsity, offering both algorithmic innovation and engineering applicability.
Guangzu Wang, Lingzhi Zhang, Tianyu Wo, Xu Wang 0007, Chunming Hu
JCC4
2025 KAIOPS: A Platform Solution of End-to-End Multi-Modal AIOps for AI Training at Scale
abstract
The resilience of large-scale AI training platforms are fundamental to enabling contemporary AI innovation and business development. However, with the rapid increase in the scale and complexity of AI model training tasks, anomalies become the norm rather than the exception at scale. Failing to handle them properly may lead to enormous resource waste and prolonged development cycles. Traditional anomaly detection methods struggle to tackle the complex temporal characteristics and extreme class imbalance inherently manifesting in training tasks, and fall short in automated solution to root cause analysis and the follow-up remediation. This paper proposes KAIOPS, an end-to-end automated platform solution for handling anomalies and engineering experience of daily operational maintenance for large-scale AI training clusters at Kuaishou. KAIOPS employs a Temporal Context Encoding mechanism to precisely capture and encode long-term trends and critical temporal context information within fault evolution. The detection model elaborates a dynamic class-weighted loss function for enhancing the detection performance. To deliver a complete end-to-end intelligent processing pipeline, KAIOPS further leverages knowledge graph and LLMs for automated root cause analysis and actionable solution generation. Extensive experiments, on the basis of data collected from Kuaishou’s production-grade training clusters, show the superior performance of our proposed approach. KAIOPS has been deployed in Kuaishou, in both testbed and production grade environments, consisting of with over 10,000 GPUs, and accelerate the reliability assurance for industry-scale model training and serving.
Zeying Wang, Penghao Zhang, Xu Wang 0007, Tianyu Wo, Chunming Hu, Chengru Song, Jin Ouyang, Renyu Yang
ASE6
2025 Kair: A Statistical and Causal Approach to Pinpointing Stragglers in Distributed Model Training
abstract
The distributed deep learning training process within large-scale clusters serves as the foundation of contemporary artificial intelligence. However, its inherent characteristics make it particularly sensitive to stragglers, specifically the presence of slow workers, which can significantly decelerate the entire procedure. Observability tools are essential for identifying stragglers within systems. However, the prevailing system profiling tools are either designed for single-node analysis, lacking visibility across multiple workers, or they recognize stragglers but only deliver high-level symptoms, providing engineers with insufficient insight into the underlying causes.We design Kair, a robust production-standard observability tool. Kair uses an innovative hierarchical approach, transitioning from statistical anomaly detection to causal inference. It employs Kolmogorov-Smirnov statistics for the identification of statistically anomalous workers and implements a causal path tracing algorithm to accurately determine the specific operations, such as computation or communication, that are responsible for the delay. Kair has been evaluated in a production cluster of 2,048 NVIDIA A800 GPUs and demonstrated high effectiveness in detecting latent stragglers at the framework level that are often overlooked by conventional tools. It offers precise suggestions that markedly reduce processing inefficiencies and engineering workload.
Yitang Yang, Jiapeng Chen, Tianyu Wo, Chunming Hu, Chengru Song, Jin Ouyang, Renyu Yang
ASE5
2025 ACbot: an IIoT platform for industrial robots
Rui Wang 0118, Xudong Mou, Tianyu Wo, Tiejun Wang 0002, Pin Liu, Jihong Yan, Xudong Liu 0001
Frontiers Comput. Sci.3
2025 A federated anti-forgetting representation method based on hybrid model architecture and gradient truncation
Hui Wang 0157, Jie Sun 0035, Tianyu Wo, Xudong Liu 0001, Suzhen Pei
Frontiers Comput. Sci.3
2024 Kale: Elastic GPU Scheduling for Online DL Model Training
abstract
Large-scale GPU clusters have been widely used for effectively training both online and offline deep learning (DL) jobs. However, elastic scheduling in most cases of resource schedulers is dedicated for offline model training where resource adjustment is planned ahead of time. The native autoscaling policy is on the basis of pre-defined threshold and, if applied directly in online model training, often suffers from belated resource adjustment, leading to diminished model accuracy. In this paper, we present Kale, a novel elastic GPU scheduling system to improve the performance of online DL model training. Through traffic forecasting and resource-throughput modeling, Kale automatically pinpoints the number of required GPUs that best accommodate the on-the-fly data samples before performing stabilized autoscaling. An advanced data shuffling strategy is further employed for balancing uneven samples among different training workers, thereby improving the runtime efficacy. Experiments show that Kale substantially outperforms the state-of-the-art solutions. Compared with the default HPA autoscaling strategy, Kale reduces the accumulated lag and downtime by 69.2% and 33.1%, respectively, whilst lowering the SLO violation rate from 19.57% to just 2.6%. Kale has been deployed at Kuaishou's production-level GPU clusters and successfully underpins real-time video recommendation and advertisement at scale.
Renyu Yang, Jin Ouyang, Weihan Jiang, Tianyu Ye, Menghao Zhang 0001, Sui Huang, Chengru Song, Di Zhang 0026, Tianyu Wo, Chunming Hu
SoCC11
2024 LDPRecover: Recovering Frequencies from Poisoning Attacks Against Local Differential Privacy
abstract
Local differential privacy (LDP), which enables an untrusted server to collect aggregated statistics from distributed users while protecting the privacy of those users, has been widely deployed in practice. However, LDP protocols for frequency estimation are vulnerable to poisoning attacks, in which an attacker can poison the aggregated frequencies by manipulating the data sent from malicious users. Therefore, it is an open challenge to recover the accurate aggregated frequencies from poisoned ones. In this work, we propose LDPRecover, a method that can recover accurate aggregated frequencies from poisoning attacks, even if the server does not learn the details of the attacks. In LDPRecover, we establish a genuine frequency estimator that theoretically guides the server to recover the frequencies aggregated from genuine users' data by eliminating the impact of malicious users' data in poisoned frequencies. Since the server has no idea of the attacks, we propose an adaptive attack to unify existing attacks and learn the statistics of the malicious data within this adaptive attack by exploiting the properties of LDP protocols. By taking the estimator and the learning statistics as constraints, we formulate the problem of recovering aggregated frequencies to approach the genuine ones as a constraint inference (CI) problem. Consequently, the server can obtain accurate aggregated frequencies by solving this problem optimally. Moreover, LDPRecover can serve as a frequency recovery paradigm that recovers more accurate aggregated frequencies by integrating attack details as new constraints in the CI problem. Our evaluation on two real-world datasets, three LDP protocols, and untargeted and targeted poisoning attacks shows that LDPRecover is both accurate and widely applicable against various poisoning attacks.
Xinyue Sun, Qingqing Ye 0001, Haibo Hu 0001, Jiawei Duan, Tianyu Wo, Jie Xu 0007, Renyu Yang
ICDE5
2024 FedFRR: Federated Forgetting-Resistant Representation Learning
abstract
Continuous learning faces the challenge of catastrophic forgetting. Our research findings indicate that in unsupervised federated continual learning (UFCL), the limited model capacity and interference among participants are the key factors contributing to this problem. Specifically, the fixed capacity of the model restricts its ability to retain historical knowledge. Besides, the indiscriminate aggregation of weights from multiple participants can cause interference, damaging the model memory. To address these challenges, we propose FedFRR, a federated anti-forgetting representation learning approach. FedFRR fits the participants’ data distribution through a weighted combination of primary network units (PNU) in the model and optimizes model memory by adjusting the structure of PNUs. Additionally, FedFRR addresses interference by truncating the PNU with less weight change, thus reducing the scope of weight aggregation. The experimental results demonstrate that FedFRR achieves state-of-the-art performance, significantly enhancing the model’s anti-forgetting ability.
Hui Wang 0157, Jie Sun 0035, Tianyu Wo, Xudong Liu 0001
ICME3
2024 PrecisionProbe: Non-intrusive Performance Analysis Tool for Deep Learning Recommendation Models
abstract
Deep learning recommendation models (DLRM) exploit user behaviors such as clicks, browse footprints, preferences, etc. for improved personalized experiences. However, in the face of the exponential growth of user data, such models require increasing GPU resources that are unaffordable and insufficient in a computing cluster. To improve GPU utilization and facilitate the advances of GPU scheduling algorithms, we present PrecisionProbe, a non-intrusive monitoring and analysis tool that can run upon Kubernetes and conduct sophisticated analytics of GPU resource utilization without altering the existing training code. PrecisionProbe captures fine-grained GPU metrics at the level of individual model layers and allows for a precise understanding of resource consumption patterns by exploring such detailed metrics. The mechanism is crucial for devising effective GPU scheduling algorithms, particularly tailored for DLRM training jobs dependent upon consumption patterns. Experimental results show that the recommendation models, as opposed to CV and NLP models, utilize less FP32 processing but have higher memory interaction frequencies. These findings indicate the unique resource needs of recommendation systems and necessitate the need of performance analytic using PrecisionProbe.
Weiyu Peng, Tianyu Wo, Renyu Yang
JCC3
2024 RESCAPE: A Resource Estimation System for Microservices with Graph Neural Network and Profile Engine
abstract
Microservice architecture has become a prevalent paradigm for constructing scalable and flexible cloud-native applications by leveraging the abundant resources of the cloud. However, the topological complexity of microservices poses significant challenges to resource management frameworks that rely on container orchestration. It is paramount to optimize resource utilization within cloud computing clusters while reducing operational costs for service providers. To this end, we present RESCAPE, a framework designed to effectively predict the resource demands of variable microservice workloads. It is instrumental for downstream optimization tasks, particularly heterogeneous resource scheduling, aiming to enhance resource utilization and efficiency. Experiments based on open-source microservice benchmarks such as DeathStarBench and HPC-AI500 demonstrate an average absolute percentage error (MAPE) of 7.9% when forecasting resource needs for the subsequent timestamp, which indicates an adequate precision for resource estimation of microservices.
Guangzu Wang, Tianyu Wo, Xu Wang 0007, Renyu Yang
JCC3
2024 CutAddPaste: Time Series Anomaly Detection by Exploiting Abnormal Knowledge
abstract
Detecting time-series anomalies is extremely intricate due to the rarity of anomalies and imbalanced sample categories, which often result in costly and challenging anomaly labeling. Most of the existing approaches largely depend on assumptions of normality, overlooking labeled abnormal samples. While anomaly assumptions based methods can incorporate prior knowledge of anomalies for data augmentation in training classifiers, the adopted random or coarse-grained augmentation approaches solely focus on pointwise anomalies and lack cutting-edge domain knowledge, making them less likely to achieve better performance. This paper introduces CutAddPaste, a novel anomaly assumption-based approach for detecting time-series anomalies. It primarily employs a data augmentation strategy to generate pseudo anomalies, by exploiting prior knowledge of anomalies as much as possible. At the core of CutAddPaste is cutting patches from random positions in temporal subsequence samples, adding linear trend terms, and pasting them into other samples, so that it can well approximate a variety of anomalies, including point and pattern anomalies. Experiments on standard benchmark datasets demonstrate that our method outperforms the state-of-the-art approaches.
Rui Wang 0118, Xudong Mou, Renyu Yang, Pin Liu, Chongwei Liu, Tianyu Wo, Xudong Liu 0001
KDD7
2024 Generating Location Traces With Semantic- Constrained Local Differential Privacy
abstract
Valuable information and knowledge can be learned from users’ location traces and support various location-based applications such as intelligent traffic control, incident response, and COVID-19 contact tracing. However, due to privacy concerns, no authority could simply collect users’ private location traces for mining or even publishing. To echo such concerns, local differential privacy (LDP) enables individual privacy by allowing each user to report a perturbed version of their data. Unfortunately, when applied to location traces, LDP cannot preserve the semantics in the context of location traces because it treats all locations (i.e., various points of interest) as equally sensitive. This results in a low utility of LDP mechanisms for collecting location traces. In this paper, we address the challenge of collecting and sharing location traces with valuable semantics while providing sufficient privacy protection for participating users. We first propose semantic-constrained local differential privacy (SLDP), a new privacy model to provide a provable mathematical privacy guarantee while preserving desirable semantics. Then, we design a location trace perturbation mechanism (LTPM) that users can use to perturb their traces in a way that satisfies SLDP. Finally, we propose a private location trace synthesis (PLTS) framework in which users use LTPM to perturb their traces before sending them to the collector, who aggregates the users’ perturbed data to generate location traces with valuable semantics. Extensive experiments on three real-world datasets demonstrate that our PLTS outperforms existing state-of-the-art methods by at least 21% in a range of real-world applications, such as spatial visiting queries and frequent pattern mining, under the same privacy leakage.
Xinyue Sun, Qingqing Ye 0001, Haibo Hu 0001, Jiawei Duan, Qiao Xue, Tianyu Wo, Weizhe Zhang, Jie Xu 0007
IEEE Trans. Inf. Forensics Secur.6
2024 PUTS: Privacy-Preserving and Utility-Enhancing Framework for Trajectory Synthesization
abstract
Vehicle trajectory data is essential for traffic management and location-based services. However, publishing real-life trajectory data has been challenging because vehicle trajectories contain users’ sensitive information. Differential privacy addresses such problems by publishing a synthetic version of the input dataset, but existing works always assume the real-world data is absolutely accurate. This assumption no longer holds in trajectory data because it typically contains errors due to inaccurate positioning services, which leads to poor performance of data synthesized by such trajectories. Even worse, existing works may generate unrealistic trajectories due to their coarse data synthesis methods, resulting in low practical utility or even inability to handle complex tasks. In this paper, we propose aPrivacy-preserving andUtility-enhancing framework forTrajectorySynthesization (PUTS). Our framework mitigates the impact of data errors in trajectories on differential privacy mechanisms, by exploiting map-matching techniques and real-world road network structure. InPUTS, a two-layer approach from path to trajectory synthesis is proposed to not only guarantee the reality of synthetic trajectories, but also scale upPUTSin real-world applications. Extensive experiments on real-world datasets show thatPUTSsignificantly outperforms existing methods in terms of utility in a range of real-world applications.
Xinyue Sun, Qingqing Ye 0001, Haibo Hu 0001, Jiawei Duan, Qiao Xue, Tianyu Wo, Jie Xu 0007
IEEE Trans. Knowl. Data Eng.6
2023 Deep Autoencoding One-Class time Series Anomaly Detection
abstract
Time-series Anomaly Detection(AD) is widely used in monitoring and security applications in various industries and has become a hot spot in the field of deep learning. Normality-representation-based methods perform well in certain scenarios but may ignore some aspects of the overall normality. Feature-extraction-based methods always take a process of pre-training, whose target differs from AD, leading to a decline in AD performance. In this paper, we propose a new AD method called deep Autoencoding One-Class (AOC), which learns features with AutoEncoder(AE). Meanwhile, the normal context vectors from AE are constrained into a hypersphere small enough, similar to one-class methods. With an objective function that optimizes the two assumptions simultaneously, AOC learns various aspects of normality, which is more effective for AD. Experiments on public datasets show that our method outperforms existing baseline approaches.
Xudong Mou, Rui Wang 0118, Tiejun Wang 0002, Jie Sun 0035, Bo Li 0005, Tianyu Wo, Xudong Liu 0001
ICASSP6
2023 FED-3DA: A Dynamic and Personalized Federated Learning Framework
abstract
In federated learning, the non-IID data generated from heterogeneous clients may reduce the global model efficiency. Previous studies use personalization as a common approach to adapt the global model to these clients (called the local model). However, client’s data distribution may change dynamically with its location or environment, which can degrade the performance of the local model, leading to a new Dynamic Personalized Federated Learning (DPFL) problem. This paper proposes a novel approach to reduce the impact of the dynamic distribution on the local model based on metalearning and distribution distance measurement named Fed-3DA. It calculates the distribution distance periodically to perceive the distribution change on the client and adjust the local model preferences from a global meta-model through the distribution representation. Our experiments on public datasets show that Fed-3DA can effectively reduce the performance fluctuation of the local model in DPFL scenarios.
Hui Wang 0157, Jie Sun 0035, Tianyu Wo, Xudong Liu 0001
ICASSP3
2023 An Approach to Workload Generation for Cloud Benchmarking: a View from Alibaba Trace
abstract
Finding performance bottlenecks through bench-marking is one of the driving forces to improve the resource provision efficiency of cloud computing. Although existing benchmarks have been designed to improve the effectiveness in system performance evaluation, the following problems still exist in these benchmarks due to insufficient consideration of the characteristics of jobs in the production environment: (i) lacking of understanding for the details of workloads composition in the production environment, which reduces the authenticity of the job. (ii) the design of workloads submission patterns lacks quantization and reproducibility, which often relies on a random setting. In our benchmarking, multiple workloads are generated by analyzing and fine-grained matching the composition of workloads in the real production, and a design of workloads submission pattern based on LSTM time series prediction is proposed to simulate the real submission behavior. We finally demonstrate the effectiveness of our work by evaluating the impact of different workloads submission patterns on system performance evaluation.
Jianyong Zhu, Xiaoqiang Yu, Jie Xu 0007, Tianyu Wo
ISADS5
2023 Fault Tolerance of Stateful Microservices for Industrial Edge Scenarios
abstract
Due to the ubiquitous increase of Industrial Internet of Things(IIoT) devices, there is a tendency to move some of the microservices-based applications from Cloud to Edge. However, edge devices are prone to node failures because of weak reliability, resulting in the loss of stateful microservices computing state, which may involve fault tolerance of stateful microservices. Moreover, the method of traditional mechanisms for microservices fault tolerance could not meet the real-time requirement. Within this context, based on stateful microservices characteristics, we propose a novel fault tolerant mechanism for IIoT Edge, which mainly consists of causal logging and distributed checkpoint algorithm. This fault recovery mechanism utilizes causal logging to record the nondeterministic events of microservices, and completes the state recovery of microservices by loading checkpoint and replaying log records, which achieves exactly-once guarantees for distributed microservices. In addition, a set of experiments was performed to evaluate the proposed mechanism by integration with Kubernetes. The results show that the proposed mechanism has less impact on service performance compared with other methods.
Yuke Jia, Tiejun Wang 0002, Tianbo Qiu, Rui Wang 0118, Tianyu Wo
JCC6
2023 HyCU: Hybrid Consistent Update for Software Defined Network
abstract
Software Defined Network (SDN) enables network operators to achieve the customization of network services, which tends to be more dynamic and fine-grained. However, the distributed nature of rule updating in SDN brings consistency problems, i.e., packets travel according to different versions of rules. It leads to the issues of blackholes, loops, congestion, and deadlock in the data plane, which may further affect the service quality of the application plane. With the emergence of new computing paradigms such as edge computing and fog computing, the heterogeneity of network devices and links, as well as the diversity of network application requirements have become increasingly prominent. Traditional update methods ignore these key factors when modeling, so they cannot cope with the increasingly complex network environment, resulting in delays or packet loss rates that do not meet service requirements. This paper proposes HyCU, which takes device performance as a constraint and optimizes updates based on flow service requirements. We conduct experiments under different scenarios and constraints over two real-world topologies with real-time running flows, demonstrating the effectiveness of HyCU.
Xudong Mou, Jie Sun 0035, Yingying Zhong, Tianyu Wo
JCC4
2023 Optimization Strategies for Data Placement in Satellite Cloud-oriented Distributed File Systems
abstract
Forming a collaborative computing network among the deployed satellites in space can process data rapidly and reduce data transmission delay by leveraging the communication, storage, and computing capacities of the satellites. Due to the dynamics of satellite orbits, the distance between satellites varies over time, bringing new challenges to the designing of dedicated distributed file systems for collaborative satellites. Traditional distributed file systems did not consider the network topology and dynamics of satellite clouds and the limited computing resources of satellite node cores. In view of the torus network topology of satellite clouds, we design a highly available distributed metadata management mechanism for satellite clouds to ensure efficient metadata reading and reliability. On this basis, a topology-aware replica placement strategy is proposed to minimize communication costs and energy consumption for replica placement. Based on the metadata strategy and replica placement strategy mentioned above, we design and implement a distributed file system for satellite clouds. Experimental results demonstrate that our proposed placement strategy can improve the data transmission performance of the replica placement by more than double compared to the random placement strategy.
Tianyu Wo, Tianyu Ye, Jiwei Zhang 0025
JCC2
2023 Deep Contrastive One-Class Time Series Anomaly Detection
abstract
The accumulation of time-series data and the absence of labels make time-series Anomaly Detection (AD) a self- supervised deep learning task. Single-normality-assumption- based methods, which reveal only a certain aspect of the whole normality, are incapable of tasks involved with a large number of anomalies. Specifically, Contrastive Learning (CL) methods distance negative pairs, many of which consist of both normal samples, thus reducing the AD performance. Existing multi-normality-assumption-based methods are usually two-staged, firstly pre-training through certain tasks whose target may differ from AD, limiting their performance. To overcome the shortcomings, a deep Contrastive One-Class Anomaly detection method of time series (COCA) is proposed by authors, following the normality assumptions of CL and one-class classification. It treats the original and reconstructed representations as the positive pair of negative-sample-free CL, namely “sequence contrast”. Next, invariance terms and variance terms compose a contrastive one-class loss function in which the loss of the assumptions is optimized by invariance terms simultaneously and the “hypersphere collapse” is prevented by variance terms. In addition, extensive experiments on two real- world time-series datasets show the superior performance of the proposed method achieves state-of-the-art. *The full version of the paper can be accessed at https://arxiv.org/abs/2207.01472
Rui Wang 0118, Chongwei Liu, Xudong Mou, Xiaohui Guo, Pin Liu, Tianyu Wo, Xudong Liu 0001
SDM7
2023 Scalable inter-domain network virtualization
Jie Sun 0035, Tianyu Wo, Xudong Liu 0001, Xudong Mou, Jinghong Lan, Jianwei Niu 0002
J. Netw. Comput. Appl.2
2023 Synthesizing Realistic Trajectory Data With Differential Privacy
abstract
Vehicle trajectory data is critical for traffic management and location-based services. However, the released trajectories raise serious privacy concerns because they contain sensitive information such as homes and workplaces. Based on differential privacy, this problem can be addressed by generating synthetic trajectories from the original sensitive data while guaranteeing personal privacy. Unfortunately, existing methods focus on synthesizing trajectory datasets that preserve summary-level statistics (e.g., the overall distribution of user movements), making these synthetic trajectories lose individual-level mobility patterns. As shown in our experiment, this results in the low performance of their synthetic datasets in real-world applications. To address these limitations, we propose a novel solution for Synthesizing Private and Realistic Trajectories, namely SPRT, whose key idea is to integrate the public geography structures of the target area into the process of private trajectory synthesis. This enables us to capture more accurate mobility patterns to synthesize realistic trajectories, which can preserve both summary-level statistics and individual-level mobility behaviors. Consequently, the synthetic trajectories generated by SPRT are more similar to real trajectories and therefore more practical. We evaluate the performance of SPRT in real-world applications by applying its synthetic data to a series of trajectory analytic tasks. The results demonstrate that our solution improves data utility by at least 37% over state-of-the-art approaches.
Xinyue Sun, Qingqing Ye 0001, Haibo Hu 0001, Yuandong Wang 0002, Kai Huang 0011, Tianyu Wo, Jie Xu 0007
IEEE Trans. Intell. Transp. Syst.6
2022 Adaptive Shapelets Preservation for Time Series Augmentation
abstract
Time series augmentation is an essential technique in training deep learning models for time series, especially achieving remarkable results in tackling the overfitting problems. However, existing methods fail to specifically protect discriminative features that contribute significantly to the classification results, so these features may be destroyed during the augmentation process. This leads to lower fidelity of augmented time series, which ultimately interferes with classification decisions during inference. To address this issue, we propose an adaptive shapelets preservation approach, named ASP. First, we exploit the saliency map to detect shapelets on the original time series that contain discriminative features. Second, we preserve them during augmentation and assign a proprietary label to each time series. It improves the fidelity of augmented time series and the confidence of their labels, thereby avoiding the risk of interfering with classification decisions. Experimental results on 128 datasets of the UCR2018 archive show that our method ASP outperforms that without augmentation on 98 datasets, and helps the classifier achieve the average accuracy improvement from 71.26% to 75.46%, which is far better than the state-of-the-art approaches.
Pin Liu, Xiaohui Guo, Bin Shi 0003, Tianyu Wo, Xudong Liu 0001
IJCNN5
2022 An Automatic Scaling System for Online Application with Microservices Architecture
abstract
Auto-scaling is an efficient technique to handle fluctuations of application workloads by acquiring or releasing resources. However, performing auto-scaling in a microservice system for online applications faces critical challenges, including unpredictably massive microservice requests, without fine-granularity performance metrics, and complex dependencies among services. In this paper, we design a cost-efficient autoscaling system, which pinpoints the scaling-needed services as quickly as possible and makes decisions on the right resource amount allocation toward them. Specifically, we first propose a multi-level microservice monitoring mechanism to capture historical and latest service-level performance metrics, and detect the over-provisioning services and under-provisioning services via jointly considering the changes of latency and throughput. For the overload anomalies, a random walk method is further adopted for detecting the root causes based on the dependency topology of microservices. When anomalies are detected, we design a threshold-based method by incorporating the ARIMI method for predicting resource usage status to allocate or recycle the right number of computation resources for them. Extensive and systematic evaluations of different algorithm modules with real-world and simulated workload data confirm the superiority of our mechanism over multiple algorithms.
Youmei Song, Kuoran Zhuang, Tianyu Wo
JCC5
2022 Two-stage Scheduling of Stream Computing for Industrial Cloud-edge Collaboration
abstract
As the Industrial Internet of Things (IIoT) develops, intelligent services applying stream computing, such as industrial robot health management, are requiring higher timeliness of data processing, which may involve scheduling of stream tasks. However, traditional scheduling methods are no longer suitable for the currently widely used cloud-edge collaboration mode, not considering the cloud-edge heterogeneity, and focusing on the scheduling of single tasks instead of the optimization of the total tasks. To improve the performance of the cloud-edge collaboration, this paper establishes a practical model for task scheduling considering respectively cloud-edge environment collaboration models. We propose a novel two-stage scheduling method for IIoT. The algorithm utilizes the idea of maximum flow to divide the task into cloud-edge deployment schemes and find the best partitioning scheme, and then deploy the operator for the edge domain based on the network topology by using dynamic programming. Experimental results show that the proposed method could reduce 7.27% the cloud-edge bandwidth usage compared with the highest greedy algorithm for traffic difference, 24.33% end-to-end latency and 11.18% back-pressure rate compared with SBON.
Tiejun Wang 0002, Xudong Mou, Rui Wang 0118, Tianyu Wo
JCC5
2022 Incentivizing Online Edge Caching via Auction - Based Subsidization
abstract
There exists a practical need for incentivizing content providers to cache contents at distributed network edges closer to users. However, this is a particularly challenging problem due to system environments that are uncertain, content placements that couple adjacent time slots, and economic properties that are desired but hard to ensure. In this paper, we present our design of an auction-based incentive mechanism for online edge caching. We formulate the long-term social cost minimization problem as a nonlinear mixed-integer program that addresses bid selections, user request dispatching, content placements, and payment determination in repetitive auctions. To solve this problem online, we devise a greedy approximation algorithm for solving each auction individually, and a lazy-replacement-based online algorithm that ties the series of auctions over time while dynamically pursuing the balance between downloading contents to new cache locations and keeping them at existing locations. We formally prove the approximation ratio for each single auction, the competitive ratio for the long-term social cost, as well as the truthfulness, the individual rationality, and the computational efficiency of our approach. Evaluations with real-world data have also validated and confirmed the practical superiority of our approach over multiple alternative algorithms.
Youmei Song, Lei Jiao 0002, Renyu Yang, Tianyu Wo, Jie Xu 0007
SECON4
2022 ScaleReactor: A graceful performance isolation agent with interference detection and investigation for container-based scale-out workloads
abstract
Summary Striking a balance between improved cluster utilization and guaranteed application QoS is a long‐standing research problem in multi‐tenants shared cluster. The typical solution is to detect performance degradation and investigate the root cause to conduct performance isolation. Existing efforts rely on lots of prior knowledge of applications and the assumption of interference‐free workload placement is possible. Performance interference is usually mitigated through application‐level approaches such as centralized rescheduling, which is usually an hindsight and a waste of resources. In this article, we present ScaleReactor, a graceful runtime agent on a per node basis that decouples the performance isolation from centralized resource management, and migrates the performance interference of scale‐out workloads in container‐based cluster using a lightweight black‐box approach. ScaleReactor analyzes the degree of contention for multi‐dimensional resources among co‐located workloads to detect the performance degradation without additional prior information, and uses correlation analysis to locate the cause of contention, while isolating resources in a graceful manner to reduce system overhead and the performance degradation of intrusive workloads. Experiments have demonstrated that ScaleReactor effectively reduces the job completion time of scale‐out applications in shared clusters, with the maximum value up to 36% and low system overhead against the existing isolation mechanism.
Jianyong Zhu, Chunming Hu, Tianyu Wo, Xiaoqiang Yu
Concurr. Comput. Pract. Exp.3
2022 Non-salient region erasure for time series augmentation
Pin Liu, Xiaohui Guo, Bin Shi 0003, Rui Wang 0118, Tianyu Wo, Xudong Liu 0001
Frontiers Comput. Sci.5
2022 Passenger Mobility Prediction via Representation Learning for Dynamic Directed and Weighted Graphs
abstract
In recent years, ride-hailing services have been increasingly prevalent, as they provide huge convenience for passengers. As a fundamental problem, the timely prediction of passenger demands in different regions is vital for effective traffic flow control and route planning. As both spatial and temporal patterns are indispensable passenger demand prediction, relevant research has evolved from pure time series to graph-structured data for modeling historical passenger demand data, where a snapshot graph is constructed for each time slot by connecting region nodes via different relational edges (origin-destination relationship, geographical distance, etc.). Consequently, the spatiotemporal passenger demand records naturally carry dynamic patterns in the constructed graphs, where the edges also encode important information about the directions and volume (i.e., weights) of passenger demands between two connected regions. aspects in the graph-structure data. representation for DDW is the key to solve the prediction problem. However, existing graph-based solutions fail to simultaneously consider those three crucial aspects of dynamic, directed, and weighted graphs, leading to limited expressiveness when learning graph representations for passenger demand prediction. Therefore, we propose a novel spatiotemporal graph attention network, namely Gallat ( G raph prediction with all at tention) as a solution. In Gallat, by comprehensively incorporating those three intrinsic properties of dynamic directed and weighted graphs, we build three attention layers to fully capture the spatiotemporal dependencies among different regions across all historical time slots. Moreover, the model employs a subtask to conduct pretraining so that it can obtain accurate results more quickly. We evaluate the proposed model on real-world datasets, and our experimental results demonstrate that Gallat outperforms the state-of-the-art approaches.
Yuandong Wang 0002, Hongzhi Yin, Tong Chen 0005, Tianyu Wo, Jie Xu 0007
ACM Trans. Intell. Syst. Technol.6
2022 QoS-Aware Co-Scheduling for Distributed Long-Running Applications on Shared Clusters
abstract
To achieve a high degree of resource utilization, production clusters need to co-schedule diverse workloads – including both batch analytic jobs with short-lived tasks and long-running applications (LRAs) that execute for a long time frame from hours to months – onto the shared resources. Microservice architecture advances the manifestation of distributed LRAs (DLRAs), comprising multiple interconnected microservices that are executed in long-lived distributed containers and serve massive user requests. Detecting and mitigating QoS violation become even more intractable due to the network uncertainties and latency propagation across dependent microservices. However, current resource managers are only responsible for resource allocation among applications/jobs but agnostic to runtime QoS such as latency at application level. The state-of-the-art QoS-aware scheduling approaches are dedicated for monolithic applications, without considering the temporal-spatio performance variability across distributed microservices. In this paper, we presentToposch, a new scheduling and execution framework to prioritize the QoS of DLRAs whilst balancing the performance of batch jobs and maintaining high cluster utilization through harvesting idle resources.Toposchtracks footprints of every single request across microservices and uses critical path analysis, based on the end-to-end latency graph, to identify microservices that have high risk of QoS violation. Based on microservice and node level risk assessment, we intervene the batch scheduling by adaptively reducing the visible resources to batch tasks and thus delaying their execution to give way to DLRAs. We propose a prediction-based vertical resource auto-scaling mechanism, with the aid of resource-performance modeling and fine-grained resource inference and access control, for prompt recovery of QoS violation. A cost-effective task preemption is leveraged to ensure a low-cost task preemption and resource reclamation during the auto-scaling.Toposchis integrated with Apache YARN and experiments show thatToposchoutperforms other baselines in terms of performance guarantee of DLRAs, at an acceptable cost of batch job slowdown. The tail latency of DLRAs is merely 1.12x of the case of executing alone on average inToposchwith a 26% JCT increase of Spark analytic jobs.
Jianyong Zhu, Renyu Yang, Tianyu Wo, Chunming Hu, Hao Peng 0001, Junqing Xiao, Albert Y. Zomaya, Jie Xu 0007
IEEE Trans. Parallel Distributed Syst.4
2021 Perph: A Workload Co-location Agent with Online Performance Prediction and Resource Inference
abstract
Striking a balance between improved cluster utilization and guaranteed application QoS is a long-standing research problem in cluster resource management. The majority of current solutions require a large number of sandboxed experimentation for different workload combinations and leverage them to predict possible interference for incoming workloads. This results in non-negligible time complexity that severely restricts its applicability to complex workload co-locations. The nature of pure offline profiling may also lead to model aging problem that drastically degrades the model precision. In this paper, we present Perph, a runtime agent on a per node basis, which decouples ML-based performance prediction and resource inference from centralized scheduler. We exploit the sensitivity of long-running applications to multi-resources for establishing a relationship between resource allocation and consequential performance. We use Online Gradient Boost Regression Tree (OGBRT) to enable the continuous model evolution. Once performance degradation is detected, resource inference is conducted to work out a proper slice of resources that will be reallocated to recover the target performance. The integration with Node Manager (NM) of Apache YARN shows that the throughput of Kafka data-streaming application is 2.0x and 1.82x times that of isolation execution schemes in native YARN and pure cgroup cpu subsystem. In TPC-C benchmarking, the throughput can also be improved by 35% and 23% respectively against YARN native and cgroup cpu subsystem.
Jianyong Zhu, Renyu Yang, Chunming Hu, Tianyu Wo, Shiqing Xue, Jin Ouyang, Jie Xu 0007
CCGRID4
2021 Gallat: A Spatiotemporal Graph Attention Network for Passenger Demand Prediction
abstract
Online ride-hailing services have become an important component of urban transportation in recent years. As a fundamental research problem for such services, the timely prediction of passenger demands in different regions is vital for effective traffic flow control. As both spatial and temporal patterns are indispensable passenger demand prediction, relevant research has evolved from pure time series to graph-structured data for modelling historical passenger demand data, where a snapshot graph is constructed for each time slot by connecting region nodes via different relational edges. Consequently, the spatiotemporal passenger demand records naturally carry dynamic patterns in the constructed graphs, where the edges also encode important information about the directions and volume (i.e., weights) of passenger demands between two connected regions. However, existing graph-based solutions fail to simultaneously consider those three crucial aspects of dynamic, directed and weighted (DDW) graphs, leading to limited expressiveness when learning graph representations for passenger demand prediction. Therefore, we propose a novel spatiotemporal graph attention network, namely Gallat (Graph prediction with all attention) as a solution. In Gallat, by comprehensively incorporating those three intrinsic properties of DDW graphs, we build three attention layers to fully capture the spatiotemporal dependencies among different regions across all historical time slots. Our experimental results on real-world datasets demonstrate that Gallat outperforms the state-of-the-art approaches.
Yuandong Wang 0002, Hongzhi Yin, Tong Chen 0005, Tianyu Wo, Jie Xu 0007
ICDE6
2021 CSMOTE: Contrastive Synthetic Minority Oversampling for Imbalanced Time Series Classification
Pin Liu, Xiaohui Guo, Rui Wang 0118, Tianyu Wo, Xudong Liu 0001
ICONIP (5)5
2021 A Cloud-edge Collaborative Architecture for Data-driven Health Condition Monitoring of Machines
abstract
With the development of the Industrial Internet of Things (IIoT), various industrial intelligent applications are emerging, especially Prognostics and Health Management (PHM). The Health Condition Monitoring(HCM) of machines is an important part of PHM and has an essential purpose to improve the intelligence level of industrial machines. However, traditional monitoring methods are no longer suitable for the current adaptability, latency, bandwidth, and privacy requirements in IIoT. In this paper, we propose a novel data-driven HCM architecture. Firstly, the Knowledge Distillation (KD) and a threshold method are respectively proposed for the cloud-edge collaborative training and inference mechanism. Secondly, a health condition representation method of multisensor signal data is proposed. Finally, experiments on the public rolling element bearing dataset show that our method can significantly improve latency and save bandwidth while ensuring accuracy.
Rui Wang 0118, Muqing Xian, Tianyu Wo
JCC4
2021 Joint optimization of cache placement and request routing in unreliable networks
Youmei Song, Tianyu Wo, Renyu Yang, Jie Xu 0007
J. Parallel Distributed Comput.2
2021 Error Bounded Line Simplification Algorithms for Trajectory Compression: An Experimental Evaluation
abstract
Nowadays, various sensors are collecting, storing, and transmitting tremendous trajectory data, and it is well known that the storage, network bandwidth, and computing resources could be heavily wasted if raw trajectory data is directly adopted. Line simplification algorithms are effective approaches to attacking this issue by compressing a trajectory to a set of continuous line segments, and are commonly used in practice. In this article, we first classify the error bounded line simplification algorithms into different categories and review each category of algorithms. We then study the data aging problem of line simplification algorithms and distance metrics from the views of aging friendliness and aging errors. Finally, we present a systematic experimental evaluation of representative error bounded line simplification algorithms, including both compression optimal and sub-optimal methods, in terms of commonly adopted perpendicular Euclidean, synchronous Euclidean, and direction-aware distances. Using real-life trajectory datasets, we systematically evaluate and analyze the performance (compression ratio, average error, running time, aging friendliness, and query friendliness) of error bounded line simplification algorithms with respect to distance metrics, trajectory sizes, and error bounds. Our study provides a full picture of error bounded line simplification algorithms, which leads to guidelines on how to choose appropriate algorithms and distance metrics for practical applications.
Xuelian Lin, Shuai Ma 0001, Yanchen Hou, Tianyu Wo
ACM Trans. Database Syst.5
2020 TOPOSCH: Latency-Aware Scheduling Based on Critical Path Analysis on Shared YARN Clusters
abstract
Balancing resource utilization and application QoS is a long-standing research topic in cluster resource management. Big data YARN clusters need to co-schedule diverse workloads on shared resources including batch processing jobs, streaming jobs, and other long-running applications such as web services, database services, etc. Current resource managers are only responsible for resource allocation among applications/jobs but completely unaware of runtime QoS requirements of interactive and latency-sensitive applications. Prior works to maximize the QoS of monolithic applications ignore inherent dependencies and temporal-spatio performance variability of components, characteristics of distributed applications primarily driven by microservices. In this paper, we present Toposch, a new resource management system to adaptively co-locate batch tasks and microservices by harvesting runtime latency. In particular, Toposch tracks full footprints of every request across microservices over time. A latency graph is periodically generated for identifying victim microservices through an end-to-end latency critical path analysis. We then exploit per-microservice and per-node risk assessment to gauge the visible resources to the capacity scheduler in YARN. Execution of batch tasks are adaptively throttled or delayed, thereby avoiding latency increase due to node over-saturation. TOPOSCH is integrated with YARN and experiments show that the latency of DLRAs can be reduced by up to 39.8% against the default capacity scheduling in YARN.
Chunming Hu, Jianyong Zhu, Renyu Yang, Hao Peng 0001, Tianyu Wo, Shiqing Xue, Xiaoqiang Yu, Jie Xu 0007, Rajiv Ranjan 0001
CLOUD5
2020 A Cost-Efficient and Low-Latency Data Hosting Scheme in JointCloud Storage
abstract
Nowadays, more and more enterprises and organizations are hosting their data into the cloud, in order to reduce the IT maintenance cost and enhance the data reliability. Generally customers put their data into a single cloud, which is subject to the vendor lock-in risk. Compared with single cloud storage, JointCloud storage can provide customers with greater flexibility and availability. In JointCloud environment, the pricing of cloud storage service and geographical distribution of cloud storage datacenters are diverse, and customers access data in different frequencies and regions. Based on comprehensive analysis of various state-of-the-art cloud vendors, this paper proposes a novel data hosting scheme, including a redundancy storage scheme using erasure coding and replication, a mathematical modeling of multi-region data placement problem, and an optimization method of global cost and latency using heuristic algorithm. We propose a dynamic adjustment method for the variations of data access pattern and a data recovery method for cloud storage service failures. We also design and implement a JointCloud storage prototype system. Finally, experiments show the effectiveness of the scheme.
Tianyu Wo, Yanlin Luo, Liyin Shao
JCC2
2020 Cloud-Edge Collaborative Industrial Robotic Intelligent Service Platform
abstract
With the development of industrial automation and intelligence deepening, industrial robots' installations increase rapidly. The traditional robot operation and management methods are no longer suitable for current requirements. Cloud computing and big data technology together are deemed to form a promising and possible solution to this challenge. From this point of view, this paper proposes a novel industrial robotics cloud platform architecture, and application orchestration and containerized software appliance technology-enabled cloud-edge collaboration mechanism is designed therein. Two exemplary and real-world case studies, supervisory control and data acquisition, and industrial robot predictive maintenance, are further conducted to demonstrate our architecture's generality and versatility.
Rui Wang 0118, Xudong Mou, Jie Sun 0035, Pin Liu, Xiaohui Guo, Tianyu Wo, Xudong Liu 0001
JCC6
2020 Performance-Aware Speculative Resource Oversubscription for Large-Scale Clusters
abstract
It is a long-standing challenge to achieve a high degree of resource utilization in cluster scheduling. Resource oversubscription has become a common practice in improving resource utilization and cost reduction. However, current centralized approaches to oversubscription suffer from the issue with resource mismatch and fail to take into account other performance requirements, e.g., tail latency. In this article we present ROSE, a new resource management platform capable of conducting performance-aware resource oversubscription. ROSE allows latency-sensitive long-running applications (LRAs) to co-exist with computation-intensive batch jobs. Instead of waiting for resource allocation to be confirmed by the centralized scheduler, job managers in ROSE can independently request to launch speculative tasks within specific machines according to their suitability for oversubscription. Node agents of those machines can however, avoid any excessive resource oversubscription by means of a mechanism for admission control using multi-resource threshold control and performance-aware resource throttle. Experiments show that in case of mixed co-location of batch jobs and latency-sensitive LRAs, the CPU utilization and the disk utilization can reach 56.34 and 43.49 percent, respectively, but the 95th percentile of read latency in YCSB workloads only increases by 5.4 percent against the case of executing the LRAs alone.
Renyu Yang, Chunming Hu, Peter Garraghan, Tianyu Wo, Zhenyu Wen, Hao Peng 0001, Jie Xu 0007
IEEE Trans. Parallel Distributed Syst.5
2019 Perphon: a ML-based Agent for Workload Co-location via Performance Prediction and Resource Inference
abstract
Cluster administrators are facing great pressures to improve cluster utilization through workload co-location. Guaranteeing performance of long-running applications (LRAs), however, is far from settled as unpredictable interference across applications is catastrophic to QoS [2]. Current solutions such as [1] usually employ sandboxed and offline profiling for different workload combinations and leverage them to predict incoming interference. However, the time complexity restricts the applicability to complex co-locations. Hence, this issue entails a new framework to harness runtime performance and mitigate the time cost with machine intelligence: i) It is desirable to explore a quantitative relationship between allocated resource and consequent workload performance, not relying on analyzing interference derived from different workload combinations. The majority of works, however, depend on offline profiling and training which may lead to model aging problem. Moreover, multi-resource dimensions (e.g., LLC contention) that are not completely included by existing works but have impact on performance interference need to be considered [3]. ii) Workload co-location also necessitates fine-grained isolation and access control mechanism. Once performance degradation is detected, dynamic resource adjustment will be enforced and application will be assigned an access to specific slices of each resources. Inferring a "just enough" amount of resource adjustment ensures the application performance can be secured whilst improving cluster utilization.
Jianyong Zhu, Renyu Yang, Chunming Hu, Tianyu Wo, Shiqing Xue, Jin Ouyang, Jie Xu 0007
SoCC4
2019 HCNet: An SDN Enabled Virtual Network Management System for Hybrid Clouds
abstract
Hybrid clouds require scalable, high performance virtual networks (VN) to facilitate the resource interoperability among multiple clouds. However, current VNs in hybrid clouds face the following challenges: (1) Management complexity. The VNs in hybrid clouds need to support inter-cloud resource deployment with multi-tenancy and full address space virtualization, and therefore are complex to manage. Meanwhile, the VNs are vulnerable to failures due to their large scale. (2) Performance limitation. Hybrid clouds rely on the Internet to support inter-cloud data transmission. Therefore without a meticulous design, the inter-cloud VNs would suffer from a performance limitation introduced by the Internet, e.g., the intercloud bandwidth provided by the Internet is much lower and more fluctuant than that of the intra-cloud networks. To the best of our knowledge, most existing VN solutions cannot solve both of these challenges simultaneously.In this paper we introduce HCNet, a hybrid cloud virtual network management system. HCNet makes use of Software Defined Networking (SDN) and: (1) supports multi-tenancy and full address space virtualization on both Layer-2 and Layer-3 networks through the help of its elaborated packet processing pipeline, (2) leverages a dynamic load balance policy atop multiple inter-cloud gateways to achieve better performance on inter-cloud virtual networking, (3) maintains master-slave backup relationship based on paxos protocol to deal with controller outages. The experiment results indicate that HCNet achieves significant inter-cloud communication performance gain with respect to packet loss. Moreover, in presence of controller failures, HCNet is able to recover in about 5 seconds.
Jie Sun 0035, Tianyu Wo, Xudong Liu 0001
ISCC3
2019 Origin-Destination Matrix Prediction via Graph Convolution: a New Perspective of Passenger Demand Modeling
abstract
Ride-hailing applications are becoming more and more popular for providing drivers and passengers with convenient ride services, especially in metropolises like Beijing or New York. To obtain the passengers' mobility patterns, the online platforms of ride services need to predict the number of passenger demands from one region to another in advance. We formulate this problem as an Origin-Destination Matrix Prediction (ODMP) problem. Though this problem is essential to large-scale providers of ride services for helping them make decisions and some providers have already put it forward in public, existing studies have not solved this problem well. One of the main reasons is that the ODMP problem is more challenging than the common demand prediction. Besides the number of demands in a region, it also requires the model to predict the destinations of them. In addition, data sparsity is a severe issue. To solve the problem effectively, we propose a unified model, Grid-Embedding based Multi-task Learning (GEML) which consists of two components focusing on spatial and temporal information respectively. The Grid-Embedding part is designed to model the spatial mobility patterns of passengers and neighboring relationships of different areas, the pre-weighted aggregator of which aims to sense the sparsity and range of data. The Multi-task Learning framework focuses on modeling temporal attributes and capturing several objectives of the ODMP problem. The evaluation of our model is conducted on real operational datasets from UCAR and Didi. The experimental results demonstrate the superiority of our GEML against the state-of-the-art approaches.
Yuandong Wang 0002, Hongzhi Yin, Hongxu Chen 0002, Tianyu Wo, Jie Xu 0007, Kai Zheng 0001
KDD4
2019 A Unified Framework with Multi-source Data for Predicting Passenger Demands of Ride Services
abstract
Ride-hailing applications have been offering convenient ride services for people in need. However, such applications still suffer from the issue of supply-demand disequilibrium, which is a typical problem for traditional taxi services. With effective predictions on passenger demands, we can alleviate the disequilibrium by pre-dispatching, dynamic pricing or avoiding dispatching cars to zero-demand areas. Existing studies of demand predictions mainly utilize limited data sources, trajectory data, or orders of ride services or both of them, which also lacks a multi-perspective consideration. In this article, we present a unified framework with a new combined model and a road-network-based spatial partition to leverage multi-source data and model the passenger demands from temporal, spatial, and zero-demand-area perspectives. In addition, our framework realizes offline training and online predicting, which can satisfy the real-time requirement more easily. We analyze and evaluate the performance of our combined model using the actual operational data from UCAR. The experimental results indicate that our model outperforms baselines on both Mean Absolute Error and Root Mean Square Error on average.
Yuandong Wang 0002, Xuelian Lin, Hua Wei 0001, Tianyu Wo, Jie Xu 0007
ACM Trans. Knowl. Discov. Data4
2018 ROSE: Cluster Resource Scheduling via Speculative Over-Subscription
abstract
A long-standing challenge in cluster scheduling is to achieve a high degree of utilization of heterogeneous resources in a cluster. In practice there exists a substantial disparity between perceived and actual resource utilization. A scheduler might regard a cluster as fully utilized if a large resource request queue is present, but the actual resource utilization of the cluster can be in fact very low. This disparity results in the formation of idle resources, leading to inefficient resource usage and incurring high operational costs and an inability to provision services. In this paper we present a new cluster scheduling system, ROSE, that is based on a multi-layered scheduling architecture with an ability to over-subscribe idle resources to accommodate unfulfilled resource requests. ROSE books idle resources in a speculative manner: instead of waiting for resource allocation to be confirmed by the centralized scheduler, it requests intelligently to launch tasks within machines according to their suitability to oversubscribe resources. A threshold control with timely task rescheduling ensures fully-utilized cluster resources without generating potential task stragglers. Experimental results show that ROSE can almost double the average CPU utilization, from 36.37% to 65.10%, compared with a centralized scheduling scheme, and reduce the workload makespan by 30.11%, with an 8.23% disk utilization improvement over other scheduling strategies.
Chunming Hu, Renyu Yang, Peter Garraghan, Tianyu Wo, Jie Xu 0007, Jianyong Zhu
ICDCS5
2018 A Context-Aware Evaluation Method of Driving Behavior
Yikai Zhai, Tianyu Wo, Xuelian Lin
PAKDD (1)2
2018 A Platform Solution of Data-Quality Improvement for Internet-of-Vehicle Services
abstract
Interconnection and intelligence have become the latest trends of the new generation of vehicle and transportation technologies. Applications built upon platforms of cloud-centered vehicle networking, i.e., Internet-of-Vehicles (IoVs), have been increasingly developed and deployed to provide data-centric services (e.g., driving assistance). Because these services are often safety critical, assuring service dependability has become an important requirement. In this paper, we propose DQI, a platform-level solution of Data-Quality Improvement designed to assure service dependability for Internet-of-Vehicle services. As an example, DQI is deployed in CarStream, an industrial system of big data processing designed for chauffeured car services. Via CarStream, over 30,000 vehicles are organized in a virtual vehicle network by sharing vehicle-status data in a near real-time manner. Such data often have low-quality issues and compromise the dependability of data-centric services. DQI includes techniques of data-quality improvement, including detecting outliers, extracting frequent patterns, and interpolating sequences. DQI enhances the dependability of data-centric services in IoVs by addressing the common data-quality requirements at the platform level. Upper-level services can benefit from DQI for data-quality improvement and reduce the complexity of service logic. We evaluate DQI by using a three-year dataset of vehicles and real applications deployed in CarStream. The result shows that compared with existing approaches, DQI can effectively restore missing data and correct anomalies with more than 30.0% improvement in precision. By studying multiple real applications, we also show that this data-quality improvement can indeed enhance the dependability of IoV services.
Tianyu Wo, Tao Xie 0001
PerCom2
2018 Preface
Tao Xie 0001, He Jiang 0001, Ge Li 0001, Tianyu Wo, Rahul Pandita, Chang Xu 0001, Lihua Xu
J. Comput. Sci. Technol.4
2017 One-Pass Error Bounded Trajectory Simplification
abstract
Nowadays, various sensors are collecting, storing and transmitting tremendous trajectory data, and it is known that raw trajectory data seriously wastes the storage, network band and computing resource. Line simplification (LS) algorithms are an effective approach to attacking this issue by compressing data points in a trajectory to a set of continuous line segments, and are commonly used in practice. However, existing LS algorithms are not sufficient for the needs of sensors in mobile devices. In this study, we first develop a one-pass error bounded trajectory simplification algorithm (OPERB), which scans each data point in a trajectory once and only once. We then propose an aggressive one-pass error bounded trajectory simplification algorithm (OPERB-A), which allows interpolating new data points into a trajectory under certain conditions. Finally, we experimentally verify that our approaches (OPERB and OPERB-A) are both efficient and effective, using four real-life trajectory datasets.
Xuelian Lin, Shuai Ma 0001, Tianyu Wo, Jinpeng Huai
Proc. VLDB Endow.4
2017 CarStream: An Industrial System of Big Data Processing for Internet-of-Vehicles
abstract
As the Internet-of-Vehicles (IoV) technology becomes an increasingly important trend for future transportation, designing large-scale IoV systems has become a critical task that aims to process big data uploaded by fleet vehicles and to provide data-driven services. The IoV data, especially high-frequency vehicle statuses (e.g., location, engine parameters), are characterized as large volume with a low density of value and low data quality. Such characteristics pose challenges for developing real-time applications based on such data. In this paper, we address the challenges in designing a scalable IoV system by describing CarStream, an industrial system of big data processing for chauffeured car services. Connected with over 30,000 vehicles, CarStream collects and processes multiple types of driving data including vehicle status, driver activity, and passenger-trip information. Multiple services are provided based on the collected data. CarStream has been deployed and maintained for three years in industrial usage, collecting over 40 terabytes of driving data. This paper shares our experiences on designing CarStream based on large-scale driving-data streams, and the lessons learned from the process of addressing the challenges in designing and maintaining CarStream.
Tianyu Wo, Xuelian Lin, Tao Xie 0001, Yaxiao Liu
Proc. VLDB Endow.2
2017 SafeDrive: Online Driving Anomaly Detection From Large-Scale Vehicle Data
abstract
Identifying driving anomalies is of great significance for improving driving safety. The development of the Internet-of-Vehicle (IoV) technology has made it feasible to acquire big data from multiple vehicle sensors, and such big data play a fundamental role in identifying driving anomalies. Existing approaches are mainly based on either rules or supervised learning. However, such approaches often require labeled data, which are typically not available in big data scenarios. In addition, because driving behaviors differ under vehicle statuses (e.g., speed and gear position), to precisely model driving behaviors needs to fuse multiple sources of sensor data. To address these issues, in this paper, we propose SafeDrive, an online and status-aware approach, which does not require labeled data. From a historical dataset, SafeDrive statistically offline derives a state graph (SG) as a behavior model. Then, SafeDrive splits the online data stream into segments and compares each segment with the SG. SafeDrive identifies a segment that significantly deviates from the SG as an anomaly. We evaluate SafeDrive on a cloud-based IoV platform with over 29 000 real connected vehicles. The evaluation results demonstrate that SafeDrive is capable of identifying a variety of driving anomalies effectively from a large-scale vehicle data stream with an overall accuracy of 93%; such identified driving anomalies can be used to timely alert drivers to correct their driving behaviors.
Chao Chen 0004, Tianyu Wo, Tao Xie 0001, Md. Zakirul Alam Bhuiyan, Xuelian Lin
IEEE Trans. Ind. Informatics3
2016 ZEST: A Hybrid Model on Predicting Passenger Demand for Chauffeured Car Service
abstract
Chauffeured car service based on mobile applications like Uber or Didi suffers from supply-demand disequilibrium, which can be alleviated by proper prediction on the distribution of passenger demand. In this paper, we propose a Zero-Grid Ensemble Spatio Temporal model (ZEST) to predict passenger demand with four predictors: a temporal predictor and a spatial predictor to model the influences of local and spatial factors separately, an ensemble predictor to combine the results of former two predictors comprehensively and a Zero-Grid predictor to predict zero demand areas specifically since any cruising within these areas costs extra waste on energy and time of driver. We demonstrate the performance of ZEST on actual operational data from ride-hailing applications with more than 6 million order records and 500 million GPS points. Experimental results indicate our model outperforms 5 other baseline models by over 10% both in MAE and sMAPE on the three-month datasets.
Hua Wei 0001, Yuandong Wang 0002, Tianyu Wo, Yaxiao Liu, Jie Xu 0007
CIKM3
2016 Analysis of students' behavior in the process of operating system experiments
abstract
Operating system (OS) experiments consolidate the understanding of the OS concepts and cultivate good engineering practices. Major challenges, however, including large class sizes, diverse software versions, and timely identification of difficulties from lab reports, hurt teaching quality. To address these, we designed an integrated environment to support OS experiments and automated the release and testing of lab code. The environment helped students to write, debug, and run code, and it collected their behavior data. These data included login hours and frequency, commands executed, files opened, and the code submission frequency. We found through analyzing the data a lack of preliminary knowledge in some students. Extra instructions and extended deadlines helped them master the subjects. For some students, the files read indicated the lack of attention to some most relevant system architecture code. This finding allowed us to provide targeted and specific help. Other data revealed an interesting “lab 2” phenomenon where the behavioral difference between the competent and average students was much more pronounced for lab 2 than that for lab 1. The data showed that login hours and frequency and log size correlated positively but submission frequency correlated negatively with grades. Analysis of students' behavior data allowed us to realize continuous improvements in the experiment process. 62% of the 152 students completed four labs and 46% all six, a significant improvement over the last semester.
Lei Wang 0126, Tianyu Wo
FIE3
2016 ScalaRDF: A Distributed, Elastic and Scalable In-Memory RDF Triple Store
abstract
The Resource Description Framework (RDF) andSPARQL query language are gaining increasing popularity andacceptance. The ever-increasing RDF data has reached a billionscale of triples, resulting in the proliferation of distributed RDFstore systems within the Semantic Web community. However, theelasticity and performance issues are still far from settled inface of data volume explosion and workload spike. In addition, providers face great pressures to provision uninterrupted reliablestorage service whilst reducing the operational costs due to avariety of system failures. Therefore, how to efficiently realizesystem fault tolerance remains an intractable problem. In this paper, we introduce ScalaRDF, a distributed and elastic in-memoryRDF triple store to provision a fault-tolerant and scalable RDFstore and query mechanism. Specifically, we describe a consistenthashing protocol that optimizes the RDF data placement, dataoperations (especially for online RDF triple update operations)and achieves an autonomously elastic data re-distribution in theevent of cluster node joining or departing, avoiding the holisticoscillation of data storage. In addition, the data store is ableto realize a rapid and transparent failover through replicationmechanism which stores in-memory data replica in the next hashhop. The experiments demonstrate that query time and updatetime are reduced by 87% and 90% respectively compared to otherapproaches. For an 18G source dataset, the data redistributiontakes at most 60 seconds when system scales out and at most 100seconds for recovery when nodes undergo crash-stop failures.
Chunming Hu, Xixu Wang, Renyu Yang, Tianyu Wo
ICPADS4
2016 Online Minimum Matching in Real-Time Spatial Data: Experiments and Analysis
abstract
Recently, with the development of mobile Internet and smartphones, the online minimum bipartite matching in real time spatial data (OMBM) problem becomes popular. Specifically, given a set of service providers with specific locations and a set of users who dynamically appear one by one, the OMBM problem is to find a maximum-cardinality matching with minimum total distance following that once a user appears, s/he must be immediately matched to an unmatched service provider, which cannot be revoked, before subsequent users arrive. To address this problem, existing studies mainly focus on analyzing the worst-case competitive ratios of the proposed online algorithms, but study on the performance of the algorithms in practice is absent. In this paper, we present a comprehensive experimental comparison of the representative algorithms of the OMBM problem. Particularly, we observe a surprising result that the simple and efficient greedy algorithm, which has been considered as the worst due to its exponential worst-case competitive ratio, is significantly more effective than other algorithms. We investigate the results and further show that the competitive ratio of the worst case of the greedy algorithm is actually just a constant, 3.195, in the average-case analysis. We try to clarify a 25-year misunderstanding towards the greedy algorithm and justify that the greedy algorithm is not bad at all. Finally, we provide a uniform implementation for all the algorithms of the OMBM problem and clarify their strengths and weaknesses, which can guide practitioners to select appropriate algorithms for various scenarios.
Yongxin Tong, Jieying She, Bolin Ding, Lei Chen 0002, Tianyu Wo, Ke Xu 0001
Proc. VLDB Endow.5
2016 MultiLanes: Providing Virtualized Storage for OS-Level Virtualization on Manycores
abstract
OS-level virtualization is often used for server consolidation in data centers because of its high efficiency. However, the sharing of storage stack services among the colocated containers incurs contention on shared kernel data structures and locks within I/O stack, leading to severe performance degradation on manycore platforms incorporating fast storage technologies (e.g., SSDs based on nonvolatile memories). This article presents MultiLanes, a virtualized storage system for OS-level virtualization on manycores. MultiLanes builds an isolated I/O stack on top of a virtualized storage device for each container to eliminate contention on kernel data structures and locks between them, thus scaling them to manycores. Meanwhile, we propose a set of techniques to tune the overhead induced by storage-device virtualization to be negligible, and to scale the virtualized devices to manycores on the host, which itself scales poorly. To reduce the contention within each single container, we further propose SFS, which runs multiple file-system instances through the proposed virtualized storage devices, distributes all files under each directory among the underlying file-system instances, then stacks a unified namespace on top of them. The evaluation of our prototype system built for Linux container (LXC) on a 32-core machine with both a RAM disk and a modern flash-based SSD demonstrates that MultiLanes scales much better than Linux in micro- and macro-benchmarks, bringing significant performance improvements, and that MultiLanes with SFS can further reduce the contention within each single container.
Junbin Kang, Chunming Hu, Tianyu Wo, Ye Zhai, Benlong Zhang, Jinpeng Huai
ACM Trans. Storage3
2015 SpanFS: A Scalable File System on Fast Storage Devices
Junbin Kang, Benlong Zhang, Tianyu Wo, Weiren Yu, Lian Du, Shuai Ma 0001, Jinpeng Huai
USENIX ATC3
2015 PARS: A Page-Aware Replication System for Efficiently Storing Virtual Machine Snapshots
abstract
Virtual machine (VM) snapshot enhances the system availability by saving the running state into stable storage during failure-free execution and rolling back to the snapshot point upon failures. Unfortunately, the snapshot state may be lost due to disk failures, so that the VM fails to be recovered. The popular distributed file systems employ replication technique to tolerate disk failures by placing redundant copies across disperse disks. However, unless user-specific personalization is provided, these systems consider the data in the file as of same importance and create identical copies of the entire file, leading to non-trivial additional storage overhead.
Lei Cui 0003, Tianyu Wo, Bo Li 0005, Jianxin Li 0002, Bin Shi 0003, Jinpeng Huai
VEE2
2014 MultiLanes: providing virtualized storage for OS-level virtualization on many cores
Junbin Kang, Benlong Zhang, Tianyu Wo, Chunming Hu, Jinpeng Huai
FAST3
2014 FENet: An SDN-based scheme for virtual network management
abstract
Virtual networking is vital to efficient resource management in Clouds, and it is in fact one of the main services provided by many Cloud Computing platforms. Virtual network management needs to meet specific requirements, including tenant isolation and adaption to virtual machines' lifecycle. Most of the existing schemes for virtual network management are based on the use of overlay networks in order to achieve a desirable degree of flexibility. However, these schemes suffer from a common limit, i.e. relatively high performance penalty due to a complicated forwarding process. We address this performance concern by developing a new management scheme, FENet, which makes use of Software-Defined Networks (SDN) to create virtual networks and manage them via the SDN controller programs. We present the design of an SDN controller, with the definition of flow entry rules based on the OpenFlow protocol and the specification of a routing algorithm. The results from our experimental evaluation show that our SDN-based prototype can control virtual network interconnections and tenant isolation appropriately. FENet achieves about 30% better network performance than the management scheme based on OpenVPN and lower latency in comparison with the traditional bridging scheme.
Tianyu Wo, Lei Cui 0003, Bin Shi 0003, Jie Xu 0007
ICPADS2
2014 Improving utilization through dynamic VM resource allocation in hybrid cloud environment
abstract
Virtualization is one of the most fascinating techniques because it can facilitate the infrastructure management and provide isolated execution for running workloads. Despite the benefits gained from virtualization and resource sharing, improved resource utilization is still far from settled due to the dynamic resource requirements and the widely-used over-provision strategy for guaranteed QoS. Additionally, with the emerging demands for big data analytic, how to effectively manage hybrid workloads such as traditional batch task and long-running virtual machine (VM) service needs to be dealt with. In this paper, we propose a system to combine long-running VM service with typical batch workload like MapReduce. The objectives are to improve the holistic cluster utilization through dynamic resource adjustment mechanism for VM without violating other batch workload executions. Furthermore, VM migration is utilized to ensure high availability and avoid potential performance degradation. The experimental results reveal that the dynamically allocated memory is close to the real usage with only 10% estimation margin, and the performance impact on VM and MapReduce jobs are both within 1%. Additionally, at most 50% increment of resource utilization could be achieved. We believe that these findings are in the right direction to solving workload consolidation issues in hybrid computing environments.
Yuda Wang, Renyu Yang, Tianyu Wo, Wenbo Jiang 0006, Chunming Hu
ICPADS3
2014 HotRestore: A Fast Restore System for Virtual Machine Cluster
Lei Cui 0003, Jianxin Li 0002, Tianyu Wo, Bo Li 0005, Renyu Yang, Yinglie Cao, Jinpeng Huai
LISA3
2014 Bounded Conjunctive Queries
abstract
A query Q is said to be effectively bounded if for all datasets D , there exists a subset D Q of D such that Q ( D ) = Q ( D Q ), and the size of DQ and time for fetching D Q are independent of the size of D . The need for studying such queries is evident, since it allows us to compute Q ( D ) by accessing a bounded dataset D Q , regardless of how big D is. This paper investigates effectively bounded conjunctive queries (SPC) under an access schema A , which specifies indices and cardinality constraints commonly used. We provide characterizations (sufficient and necessary conditions) for determining whether an SPC query Q is effectively bounded under A . We study several problems for deciding whether Q is bounded, and if not, for identifying a minimum set of parameters of Q to instantiate and make Q bounded. We show that these problems range from quadratic-time to NP-complete, and develop efficient (heuristic) algorithms for them. We also provide an algorithm that, given an effectively bounded SPC query Q and an access schema A , generates a query plan for evaluating Q by accessing a bounded amount of data in any (possibly big) dataset. We experimentally verify that our algorithms substantially reduce the cost of query evaluation.
Yang Cao 0012, Wenfei Fan, Tianyu Wo, Wenyuan Yu
Proc. VLDB Endow.3
2014 Strong simulation: Capturing topology in graph pattern matching
abstract
Graph pattern matching is finding all matches in a data graph for a given pattern graph and is often defined in terms of subgraph isomorphism, an NP -complete problem. To lower its complexity, various extensions of graph simulation have been considered instead. These extensions allow graph pattern matching to be conducted in cubic time. However, they fall short of capturing the topology of data graphs, that is, graphs may have a structure drastically different from pattern graphs they match, and the matches found are often too large to understand and analyze. To rectify these problems, this article proposes a notion of strong simulation , a revision of graph simulation for graph pattern matching. (1) We identify a set of criteria for preserving the topology of graphs matched. We show that strong simulation preserves the topology of data graphs and finds a bounded number of matches. (2) We show that strong simulation retains the same complexity as earlier extensions of graph simulation by providing a cubic-time algorithm for computing strong simulation. (3) We present the locality property of strong simulation which allows us to develop an effective distributed algorithm to conduct graph pattern matching on distributed graphs. (4) We experimentally verify the effectiveness and efficiency of these algorithms using both real-life and synthetic data.
Shuai Ma 0001, Yang Cao 0012, Wenfei Fan, Jinpeng Huai, Tianyu Wo
ACM Trans. Database Syst.5
2013 An Analysis of Performance Interference Effects on Energy-Efficiency of Virtualized Cloud Environments
abstract
Co-allocated workloads in a virtualized computing environment often have to compete for resources, thereby suffering from performance interference. While this phenomenon has a direct impact on the Quality of Service provided to customers, it also changes the patterns of resource utilization and reduces the amount of work per Watt consumed. Unfortunately, there has been only limited research into how performance interference affects energy-efficiency of servers in such environments. In reality, there is a highly dynamic and complicated correlation among resource utilization, performance interference and energy-efficiency. This paper presents a comprehensive analysis that quantifies the negative impact of performance interference on the energy-efficiency of virtualized servers. Our analysis methodology takes into account the heterogeneous workload characteristics identified from a real Cloud environment. In particular, we investigate the impact due to different workload type combinations and develop a method for approximating the levels of performance interference and energy-efficiency degradation. The proposed method is based on profiles of pair combinations of existing workload types and the patterns derived from the analysis. Our experimental results reveal a non-linear relationship between the increase in interference and the reduction in energy-efficiency as well as an average precision within +/-5% of error margin for the estimation of both parameters. These findings provide vital information for research into dynamic trade-offs between resource utilization, performance, and energy-efficiency of a data center.
Renyu Yang, Ismael Solís Moreno, Jie Xu 0007, Tianyu Wo
CloudCom (1)4
2013 Improved energy-efficiency in cloud datacenters with interference-aware virtual machine placement
abstract
Virtualization is one of the main technologies used for improving resource efficiency in datacenters; it allows the deployment of co-existing computing environments over the same hardware infrastructure. However, the co-existing of environments — along with management inefficiencies — often creates scenarios of high-competition for resources between running workloads, leading to performance degradation. This phenomenon is known as Performance Interference, and introduces a non-negligible overhead that affects both a datacenter's Quality of Service and its energy-efficiency. This paper introduces a novel approach to workload allocation that improves energy-efficiency in Cloud datacenters by taking into account their workload heterogeneity. We analyze the impact of performance interference on energy-efficiency using workload characteristics identified from a real Cloud environment, and develop a model that implements various decision-making techniques intelligently to select the best workload host according to its internal interference level. Our experimental results show reductions in interference by 27.5% and increased energy-efficiency up to 15% in contrast to current mechanisms for workload allocation.
Ismael Solís Moreno, Renyu Yang, Jie Xu 0007, Tianyu Wo
ISADS4
2013 VMScatter: migrate virtual machines to many hosts
abstract
Live virtual machine migration is a technique often used to migrate an entire OS with running applications in a non-disruptive fashion. Prior works concerned with one-to-one live migration with many techniques have been proposed such as pre-copy, post-copy and log/replay. In contrast, we propose VMScatter, a one-to-many migration method to migrate virtual machines from one to many other hosts simultaneously. First, by merging the identical pages within or across virtual machines, VMScatter multicasts only a single copy of these pages to associated target hosts for avoiding redundant transmission. This is impactful practically when the same OS and similar applications running in the virtual machines where there are plenty of identical pages. Second, we introduce a novel grouping algorithm to decide the placement of virtual machines, distinguished from the previous schedule algorithms which focus on the workload for load balance or power saving, we also focus on network traffic, which is a critical metric in data-intensive data centers. Third, we schedule the multicast sequence of packets to reduce the network overhead introduced by joining or quitting the multicast groups of target hosts. Compared to traditional live migration technique in QEMU/KVM, VMScatter reduces 74.2% of the total transferred data, 69.1% of the total migration time and achieves the network traffic reduction from 50.1% to 70.3%.
Lei Cui 0003, Jianxin Li 0002, Bo Li 0005, Jinpeng Huai, Chunming Hu, Tianyu Wo, Hussain Al-Aqrabi, Lu Liu 0001
VEE6
2013 CyberLiveApp: A secure sharing and migration approach for live virtual desktop applications in a cloud environment
Jianxin Li 0002, Lu Liu 0001, Tianyu Wo
Future Gener. Comput. Syst.4
2012 A Virtual File System for Streaming Loading of Virtual Software on Windows NT
Yabing Cui, Chunming Hu, Tianyu Wo
GPC3
2012 A Remote USB Architecture for Virtual Machine Oriented Device Sharing and Transparent Mgration
abstract
IaaS cloud environments which are driven by virtualization technologies enable managing applications and resources in a cost-efficient way and become the main operating environments of modern data centers. While, in such environments, how to use remote peripheral devices in a virtual machine (VM) becomes a key research problem, and the problem is aggravated when facing VM migration. The state of the art migration technologies lack for the consideration of peripheral devices, which can result in data loss. To address these two problems, the paper presents a remote USB architecture, which consists of two parts: the virtual machine oriented USB device sharing (VMDS) and transparent virtual USB device migration (TVDM). VMDS is used to share locally attached USB devices to remote virtual machines, and TVDM supports continuous accessing to remote devices during virtual machine live migration. A system based on Linux and KVM is implemented to demonstrate the ideas. Experimental evaluations illustrate the system's excellent usability and performance.
Ye Jiao, Tianyu Wo, Bo Li 0005
ICPADS2
2012 iROW: An Efficient Live Snapshot System for Virtual Machine Disk
abstract
The high-availiablity of mission-critical data and services hosted in a virtual machine (VM) is one of the top concerns in a cloud computing environment. The live disk snapshot is an emerging technology to save the whole state and the data of a VM at a specific point of time, and be used for quick disaster recovery. However, the existing VM disk snapshot systems suffer from long operation time and I/O performance degradation problems during snapshots creating and managing, and thereby affecting the performance of the VM and its services. To address such issues, we designed an efficient VM disk snapshot system, named iROW (improved Redirect-on-Write). In iROW, a bitmap based light-weight index scheme is adopted to replace the existing multi-level index tree structure to reduce query cost. Additionally, through a combination of Redirect-on-Write (ROW) and Copy-on-Demand (COD) schema to avoid extra copy operation on the first write after snapshot with Copy-on-Write (COW) schema, and the file fragmentation problem caused by ROW snapshot after long-term using. Finally, iROW gives a unified disk space allocation function by the host machine's file system. We have implemented iROW in qemu-kvm 0.12.5 and conducted some experiments. The implementation of iROW completely obey the interfaces of the block device driver in QEMU, so it is transparent to the upper system or applications and original disk image formats can be also supported. The experimental results show that iROW has obvious performance advantages in snapshot creating and management operations. Compared with the existing qcow2 disk image in KVM, when the VM disk size is 50GB, and the cluster size is 64KB (the default cluster size of qcow2), the snapshot creation and rollback time is only about 6% and 3% of original qcow2's. With the increasing of the VM disk size, iROW has more performance advantages on snapshot creation and rollback operations. In addition, the I/O performance of iROW is better than qcow2. When the cluster size is 64 KB, typically the iROW's performance loss is 10% less than qcow2's, and its first write performance after snapshot creation is about 250% of qcow2's.
Jianxin Li 0002, Lei Cui 0003, Bo Li 0005, Tianyu Wo
ICPADS5
2012 Distributed graph pattern matching
abstract
Graph simulation has been adopted for pattern matching to reduce the complexity and capture the need of novel applications. With the rapid development of the Web and social networks, data is typically distributed over multiple machines. Hence a natural question raised is how to evaluate graph simulation on distributed data. To our knowledge, no such distributed algorithms are in place yet. This paper settles this question by providing evaluation algorithms and optimizations for graph simulation in a distributed setting. (1) We study the impacts of components and data locality on the evaluation of graph simulation. (2) We give an analysis of a large class of distributed algorithms, captured by a message-passing model, for graph simulation. We also identify three complexity measures: visit times, makespan and data shipment, for analyzing the distributed algorithms, and show that these measures are essentially controversial with each other. (3) We propose distributed algorithms and optimization techniques that exploit the properties of graph simulation and the analyses of distributed algorithms. (4) We experimentally verify the effectiveness and efficiency of these algorithms, using both real-life and synthetic data.
Shuai Ma 0001, Yang Cao 0012, Jinpeng Huai, Tianyu Wo
WWW4
2012 Complexity of synthesis of composite service with correctness guarantee
Ting Deng, Jinpeng Huai, Tianyu Wo
Sci. China Inf. Sci.3
2012 CyberGuarder: A virtualization security assurance architecture for green cloud computing
Jianxin Li 0002, Bo Li 0005, Tianyu Wo, Chunming Hu, Jinpeng Huai, Lu Liu 0001
Future Gener. Comput. Syst.3
2011 Soft-Union: An Overlay Based Efficient Software P2P Distribution Scheme
abstract
Recent years, SaaS (Software as a Service) has become an innovative software delivery model. The traditional centralized software distribution model failed to meet the rapid deployment requirement of distribute software in large scale environment. In this paper, we present Soft-Union: a novel peer-to-peer based software distribution scheme. Soft-Union organizes the nodes sharing the same software into an overlay to reduce the query time. Furthermore, it incorporates a distributed mechanism for meta-information placement mechanism, and a search protocol to balance the query load and shorter the query time of the block meta-information. The experimental results validate that Soft-Union significantly reduces the block query time comparing with flooding and BT-like software distribution solutions.
Chunming Hu, Tianyu Wo, Jianxin Li 0002, Weiji Zeng
IEEE CLOUD3
2011 VirtualRank: A Prediction Based Load Balancing Technique in Virtual Computing Environment
abstract
This paper presents Virtual Rank, a load balancing technique which is on the basis of virtual machine migration. Virtual Rank proposes a solution that determines when to migrate virtual machines, and where to migrate. Most of the traditional load balancing techniques are based on threshold, whereas Virtual Rank predicts load tendency in the upcoming time slots. It ensures a small transient spike which does not trigger needless virtual machine(VM) migration. After triggering migration, the technique selects the potential migration target applying the Markov stochastic process. Finally the weighted probability method is applied to confirm the final migration target. It resolves the accumulation conflicts, as well as increases the stability. We implement our techniques in virtual computing environment iVic and conduct a detailed evaluation using a mix of CPU, network applications. We demonstrate that in different scale virtual network, Virtual Rank achieves better load balancing performance, compared with traditional methods.
Qingyi Gao, Ting Deng, Tianyu Wo
SERVICES4
2011 Capturing Topology in Graph Pattern Matching
abstract
Graph pattern matching is often defined in terms of subgraph isomorphism, an np-complete problem. To lower its complexity, various extensions of graph simulation have been considered instead. These extensions allow pattern matching to be conducted in cubic-time. However, they fall short of capturing the topology of data graphs, i.e. , graphs may have a structure drastically different from pattern graphs they match, and the matches found are often too large to understand and analyze. To rectify these problems, this paper proposes a notion of strong simulation , a revision of graph simulation, for graph pattern matching. (1) We identify a set of criteria for preserving the topology of graphs matched. We show that strong simulation preserves the topology of data graphs and finds a bounded number of matches. (2) We show that strong simulation retains the same complexity as earlier extensions of simulation, by providing a cubic-time algorithm for computing strong simulation. (3) We present the locality property of strong simulation, which allows us to effectively conduct pattern matching on distributed graphs. (4) We experimentally verify the effectiveness and efficiency of these algorithms, using real-life data and synthetic data.
Shuai Ma 0001, Yang Cao 0012, Wenfei Fan, Jinpeng Huai, Tianyu Wo
Proc. VLDB Endow.5
2010 CloudView: Describe and Maintain Resource View in Cloud
abstract
Resource view is the user defined table to provide specific view on resource status in cloud computing environment. It provides a convenient way to retrieve resource data for applications at infrastructure, platform and service layers. But the description and maintenance of these diverse resource views are inconvenient and dramatically difficult due to massive, heterogeneous and dynamic characteristics of the cloud resources involved. In this paper we present a resource view description scheme RQL and the corresponding system Cloud View to address these difficulties. RQL provides users a scheme to specify the data processing flow from resource raw data collected to resource view data objected. By constructing data processing a cyclic graph based on view definitions and using basic routines, view maintenance mechanism update user defined resource views automatically and periodically. Cloud View use a centralized scheduler to distribute maintenance jobs to a set of scalable worker nodes. It leverages distributed key-value database to store view data. Compared to related resource monitoring and discovering systems, Cloud View is flexible in application oriented view description and maintenance. Experiments show it updates typical user defined views with desired performance.
Dehui Zhou, Tianyu Wo, Junbin Kang
CloudCom3
2010 Resilient Virtual Network Service Provision in Network Virtualization Environments
abstract
Network Virtualization has recently emerged to provide scalable, customized and on-demand virtual network services over a shared substrate network. How to provide VN services with resiliency guarantees against network failures has become a critical issue, meanwhile the service resource usages should be minimized under the strict constraints such as link bandwidth capability and service resiliency guarantees etc. In this paper, we present a resource allocation algorithm to balance the tradeoff between service resource consumptions and service resiliency. By exploiting a heuristic VN mapping scheme and a restoration path selection scheme based on intelligent bandwidth sharing, the algorithm simultaneously makes cost-effective usage of network resources and protects VN services against network failures. We perform evaluations and find that the algorithm is near optimal in terms of network resource usage, especially the additional restoration bandwidth cost for resiliency protection.
Jianxin Li 0002, Tianyu Wo, Chunming Hu, Wantao Liu
ICPADS3
2010 A VMM-Based System Call Interposition Framework for Program Monitoring
abstract
System call interposition is a powerful method for regulating and monitoring program behavior. A wide variety of security tools have been developed which use this technique. However, traditional system call interposition techniques are vulnerable to kernel attacks and have some limitations on effectiveness and transparency. In this paper, we propose a novel approach named VSyscall, which leverages virtualization technology to enable system call interposition outside the operating system. A system call correlating method is proposed to identify the coherent system calls belonging to the same process from the system call sequence. We have developed a prototype of VSyscall and implemented it in two mainstream virtual machine monitors, Qemu and KVM, respectively. We also evaluate the effectiveness and performance overhead of our approach by comprehensive experiments. The results show that VSyscall achieves effectiveness with a small overhead, and our experiments with six real-world applications indicate its practicality.
Bo Li 0005, Jianxin Li 0002, Tianyu Wo, Chunming Hu
ICPADS3
2010 A Prefetching Framework for the Streaming Loading of Virtual Software
abstract
In recent years, the Software as a Service, largely enabled by the Internet, has become an innovative software delivery model. During the streaming execution of virtualization software, the execution will wait until the missing data was downloaded, which greatly influences the user experience. In this paper, we present a block-level prefetching framework for streaming delivery of software based on N-Gram prediction model and an incremental data mining algorithm. The prefetching framework uses the historical block access logs for data mining, then dynamically updates and polishes the prefetching rules. The experimental results show that this prefetching framework achieves a launch time reduced by 10% to 50%, as well as hit rate between 81% and 97%.
Junbin Kang, Chunming Hu, Tianyu Wo, Haibing Zheng, Bo Li 0005
ICPADS4
2009 An Efficient Resource Management System for On-Line Virtual Cluster Provision
abstract
As a prevalent paradigm for flexible, scalable and on-demand provisions of computing services, Cloud computing can be an alternative platform for scientific computing. In this paper, we propose an efficient resource management system for on-line virtual clusters provision, aiming to provide immediately-available virtual clusters for academic users. Particularly, we investigated two crucial problems: efficient VM image management and intelligent resource mapping, either of them has remarkable impact on the performance of the system. VM image management includes image preparation and local image management on physical resources. A resource mapping refers to a mapping from userpsilas resource constraints to specific physical resources. We explore how to simplify VM image management and reduce image preparation overhead by the multicast file transferring and image caching/reusing. Additionally, the Load-Aware Mapping, a novel resource mapping strategy, is proposed in order to further reduce deploying overhead and make efficient use of resources. The strategy takes account of both image cache and VM load distribution information. System evaluation is conducted through various real stress workloads, and results show that our approaches are effective comparing to other common solutions.
Tianyu Wo
IEEE CLOUD2
2009 EnaCloud: An Energy-Saving Application Live Placement Approach for Cloud Computing Environments
abstract
With the increasing prevalence of large scale cloud computing environments, how to place requested applications into available computing servers regarding to energy consumption has become an essential research problem, but existing application placement approaches are still not effective for live applications with dynamic characters. In this paper, we proposed a novel approach named EnaCloud, which enables application live placement dynamically with consideration of energy efficiency in a cloud platform. In EnaCloud, we use a Virtual Machine to encapsulate the application, which supports applications scheduling and live migration to minimize the number of running machines, so as to save energy. Specially, the application placement is abstracted as a bin packing problem, and an energy-aware heuristic algorithm is proposed to get an appropriate solution. In addition, an over-provision approach is presented to deal with the varying resource demands of applications. Our approach has been successfully implemented as useful components and fundamental services in the iVIC platform. Finally, we evaluate our approach by comprehensive experiments based on virtual machine monitor Xen and the results show that it is feasible.
Bo Li 0005, Jianxin Li 0002, Jinpeng Huai, Tianyu Wo, Qin Li 0014
IEEE CLOUD4
2009 Enhancing Reliability for Virtual Machines via Continual Migration
abstract
Our approach is to design and implement a continual migration strategy for virtual machines to achieve automatic failure recovery. By continually and transparently propagating virtual machine's state to a backup host via live migration techniques, trivial applications encapsulated in the virtual machine can be recovered from hardware failures with minimal downtime while no modifications are required. Deployment agility is considered so that initiating continual migration can be done with no time consuming preparations. Experimental results show that virtual machine in a continual migration system can be recovered in less than one second after a failure is detected, while performance impact to the protected virtual machine can be reduced to 30%.
Wenchao Cui, Dian-fu Ma, Tianyu Wo, Qin Li 0014
ICPADS3
2008 CROWN: A Service-Oriented Grid Middleware System: Experience and Applications
abstract
Grid computing has emerged as a new paradigm of distributed computing technology on large-scale resource sharing and coordinated problem solving. Based on a proposed Web service-based grid architecture, we have designed a service grid middleware system called CROWN which aims to promote the utilization of valuable resources and cooperation of researchers nationwide and world-wide. To address the issues of CROWN resource management, we proposed some key technologies including trustworthy remote and hot service deployment, overlay-based distributed resource organization, resource scheduling and load balance, and federation-based virtual organization management. A status of the wide-area CROWN testbed is also introduced in this paper. Three typical applications including AREM, MDP and gViz are deployed on the CROWN testbed. Experience of CROWN testbed deployment and application development shows that the middleware can support the typical scenarios such as computing-intensive applications and data-intensive applications etc.
Jinpeng Huai, Chunming Hu, Tianyu Wo, Jianxin Li 0002
ISORC3
2007 BBCLB: A Bulletin-Board based Cooperative Load Balance Strategy for Service Grid
abstract
Although many efforts have been put on the load balance in network and job scheduling systems, most of them, however, can not be applied in the service grid environment directly since they are often designed for a homogeneous system with limited scalability. It is still a challenge problem to balance the load among service grid nodes which are often highly dynamic, heterogeneous and linked by wide-area network. In this paper, we present a load balance strategy using several bulletin-boards as load intermediates among grid nodes. A modified thresholds based load transfer algorithm has been applied with a non-preemptive selection policy. Based on the strategy above, a load balance system is realized in CROWN, a service oriented grid middleware, and deployed in the CROWN testbed. The performance evaluations have shown that our strategy can effectively balance the load of service invocation, and improve the system throughput.
Tianyu Wo, Chunming Hu, Jinpeng Huai
CCGRID1
2006 CROWN: A service grid middleware with trust management mechanism
Jinpeng Huai, Chunming Hu, Jianxin Li 0002, Hailong Sun 0001, Tianyu Wo
Sci. China Ser. F Inf. Sci.5