Jie Xu 0007

dblp:37/5126-7 · DBLP profile ↗
← Back
122ranked-venue papers
13as first author
26since 2021 · last 2026
0000-0003-4102-233XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 36 · 10 first-author · 7 since 2021Software engineering, systems software and programming languages · 19 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 17 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 since 2021Artificial intelligence and machine learning · 8 · 2 since 2021Computer networks · 8 · 1 since 2021Security and privacy · 8 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 SynergyScale: Optimizing Offloading and Task Partitioning for Efficient Model Training
Jie Xu 0007, Zheng Wang 0001
IEEE Trans. Parallel Distributed Syst.2
2026 A Provably Cost-Efficient Approach to Deploying MoE Inference Models at the Network Edge
Chao Wang 0153, Danyang Zheng 0001, Huanlai Xing, Chen Yang 0043, Xiaojun Cao, Jie Xu 0007, Fei Teng 0001
IEEE Trans. Serv. Comput.6
2025 Tuning LLM-based Code Optimization via Meta-Prompting: An Industrial Perspective
abstract
There is a growing interest in leveraging multiple large language models (LLMs) for automated code optimization. However, industrial platforms deploying multiple LLMs face a critical challenge: prompts optimized for one LLM often fail with others, requiring expensive model-specific prompt engineering. This cross-model prompt engineering bottleneck severely limits the practical deployment of multi-LLM systems in production environments. We introduce Meta-Prompted Code Optimization (Mpco), a framework that automatically generates high-quality, task-specific prompts across diverse LLMs while maintaining industrial efficiency requirements. Mpco leverages meta-prompting to dynamically synthesize context-aware optimization prompts by integrating project metadata, task requirements, and LLM-specific contexts. It is an essential part of the ARTEMIS code optimization platform for automated validation and scaling.Our comprehensive evaluation on five real-world codebases with 366 hours of runtime benchmarking demonstrates Mpco’s effectiveness: it achieves overall performance improvements up to 19.06% with the best statistical rank across all systems compared to baseline methods. Analysis shows that 96% of the top-performing optimizations stem from meaningful edits. Through systematic ablation studies and meta-prompter sensitivity analysis, we identify that comprehensive context integration is essential for effective meta-prompting and that major LLMs can serve effectively as meta-prompters, providing actionable insights for industrial practitioners.
Jingzhi Gong, Rafail Giavrimis, Paul Brookes, Vardan Voskanyan 0001, Fan Wu 0009, Mari Ashiga, Matthew Truscott, Michail Basios, Leslie Kanthan, Jie Xu 0007, Zheng Wang 0001
ASE10
2025 Feature selection based on fuzzy joint entropy and feature interaction for label distribution learning
Dayong Deng, Jie Xu 0007, Zhixuan Deng, Jihong Wan, Deyou Xia, Zhenxin Cao, Tianrui Li 0001
Inf. Process. Manag.2
2025 Modeling Temporal Dependencies Within the Target for Long-Term Time Series Forecasting
abstract
Long-term time series forecasting (LTSF) is a critical task across diverse domains. Despite significant advancements in LTSF research, we identify a performance bottleneck in existing LTSF methods caused by the inadequate modeling of Temporal Dependencies within the Target (TDT). To address this issue, we propose a novel and generic temporal modeling framework, Temporal Dependency Alignment (TDAlign), that equips existing LTSF methods with TDT learning capabilities. TDAlign introduces two key innovations: 1) a loss function that aligns the change values between adjacent time steps in the predictions with those in the target, ensuring consistency with variation patterns, and 2) an adaptive loss balancing strategy that seamlessly integrates the new loss function with existing LTSF methods without introducing additional learnable parameters. As a plug-and-play framework, TDAlign enhances existing methods with minimal computational overhead, featuring only linear time complexity and constant space complexity relative to the prediction length. Extensive experiments on six strong LTSF baselines across seven real-world datasets demonstrate the effectiveness and flexibility of TDAlign. On average, TDAlign reduces baseline prediction errors by1.47%to9.19%and change value errors by4.57%to15.78%, highlighting its substantial performance improvements.
Minbo Ma, Ji Zhang 0012, Jie Xu 0007, Tianrui Li 0001
IEEE Trans. Knowl. Data Eng.5
2025 Looking Back on Recovery Blocks and Conversations
abstract
Our 1975 paper “System Structure for Software Fault Tolerance” introduced “Recovery Blocks” (a backward error recovery strategy for use in isolated processes), the “Domino Effect” (the problem that a single error could cause multiple interacting processes with uncoordinated error recovery strategies to lose all their recovery capability) and “Conversations” (an error recovery strategy for interacting processes motivated by the danger of the domino effect). This retrospective account describes how these ideas were developed by the Newcastle group and its collaborators, and what further research ensued. A tentative assessment is then provided of the impact of this research.
Brian Randell, Jie Xu 0007
IEEE Trans. Software Eng.2
2024 LDPRecover: Recovering Frequencies from Poisoning Attacks Against Local Differential Privacy
abstract
Local differential privacy (LDP), which enables an untrusted server to collect aggregated statistics from distributed users while protecting the privacy of those users, has been widely deployed in practice. However, LDP protocols for frequency estimation are vulnerable to poisoning attacks, in which an attacker can poison the aggregated frequencies by manipulating the data sent from malicious users. Therefore, it is an open challenge to recover the accurate aggregated frequencies from poisoned ones. In this work, we propose LDPRecover, a method that can recover accurate aggregated frequencies from poisoning attacks, even if the server does not learn the details of the attacks. In LDPRecover, we establish a genuine frequency estimator that theoretically guides the server to recover the frequencies aggregated from genuine users' data by eliminating the impact of malicious users' data in poisoned frequencies. Since the server has no idea of the attacks, we propose an adaptive attack to unify existing attacks and learn the statistics of the malicious data within this adaptive attack by exploiting the properties of LDP protocols. By taking the estimator and the learning statistics as constraints, we formulate the problem of recovering aggregated frequencies to approach the genuine ones as a constraint inference (CI) problem. Consequently, the server can obtain accurate aggregated frequencies by solving this problem optimally. Moreover, LDPRecover can serve as a frequency recovery paradigm that recovers more accurate aggregated frequencies by integrating attack details as new constraints in the CI problem. Our evaluation on two real-world datasets, three LDP protocols, and untargeted and targeted poisoning attacks shows that LDPRecover is both accurate and widely applicable against various poisoning attacks.
Xinyue Sun, Qingqing Ye 0001, Haibo Hu 0001, Jiawei Duan, Tianyu Wo, Jie Xu 0007, Renyu Yang
ICDE6
2024 MatchCom: Stable Matching-Based Software Services Composition in Cloud Computing Environments
Renyu Yang, Rajiv Ranjan 0001, Rami Bahsoon, Jie Xu 0007, Rajkumar Buyya
ICWE5
2024 Generating Location Traces With Semantic- Constrained Local Differential Privacy
abstract
Valuable information and knowledge can be learned from users’ location traces and support various location-based applications such as intelligent traffic control, incident response, and COVID-19 contact tracing. However, due to privacy concerns, no authority could simply collect users’ private location traces for mining or even publishing. To echo such concerns, local differential privacy (LDP) enables individual privacy by allowing each user to report a perturbed version of their data. Unfortunately, when applied to location traces, LDP cannot preserve the semantics in the context of location traces because it treats all locations (i.e., various points of interest) as equally sensitive. This results in a low utility of LDP mechanisms for collecting location traces. In this paper, we address the challenge of collecting and sharing location traces with valuable semantics while providing sufficient privacy protection for participating users. We first propose semantic-constrained local differential privacy (SLDP), a new privacy model to provide a provable mathematical privacy guarantee while preserving desirable semantics. Then, we design a location trace perturbation mechanism (LTPM) that users can use to perturb their traces in a way that satisfies SLDP. Finally, we propose a private location trace synthesis (PLTS) framework in which users use LTPM to perturb their traces before sending them to the collector, who aggregates the users’ perturbed data to generate location traces with valuable semantics. Extensive experiments on three real-world datasets demonstrate that our PLTS outperforms existing state-of-the-art methods by at least 21% in a range of real-world applications, such as spatial visiting queries and frequent pattern mining, under the same privacy leakage.
Xinyue Sun, Qingqing Ye 0001, Haibo Hu 0001, Jiawei Duan, Qiao Xue, Tianyu Wo, Weizhe Zhang, Jie Xu 0007
IEEE Trans. Inf. Forensics Secur.8
2024 PUTS: Privacy-Preserving and Utility-Enhancing Framework for Trajectory Synthesization
abstract
Vehicle trajectory data is essential for traffic management and location-based services. However, publishing real-life trajectory data has been challenging because vehicle trajectories contain users’ sensitive information. Differential privacy addresses such problems by publishing a synthetic version of the input dataset, but existing works always assume the real-world data is absolutely accurate. This assumption no longer holds in trajectory data because it typically contains errors due to inaccurate positioning services, which leads to poor performance of data synthesized by such trajectories. Even worse, existing works may generate unrealistic trajectories due to their coarse data synthesis methods, resulting in low practical utility or even inability to handle complex tasks. In this paper, we propose aPrivacy-preserving andUtility-enhancing framework forTrajectorySynthesization (PUTS). Our framework mitigates the impact of data errors in trajectories on differential privacy mechanisms, by exploiting map-matching techniques and real-world road network structure. InPUTS, a two-layer approach from path to trajectory synthesis is proposed to not only guarantee the reality of synthetic trajectories, but also scale upPUTSin real-world applications. Extensive experiments on real-world datasets show thatPUTSsignificantly outperforms existing methods in terms of utility in a range of real-world applications.
Xinyue Sun, Qingqing Ye 0001, Haibo Hu 0001, Jiawei Duan, Qiao Xue, Tianyu Wo, Jie Xu 0007
IEEE Trans. Knowl. Data Eng.7
2024 Hawk: Rapid Android Malware Detection Through Heterogeneous Graph Attention Networks
abstract
Android is undergoing unprecedented malicious threats daily, but the existing methods for malware detection often fail to cope with evolving camouflage in malware. To address this issue, we present Hawk, a new malware detection framework for evolutionary Android applications. We model Android entities and behavioral relationships as a heterogeneous information network (HIN), exploiting its rich semantic meta-structures for specifying implicit higher order relationships. An incremental learning model is created to handle the applications that manifest dynamically, without the need for reconstructing the whole HIN and the subsequent embedding model. The model can pinpoint rapidly the proximity between a new application and existing in-sample applications and aggregate their numerical embeddings under various semantics. Our experiments examine more than 80 860 malicious and 100 375 benign applications developed over a period of seven years, showing that Hawk achieves the highest detection accuracy against baselines and takes only 3.5 ms on average to detect an out-of-sample application, with the accelerated training time of 50× faster than the existing approach.
Yiming Hei, Renyu Yang, Hao Peng 0001, Jianwei Liu 0001, Hong Liu 0006, Jie Xu 0007, Lichao Sun 0001
IEEE Trans. Neural Networks Learn. Syst.8
2023 An Approach to Workload Generation for Cloud Benchmarking: a View from Alibaba Trace
abstract
Finding performance bottlenecks through bench-marking is one of the driving forces to improve the resource provision efficiency of cloud computing. Although existing benchmarks have been designed to improve the effectiveness in system performance evaluation, the following problems still exist in these benchmarks due to insufficient consideration of the characteristics of jobs in the production environment: (i) lacking of understanding for the details of workloads composition in the production environment, which reduces the authenticity of the job. (ii) the design of workloads submission patterns lacks quantization and reproducibility, which often relies on a random setting. In our benchmarking, multiple workloads are generated by analyzing and fine-grained matching the composition of workloads in the real production, and a design of workloads submission pattern based on LSTM time series prediction is proposed to simulate the real submission behavior. We finally demonstrate the effectiveness of our work by evaluating the impact of different workloads submission patterns on system performance evaluation.
Jianyong Zhu, Xiaoqiang Yu, Jie Xu 0007, Tianyu Wo
ISADS4
2023 Affinity-aware resource provisioning for long-running applications in shared clusters
abstract
Resource provisioning plays a pivotal role in determining the right amount of infrastructure resource to run applications and reduce the monetary cost. A significant portion of production clusters is now dedicated to long-running applications (LRAs), which are typically in the form of microservices and executed in the order of hours or even months. It is therefore practically important to plan ahead the placement of LRAs in a shared cluster for the minimized number of compute nodes required by them. Existing works on LRA scheduling are often application-agnostic, without particularly addressing the constraining requirements imposed by LRAs, such as co-location affinity constraints and time-varying resource requirements. In this paper, we present an affinity-aware resource provisioning approach for deploying large-scale LRAs in a shared cluster subject to multiple constraints, with the objective of minimizing the number of compute nodes in use. We investigate a broad range of solution algorithms which fall into three main categories: Application-Centric, Node-Centric, and Multi-Node approaches, and tune them for typical large-scale real-world scenarios. Experimental studies driven by the Alibaba Tianchi dataset show that our algorithms can achieve competitive scheduling effectiveness and running time, as compared with the heuristics used by the latest work including Medea and LraSched. Best results are obtained by the Application-Centric algorithms, if the algorithm's running time is of primary concern, and by Multi-Node algorithms, if the solution quality is of primary concern.
Clément Mommessin, Renyu Yang, Natalia V. Shakhlevich, Junqing Xiao, Jie Xu 0007
J. Parallel Distributed Comput.7
2023 Synthesizing Realistic Trajectory Data With Differential Privacy
abstract
Vehicle trajectory data is critical for traffic management and location-based services. However, the released trajectories raise serious privacy concerns because they contain sensitive information such as homes and workplaces. Based on differential privacy, this problem can be addressed by generating synthetic trajectories from the original sensitive data while guaranteeing personal privacy. Unfortunately, existing methods focus on synthesizing trajectory datasets that preserve summary-level statistics (e.g., the overall distribution of user movements), making these synthetic trajectories lose individual-level mobility patterns. As shown in our experiment, this results in the low performance of their synthetic datasets in real-world applications. To address these limitations, we propose a novel solution for Synthesizing Private and Realistic Trajectories, namely SPRT, whose key idea is to integrate the public geography structures of the target area into the process of private trajectory synthesis. This enables us to capture more accurate mobility patterns to synthesize realistic trajectories, which can preserve both summary-level statistics and individual-level mobility behaviors. Consequently, the synthetic trajectories generated by SPRT are more similar to real trajectories and therefore more practical. We evaluate the performance of SPRT in real-world applications by applying its synthetic data to a series of trajectory analytic tasks. The results demonstrate that our solution improves data utility by at least 37% over state-of-the-art approaches.
Xinyue Sun, Qingqing Ye 0001, Haibo Hu 0001, Yuandong Wang 0002, Kai Huang 0011, Tianyu Wo, Jie Xu 0007
IEEE Trans. Intell. Transp. Syst.7
2023 A Neural Expectation-Maximization Framework for Noisy Multi-Label Text Classification
abstract
Multi-label text classification (MLTC) has a wide range of real-world applications. Neural networks recently promoted the performance of MLTC models. Training these neural-network models relies on sufficient accurately labelled data. However, manually annotating large-scale multi-label text classification datasets is expensive and impractical for many applications. Weak supervision techniques have thus been developed to reduce the cost of annotating text corpus. However, these techniques introduce noisy labels into the training data and may degrade the model performance. This paper aims to deal with such noise-label problems in MLTC in both single-instance and multi-instance settings. We build a novel Neural Expectation-Maximization Framework (nEM) that combines neural networks with probabilistic modelling. The nEM framework produces text representations using neural-network text encoders and is optimized with the Expectation-Maximization algorithm. It naturally considers the noisy labels during learning by iteratively updating the model parameters and estimating the distribution of the ground-truth labels. We evaluate our nEM framework in multi-instance noisy MLTC on a benchmark relation extraction dataset constructed by distant supervision and in single-instance noisy MLTC on synthetic noisy datasets constructed by keywords supervision and label flipping. The experimental results demonstrate that nEM significantly improves upon baseline models in both single-instance and multi-instance noisy MLTC tasks. The experiment analysis suggests that our nEM framework efficiently reduces the noisy labels in MLTC datasets and significantly improves model performance.
Junfan Chen 0001, Richong Zhang, Jie Xu 0007, Chunming Hu, Yongyi Mao
IEEE Trans. Knowl. Data Eng.3
2023 Janus: Latency-Aware Traffic Scheduling for IoT Data Streaming in Edge Environments
abstract
This article focuses on a simple, yet fundamental question of distributed edge computing: “how to handle IoT traffic with different levels of sensitivity and criticality by satisfying the application-specific latency constraints?” This question arises in the practical deployment of edge computing, where user data can arrive at a much faster rate than that they can be processed by an edge node. Addressing this question is critical for meeting the latency requirement for latency-sensitive applications, but existing approaches are inadequate to the problem. We presentJanus, a multi-level traffic scheduling system for managing multiple data streams with various degrees of latency constraints. At the edge node level,Janususes multi-level queues to manage data streams with different latency constraints. It then allocates the output bandwidth of the edge node according to the requirements of applications in different priority queues, aiming to reduce the queuing and processing delay of latency-sensitive streams while maximizing the edge-node throughput. At the network level,Janusactively redirects incoming data streams to the less-loaded ones to achieve better network-wide load balance and improve the overall throughput. Experiments show thatJanusreduces the latency to only 16.6% of a non-priority based solution and improves the throughput by 1.7x of a state-of-the-art priority-aware data stream scheduling approach.
Zhenyu Wen, Renyu Yang, Bin Qian 0002, Yubo Xuan, Lingling Lu, Zheng Wang 0001, Hao Peng 0001, Jie Xu 0007, Albert Y. Zomaya, Rajiv Ranjan 0001
IEEE Trans. Serv. Comput.8
2023 RoSGAS: Adaptive Social Bot Detection with Reinforced Self-supervised GNN Architecture Search
abstract
Social bots are referred to as the automated accounts on social networks that make attempts to behave like humans. While Graph Neural Networks (GNNs) have been massively applied to the field of social bot detection, a huge amount of domain expertise and prior knowledge is heavily engaged in the state-of-the-art approaches to design a dedicated neural network architecture for a specific classification task. Involving oversized nodes and network layers in the model design, however, usually causes the over-smoothing problem and the lack of embedding discrimination. In this article, we propose RoSGAS , a novel R einf o rced and S elf-supervised G NN A rchitecture S earch framework to adaptively pinpoint the most suitable multi-hop neighborhood and the number of layers in the GNN architecture. More specifically, we consider the social bot detection problem as a user-centric subgraph embedding and classification task. We exploit the heterogeneous information network to present the user connectivity by leveraging account metadata, relationships, behavioral features, and content features. RoSGAS uses a multi-agent deep reinforcement learning (RL), 31 pages. mechanism for navigating the search of optimal neighborhood and network layers to learn individually the subgraph embedding for each target user. A nearest neighbor mechanism is developed for accelerating the RL training process, and RoSGAS can learn more discriminative subgraph embedding with the aid of self-supervised learning. Experiments on five Twitter datasets show that RoSGAS outperforms the state-of-the-art approaches in terms of accuracy, training efficiency, and stability and has better generalization when handling unseen samples.
Yingguang Yang, Renyu Yang, Zhiqin Yang, Yue Wang 0129, Jie Xu 0007, Haiyong Xie 0001
ACM Trans. Web7
2022 ContrastNet: A Contrastive Learning Framework for Few-Shot Text Classification
abstract
Few-shot text classification has recently been promoted by the meta-learning paradigm which aims to identify target classes with knowledge transferred from source classes with sets of small tasks named episodes. Despite their success, existing works building their meta-learner based on Prototypical Networks are unsatisfactory in learning discriminative text representations between similar classes, which may lead to contradictions during label prediction. In addition, the task-level and instance-level overfitting problems in few-shot text classification caused by a few training examples are not sufficiently tackled. In this work, we propose a contrastive learning framework named ContrastNet to tackle both discriminative representation and overfitting problems in few-shot text classification. ContrastNet learns to pull closer text representations belonging to the same class and push away text representations belonging to different classes, while simultaneously introducing unsupervised contrastive regularization at both task-level and instance-level to prevent overfitting. Experiments on 8 few-shot text classification datasets show that ContrastNet outperforms the current state-of-the-art models.
Junfan Chen 0001, Richong Zhang, Yongyi Mao, Jie Xu 0007
AAAI4
2022 AIACC-Training: Optimizing Distributed Deep Learning Training through Multi-streamed and Concurrent Gradient Communications
abstract
There is a growing interest in training deep neural networks (DNNs) in a GPU cloud environment. This is typically achieved by running parallel training workers on multiple GPUs across computing nodes. Under such a setup, the communication overhead is often responsible for long training time and poor scalability. This paper presents AIACC-Training, a unified communication framework designed for the distributed training of DNNs in a GPU cloud environment. AIACC-Training permits a training worker to participate in multiple gradient communication operations simultaneously to improve network bandwidth utilization and reduce communication latency. It employs auto-tuning techniques to dynamically determine the right communication parameters based on the input DNN workloads and the underlying network infrastructure. AIACC-Training has been deployed to production at Alibaba GPU Cloud with 3000+ GPUs executing AIACC-Training optimized code at any time. Experiments performed on representative DNN workloads show that AIACC-Training outperforms existing solutions, improving the training throughput and scalability by a large margin.
Lixiang Lin, Shenghao Qiu, Liang You, Long Xin, Jie Xu 0007, Zheng Wang 0001
ICDCS7
2022 STRONGHOLD: Fast and Affordable Billion-Scale Deep Learning Model Training
abstract
Deep neural networks (DNNs) with billion-scale parameters have demonstrated impressive performance in solving many tasks. Unfortunately, training a billion-scale DNN is out of the reach of many data scientists because it requires high-performance GPU servers that are too expensive to purchase and maintain. We present STRONGHOLD, a novel approach for enabling large DNN model training with no change to the user code. STRONGHOLD scales up the largest trainable model size by dynamically offloading data to the CPU RAM and enabling the use of secondary storage. It automatically determines the minimum amount of data to be kept in the GPU memory to minimize GPU memory usage. Compared to state-of-the-art offloading-based solutions, STRONGHOLD improves the trainable model size by 1.9x~6. Sx on a 32GB V100 GPU, with 1.2x~3.7x improvement on the training throughput. It has been deployed into production to successfully support large-scale DNN training.
Wei Wang 0225, Shenghao Qiu, Renyu Yang, Songfang Huang, Jie Xu 0007, Zheng Wang 0001
SC6
2022 Incentivizing Online Edge Caching via Auction - Based Subsidization
abstract
There exists a practical need for incentivizing content providers to cache contents at distributed network edges closer to users. However, this is a particularly challenging problem due to system environments that are uncertain, content placements that couple adjacent time slots, and economic properties that are desired but hard to ensure. In this paper, we present our design of an auction-based incentive mechanism for online edge caching. We formulate the long-term social cost minimization problem as a nonlinear mixed-integer program that addresses bid selections, user request dispatching, content placements, and payment determination in repetitive auctions. To solve this problem online, we devise a greedy approximation algorithm for solving each auction individually, and a lazy-replacement-based online algorithm that ties the series of auctions over time while dynamically pursuing the balance between downloading contents to new cache locations and keeping them at existing locations. We formally prove the approximation ratio for each single auction, the competitive ratio for the long-term social cost, as well as the truthfulness, the individual rationality, and the computational efficiency of our approach. Evaluations with real-world data have also validated and confirmed the practical superiority of our approach over multiple alternative algorithms.
Youmei Song, Lei Jiao 0002, Renyu Yang, Tianyu Wo, Jie Xu 0007
SECON5
2022 Passenger Mobility Prediction via Representation Learning for Dynamic Directed and Weighted Graphs
abstract
In recent years, ride-hailing services have been increasingly prevalent, as they provide huge convenience for passengers. As a fundamental problem, the timely prediction of passenger demands in different regions is vital for effective traffic flow control and route planning. As both spatial and temporal patterns are indispensable passenger demand prediction, relevant research has evolved from pure time series to graph-structured data for modeling historical passenger demand data, where a snapshot graph is constructed for each time slot by connecting region nodes via different relational edges (origin-destination relationship, geographical distance, etc.). Consequently, the spatiotemporal passenger demand records naturally carry dynamic patterns in the constructed graphs, where the edges also encode important information about the directions and volume (i.e., weights) of passenger demands between two connected regions. aspects in the graph-structure data. representation for DDW is the key to solve the prediction problem. However, existing graph-based solutions fail to simultaneously consider those three crucial aspects of dynamic, directed, and weighted graphs, leading to limited expressiveness when learning graph representations for passenger demand prediction. Therefore, we propose a novel spatiotemporal graph attention network, namely Gallat ( G raph prediction with all at tention) as a solution. In Gallat, by comprehensively incorporating those three intrinsic properties of dynamic directed and weighted graphs, we build three attention layers to fully capture the spatiotemporal dependencies among different regions across all historical time slots. Moreover, the model employs a subtask to conduct pretraining so that it can obtain accurate results more quickly. We evaluate the proposed model on real-world datasets, and our experimental results demonstrate that Gallat outperforms the state-of-the-art approaches.
Yuandong Wang 0002, Hongzhi Yin, Tong Chen 0005, Tianyu Wo, Jie Xu 0007
ACM Trans. Intell. Syst. Technol.7
2022 QoS-Aware Co-Scheduling for Distributed Long-Running Applications on Shared Clusters
abstract
To achieve a high degree of resource utilization, production clusters need to co-schedule diverse workloads – including both batch analytic jobs with short-lived tasks and long-running applications (LRAs) that execute for a long time frame from hours to months – onto the shared resources. Microservice architecture advances the manifestation of distributed LRAs (DLRAs), comprising multiple interconnected microservices that are executed in long-lived distributed containers and serve massive user requests. Detecting and mitigating QoS violation become even more intractable due to the network uncertainties and latency propagation across dependent microservices. However, current resource managers are only responsible for resource allocation among applications/jobs but agnostic to runtime QoS such as latency at application level. The state-of-the-art QoS-aware scheduling approaches are dedicated for monolithic applications, without considering the temporal-spatio performance variability across distributed microservices. In this paper, we presentToposch, a new scheduling and execution framework to prioritize the QoS of DLRAs whilst balancing the performance of batch jobs and maintaining high cluster utilization through harvesting idle resources.Toposchtracks footprints of every single request across microservices and uses critical path analysis, based on the end-to-end latency graph, to identify microservices that have high risk of QoS violation. Based on microservice and node level risk assessment, we intervene the batch scheduling by adaptively reducing the visible resources to batch tasks and thus delaying their execution to give way to DLRAs. We propose a prediction-based vertical resource auto-scaling mechanism, with the aid of resource-performance modeling and fine-grained resource inference and access control, for prompt recovery of QoS violation. A cost-effective task preemption is leveraged to ensure a low-cost task preemption and resource reclamation during the auto-scaling.Toposchis integrated with Apache YARN and experiments show thatToposchoutperforms other baselines in terms of performance guarantee of DLRAs, at an acceptable cost of batch job slowdown. The tail latency of DLRAs is merely 1.12x of the case of executing alone on average inToposchwith a 26% JCT increase of Spark analytic jobs.
Jianyong Zhu, Renyu Yang, Tianyu Wo, Chunming Hu, Hao Peng 0001, Junqing Xiao, Albert Y. Zomaya, Jie Xu 0007
IEEE Trans. Parallel Distributed Syst.9
2021 Perph: A Workload Co-location Agent with Online Performance Prediction and Resource Inference
abstract
Striking a balance between improved cluster utilization and guaranteed application QoS is a long-standing research problem in cluster resource management. The majority of current solutions require a large number of sandboxed experimentation for different workload combinations and leverage them to predict possible interference for incoming workloads. This results in non-negligible time complexity that severely restricts its applicability to complex workload co-locations. The nature of pure offline profiling may also lead to model aging problem that drastically degrades the model precision. In this paper, we present Perph, a runtime agent on a per node basis, which decouples ML-based performance prediction and resource inference from centralized scheduler. We exploit the sensitivity of long-running applications to multi-resources for establishing a relationship between resource allocation and consequential performance. We use Online Gradient Boost Regression Tree (OGBRT) to enable the continuous model evolution. Once performance degradation is detected, resource inference is conducted to work out a proper slice of resources that will be reallocated to recover the target performance. The integration with Node Manager (NM) of Apache YARN shows that the throughput of Kafka data-streaming application is 2.0x and 1.82x times that of isolation execution schemes in native YARN and pure cgroup cpu subsystem. In TPC-C benchmarking, the throughput can also be improved by 35% and 23% respectively against YARN native and cgroup cpu subsystem.
Jianyong Zhu, Renyu Yang, Chunming Hu, Tianyu Wo, Shiqing Xue, Jin Ouyang, Jie Xu 0007
CCGRID7
2021 Gallat: A Spatiotemporal Graph Attention Network for Passenger Demand Prediction
abstract
Online ride-hailing services have become an important component of urban transportation in recent years. As a fundamental research problem for such services, the timely prediction of passenger demands in different regions is vital for effective traffic flow control. As both spatial and temporal patterns are indispensable passenger demand prediction, relevant research has evolved from pure time series to graph-structured data for modelling historical passenger demand data, where a snapshot graph is constructed for each time slot by connecting region nodes via different relational edges. Consequently, the spatiotemporal passenger demand records naturally carry dynamic patterns in the constructed graphs, where the edges also encode important information about the directions and volume (i.e., weights) of passenger demands between two connected regions. However, existing graph-based solutions fail to simultaneously consider those three crucial aspects of dynamic, directed and weighted (DDW) graphs, leading to limited expressiveness when learning graph representations for passenger demand prediction. Therefore, we propose a novel spatiotemporal graph attention network, namely Gallat (Graph prediction with all attention) as a solution. In Gallat, by comprehensively incorporating those three intrinsic properties of DDW graphs, we build three attention layers to fully capture the spatiotemporal dependencies among different regions across all historical time slots. Our experimental results on real-world datasets demonstrate that Gallat outperforms the state-of-the-art approaches.
Yuandong Wang 0002, Hongzhi Yin, Tong Chen 0005, Tianyu Wo, Jie Xu 0007
ICDE7
2021 Joint optimization of cache placement and request routing in unreliable networks
Youmei Song, Tianyu Wo, Renyu Yang, Jie Xu 0007
J. Parallel Distributed Comput.5
2020 TOPOSCH: Latency-Aware Scheduling Based on Critical Path Analysis on Shared YARN Clusters
abstract
Balancing resource utilization and application QoS is a long-standing research topic in cluster resource management. Big data YARN clusters need to co-schedule diverse workloads on shared resources including batch processing jobs, streaming jobs, and other long-running applications such as web services, database services, etc. Current resource managers are only responsible for resource allocation among applications/jobs but completely unaware of runtime QoS requirements of interactive and latency-sensitive applications. Prior works to maximize the QoS of monolithic applications ignore inherent dependencies and temporal-spatio performance variability of components, characteristics of distributed applications primarily driven by microservices. In this paper, we present Toposch, a new resource management system to adaptively co-locate batch tasks and microservices by harvesting runtime latency. In particular, Toposch tracks full footprints of every request across microservices over time. A latency graph is periodically generated for identifying victim microservices through an end-to-end latency critical path analysis. We then exploit per-microservice and per-node risk assessment to gauge the visible resources to the capacity scheduler in YARN. Execution of batch tasks are adaptively throttled or delayed, thereby avoiding latency increase due to node over-saturation. TOPOSCH is integrated with YARN and experiments show that the latency of DLRAs can be reduced by up to 39.8% against the default capacity scheduling in YARN.
Chunming Hu, Jianyong Zhu, Renyu Yang, Hao Peng 0001, Tianyu Wo, Shiqing Xue, Xiaoqiang Yu, Jie Xu 0007, Rajiv Ranjan 0001
CLOUD8
2020 Parallel Interactive Networks for Multi-Domain Dialogue State Generation
abstract
The dependencies between system and user utterances in the same turn and across different turns are not fully considered in existing multidomain dialogue state tracking (MDST) models.In this study, we argue that the incorporation of these dependencies is crucial for the design of MDST and propose Parallel Interactive Networks (PIN) to model these dependencies.Specifically, we integrate an interactive encoder to jointly model the in-turn dependencies and cross-turn dependencies.The slot-level context is introduced to extract more expressive features for different slots.And a distributed copy mechanism is utilized to selectively copy words from historical system utterances or historical user utterances.Empirical studies demonstrated the superiority of the proposed PIN model.
Junfan Chen 0001, Richong Zhang, Yongyi Mao, Jie Xu 0007
EMNLP (1)4
2020 Improving Policy Generalization for Teacher-Student Reinforcement Learning
Xudong Gong, Hongda Jia, Xing Zhou 0004, Bo Ding 0001, Jie Xu 0007
KSEM (2)6
2020 Integrating clustering and regression for workload estimation in the cloud
abstract
Abstract Workload prediction has been widely researched in the literature. However, existing techniques are per‐job based and useful for service‐like tasks whose workloads exhibit seasonality and trend. But cloud jobs have many different workload patterns and some do not exhibit recurring workload patterns. We consider job‐pool‐based workload estimation, which analyzes the characteristics of existing tasks' workloads to estimate the currently running tasks' workload. First cluster existing tasks based on their workloads. For a new task J, collect the initial workload of J and determine which cluster J may belong to, then use the cluster's characteristics to estimate J′s workload. Based on the Google dataset, the algorithm is experimentally evaluated and its effectiveness is confirmed. However, the workload patterns of some tasks do have seasonality and trend, and conventional per‐job‐based regression methods may yield better workload prediction results. Also, in some cases, some new tasks may not follow the workload patterns of existing tasks in the pool. Thus, develop an integrated scheme which combines clustering and regression and utilize the best of them for workload prediction. Experimental study shows that the combined approach can further improve the accuracy of workload prediction.
Yongjia Yu, Vasu Jindal, I-Ling Yen, Farokh B. Bastani, Jie Xu 0007, Peter Garraghan
Concurr. Comput. Pract. Exp.5
2020 GA-Par: Dependable Microservice Orchestration Framework for Geo-Distributed Clouds
abstract
Recent advances in composing Cloud applications have been driven by deployments of inter-networking heterogeneous microservices across multiple Cloud datacenters. System dependability has been of the upmost importance and criticality to both service vendors and customers. Security, a measurable attribute, is increasingly regarded as the representative example of dependability. Literally, with the increment of microservice types and dynamicity, applications are exposed to aggravated internal security threats and externally environmental uncertainties. Existing work mainly focuses on the QoS-aware composition of native VM-based Cloud application components, while ignoring uncertainties and security risks among interactive and interdependent container-based microservices. Still, orchestrating a set of microservices across datacenters under those constraints remains computationally intractable. This paper describes a new dependable microservice orchestration framework GA-Par to effectively select and deploy microservices whilst reducing the discrepancy between user security requirements and actual service provision. We adopt a hybrid (both whitebox and blackbox based) approach to measure the satisfaction of security requirement and the environmental impact of network QoS on system dependability. Due to the exponential grow of solution space, we develop a parallel Genetic Algorithm framework based on Spark to accelerate the operations for calculating the optimal or near-optimal solution. Large-scale real world datasets are utilized to validate models and orchestration approach. Experiments show that our solution outperforms the greedy-based security aware method with 42.34 percent improvement. GA-Par is roughly 4× faster than a Hadoop-based genetic algorithm solver and the effectiveness can be constantly guaranteed under different application scales.
Zhenyu Wen, Tao Lin 0004, Renyu Yang, Shouling Ji, Rajiv Ranjan 0001, Alexander B. Romanovsky, Chang-Ting Lin, Jie Xu 0007
IEEE Trans. Parallel Distributed Syst.8
2020 Performance-Aware Speculative Resource Oversubscription for Large-Scale Clusters
abstract
It is a long-standing challenge to achieve a high degree of resource utilization in cluster scheduling. Resource oversubscription has become a common practice in improving resource utilization and cost reduction. However, current centralized approaches to oversubscription suffer from the issue with resource mismatch and fail to take into account other performance requirements, e.g., tail latency. In this article we present ROSE, a new resource management platform capable of conducting performance-aware resource oversubscription. ROSE allows latency-sensitive long-running applications (LRAs) to co-exist with computation-intensive batch jobs. Instead of waiting for resource allocation to be confirmed by the centralized scheduler, job managers in ROSE can independently request to launch speculative tasks within specific machines according to their suitability for oversubscription. Node agents of those machines can however, avoid any excessive resource oversubscription by means of a mechanism for admission control using multi-resource threshold control and performance-aware resource throttle. Experiments show that in case of mixed co-location of batch jobs and latency-sensitive LRAs, the CPU utilization and the disk utilization can reach 56.34 and 43.49 percent, respectively, but the 95th percentile of read latency in YCSB workloads only increases by 5.4 percent against the case of executing the LRAs alone.
Renyu Yang, Chunming Hu, Peter Garraghan, Tianyu Wo, Zhenyu Wen, Hao Peng 0001, Jie Xu 0007
IEEE Trans. Parallel Distributed Syst.8
2019 Perphon: a ML-based Agent for Workload Co-location via Performance Prediction and Resource Inference
abstract
Cluster administrators are facing great pressures to improve cluster utilization through workload co-location. Guaranteeing performance of long-running applications (LRAs), however, is far from settled as unpredictable interference across applications is catastrophic to QoS [2]. Current solutions such as [1] usually employ sandboxed and offline profiling for different workload combinations and leverage them to predict incoming interference. However, the time complexity restricts the applicability to complex co-locations. Hence, this issue entails a new framework to harness runtime performance and mitigate the time cost with machine intelligence: i) It is desirable to explore a quantitative relationship between allocated resource and consequent workload performance, not relying on analyzing interference derived from different workload combinations. The majority of works, however, depend on offline profiling and training which may lead to model aging problem. Moreover, multi-resource dimensions (e.g., LLC contention) that are not completely included by existing works but have impact on performance interference need to be considered [3]. ii) Workload co-location also necessitates fine-grained isolation and access control mechanism. Once performance degradation is detected, dynamic resource adjustment will be enforced and application will be assigned an access to specific slices of each resources. Inferring a "just enough" amount of resource adjustment ensures the application performance can be secured whilst improving cluster utilization.
Jianyong Zhu, Renyu Yang, Chunming Hu, Tianyu Wo, Shiqing Xue, Jin Ouyang, Jie Xu 0007
SoCC7
2019 Uncover the Ground-Truth Relations in Distant Supervision: A Neural Expectation-Maximization Framework
abstract
Junfan Chen, Richong Zhang, Yongyi Mao, Hongyu Guo, Jie Xu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Junfan Chen 0001, Richong Zhang, Yongyi Mao, Jie Xu 0007
EMNLP/IJCNLP (1)5
2019 Origin-Destination Matrix Prediction via Graph Convolution: a New Perspective of Passenger Demand Modeling
abstract
Ride-hailing applications are becoming more and more popular for providing drivers and passengers with convenient ride services, especially in metropolises like Beijing or New York. To obtain the passengers' mobility patterns, the online platforms of ride services need to predict the number of passenger demands from one region to another in advance. We formulate this problem as an Origin-Destination Matrix Prediction (ODMP) problem. Though this problem is essential to large-scale providers of ride services for helping them make decisions and some providers have already put it forward in public, existing studies have not solved this problem well. One of the main reasons is that the ODMP problem is more challenging than the common demand prediction. Besides the number of demands in a region, it also requires the model to predict the destinations of them. In addition, data sparsity is a severe issue. To solve the problem effectively, we propose a unified model, Grid-Embedding based Multi-task Learning (GEML) which consists of two components focusing on spatial and temporal information respectively. The Grid-Embedding part is designed to model the spatial mobility patterns of passengers and neighboring relationships of different areas, the pre-weighted aggregator of which aims to sense the sparsity and range of data. The Multi-task Learning framework focuses on modeling temporal attributes and capturing several objectives of the ODMP problem. The evaluation of our model is conducted on real operational datasets from UCAR and Didi. The experimental results demonstrate the superiority of our GEML against the state-of-the-art approaches.
Yuandong Wang 0002, Hongzhi Yin, Hongxu Chen 0002, Tianyu Wo, Jie Xu 0007, Kai Zheng 0001
KDD5
2019 Mitigating stragglers to avoid QoS violation for time-critical applications through dynamic server blacklisting
Xue Ouyang 0003, Jie Xu 0007
Future Gener. Comput. Syst.3
2019 A Unified Framework with Multi-source Data for Predicting Passenger Demands of Ride Services
abstract
Ride-hailing applications have been offering convenient ride services for people in need. However, such applications still suffer from the issue of supply-demand disequilibrium, which is a typical problem for traditional taxi services. With effective predictions on passenger demands, we can alleviate the disequilibrium by pre-dispatching, dynamic pricing or avoiding dispatching cars to zero-demand areas. Existing studies of demand predictions mainly utilize limited data sources, trajectory data, or orders of ride services or both of them, which also lacks a multi-perspective consideration. In this article, we present a unified framework with a new combined model and a road-network-based spatial partition to leverage multi-source data and model the passenger demands from temporal, spatial, and zero-demand-area perspectives. In addition, our framework realizes offline training and online predicting, which can satisfy the real-time requirement more easily. We analyze and evaluate the performance of our combined model using the actual operational data from UCAR. The experimental results indicate that our model outperforms baselines on both Mean Absolute Error and Root Mean Square Error on average.
Yuandong Wang 0002, Xuelian Lin, Hua Wei 0001, Tianyu Wo, Jie Xu 0007
ACM Trans. Knowl. Discov. Data7
2019 Straggler Root-Cause and Impact Analysis for Massive-scale Virtualized Cloud Datacenters
abstract
Increased complexity and scale of virtualized distributed systems has resulted in the manifestation of emergent phenomena substantially affecting overall system performance. This phenomena is known as “Long Tail”, whereby a small proportion of task stragglers significantly impede job completion time. While work focuses on straggler detection and mitigation, there is limited work that empirically studies straggler root-cause and quantifies its impact upon system operation. Such analysis is critical to ascertain in-depth knowledge of straggler occurrence for focusing developmental and research efforts towards solving the Long Tail challenge. This paper provides an empirical analysis of straggler root-cause within virtualized Cloud datacenters; we analyze two large-scale production systems to quantify the frequency and impact stragglers impose, and propose a method for conducting root-cause analysis. Results demonstrate approximately 5 percent of task stragglers impact 50 percent of total jobs for batch processes, and 53 percent of stragglers occur due to high server resource utilization. We leverage these findings to propose a method for extreme straggler detection through a combination of offline execution patterns modeling and online analytic agents to monitor tasks at runtime. Experiments show the approach is capable of detecting stragglers less than 11 percent into their execution lifecycle with 95 percent accuracy for short duration jobs.
Peter Garraghan, Xue Ouyang 0003, Renyu Yang, David McKee 0001, Jie Xu 0007
IEEE Trans. Serv. Comput.5
2018 ROSE: Cluster Resource Scheduling via Speculative Over-Subscription
abstract
A long-standing challenge in cluster scheduling is to achieve a high degree of utilization of heterogeneous resources in a cluster. In practice there exists a substantial disparity between perceived and actual resource utilization. A scheduler might regard a cluster as fully utilized if a large resource request queue is present, but the actual resource utilization of the cluster can be in fact very low. This disparity results in the formation of idle resources, leading to inefficient resource usage and incurring high operational costs and an inability to provision services. In this paper we present a new cluster scheduling system, ROSE, that is based on a multi-layered scheduling architecture with an ability to over-subscribe idle resources to accommodate unfulfilled resource requests. ROSE books idle resources in a speculative manner: instead of waiting for resource allocation to be confirmed by the centralized scheduler, it requests intelligently to launch tasks within machines according to their suitability to oversubscribe resources. A threshold control with timely task rescheduling ensures fully-utilized cluster resources without generating potential task stragglers. Experimental results show that ROSE can almost double the average CPU utilization, from 36.37% to 65.10%, compared with a centralized scheduling scheme, and reduce the workload makespan by 30.11%, with an 8.23% disk utilization improvement over other scheduling strategies.
Chunming Hu, Renyu Yang, Peter Garraghan, Tianyu Wo, Jie Xu 0007, Jianyong Zhu
ICDCS6
2018 Context-Aware Location Annotation on Mobility Records Through User Grouping
Hua Wei 0001, Xuelian Lin, Fei Wu 0007, Zhenhui Li, Kaiheng Chen, Yuandong Wang 0002, Jie Xu 0007
PAKDD (3)8
2018 Adaptive Speculation for Efficient Internetware Application Execution in Clouds
abstract
Modern Cloud computing systems are massive in scale, featuring environments that can execute highly dynamic Internetware applications with huge numbers of interacting tasks. This has led to a substantial challenge—the straggler problem, whereby a small subset of slow tasks significantly impede parallel job completion. This problem results in longer service responses, degraded system performance, and late timing failures that can easily threaten Quality of Service (QoS) compliance. Speculative execution (or speculation) is the prominent method deployed in Clouds to tolerate stragglers by creating task replicas at runtime. The method detects stragglers by specifying a predefined threshold to calculate the difference between individual tasks and the average task progression within a job. However, such a static threshold debilitates speculation effectiveness as it fails to capture the intrinsic diversity of timing constraints in Internetware applications, as well as dynamic environmental factors, such as resource utilization. By considering such characteristics, different levels of strictness for replica creation can be imposed to adaptively achieve specified levels of QoS for different applications. In this article, we present an algorithm to improve the execution efficiency of Internetware applications by dynamically calculating the straggler threshold, considering key parameters including job QoS timing constraints, task execution progress, and optimal system resource utilization. We implement this dynamic straggler threshold into the YARN architecture to evaluate it’s effectiveness against existing state-of-the-art solutions. Results demonstrate that the proposed approach is capable of reducing parallel job response time by up to 20% compared to the static threshold, as well as a higher speculation success rate, achieving up to 66.67% against 16.67% in comparison to the static method.
Xue Ouyang 0003, Peter Garraghan, Bernhard Primas, David McKee 0001, Paul Townend, Jie Xu 0007
ACM Trans. Internet Techn.6
2018 Holistic Virtual Machine Scheduling in Cloud Datacenters towards Minimizing Total Energy
abstract
Energy consumed by Cloud datacenters has dramatically increased, driven by rapid uptake of applications and services globally provisioned through virtualization. By applying energy-aware virtual machine scheduling, Cloud providers are able to achieve enhanced energy efficiency and reduced operation cost. Energy consumption of datacenters consists of computing energy and cooling energy. However, due to the complexity of energy and thermal modeling of realistic Cloud datacenter operation, traditional approaches are unable to provide a comprehensive in-depth solution for virtual machine scheduling which encompasses both computing and cooling energy. This paper addresses this challenge by presenting an elaborate thermal model that analyzes the temperature distribution of airflow and server CPU. We propose GRANITE - a holistic virtual machine scheduling algorithm capable of minimizing total datacenter energy consumption. The algorithm is evaluated against other existing workload scheduling algorithms MaxUtil, TASA, IQR and Random using real Cloud workload characteristics extracted from Google datacenter tracelog. Results demonstrate that GRANITE consumes 4.3-43.6 percent less total energy in comparison to the state-of-the-art, and reduces the probability of critical temperature violation by 99.2 with 0.17 percent SLA violation rate as the performance penalty.
Xiang Li 0017, Peter Garraghan, Xiaohong Jiang 0002, Zhaohui Wu 0001, Jie Xu 0007
IEEE Trans. Parallel Distributed Syst.5
2017 A Framework and Task Allocation Analysis for Infrastructure Independent Energy-Efficient Scheduling in Cloud Data Centers
abstract
Cloud computing represents a paradigm shift in provisioning on-demand computational resources underpinned by data center infrastructure, which now constitutes 1.5% of worldwide energy consumption. Such consumption is not merely limited to operating IT devices, but encompasses cooling systems representing 40% total data center energy usage. Given the substantive complexity and heterogeneity of data center operation spanning both computing and cooling components, obtaining analytical models for optimizing data center energy-efficiency is an inherently difficult challenge. Specifically, difficulties arise pertaining to the non-intuitive relationship between computing and cooling energy in the data center, computationally complex energy modeling, as well as cooling models restricted to a specific class of data center facility geometry - all of which arise from the interdisciplinary nature of this research domain.In this paper we propose a framework for energy-efficient scheduling to alleviate these challenges. It is applicable to any type of data center infrastructure and does not require complex modeling of energy.Instead, the concept of a target workload distribution is proposed. If the workload is assigned to nodes according to the target workload distribution, then the energy consumption is minimized. The exact target workload distribution is unknown, but an approximated distribution is delivered by the framework. The scheduling objective is to assign workload to nodes such that the workload distribution becomes as similar as possible to the target distribution in order to reduce energy consumption.Several mathematically sound algorithms have been designed to address this novel type of scheduling problem. Simulation results demonstrate that our algorithms reduce the relative deviation by at least 16.9% and the relative variance by at least 22.67% in comparison to (asymmetric) load balancing algorithms.
Bernhard Primas, Peter Garraghan, David McKee 0001, Jon Summers, Jie Xu 0007
CloudCom5
2017 ML-NA: A Machine Learning Based Node Performance Analyzer Utilizing Straggler Statistics
abstract
Current Cloud clusters often consist of heterogeneous machine nodes, which can trigger performance challenges such as the task straggler problem, whereby a small subset of parallel tasks running abnormally slower than the other sibling ones. The straggler problem leads to extended job response and deteriorates system throughput. Poor performance nodes are more likely to engender stragglers, and can undermine straggler mitigation effectiveness. For example, as the dominant mechanism for straggler alleviation, speculative execution functions by creating redundant task replicas on other machine nodes as soon as a straggler is detected. When speculative copies are assigned onto the poor performance nodes, it is hard for them to catch up with the stragglers compared to replicas run on fast nodes. And due to the fact that the performance heterogeneity is caused not only by static attribute variations such as physical capacity, but also dynamic characteristic uctuations such as contention level, analyzing node performance is important yet challenging. In this paper we develop ML-NA, a Machine Learning based Node performance Analyzer. By leveraging historical parallel tasks execution log data, ML-NA classies cluster nodes into different categories and predicts their performance in the near future as a scheduling guide to improve speculation effectiveness and minimize task straggler generation. We consider MapReduce as a representative framework to perform our analysis, and use the published OpenCloud trace as a case study to train and to evaluate our model. Results show that ML-NA can predict node performance categories with an average accuracy up to 92.86%.
Xue Ouyang 0003, Renyu Yang, Guogui Yang, Paul Townend, Jie Xu 0007
ICPADS6
2017 Practical Homomorphic Encryption Over the Integers for Secure Computation in the Cloud
James Dyer, Martin E. Dyer, Jie Xu 0007
IMACC3
2017 Mitigate data skew caused stragglers through ImKP partition in MapReduce
abstract
Speculative execution is the mechanism adopted by current MapReduce framework when dealing with the straggler problem, and it functions through creating redundant copies for identified stragglers. The result of the quicker task will be adopted to improve the overall job execution performance. Although proved to be effective for contention caused stragglers, speculative execution can easily meet its bottleneck when mitigating data skew caused stragglers due to its replication nature: the identical unbalanced input data will lead to a slow speculative task. The Map inputs are typically even in size according to the HDFS block configuration, therefore the skew caused stragglers happen mainly in the Reduce phase because of the unknown intermediate key distribution. In this paper, we focus on mitigating data skew caused Reduce stragglers, propose ImKP, an Intermediate Key Pre-processing framework that enables the even distributed partition for Reduce inputs. A group based ranking technique has been developed that dramatically decreases the pre-processing time, and ImKP manages to eliminate this timing overhead through parallelizing the pre-processing with the file uploading procedure (from local file system to HDFS). For jobs that take input directly from HDFS, ImKP minimizes the overhead by storing themapping result on every node within the cluster for reuse. Experiments are conducted on different datasets with various workloads. Results show that, compared to the popular hash partition, ImKP can dramatically decrease Reduce skew, achieving a 99.8% reduction in the coefficient of variation of the input sizes in average, and improve up to 29.37% job response performance.
Xue Ouyang 0003, Huan Zhou 0006, Stephen J. Clement, Paul Townend, Jie Xu 0007
IPCCC5
2017 Massive-Scale Automation in Cyber-Physical Systems: Vision & Challenges
abstract
The next era of computing is the evolution of the Internet of Things (IoT) and Smart Cities with development of the Internet of Simulation (IoS). The existing technologies of Cloud, Edge, and Fog computing as well as HPC being applied to the domains of Big Data and deep learning are not adequate to handle the scale and complexity of the systems required to facilitate a fully integrated and automated smart city. This integration of existing systems will create an explosion of data streams at a scale not yet experienced. The additional data can be combined with simulations as services (SIMaaS) to provide a shared model of reality across all integrated systems, things, devices, and individuals within the city. There are also numerous challenges in managing the security and safety of the integrated systems. This paper presents an overview of the existing state-of-the-art in automating, augmenting, and integrating systems across the domains of smart cities, autonomous vehicles, energy efficiency, smart manufacturing in Industry 4.0, and healthcare. Additionally the key challenges relating to Big Data, a model of reality, augmentation of systems, computation, and security are examined.
David McKee 0001, Stephen J. Clement, Jaber Almutairi, Jie Xu 0007
ISADS4
2017 Reliable Computing Service in Massive-Scale Systems through Rapid Low-Cost Failover
abstract
Large-scale distributed systems deployed as Cloud datacenters are capable of provisioning service to consumers with diverse business requirements. Providers face pressure to provision uninterrupted reliable services while reducing operational costs due to significant software and hardware failures. A widely adopted means to achieve such a goal is using redundant system components to implement user-transparent failover, yet its effectiveness must be balanced carefully without incurring heavy overhead when deployed—an important practical consideration for complex large-scale systems. Failover techniques developed for Cloud systems often suffer serious limitations, including mandatory restart leading to poor cost-effectiveness, as well as solely focusing on crash failures, omitting other important types, such as timing failures and simultaneous failures. This paper addresses these limitations by presenting a new approach to user-transparent failover for massive-scale systems. The approach uses soft-state inference to achieve rapid failure recovery and avoid unnecessary restart, with minimal system resource overhead. It also copes with different failures, including correlated and simultaneous events. The proposed approach was implemented, deployed and evaluated within Fuxi system, the underlying resource management system used within Alibaba Cloud. Results demonstrate that our approach tolerates complex failure scenarios while incurring at worst 228.5 microsecond instance overhead with 1.71 percent additional CPU usage.
Renyu Yang, Peter Garraghan, Yihui Feng, Jin Ouyang, Jie Xu 0007, Zhuo Zhang 0015
IEEE Trans. Serv. Comput.6
2016 Straggler Detection in Parallel Computing Systems through Dynamic Threshold Calculation
abstract
Cloud computing systems face the substantial challenge of the Long Tail problem: a small subset of straggling tasks significantly impede parallel jobs completion. This behavior results in longer service response times and degraded system utilization. Speculative execution, which create task replicas at runtime, is a typical method deployed in large-scale distributed systems to tolerate stragglers. This approach defines stragglers by specifying a static threshold value, which calculates the temporal difference between an individual task and the average task progression for a job. However, specifying static threshold debilitates speculation effectiveness as it fails to consider the intrinsic diversity of job timing constraints within modern day Cloud computing systems. Capturing such heterogeneity enables the ability to impose different levels of strictness for replica creation while achieving specified levels of QoS for different application types. Furthermore, a static threshold also fails to consider system environmental constraints in terms of replication overheads and optimal system resource usage. In this paper we present an algorithm for dynamically calculating a threshold value to identify task stragglers, considering key parameters including job QoS timing constraints, task execution characteristics, and optimal system resource utilization. We study and demonstrate the effectiveness of our algorithm through simulating a number of different operational scenarios based on real production cluster data against state-of-the-art solutions. Results demonstrate that our approach is capable of creating 58.62% less replicas under high resource utilization while reducing response time up to 17.86% for idle periods compared to a static threshold.
Xue Ouyang 0003, Peter Garraghan, David McKee 0001, Paul Townend, Jie Xu 0007
AINA5
2016 ZEST: A Hybrid Model on Predicting Passenger Demand for Chauffeured Car Service
abstract
Chauffeured car service based on mobile applications like Uber or Didi suffers from supply-demand disequilibrium, which can be alleviated by proper prediction on the distribution of passenger demand. In this paper, we propose a Zero-Grid Ensemble Spatio Temporal model (ZEST) to predict passenger demand with four predictors: a temporal predictor and a spatial predictor to model the influences of local and spatial factors separately, an ensemble predictor to combine the results of former two predictors comprehensively and a Zero-Grid predictor to predict zero demand areas specifically since any cruising within these areas costs extra waste on energy and time of driver. We demonstrate the performance of ZEST on actual operational data from ride-hailing applications with more than 6 million order records and 500 million GPS points. Experimental results indicate our model outperforms 5 other baseline models by over 10% both in MAE and sMAPE on the three-month datasets.
Hua Wei 0001, Yuandong Wang 0002, Tianyu Wo, Yaxiao Liu, Jie Xu 0007
CIKM5
2016 Holistic data centres: Next generation data and thermal energy infrastructures
abstract
Digital infrastructure is becoming more distributed and requiring more power for operation. At the same time, many countries are working to de-carbonise their energy, which will require electrical generation of heat for populated areas. What if this heat generation was combined with digital processing?
Paul Townend, Jie Xu 0007, Jon Summers, Daniel Ruprecht, Harvey Thompson
IPCCC2
2016 SEED: A Scalable Approach for Cyber-Physical System Simulation
abstract
Simulation is critical when studying real operational behavior of increasingly complex Cyber-Physical Systems, forecasting future behavior, and experimenting with hypothetical scenarios. A critical aspect of simulation is the ability to evaluate large-scale systems within a reasonable time frame while modeling complex interactions between millions of components. However, modern simulations face limitations in provisioning this functionality for CPSs in terms of balancing simulation complexity with performance, resulting in substantial operational costs required for completing simulation execution. Moreover, users are required to have expertise in modeling and configuring simulations to infrastructure which is time consuming. In this paper we present Simulation EnvironmEnt Distributor (SEED), a novel approach for simulating large-scale CPSs across a loosely-coupled distributed system requiring minimal user configuration. This is achieved through automated simulation partitioning and instantiation while enforcing tight event messaging across the system. SEED operates efficiently within both small and large-scale OTS hardware, agnostic of cluster heterogeneity and OS running, and is capable of simulating the full system and network stack of a CPS. Our approach is validated through experiments conducted in a cluster to simulate CPS operation. Results demonstrate that SEED is capable of simulating CPSs containing 2,000,000 tasks across 2,000 nodes with only 6.89× slow down relative to real time, and executes effectively across distributed infrastructure.
Peter Garraghan, David McKee 0001, Xue Ouyang 0003, David Webster, Jie Xu 0007
IEEE Trans. Serv. Comput.5
2015 Workload Estimation for Improving Resource Management Decisions in the Cloud
abstract
In cloud computing, good resource management can benefit both cloud users as well as cloud providers. Workload prediction is a crucial step towards achieving good resource management. While it is possible to estimate the workloads of long-running tasks based on the periodicity in their historical workloads, it is difficult to do so for tasks which do not have such recurring workload patterns. In this paper, we present an innovative clustering based resource estimation approach which groups tasks that have similar characteristics into the same cluster. The historical workload data for tasks in a cluster are used to estimate the resources needed by new tasks based on the cluster(s) to which they belong. In particular, for a new task T, we measure T's initial workload and predict to which cluster(s) it may belong. Then, the workload information of the cluster(s) is used to estimate the workload of T. The approach is experimentally evaluated using Google dataset, including resource usage data of over half a million tasks. We develop a workload model based on the dataset which is then used to estimate the workload patterns of several randomly selected tasks from the trace log. The results confirm the effectiveness of this cluster-based method for estimating the resources required by each task.
Jemishkumar Patel, Vasu Jindal, I-Ling Yen, Farokh B. Bastani, Jie Xu 0007, Peter Garraghan
ISADS5
2015 Timely Long Tail Identification through Agent Based Monitoring and Analytics
abstract
The increasing complexity and scale of distributed systems has resulted in the manifestation of emergent behavior which substantially affects overall system performance. A significant emergent property is that of the "Long Tail", whereby a small proportion of task stragglers significantly impact job execution completion times. To mitigate such behavior, straggling tasks occurring within the system need to be accurately identified in a timely manner. However, current approaches focus on mitigation rather than identification, which typically identify stragglers too late in the execution lifecycle. This paper presents a method and tool to identify Long Tail behavior within distributed systems in a timely manner, through a combination of online and offline analytics. This is achieved through historical analysis to profile and model task execution patterns, which then inform online analytic agents that monitor task execution at runtime. Furthermore, we provide an empirical analysis of two large-scale production Cloud data enters that demonstrate the challenge of data skew within modern distributed systems, this analysis shows that approximately 5% of task stragglers caused by data skew impact 50% of the total jobs for batch processes. Our results demonstrate that our approach is capable of identifying task stragglers less than 11% into their execution lifecycle with 98% accuracy, signifying significant improvement over current state-of-the-art practice and enables far more effective mitigation strategies in large-scale distributed systems worldwide.
Peter Garraghan, Xue Ouyang 0003, Paul Townend, Jie Xu 0007
ISORC4
2015 DIVIDER: Modelling and Evaluating Real-Time Service-Oriented Cyberphysical Co-Simulations
abstract
The ability to reliably distribute simulations across a distributed system and seamlessly integrate them as a workflow regardless of their level of abstraction is critical to improving the quality of product manufacturing. This paper presents the DIVIDER architecture for managing and maintaining real-time performance simulations integrated through SOAs. The described approach captures features present in complex workflow patterns such as asynchronous arbitrary cycles and estimates the worst case execution time in the context of the interfering execution environment.
David McKee 0001, David Webster, Jie Xu 0007, David Battersby
ISORC3
2014 FENet: An SDN-based scheme for virtual network management
abstract
Virtual networking is vital to efficient resource management in Clouds, and it is in fact one of the main services provided by many Cloud Computing platforms. Virtual network management needs to meet specific requirements, including tenant isolation and adaption to virtual machines' lifecycle. Most of the existing schemes for virtual network management are based on the use of overlay networks in order to achieve a desirable degree of flexibility. However, these schemes suffer from a common limit, i.e. relatively high performance penalty due to a complicated forwarding process. We address this performance concern by developing a new management scheme, FENet, which makes use of Software-Defined Networks (SDN) to create virtual networks and manage them via the SDN controller programs. We present the design of an SDN controller, with the definition of flow entry rules based on the OpenFlow protocol and the specification of a routing algorithm. The results from our experimental evaluation show that our SDN-based prototype can control virtual network interconnections and tenant isolation appropriately. FENet achieves about 30% better network performance than the management scheme based on OpenVPN and lower latency in comparison with the traditional bridging scheme.
Tianyu Wo, Lei Cui 0003, Bin Shi 0003, Jie Xu 0007
ICPADS5
2014 Fault-Tolerant Dynamic Deduplication for Utility Computing
abstract
Utility computing is an increasingly important paradigm, whereby computing resources are provided on-demand as utilities. An important component of utility computing is storage, data volumes are growing rapidly, and mechanisms to mitigate this growth need to be developed. Data deduplication is a promising technique for drastically reducing the amount of data stored in such system systems, however, current approachs are static in nature, using an amount of redundancy fixed at design time. This is inappropriate for truly dynamic modern systems. We propose a real-time adaptive deduplication system for Cloud and Utility computing that monitors in real-time for changing system, user, and environmental behaviour in order to fulfill a balance between changing storage efficiency, performance, and fault tolerance requirements. We evaluate our system through simulation, with experimental results showing that our system is both efficient and sclable. We also perform experimentation to evaluate the fault tolerance of the system by measuring Mean Time to Repair (MTTR), and using these values to calculate availability of the system. The results show that higher replication levels result in higher system availability, however, the number of files in the system also effects recovery time. We show that the tradeoff between replication levels and recovery time when the system overloads needs further investigation.
Waraporn Leesakul, Paul Townend, Peter Garraghan, Jie Xu 0007
ISORC4
2014 M-VCR: Multi-View Consensus Recognition for Real-Time Experimentation
abstract
A major application area in the computer vision domain is gesture recognition, requiring real-time image classification to respond to human interactions. However, current state-of-the-art high-quality algorithms for image classification do not meet many dynamic real-time requirements. This paper presents the development of M-VCR - a novel approach for improving the reliability of real-time image classification. M-VCR increases the quality of classifications under real-time constraints through the adoption of fast classification algorithms, although these algorithms individually produce lower quality results, utilisation under a 'consensus' approach can achieve results equivalent to those of much higher-quality algorithms. The proposed approach also allows for different algorithms to be utilised in parallel, building on the fault tolerance technique of N-versioning. A significant improvement in image classification is experimentally demonstrated for both the SURF and MSER feature detectors through our integration consensus approach. This improvement is delivered entirely through the integration method without requiring modification of the source algorithms being used.
David McKee 0001, Paul Townend, David Webster, Jie Xu 0007
ISORC4
2014 Towards a Virtual Integration Design and Analysis Enviroment for Automotive Engineering
abstract
As the automotive industry moves towards reduced physical prototyping it is becoming more dependent on distributed simulations. However, the current technologies do not fully enable real-time distributed simulations which involve both virtual and physical components. This paper considers the current approaches to real-time distributed simulation and proposes the use of service-orientation. The highlights and current shortfalls of current research in real-time service orientation are then identified. Finally a key area of research is focused upon requiring a fundamental change of the understanding of quality of service and capability that is necessary to enable dependable real-time service orientated architectures.
David McKee 0001, David Webster, Paul Townend, Jie Xu 0007, David Battersby
ISORC4
2014 A Broker-Based Self-organizing Mechanism for Cloud-Market
Jie Xu 0007, Jian Cao 0001
NPC1
2014 Fuxi: a Fault-Tolerant Resource Management and Job Scheduling System at Internet Scale
abstract
Scalability and fault-tolerance are two fundamental challenges for all distributed computing at Internet scale. Despite many recent advances from both academia and industry, these two problems are still far from settled. In this paper, we present Fuxi, a resource management and job scheduling system that is capable of handling the kind of workload at Alibaba where hundreds of terabytes of data are generated and analyzed everyday to help optimize the company's business operations and user experiences. We employ several novel techniques to enable Fuxi to perform efficient scheduling of hundreds of thousands of concurrent tasks over large clusters with thousands of nodes: 1) an incremental resource management protocol that supports multi-dimensional resource allocation and data locality; 2) user-transparent failure recovery where failures of any Fuxi components will not impact the execution of user jobs; and 3) an effective detection mechanism and a multi-level blacklisting scheme that prevents them from affecting job execution. Our evaluation results demonstrate that 95% and 91% scheduled CPU/memory utilization can be fulfilled under synthetic workloads, and Fuxi is capable of achieving 2.36T-B/minute throughput in GraySort. Additionally, the same Fuxi job only experiences approximately 16% slowdown under a 5% fault-injection rate. The slowdown only grows to 20% when we double the fault-injection rate to 10%. Fuxi has been deployed in our production environment since 2009, and it now manages hundreds of thousands of server nodes.
Zhuo Zhang 0015, Yangyu Tao, Renyu Yang, Jie Xu 0007
Proc. VLDB Endow.6
2014 Analysis, Modeling and Simulation of Workload Patterns in a Large-Scale Utility Cloud
abstract
Understanding the characteristics and patterns of workloads within a Cloud computing environment is critical in order to improve resource management and operational conditions while Quality of Service (QoS) guarantees are maintained. Simulation models based on realistic parameters are also urgently needed for investigating the impact of these workload characteristics on new system designs and operation policies. Unfortunately there is a lack of analyses to support the development of workload models that capture the inherent diversity of users and tasks, largely due to the limited availability of Cloud tracelogs as well as the complexity in analyzing such systems. In this paper we present a comprehensive analysis of the workload characteristics derived from a production Cloud data center that features over 900 users submitting approximately 25 million tasks over a time period of a month. Our analysis focuses on exposing and quantifying the diversity of behavioral patterns for users and tasks, as well as identifying model parameters and their values for the simulation of the workload created by such components. Our derived model is implemented by extending the capabilities of the CloudSim framework and is further validated through empirical comparison and statistical hypothesis tests. We illustrate several examples of this work's practical applicability in the domain of resource management and energy-efficiency.
Ismael Solís Moreno, Peter Garraghan, Paul Townend, Jie Xu 0007
IEEE Trans. Cloud Comput.4
2013 An Analysis of Performance Interference Effects on Energy-Efficiency of Virtualized Cloud Environments
abstract
Co-allocated workloads in a virtualized computing environment often have to compete for resources, thereby suffering from performance interference. While this phenomenon has a direct impact on the Quality of Service provided to customers, it also changes the patterns of resource utilization and reduces the amount of work per Watt consumed. Unfortunately, there has been only limited research into how performance interference affects energy-efficiency of servers in such environments. In reality, there is a highly dynamic and complicated correlation among resource utilization, performance interference and energy-efficiency. This paper presents a comprehensive analysis that quantifies the negative impact of performance interference on the energy-efficiency of virtualized servers. Our analysis methodology takes into account the heterogeneous workload characteristics identified from a real Cloud environment. In particular, we investigate the impact due to different workload type combinations and develop a method for approximating the levels of performance interference and energy-efficiency degradation. The proposed method is based on profiles of pair combinations of existing workload types and the patterns derived from the analysis. Our experimental results reveal a non-linear relationship between the increase in interference and the reduction in energy-efficiency as well as an average precision within +/-5% of error margin for the estimation of both parameters. These findings provide vital information for research into dynamic trade-offs between resource utilization, performance, and energy-efficiency of a data center.
Renyu Yang, Ismael Solís Moreno, Jie Xu 0007, Tianyu Wo
CloudCom (1)3
2013 An Analysis of the Server Characteristics and Resource Utilization in Google Cloud
abstract
Understanding the resource utilization and server characteristics of large-scale systems is crucial if service providers are to optimize their operations whilst maintaining Quality of Service. For large-scale data enters, identifying the characteristics of resource demand and the current availability of such resources, allows system managers to design and deploy mechanisms to improve data enter utilization and meet Service Level Agreements with their customers, as well as facilitating business expansion. In this paper, we present a large-scale analysis of server resource utilization and a characterization of a production Cloud data enter using the most recent data enter trace logs made available by Google. We present their statistical properties, and a comprehensive coarse-grain analysis of the data, including submission rates, server classification, and server resource utilization. Additionally, we perform a fine-grained analysis to quantify the resource utilization of servers wasted due to the early termination of tasks. Our results show that data enter resource utilization remains relatively stable at between 40 - 60%, that the degree of correlation between server utilization and Cloud workload environment varies by server architecture, and that the amount of resource utilization wasted varies between 4.53 - 14.22% for different server architectures. This provides invaluable real-world empirical data for Cloud researchers in many subject areas.
Peter Garraghan, Paul Townend, Jie Xu 0007
IC2E3
2013 Improved energy-efficiency in cloud datacenters with interference-aware virtual machine placement
abstract
Virtualization is one of the main technologies used for improving resource efficiency in datacenters; it allows the deployment of co-existing computing environments over the same hardware infrastructure. However, the co-existing of environments — along with management inefficiencies — often creates scenarios of high-competition for resources between running workloads, leading to performance degradation. This phenomenon is known as Performance Interference, and introduces a non-negligible overhead that affects both a datacenter's Quality of Service and its energy-efficiency. This paper introduces a novel approach to workload allocation that improves energy-efficiency in Cloud datacenters by taking into account their workload heterogeneity. We analyze the impact of performance interference on energy-efficiency using workload characteristics identified from a real Cloud environment, and develop a model that implements various decision-making techniques intelligently to select the best workload host according to its internal interference level. Our experimental results show reductions in interference by 27.5% and increased energy-efficiency up to 15% in contrast to current mechanisms for workload allocation.
Ismael Solís Moreno, Renyu Yang, Jie Xu 0007, Tianyu Wo
ISADS3
2013 An evaluation framework for assessing the dependability of Dynamic Binding in Service-Oriented Computing
abstract
Service-Oriented Computing (SOC) provides a flexible framework in which applications may be built up from services, often distributed across a network. One of the promises of SOC is that of Dynamic Binding where abstract consumer requests are bound to concrete service instances at runtime, thereby offering a high level of flexibility and adaptability. Existing research has so far focused mostly on the design and implementation of dynamic binding operations and there is little research into a comprehensive evaluation of dynamic binding systems, especially in terms of system failure and dependability. In this paper, we present a novel, extensible evaluation framework that allows for the testing and assessment of a Dynamic Binding System (DBS). Based on a fault model specially built for DBS's, we are able to insert selectively the types of fault that would affect a DBS and observe its behavior. By treating the DBS as a black box and distributing the components of the evaluation framework we are not restricted to the implementing technologies of the DBS, nor do we need to be co-located in the same environment as the DBS under test. We present the results of a series of experiments, with a focus on the interactions between a real-life DBS and the services it employs. The results on the NECTISE Software Demonstrator (NSD) system show that our proposed method and testing framework is able to trigger abnormal behavior of the NSD due to interaction faults and generate important information for improving both dependability and performance of the system under test.
Anthony Sargeant, Paul Townend, Jie Xu 0007, Karim Djemame
ISORC3
2013 A novel intrusion severity analysis approach for Clouds
Junaid Arshad, Paul Townend, Jie Xu 0007
Future Gener. Comput. Syst.3
2013 Clouds and service-oriented architectures
Lu Liu 0001, Jie Xu 0007
Future Gener. Comput. Syst.2
2013 Internet-based Virtual Computing Environment: Beyond the data center as a computer
Xicheng Lu, Huaimin Wang 0001, Ji Wang 0001, Jie Xu 0007, Dongsheng Li 0001
Future Gener. Comput. Syst.4
2012 Neural Network-Based Overallocation for Improved Energy-Efficiency in Real-Time Cloud Environments
abstract
This paper introduces a dynamic resource provisioning mechanism for over allocating the capacity of Cloud data centers based on customer resource utilization patterns. The proposed mechanism reduces the impact on Real-Time constraints while improvements on the overall energy-efficiency are sought. The main idea is to exploit the resource utilization patterns of each customer for smartly under allocating resources to the requested Virtual Machines. This reduces the waste produced by frequent overestimations and increases the data center availability. Consequently, it creates the opportunity to host additional Virtual Machines in the same computing infrastructure improving its energy-efficiency. In order to mitigate the negative effect on deadlines, the proposed over allocation service implements a multiplayer Neural Network to anticipate the resource usage patterns based on historical data. Additionally, a compensation mechanism for adjusting the resource allocation in cases of unexpected higher demand is also described. The experiments contrast the proposed approach against traditional "Dynamic Resource Resizing" energy-aware mechanisms and also to our previous work that implements Low-Pass-Filter as predictor. Results demonstrate meaningful improvements in energy-efficiency while time constraints are slightly affected.
Ismael Solís Moreno, Jie Xu 0007
ISORC2
2012 Interface Refactoring in Performance-Constrained Web Services
abstract
This paper presents the development of REF-WS an approach to enable a Web Service provider to reliably evolve their service through the application of refactoring transformations. REF-WS is intended to aid service providers, particularly in a reliability and performance constrained domain as it permits upgraded 'non-backwards compatible' services to be deployed into a performance constrained network where existing consumers depend on an older version of the service interface. In order for this to be successful, the refactoring and message mediation needs to occur without affecting functional compatibility with the services' consumers, and must operate within the performance overhead expected of the original service, introducing as little latency as possible. Furthermore, compared to a manually programmed solution, the presented approach enables the service developer to apply and parameterize refactorings with a level of confidence that they will not produce an invalid or 'corrupt' transformation of messages. This is achieved through the use of preconditions for the defined refactorings.
David Webster, Paul Townend, Jie Xu 0007
ISORC3
2012 Dynamic Authentication for Cross-Realm SOA-Based Business Processes
abstract
Modern distributed applications are embedding an increasing degree of dynamism, from dynamic supply-chain management, enterprise federations, and virtual collaborations to dynamic resource acquisitions and service interactions across organizations. Such dynamism leads to new challenges in security and dependability. Collaborating services in a system with a Service-Oriented Architecture (SOA) may belong to different security realms but often need to be engaged dynamically at runtime. If their security realms do not have a direct cross-realm authentication relationship, it is technically difficult to enable any secure collaboration between the services. A potential solution to this would be to locate intermediate realms at runtime, which serve as an authentication path between the two separate realms. However, the process of generating an authentication path for two distributed services can be highly complicated. It could involve a large number of extra operations for credential conversion and require a long chain of invocations to intermediate services. In this paper, we address this problem by designing and implementing a new cross-realm authentication protocol for dynamic service interactions, based on the notion of service-oriented multiparty business sessions. Our protocol requires neither credential conversion nor establishment of any authentication path between the participating services in a business session. The correctness of the protocol is formally analyzed and proven, and an empirical study is performed using two production-quality Grid systems, Globus 4 and CROWN. The experimental results indicate that the proposed protocol and its implementation have a sound level of scalability and impose only a limited degree of performance overhead, which is for example comparable with those security-related overheads in Globus 4.
Jie Xu 0007, Dacheng Zhang, Lu Liu 0001, Xianxian Li
IEEE Trans. Serv. Comput.1
2011 Monitoring Resources Allocation for Service Composition Under Different Monitoring Mechanisms
abstract
As availability of web service has become a great concern in SOA, monitoring mechanism is often deployed to detect and recover failures for service composition. While monitoring mechanism could improve the availability to an extent, it may cost more resources and increase the response time perceived by end users. To decrease the overall usage of monitoring resources, this paper proposes to select some services bringing the highest availability improvement to the composition and allocate monitors on them while leaving others unmonitored. This paper first researched two common monitoring mechanisms in service composition and analyzed their different impact on the service composition QoS values. Then continuous-time Markov chain and discrete-time Markov chain were employed to build the availability model related to the monitoring rate or service pool size according to different kind of monitoring mechanisms. Based on these models, two algorithms were proposed, for two monitoring mechanisms respectively, to allocate monitors in the composition aiming at minimizing the overall number of monitors while making sure the composition availability could meet certain requirements. The monitor allocation algorithm could be used to get the overall number of monitors and those services to monitor in different scenarios. Empirical studies results showed that it was feasible to monitor only some services in the composition to meet certain availability requirement. Monitors allocation decreased the overall number of monitors in the service composition and also decreased the mean response time comparing with the scenario that all services were monitored.
Pan He, Kaigui Wu, Jie Xu 0007
CISIS4
2011 QoS Driven Web Services Evolution
abstract
The loose coupling and on-demand integration are the fundamental characteristics of the Service Oriented Architecture (SOA), which have enforced rapid development of Web services. However, nonfunctional quality of service (QoS) attributes may evolve due to the changes of network conditions and locations of the service users. Some real services may update their QoS properties on-the-fly, others may turn to unavailable. Thus, addressing the problem of uninformed QoS evolution of Web services has become a significant research issue. This paper proposes a dynamic evolution framework of Web services, which uses the Collaborative Filtering (CF) to predict the QoS values and enables the evolution of Web services. In this framework, the QoS values of current users can be predicted using the past QoS data of similar users. There is no extra Web services invocation. About 1.5 millions real-world QoS data are used for evaluation and the experimental results show that it is a feasible and supplementary manner in dynamic evolution of the Web Services.
Kaigui Wu, Jie Xu 0007
CISIS3
2011 Distributed service integration for disaster monitoring sensor systems
abstract
Sensor networks have the potential to revolutionise the capture, processing and communication of critical data for use of disaster rescue and relief. In order to provide a dependable rescue capability through dynamically integrating newly developed and legacy sensor systems with other systems and computing, new methodologies are required for the dependable integration of services in heterogeneous environments. In this study, the authors present a new architectural model which can proactively self-adapt to changes and evolution occurring in the provision of search and rescue capabilities in a dynamic environment. This performance and reliability of the approach has been evaluated using simulations in a dynamic environment and demonstrated through developing and testing a demonstration system for a scenario of disaster area monitoring.
Lu Liu 0001, Nick Antonopoulos, Jie Xu 0007, David Webster, Kaigui Wu
IET Commun.3
2010 Personalized Context-Aware QoS Prediction for Web Services Based on Collaborative Filtering
Kaigui Wu, Jie Xu 0007, Pan He
ADMA (2)3
2010 A New Method for Formalizing Optimistic Fair Exchange Protocols
Ming Chen 0009, Kaigui Wu, Jie Xu 0007, Pan He
ICICS3
2010 Towards Building Efficient Content-Based Publish/Subscribe Systems over Structured P2P Overlays
abstract
In this paper, we introduce a generic model to deal with the event matching problem of content-based publish/subscribe systems over structured P2P overlays. In this model, we claim that there are three methods (event-oriented, subscription-oriented and hybrid) to make all the matched pairs (event, subscription) meet in a system. By theoretically analyzing the inherent problem of both event-oriented and subscription-oriented methods, we propose PEM (Popularity-based Event Matching), a variant of hybrid method. PEM can achieve better trade-off between event processing load and subscription storage load of a system. PEM has been verified through both mathematical and simulation-based evaluation.
Shengdong Zhang, Ji Wang 0001, Rui Shen 0003, Jie Xu 0007
ICPP4
2010 State-Based Search Strategy in Unstructured P2P
abstract
Efficient resource search in large-scale unstructured peer-to-peer (P2P) systems remains a fundamental challenge. In order to improve search performance, interest-based search is a good way to tackle the challenge. However, the existing interest-based search algorithms pay attention to user interest model and search history, but ignore some influence factors of improving performance. In this paper, through analysis of node's multi-dimensional quality-of-service (QoS) for search, it is found that node state's information is essential for improving search performance. So a State-Based Search (SBS) strategy is presented by improving Sripanidkulchai's interest-based search model, where Node's state evaluation is based on node's QoS fuzzification. The SBS enhances the methods of shortcuts adding, ranking and selecting, and state's life-cycle is predicted by grey theory. The experimental results show that the SBS can improve performance by reducing search response time and achieving load balance.
Kaigui Wu, Changze Wu, Lu Liu 0001, Jie Xu 0007
ISORC4
2010 Editorial: Special issue on dependable peer-to-peer systems
Lu Liu 0001, Jie Xu 0007
Peer-to-Peer Netw. Appl.2
2009 A Scenario-Based Architecture Evaluation Framework for Network Enabled Capability
abstract
The vision of service-oriented computing is one of loosely coupled services that create agile applications to encapsulate business objectives and processes. The potential of services to form complex systems of systems, with emergent behaviour, necessitates the need to understand how we can we develop sufficient confidence in their qualities. In this paper we argue that successful development and evolution of service oriented computing is dependent on making informed decisions at the architectural level. This poses a number of challenges concerning how to evaluate such system architectures and the measurements and metrics that are important in their assessment. This paper describes research-in-progress into the development of a scenario-based architectural evaluation framework that allows architectural-level reasoning across systems of systems in the context of Network Enabled Capability: a UK Ministry endeavour designed to achieve enhanced [military] effect through the physical networking and coherent integration of existing and future resources.
Colin C. Venters, Duncan Russell, Lu Liu 0001, Zongyang Luo, David Webster, Jie Xu 0007
COMPSAC (2)6
2009 A Practical Model for Conceptual Comparison Using a Wiki
abstract
One of the key concerns in the conceptualisation of a single object is understanding the context under which that object exists (or can exist). This contextual understanding should provide us with clear conceptual identification of an object including implicit situational information and detail of surrounding objects. For example in learning terms, a learner should be aware of concepts related to the context of their field of study and a surrounding cloud of contextually related concepts. This paper explores the use of an evolving community maintained knowledge-base (that of Wikipedia) in order to prioritise concepts that are semantically relevant to the user's interest space.
David Webster, Jie Xu 0007, Darren Mundy
ICALT2
2009 Quantification of Security for Compute Intensive Workloads in Clouds
abstract
Cloud computing is a promising technology to facilitate development of large-scale, on-demand, flexible computing infrastructures. However, improving dependability of cloud computing is critical for realization of its potential. In this paper, we describe our efforts to quantify security for Clouds to facilitate provision of assurance for quality of service, one of the factors contributing to dependability. This has profound implications for delivering customized security solutions such as effective intrusion prevention and detection which is the overall objective of our research. In order to demonstrate the applicability of our research, we have incorporated these requirements in the resource acquisition phase for Clouds. We also present experiments to demonstrate the effectiveness of our approach to address the random migration problem for virtualized computing environments.
Junaid Arshad, Paul Townend, Jie Xu 0007
ICPADS3
2009 Delivering Sustainable Capability on Evolutionary Service-oriented Architecture
abstract
Network enabled capability (NEC) is the U.K. Ministry of Defencepsilas response to the quickly changing conflict environment in which its forces must operate. In NEC, systems need to be integrated in context, to assist in human activity and provide dependable inter-operation. In order to provide reliable and sustainable military capability, fast paced changes must be conducted without halting the operation of a capability. In this paper we present the concept of evolutionary service-oriented architecture for delivering sustainable capability. The reliability of the architecture is evaluated by simulations using a computer-based model. The simulation results indicate that the evolutionary service-oriented architecture can provide higher reliability and sustainability in the provision of capability in a dynamic environment.
Lu Liu 0001, Duncan Russell, David Webster, Zongyang Luo, Colin C. Venters, Jie Xu 0007, John K. Davies
ISORC6
2009 Efficient resource discovery in self-organized unstructured peer-to-peer networks
abstract
Abstract In unstructured peer‐to‐peer (P2P) networks, two autonomous peer nodes can be connected if users in those nodes are interested in each other's data. Owing to the similarity between P2P networks and social networks, where peer nodes can be regarded as people and connections can be regarded as relationships, social strategies are useful for improving the performance of resource discovery by self‐organizing autonomous peers on unstructured P2P networks. In this paper, we present an efficient social‐like peer‐to‐peer (ESLP) method for resource discovery by mimicking different human behaviours in social networks. ESLP has been simulated in a dynamic environment with a growing number of peer nodes. From the simulation results and analysis, ESLP achieved better performance than current methods. Copyright © 2008 John Wiley & Sons, Ltd.
Lu Liu 0001, Nick Antonopoulos, Stephen Mackin, Jie Xu 0007, Duncan Russell
Concurr. Comput. Pract. Exp.4
2009 Context-aware trust negotiation in peer-to-peer service collaborations
Jianxin Li 0002, Dacheng Zhang, Jinpeng Huai, Jie Xu 0007
Peer-to-Peer Netw. Appl.4
2009 Efficient and scalable search on scale-free P2P networks
Lu Liu 0001, Jie Xu 0007, Duncan Russell, Paul Townend, David Webster
Peer-to-Peer Netw. Appl.2
2008 Self-Organization of Autonomous Peers with Human Strategies
abstract
Similarly to social networks where people are connected by their social relationships, two autonomous peer nodes can be connected in unstructured peer-to-peer (P2P) networks if users in those nodes are interested in each other's data. The similarity between P2P networks and social networks, where peer nodes are people and connections are relationships, leads us to believe that human strategies in social networks are useful for improving the performance of resource discovery by self-organising autonomous peers on unstructured P2P networks. In this paper, we present an efficient social-like peer-to-peer (ESLP) model for resource discovery by mimicking different human behaviours in social networks.
Lu Liu 0001, Jie Xu 0007, Duncan Russell, Nick Antonopoulos
ICIW2
2008 On Dynamic Replication Strategies in Data Service Grids
abstract
Service oriented architecture (SOA) allows multiple and heterogeneous data resources to be integrated within a single service while hiding the implementation details and formats of data resources from users of the service. However, data sources for a service are often distributed geographically and connected with long-latency networks; time and bandwidth consumption of data transportation may have an impact on the system performance. Dynamic data replication is a practical solution to this problem. By replicating data copies to appropriate sites, this approach aims to reduce time and bandwidth consumptions over networks. Existing strategies for dynamic replication are typically based on so-called single-location algorithms for identifying a single site for data replication. In this paper we discuss the issues with single-location strategies in large-scale data integration applications, and examine potential multiple-location schemes. Dynamic multiple-location replication is NP-complete in nature. We therefore transform the multiple-location problem into several classical mathematical problems with different parameter settings, for which efficient approximation algorithms exist. Experimental results indicate that unlike single-location strategies our multiple-location schemes are efficient with respect to access latency and bandwidth consumption, especially when the requesters of a data set are distributed over a large scale of locations.
Xiaohua Dong, Zhongfu Wu, Dacheng Zhang, Jie Xu 0007
ISORC5
2008 Scenario Based Evaluation
abstract
The concept of a scenario has long been utilized in military procurement as a means of evaluating capability in an operational context. With the advent of initiatives such as the USA Department of Defence's Network-Centric Operations and the UK Ministry of Defence's Network Enabled Capability, research into the application of service-oriented architectures (SOA) as a means of delivering capability is being undertaken. A promising output of this research is the definition of scenario based evaluation models that can be applied to SOA. This paper examines the application of a military scenario framework to more conventional software evaluation situations by means of case studies. This demonstrates that similar techniques can be used in their construction and thus provides a basis for the application of scenario based evaluation models in future research.
Nik Looker, David Webster, Duncan Russell, Jie Xu 0007
ISORC4
2008 Service-Oriented Integration of Systems for Military Capability
abstract
Service oriented architecture (SOA) is becoming established in computing as a means to integrate processing and data across organisations. This paper proposes that system-level integration can benefit from service oriented architectural descriptions and loose coupling between the problem domain requirements and different system solutions. The problem domain is exemplified as military capability, from the UK Ministry of Defence (MoD), in particular, network enabled capability (NEC). Representations of military capability in the problem domain can be described in terms of processes. The processes are sequences of functions that can be described as services. Then different types of system solutions can implement the described services. Firstly, the paper presents an overview of conceptual SOA and in the context of military capability compares three levels of service integration: business services, systems services and computing services. Secondly, the paper presents a framework for evaluating the performance and effectiveness of service integration to compare different solutions in delivering military capability.
Duncan Russell, Nik Looker, Lu Liu 0001, Jie Xu 0007
ISORC4
2008 Special Issue: Selected Papers from the U.K. e-Science All Hands Meeting in 2006
abstract
U.K. e-Science All Hands Meetings provide a premier forum for the exchange of ideas and the dissemination of research findings from a broad range of e-Science research activities and projects across all disciplines.Since 2001, the U.K. e-Science Core Programme has funded more than 100 research projects to develop e-Science techniques and applications.The year 2006 was a particularly important year for the U.K. e-Science Programme, as many of the first set of e-Science pilot projects had completed with significant outputs, and the Programme itself was moving into a new phase towards the establishment of a U.K. national e-Infrastructure for research and innovation.In response to this important transition, the fifth All Hands Meeting (AHM 2006) was held on 18th-21st September 2006 at the East Midlands Conference Centre in Nottingham, U.K. The theme of the conference-Achievements, Challenges and New Opportunities-reflected precisely our willingness to review past achievements, highlight current challenges, and identify future opportunities for e-Science.AHM 2006 attracted approximately 630 delegates from both the U.K. and overseas.Submissions of papers and proposals for workshops and sessions reached a record level: 84 papers were accepted out of 128 submitted, 10 out of 25 workshop proposals were selected, as were 5 out of 13 BoF session proposals.This special issue of Concurrency and Computation: Practice and Experience contains extended versions of the seven best papers, which were selected from the 84 quality papers and accepted after an additional peer review process.Multidisciplinary research is a key driver for e-Science innovation.The accepted papers in this special issue cover a diverse range of disciplines, including computer science, physics, medicine, chemistry and the social sciences, which represent some of the most innovative and exciting research projects within the U.K. e-Science community.In particular, Nicholson et al. examine different data replication methods in the large hadron collider (LHC).The results from the simulation of the LHC computing grid show that dynamic replication does give an improved job throughput.For the complex Grid system assessed, simple replication strategies, such as, least recently used (LRU) and least frequently used (LFU) are as effective as more advanced economic models.Periorellis et al. introduce the GOLD e-Science pilot project (Grid-based information models to support the rapid innovation of new high value-added chemicals), which aims at carrying out research into enabling technology to support the formation, operation and termination of virtual organizations.This paper discusses the issues with the GOLD infrastructure such as trust, security, contract monitoring and enforcement, information management and coordination.
Jie Xu 0007
Concurr. Comput. Pract. Exp.1
2008 FT-Grid: a system for achieving fault tolerance in grids
abstract
Abstract The FT‐Grid system introduces a fault‐tolerance framework that allows faults occurring in service‐oriented systems to be tolerated, thus increasing the dependability of such systems. This paper presents the design, development and evaluation of FT‐Grid. We show empirical evidence of the dependability benefits offered by FT‐Grid by performing an experimental dependability analysis using fault‐injection testing performed with the WS‐FIT tool. We then illustrate a potential problem with voting‐based fault‐tolerance schemes in the service‐oriented paradigm—namely that individual channels within a fault‐tolerant system, supposed to be independent of each other, may in fact invoke common services as part of their workflow, thus increasing the potential for common‐mode failure of those channels. We propose a solution to this issue by using the technique of provenance to provide FT‐Grid with topological awareness. We implement a large experimental system, and—with the use of the Provenance Recording for Services system developed as part of the PASOA project at the University of Southampton—perform a large number of experiments that show that a topologically aware FT‐Grid system serves as a much more dependable system than any other configuration tested, while imposing a negligible timing overhead. Copyright © 2007 John Wiley & Sons, Ltd.
Jie Xu 0007, Paul Townend, Nik Looker, Paul Groth
Concurr. Comput. Pract. Exp.1
2007 Dependability Assessment of Grid Middleware
abstract
Dependability is a key factor in any software system due to the potential costs in both time and money a failure may cause. Given the complexity of grid applications that rely on dependable grid middleware, tools for the assessment of grid middleware are highly desirable. Our past research, based around our fault injection technology (FIT) framework and its implementation, WS-FIT, has demonstrated that network level fault injection can be a valuable tool in assessing the dependability of traditional Web services. Here we apply our FIT framework to globus grid middleware using grid-FIT, our new implementation of the FIT framework, to obtain middleware dependability assessment data. We conclude by demonstrating that grid-FIT can be applied to globus grid systems to assess dependability as part of a fault removal mechanism and thus allow middleware dependability to be increased.
Nik Looker, Jie Xu 0007
DSN2
2007 Dynamic Cross-Realm Authentication for Multi-Party Service Interactions
abstract
Modern distributed applications are embedding an increasing degree of dynamism, from dynamic supply-chain management, enterprise federations, and virtual collaborations to dynamic service interactions across organisations. Such dynamism leads to new security challenges. Collaborating services may belong to different security realms but often have to be engaged dynamically at run time. If their security realms do not have in place a direct cross-realm authentication relationship, it is technically difficult to enable any secure collaboration between the services. A typical solution to this is to locate at run time intermediate realms that serve as an authentication-path between the two separate realms. However, the process of generating an authentication path for two distributed services can be very complex. It could involve a large number of extra operations for credential conversion and require a long chain of invocations to intermediate services. In this paper, we address this problem by presenting a new cross-realm authentication protocol for dynamic service interactions, based on the notion of multi-party business sessions. Our protocol requires neither credential conversion nor establishment of any authentication path between session members. The correctness of the protocol is analysed, and a comprehensive empirical study is performed using two production quality grid systems, Globus 4 and CROWN. The experimental results indicate that our protocol and its implementation have a sound level of scalability and impose only a limited degree of performance overhead, which is comparable with those security-related overheads in Globus 4.
Dacheng Zhang, Jie Xu 0007, Xianxian Li
DSN2
2006 TOWER: Practical Trust Negotiation Framework for Grids
abstract
In order to establish trust relationship between service requesters and providers in an open decentralized environment, we propose a novel trust negotiation framework, TOWER, which integrates distributed trust chain construction of trust management and aims to enhance the grid security infrastructure. Our approach leverages attribute-based credentials to support flexible delegation, and dynamically constructs trust chains. A novel TRust chAin based Negotiation Strategy (TRANS) is proposed to establish trust relationship on the fly by gradually disclosing credentials according to various access control policies. Our approach has been successfully implemented as useful components and fundamental security services in the CROWN Grid, and techniques such as trust tickets and policy caching that can greatly increase service efficiency are used. Finally, we evaluate our approach by comprehensive experiments and the results show that it is feasible.
Jianxin Li 0002, Jinpeng Huai, Jie Xu 0007, Yanmin Zhu 0006
e-Science3
2006 A Practical Approach to Secure Web Services
abstract
Web services provide the potential to offer interoperability of distributed business-to-business application integration between autonomous organisations, regardless of platforms, operating systems or languages. For both user and vendor organisations, this raises immediate problems of trust, security, privacy and prevention of malicious attacks. Until these problems are addressed and solved properly, the use of Web services will be severely restricted because no-one will trust them. We describe in this paper a service-oriented architecture and an attack-tolerant information retrieval (ATIR) service which tackles certain classes of privacy problems. In particular, we address the problem of protecting a user against malicious attacks upon an information service when the user retrieves some information from the service. Although there have been many theoretical solutions to certain aspects of this problem, the results have yet to be adapted to real systems. We report our experience of integrating the ATIR service with Taverna, a popular workflow system used amongst the UK e-science/grid computing community, to support secure information retrieval in the biology context. Performance studies show that the overhead of ATIR server-side processing is trivial (<5%) in comparison with the total processing time of the integrated Taverna. Our experimental results also show that the major processing overhead is caused by the Taverna enactor operations which consume no less than 50% of the total processing time.
Jie Xu 0007, Erica Y. Yang, Keith H. Bennett
ISORC1
2006 A Dynamic Shadow Approach to Fault-Tolerant Mobile Agents in an Autonomic Environment
Jie Xu 0007, Simon Pears
Real Time Syst.1
2005 A Comparison of Network Level Fault Injection with Code Insertion
abstract
This paper describes our research into the application of fault injection to Simple Object Access Protocol (SOAP) based service oriented-architectures (SOA). We show that our previously devised WS-FIT method, when combined with parameter perturbation, gives comparable performance to code insertion techniques with the benefit that it is less invasive. Finally we demonstrate that this technique can be used to compliment certification testing of a production system by strategic instrumentation of selected servers in a system.
Nik Looker, Malcolm Munro, Jie Xu 0007
COMPSAC (1)3
2005 Increasing Web Service Dependability Through Consensus Voting
abstract
This paper demonstrates our Web Service based N-Version model, WS-FTM (Web Service-Fault Tolerance Mechanism), which applies this well proven technique to the domain of Web Services to increase system dependability. WS-FTM achieves transparent usage of replicated Web Services by use of a modified stub. The stub is created using tools included in WS-FTM. Our initial implementation includes a simple consensus voter that allows generic result comparison. Finally we show, through the use of a non-trivial example, that WS-FTM can be used to increase the reliability of a Web Service system.
Nik Looker, Malcolm Munro, Jie Xu 0007
COMPSAC (2)3
2005 Securing instance-level interactions in Web services
abstract
The Web service technology enables dynamic service composition, resource utilisation and application integration in a heterogeneous computing environment. Web services can be used to compose and perform flexible and complex business flows. In practice, a Web service may create multiple service instances working for different business flows or business sessions, whilst the service instances within a business session may be created by different Web services, often designed, implemented and maintained by different organisations across different security domains. This introduces new challenges to existing security systems and solutions. For many applications ensuring security only at the level of Web services is not enough for a fine-grained level of control for multi-party collaborations because interactions amongst Web services in fact happen at the level of service instances. In this paper, we address the problem of how to secure instance-level interactions in Web services, and discuss different schemes for identifying and authenticating service instances. We present an experimental system and analyse some performance results. The experimental system implements instance-level communication control and instance authentication. The experimental results demonstrate that the overhead of execution time introduced by instance authentication is proportional to the number of the session partners within a business session.
Dacheng Zhang, Jie Xu 0007
ISADS2
2005 A Provenance-Aware Weighted Fault Tolerance Scheme for Service-Based Applications
abstract
Service-orientation has been proposed as a way of facilitating the development and integration of increasingly complex and heterogeneous system components. However, there are many new challenges to the dependability community in this new paradigm, such as how individual channels within fault-tolerant systems may invoke common services as part of their workflow, thus increasing the potential for common-mode failure. We propose a scheme that - for the first time - links the technique of provenance with that of multi-version fault tolerance. We implement a large test system and perform experiments with a single-version system, a traditional MVD system, and a provenance-aware MVD system, and compare their results. We show that for this experiment, our provenance-aware scheme results in a much more dependable system than either of the other systems tested, whilst imposing a negligible timing overhead.
Paul Townend, Paul Groth, Jie Xu 0007
ISORC3
2004 Dynamic Data Integration Using Web Services
abstract
We address the problem of large-scale data integration, where the data sources are unknown at design time, are from autonomous organisations, and may evolve. Experiments are described involving a demonstrator system in the field of health services data integration within the UK. Current Web services technology has been used extensively and largely successfully in these distributed prototype systems. The work shows that Web services provide a good infrastructure layer, but integration demands a higher level "broker" architectural layer; the paper identifies eight specific requirements for such an architecture that have emerged from the experiments, derived from an analysis of shortcomings which are collectively due to the static nature of the initial prototype. The way in which these are being met in the current version in order to achieve a more dynamic integration is described.
Fujun Zhu, Mark Turner 0001, Ioannis Kotsiopoulos, Keith H. Bennett, Michelle Russell, David Budgen, Pearl Brereton, John A. Keane, Paul J. Layzell 0001, Michael Rigby, Jie Xu 0007
ICWS11
2004 Multi-Party Authentication for Web Services: Protocols, Implementation and Evaluation
abstract
The Web service technology allows the dynamic composition of a workflow (or a business flow) by composing a set of existing Web services scattered across the Internet. While a given Web service may have multiple service instances taking pan in several workflows simultaneously, a workflow often involves a set of service instances that belong to different Web services. In order to establish trust relationships amongst service instances, new security protocols are urgently needed. Hada and Maruyabma [2002] presented a session-oriented, multi-party authentication protocol to resolve this problem. Within a session their protocol provides a commonly shared session secret for all the service instances, thereby distinguishing the instances from those of other sessions. However, individual instances cannot be distinguished and identified using the session secret. This leads to vulnerable session management and poor threat containment. In this paper we present a new protocol design for multiparty authentication in which each service instance of a given session is provided with a unique identifier. The coordinated atomic action scheme is exploited for achieving an improved level of threat containment. We evaluate the scalability of our design by means of both experiments and an analytical model. The result shows that time consumed by the authentication process increases linearly with an increase in the number of session participants.
Dacheng Zhang, Jie Xu 0007
ISORC2
2004 Dependability in Web Services
Jie Xu 0007, Graeme Dixon, Priya Narasimhan, Luigi Romano, Rob Smith
SRDS1
2003 A Broker Architecture for Integrating Data Using a Web Services Environment
Keith H. Bennett, Nicolas E. Gold, Paul J. Layzell 0001, Fujun Zhu, Pearl Brereton, David Budgen, John A. Keane, Ioannis Kotsiopoulos, Mark Turner 0001, Jie Xu 0007, Orouba Almilaji, Jung-Ching Chen, Ali Owrak
ICSOC10
2003 Mobile Agent Fault Tolerance for Information Retrieval Applications: An Exception Handling Approach
abstract
Maintaining mobile agent availability in the presence of agent server crashes is a challenging issue since developers normally have no control over remote agent servers. A popular technique is that a mobile agent injects a replica into stable storage upon its arrival at each agent server. However, a server crash leaves the replica unavailable, for an unknown time period, until the agent server is back online. This paper uses exception handling to maintain the availability, of mobile agents in the presence of agent server crash failures. Two exception handler designs are proposed. The first exists at the agent server that created the mobile agent. The second operates at the previous agent server visited by the mobile agent. Initial performance results demonstrate that although the second design is slower it offers the smaller trip time increase in the presence of agent server crashes.
Simon Pears, Jie Xu 0007, Cornelia Boldyreff
ISADS2
2003 A Dynamic Shadow Approach for Mobile Agents to Survive Crash Failures
abstract
Fault tolerance schemes for mobile agents to survive agent server crash failures are complex since developers normally have no control over remote agent servers. Some solutions inject a replica into stable storage upon its arrival at an agent server. However in the event of an agent server crash the replica is unavailable until the agent server recovers. This paper presents a failure model and a revised exception handling framework for mobile agent systems. An exception handler design is presented for mobile agents to survive agent server crash failures. A replica mobile agent operates at the agent server visited prior to its master's current location. If a master crashes its replica is available as a replacement. Experimental evaluation is performed and performance results are used to suggest some useful design guidelines.
Simon Pears, Jie Xu 0007, Cornelia Boldyreff
ISORC2
2003 Assessing the Dependability of OGSA Middleware by Fault Injection
abstract
This paper presents our research on devising a dependability assessment method for the upcoming OGSA 3.0 middleware using network level fault injection. We compare existing DCE middleware dependability testing research with the requirements of testing OGSA middleware and derive a new method and fault model. From this we have implemented an extendable fault injector framework and undertaken some proof of concept experiments with a simulated OGSA middleware system based around Apache SOAP and Apache Tomcat. We also present results from our initial experiments, which uncovered a discrepancy with our simulated OGSA system. We finally detail future research, including plans to adapt this fault injector framework from the stateless environment of a standard Web service to the stateful environment of an OGSA service.
Nik Looker, Jie Xu 0007
SRDS2
2002 Private Information Retrieval in the Presence of Malicious Failures
abstract
In the application domain of online information services such as online census information, health records and real-time stock quotes, there are at least two fundamental challenges: the protection of users' privacy and the assurance of service availability. We present a fault-tolerant scheme for private information retrieval (FT-PIR) that protects users' privacy and ensure service provision in the presence of malicious server failures. An error detection algorithm is introduced into this scheme to detect the corrupted results from servers. The analytical and experimental results show that the FT-PIR scheme can tolerate malicious server failures effectively and prevent any information of users front being leaked to attackers. This new scheme does not rely on any unproven cryptographic premise and the availability of tamperproof hardware. An implementation of the FT-PIR scheme on a distributed database system suggests just a modest level of performance overhead.
Erica Y. Yang, Jie Xu 0007, Keith H. Bennett
COMPSAC2
2002 A Fault-Tolerant Approach to Secure Information Retrieval
abstract
Several private information retrieval (PIR) schemes were proposed to protect users' privacy when sensitive information stored in database servers is retrieved. However, existing PIR schemes assume that any attack to the servers does not change the information stored and any computational results. We present a novel fault-tolerant PIR scheme (called FT-PIR) that protects users' privacy and at the same time ensures service availability in the presence of malicious server faults. Our scheme neither relies on any unproven cryptographic assumptions nor the availability of tamper-proof hardware. A probabilistic verification function is introduced into the scheme to detect corrupted results. Unlike previous PIR research that attempted mainly to demonstrate the theoretical feasibility of PIR, we have actually implemented both a PIR scheme and our FT-PIR scheme in a distributed database environment. The experimental and analytical results show that only modest performance overhead is introduced by FT-PIR while comparing with PIR in the fault-free cases. The FT-PIR scheme tolerates a variety of server faults effectively. In certain fail-stop fault scenarios, FT-PIR performs even better than PIR. It was observed that 35.82% less processing time was actually needed for FT-PIR to tolerate one server fault.
Erica Y. Yang, Jie Xu 0007, Keith H. Bennett
SRDS2
2002 Rigorous Development of an Embedded Fault-Tolerant System Based on Coordinated Atomic Actions
abstract
Describes our experience using coordinated atomic (CA) actions as a system structuring tool to design and validate a sophisticated and embedded control system for a complex industrial application that has high reliability and safety requirements. Our study is based on an extended production cell model, the specification and simulator for which were defined and developed by FZI (Forschungszentrum Informatik, Germany). This "fault-tolerant production cell" represents a manufacturing process involving redundant mechanical devices (provided in order to enable continued production in the presence of machine faults). The challenge posed by the model specification is to design a control system that maintains specified safety and liveness properties even in the presence of a large number and variety of device and sensor failures. Based on an analysis of such failures, we provide details of: (1) a design for a control program that uses CA actions to deal with both safety-related and fault tolerance concerns and (2) the formal verification of this design based on the use of model checking. We found that CA action structuring facilitated both the design and verification tasks by enabling the various safety problems (involving possible clashes of moving machinery) to be treated independently. Even complex situations involving the concurrent occurrence of any pairs of the many possible mechanical and sensor failures can be handled simply yet appropriately. The formal verification activity was performed in parallel with the design activity, and the interaction between them resulted in a combined exercise in "design for validation"; formal verification was very valuable in identifying some very subtle residual bugs in early versions of our design which would have been difficult to detect otherwise.
Jie Xu 0007, Brian Randell, Alexander B. Romanovsky, Robert J. Stroud, Avelino Francisco Zorzo, Ercument Canver, Friedrich W. von Henke
IEEE Trans. Computers1
2001 An Architectural Model for Service-Based Flexible Software
abstract
The urgent need to change software easily to meet evolving business requirements requires a radical shift in the development of software, with a more demand-centric view leading to software which will be delivered as a service, within the framework of an open marketplace. We describe a service architecture and its rationale, in which components may be bound instantly, just at the time they are needed and then the binding may, be disengaged. This allows highly flexible software services to be evolved in "internet time". The paper focuses on early results: some of the aims have been demonstrated and amplified through an experimental implementation based on e-Speak, an existing and available technology. It is concluded that technology such as e-Speak provides a useful infrastructure that rapidly enabled us to demonstrate the basic operation and viability of our approach.
Keith H. Bennett, Jie Xu 0007, Malcolm Munro, Zhuang Hong, Paul J. Layzell 0001, Nicolas E. Gold, David Budgen, Pearl Brereton
COMPSAC2
2001 A comparative study of exception handling mechanisms for building dependable object-oriented software
Alessandro F. Garcia 0001, Cecília M. F. Rubira, Alexander B. Romanovsky, Jie Xu 0007
J. Syst. Softw.4
2000 Concurrent Exception Handling and Resolution in Distributed Object Systems
abstract
We address the problem of how to handle exceptions in distributed object systems. In a distributed computing environment, exceptions may be raised simultaneously in different processing nodes and thus need to be treated in a coordinated manner. Mishandling concurrent exceptions can lead to catastrophic consequences. We take two kinds of concurrency into account: 1) Several objects are designed collectively and invoked concurrently to achieve a global goal and 2) multiple objects (or object groups) that are designed independently compete for the same system resources. We propose a new distributed algorithm for resolving concurrent exceptions and show that the algorithm works correctly even in complex nested situations, and is an improvement over previous proposals in that it requires only O(n/sub max/N/sup 2/) messages, thereby permitting quicker response to exceptions.
Jie Xu 0007, Alexander B. Romanovsky, Brian Randell
IEEE Trans. Parallel Distributed Syst.1
1999 Using Coordinated Atomic Actions to Design Safety-Critical Systems: a Production Cell Case Study
abstract
Coordinated Atomic actions (CA actions) are a unified approach to structuring complex concurrent activities and supporting error recovery between multiple interacting objects in object-oriented systems. This paper explains how we have used the CA action concept to design and implement a safety-critical application. We have used the Production Cell model that was developed in the Forschungszentrum Informatik (FZI), Karlsruhe, Germany, to present a realistic industry-oriented problem, where safety requirements play a significant role. Our design consists of two levels: the first level deals with the scheduling of CA actions, and the second level deals with the interactions between devices. Both the scheduling mechanism and the device interactions are enclosed by CA actions. Exception handling and error recovery are incorporated into CA actions in order to satisfy high safety and fault tolerance requirements. A controlling program based on our design was developed in the Java language and used to drive a graphical simulator provided by the FZI. Copyright © 1999 John Wiley & Sons, Ltd.
Avelino Francisco Zorzo, Alexander B. Romanovsky, Jie Xu 0007, Brian Randell, Robert J. Stroud, Ian Welch
Softw. Pract. Exp.3
1998 Coordinated Exception Handling in Distributed Object Systems: From Model to System Implementation
abstract
Exception handling in concurrent and distributed programs is a difficult task though it is often necessary. In many cases traditional exception mechanisms for sequential programs are no longer appropriate. One major difficulty is that the process of handling an exception may need to involve multiple concurrent components that are cooperating in pursuit of some global goal. Another complication is that several exceptions may be raised concurrently in different nodes of a distributed environment. Existing proposals and actual concurrent languages either ignore these difficulties or only cope with a limited form of them. The paper attempts a general solution, developed especially for distributed object systems, starting from a conceptual model, together with algorithms for coordinating concurrent components and resolving multiple exceptions, through to an actual system implementation. An industrial production cell is chosen as a case study to demonstrate the usefulness of the proposed model and algorithms. A system that supports coordinated atomic actions and exception resolution is implemented in distributed Ada 95 and examined through several performance-related experiments.
Jie Xu 0007, Alexander B. Romanovsky, Brian Randell
ICDCS1
1998 Exception Handling in Object-Oriented Real-Time Distributed Systems
abstract
Exception handling in a complex concurrent and distributed system (e.g. one involving cooperating rather than just competing activities) is often a necessary, but a very difficult, task. No widely accepted models or approaches exist in this area. The object-oriented paradigm, for all its structuring benefits, and real-time requirements each add further difficulties to the design and implementation of exception handling in such systems. In this paper, we develop a general structuring framework based on the coordinated atomic (CA) action concept for handling exceptions in an object-oriented distributed system, in which exceptions in both the value and the time domain are taken into account. In particular, we attempt to attack several difficult problems related to real-time system design and error recovery, including action-level timing constraints, time-triggered CA actions, and time-dependent exception handling. The proposed framework is then demonstrated and assessed using an industrial real-time application-the Production Cell III case study.
Alexander B. Romanovsky, Jie Xu 0007, Brian Randell
ISORC2
1997 Implementation of blocking coordinated atomic actions based on forward error recovery
Alexander B. Romanovsky, Brian Randell, Robert J. Stroud, Jie Xu 0007, Avelino Francisco Zorzo
J. Syst. Archit.4
1996 Exception Handling and Resolution in Distributed Object-oriented Systems
abstract
We address the problem of how to handle exceptions in distributed object-oriented systems. In a distributed computing environment exceptions may be raised simultaneously and thus need to be treated in a coordinated manner. We take two kinds of concurrency into account: 1) several objects are designed collectively and invoked concurrently to achieve a global goal, and 2) concurrent objects or object groups that are designed independently compete for the same system resources. We propose a new distributed algorithm for resolving concurrent exceptions and show that the algorithm works correctly even in complex nested situations, and is an improvement over previous proposals in that it requires only O(N/sup 2/) messages, and is fully object-oriented.
Jie Xu 0007, Alexander B. Romanovsky, Brian Randell
ICDCS1
1996 Roll-forward error recovery in embedded real-time systems
abstract
Roll-forward checkpointing schemes are developed in order to avoid rollback in the presence of independent faults and to increase the possibility that a task completes within a tight deadline. However, despite of the adoption of roll-forward recovery, these schemes are not necessarily appropriate for time-critical applications because interactions with the external environment and communications between processes must be deferred during checkpoint validation steps (typically, two checkpoint intervals) until the fault-free processors are identified. The deadlines on providing services may thus be violated. In this paper we present and discuss two alternative roll-forward recovery schemes, especially for time-critical and interaction-intensive applications, that deliver correct, timely results even when checkpoint validation is required.
Jie Xu 0007, Brian Randell
ICPADS1
1995 Sequentially t-Diagnosable Systems: A Characterization and Its Applications
abstract
In the system level fault diagnosis area, the fundamental problem of characterizing sequentially t-diagnosable systems in the PMC model has remained open for more than two decades. We resolve this problem by providing a complete characterization of such systems. Our solution to the characterization problem leads to the correct identification of optimal sequentially t-diagnosable D/sub /spl delta/,k/ systems. Given a set of n units where n=2t+1, an optimal D/sub /spl delta/,k/ system can be constructed with just n([(t+2)/3]) tests, rather than n([t/2]+1) tests-a previously misjudged bound. An efficient algorithm for identifying the set of faulty units in a sequentially t-diagnosable D/sub /spl delta/,k/ system is given along the line of the proposed characterization, which is linear with respect to the number of tests in the system.>
Jie Xu 0007, Shi-ze Huang
IEEE Trans. Computers1