EDBT 2026 Demo / reviewers in the wild / expert
Paul Townend
dblp:48/917
· DBLP profile ↗
28ranked-venue papers
2as first author
6since 2021 · last 2026
0000-0002-9698-8361ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 3 since 2021Computer networks · 5 · 1 first-authorSoftware engineering, systems software and programming languages · 2Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient retraining of machine learning algorithms in cloud management systemsabstractCloud management systems performing capacity autoscaling, application orchestration, server consolidation, and service differentiation increasingly rely on machine learning (ML) models for predictive decision-making. However, changes in user behavior, software updates, and hardware upgrades cause monitoring data to deviate from the training distribution, leading to model performance degradation. This phenomenon—known as concept drift poses a significant challenge to maintaining prediction accuracy in dynamic cloud environments. In this work, we propose a hybrid concept drift detection approach that combines statistical data monitoring with model performance metrics to inform efficient model retraining. The proposed method minimizes false positives, reduces adaptation delays, and avoids unnecessary retraining during stable periods. We conduct extensive experiments on both synthetic and real-world cloud workload datasets collected from multiple data centers, evaluating the hybrid approach against five established drift detection algorithms. To ensure statistical rigor, all experiments are repeated ten times, and the results are reported with 95% confidence intervals and significance tests. The results show that our proposed approach and implemented methods improve drift detection efficiency by eliminating unnecessary retraining and lead to more than 60% improvements over the baseline prediction accuracy. We Edited the abstract and removed the inconsistencies Lidia Kidane, Paul Townend, Thijs Metsch, Erik Elmroth |
Future Gener. Comput. Syst. | 2 |
| 2026 | GraphOpticon: A Global proactive horizontal autoscaler for improved service performance & resource consumptionabstractThe increasing complexity of distributed computing environments necessitates efficient resource management strategies to optimize performance and minimize resource consumption. Although proactive horizontal autoscaling dynamically adjusts computational resources based on workload predictions, existing approaches primarily focus on improving workload resource consumption, often neglecting the overhead introduced by the autoscaling system itself. This could have dire ramifications on resource efficiency, since many prior solutions rely on multiple forecasting models per compute node or group of pods, leading to significant resource consumption associated with the autoscaling system. To address this, we propose GraphOpticon, a novel proactive horizontal autoscaling framework that leverages a singular global forecasting model based on Spatiotemporal Graph Neural Networks. The experimental results demonstrate that GraphOpticon is capable of providing improved service performance, and resource consumption (caused by the workloads involved and the autoscaling system itself). As a matter of fact, GraphOpticon manages to consistently outperform other contemporary horizontal autoscaling solutions, such as Kubernetes’ Horizontal Pod Autoscaler, with improvements of 6.62% in median execution time, 7.62% in tail latency, and 6.77% in resource consumption, among others. Theodoros Theodoropoulos, Yashwant Singh Patel, Uwe Zdun, Paul Townend, Ioannis Korontanis, Antonios Makris, Konstantinos Tserpes |
Future Gener. Comput. Syst. | 4 |
| 2025 | Balancing Compression and Prediction: A Hybrid Autoencoder-LSTM Framework for Cloud WorkloadsabstractAccurate future workload prediction is an essential step for proactive resource allocation and efficient provisioning in cloud computing environments. Deep learning strategies have proven successful for this task, but they face challenges due to the high dimensionality of monitoring data, extensive preprocessing requirements, and computational overhead. In this paper, we propose a hybrid framework that integrates autoencoders for workload compression with Long Short-Term Memory (LSTM) networks for time-series forecasting. Unlike prior studies, our approach systematically analyzes the trade-off between compression ratio and predictive accuracy, demonstrating how dimensionality reduction can improve both scalability and robustness. Thereby reducing the computational burden associated with processing massive-scale monitoring data. Experiments conducted on both synthetic and real-world datasets demonstrate that the proposed method achieves up to 60% data compression with minimal reconstruction loss, while also improving prediction accuracy compared to baseline LSTM models. We evaluate the overall performance of the framework using various metrics, including data reduction ratio, prediction accuracy, and the effects of different compression stages on predictive performance. Additionally, we quantify the computational savings in terms of CPU usage, memory footprint, and training/inference times, confirming the framework’s feasibility for real-world deployment. These results underscore the potential of integrating compression and prediction to achieve scalable, accurate, and resource-efficient management of cloud workloads. Lidia Kidane, Paul Townend, Thijs Metsch, Erik Elmroth |
BDCAT | 2 |
| 2025 | GreenContinuum: A Formal Model of a Smart Grid-Aware Edge-Cloud Continuum for Carbon and Energy ManagementabstractThe Edge-Cloud Continuum is a large-scale, loosely coupled system consisting of multiple stakeholders, regions, dynamic infrastructures, and conflicting objectives. With surging growth and demand, the Continuum's energy and carbon footprint have massively increased, resulting in great operational expense, environmental impact, and strain on power grids. Methods to mitigate this face significant challenges: Quality of Service (QoS) guarantees must be balanced against not only carbon emissions, but the loadings, capacities, and QoS of the (smart) grids that power the underlying infrastructure. Integrated models to enable reasoning across both a Continuum and its associated Smart Grids are therefore required. This work presents a formal model to reason across the integration of Smart Grids and the Edge-Cloud Continuum. Firstly, we identify the components, interactions, and properties crucial to mitigating cross-Continuum energy and carbon footprint while maintaining user, provider, and power grid QoS. We then present associated mathematical models to enable a model-based simulation to be developed based on our work. We present this simulation (all code is available for download) and use a simple scheduling algorithm to demonstrate the feasibility of utilizing knowledge from both the Smart Grid and Edge-Cloud Continuum for carbon and energy management, showing that significant savings are possible. Rohail Gulbaz, Paul Townend, Per-Olov Östberg |
CloudCom | 2 |
| 2025 | A Decentralized Microservice Scheduling Approach Using Service Mesh in Cloud-Edge SystemsabstractAs microservice-based systems scale across the cloud-edge continuum, traditional centralized scheduling mechanisms increasingly struggle with latency, coordination overhead, and fault tolerance. This paper presents a new architectural direction: leveraging service mesh sidecar proxies as decentralized, in-situ schedulers to enable scalable, low-latency coordination in large-scale, cloud-native environments. We propose embedding lightweight, autonomous scheduling logic into each sidecar, allowing scheduling decisions to be made locally without centralized control. This approach leverages the growing maturity of service mesh infrastructures, which support programmable distributed traffic management. We describe the design of such an architecture and present initial results demonstrating its scalability potential in terms of response time and latency under varying request rates. Rather than delivering a finalized scheduling algorithm, this paper presents a system-level architectural direction and preliminary evidence to support its scalability potential. Yangyang Wen, Paul Townend, Per-Olov Östberg, Abel Souza, Clément Courageux-Sudan |
JCC | 2 |
| 2025 | Towards Adaptive Rule Replacement for Mitigating Inference Attacks in Serverless SDN FrameworkabstractIn the rapidly evolving landscape of Software-Defined Networking (SDN), the enhancement of security measures against sophisticated cyber threats is paramount. Among these threats, inference attacks pose a significant risk by allowing adversaries to deduce the configurations and policies of SDN switches, thereby undermining the integrity and confidentiality of the network infrastructure. To address this critical issue, we introduce a novel dynamic rule replacement policy for SDN switches, leveraging the capabilities of a Support Vector Machine (SVM) for its implementation. Our approach utilizes a comprehensive set of statistical features, including duration analysis of flow rules, dispersion of packet match fields, and frequency of packet arrivals to identify patterns indicative of potential inference attacks. By dynamically adjusting the rules within SDN switches based on the analysis of these features, our policy significantly enhances the resilience of the network against such attacks. To accelerate the innovation and development of network services, this study proposes an integrated SDN architecture deployed over a serverless framework. This work serves as a starting point to enable researchers to realize the concept of modular serverless functions over traditional SDN environments. We show during inference attacks how a serverless framework improves the latency and resource utilization of the network compared to a traditional SDN framework. This study demonstrates an improvement in preventing inference attacks without compromising the performance and efficiency of the SDN infrastructure. Ankur Mudgal, Munesh Singh, Abhishek Verma 0003, Kshira Sagar Sahoo, Paul Townend, Monowar Bhuyan |
NOMS | 5 |
| 2019 | Guest Editor's Introduction: Special Section on Virtualization and Services for Cloud-Based Application SystemsabstractCloud-based application systems are rapidly deployed worldwide in production use via virtualization and services computing technologies. The scaling demands for these application capabilities to the cloud providers, compound with differentiated requirements on the quality of services, have brought severe technical challenges. This special section focuses on the techniques of virtualization and services for cloud-based application systems, mainly including multi-scale resource management and sharing, elastic scheduling and allocation of computing and network resources, monitoring and diagnosis for cloud-based services, and cloud-based mobile systems. The articles of this special section illustrate recent advances in virtualization and services provisioning for cloud-based application systems. We expect that this special section will provide an integrated view of the state-of-the-art techniques, identify new challenges as well as opportunities, and promote collaboration among researchers in this field. We received 34 submissions and we finally accepted 4 articles. The acceptance rate is as low as 11.8 percent. Yiming Zhang 0003, Rong Chang 0001, Paul Townend |
IEEE Trans. Serv. Comput. | 3 |
| 2019 | Errata to "Guest Editor's Introduction: Special Section on Virtualization and Services for Cloud-Based Application Systems"abstractPresents corrections to author affiliation information in the paper, “Guest Editor’s Introduction: Special Section on Virtualization and Services for Cloud-Based Application Systems,” (Zhange, Y. et al), IEEE Trans. Serv. Comput., vol. 12, no. 1, pp. 88–90, Jan./Feb. 2019. Yiming Zhang 0003, Rong Chang 0001, Paul Townend |
IEEE Trans. Serv. Comput. | 3 |
| 2018 | Improving content-based image retrieval for heterogeneous datasets using histogram-based descriptors
Carolina Reta, Ismael Solís Moreno, Jose Antonio Cantoral-Ceballos, Rogelio Alvarez-Vargas, Paul Townend |
Multim. Tools Appl. | 5 |
| 2018 | Adaptive Speculation for Efficient Internetware Application Execution in CloudsabstractModern Cloud computing systems are massive in scale, featuring environments that can execute highly dynamic Internetware applications with huge numbers of interacting tasks. This has led to a substantial challenge—the straggler problem, whereby a small subset of slow tasks significantly impede parallel job completion. This problem results in longer service responses, degraded system performance, and late timing failures that can easily threaten Quality of Service (QoS) compliance. Speculative execution (or speculation) is the prominent method deployed in Clouds to tolerate stragglers by creating task replicas at runtime. The method detects stragglers by specifying a predefined threshold to calculate the difference between individual tasks and the average task progression within a job. However, such a static threshold debilitates speculation effectiveness as it fails to capture the intrinsic diversity of timing constraints in Internetware applications, as well as dynamic environmental factors, such as resource utilization. By considering such characteristics, different levels of strictness for replica creation can be imposed to adaptively achieve specified levels of QoS for different applications. In this article, we present an algorithm to improve the execution efficiency of Internetware applications by dynamically calculating the straggler threshold, considering key parameters including job QoS timing constraints, task execution progress, and optimal system resource utilization. We implement this dynamic straggler threshold into the YARN architecture to evaluate it’s effectiveness against existing state-of-the-art solutions. Results demonstrate that the proposed approach is capable of reducing parallel job response time by up to 20% compared to the static threshold, as well as a higher speculation success rate, achieving up to 66.67% against 16.67% in comparison to the static method. Xue Ouyang 0003, Peter Garraghan, Bernhard Primas, David McKee 0001, Paul Townend, Jie Xu 0007 |
ACM Trans. Internet Techn. | 5 |
| 2017 | ML-NA: A Machine Learning Based Node Performance Analyzer Utilizing Straggler StatisticsabstractCurrent Cloud clusters often consist of heterogeneous machine nodes, which can trigger performance challenges such as the task straggler problem, whereby a small subset of parallel tasks running abnormally slower than the other sibling ones. The straggler problem leads to extended job response and deteriorates system throughput. Poor performance nodes are more likely to engender stragglers, and can undermine straggler mitigation effectiveness. For example, as the dominant mechanism for straggler alleviation, speculative execution functions by creating redundant task replicas on other machine nodes as soon as a straggler is detected. When speculative copies are assigned onto the poor performance nodes, it is hard for them to catch up with the stragglers compared to replicas run on fast nodes. And due to the fact that the performance heterogeneity is caused not only by static attribute variations such as physical capacity, but also dynamic characteristic uctuations such as contention level, analyzing node performance is important yet challenging. In this paper we develop ML-NA, a Machine Learning based Node performance Analyzer. By leveraging historical parallel tasks execution log data, ML-NA classies cluster nodes into different categories and predicts their performance in the near future as a scheduling guide to improve speculation effectiveness and minimize task straggler generation. We consider MapReduce as a representative framework to perform our analysis, and use the published OpenCloud trace as a case study to train and to evaluate our model. Results show that ML-NA can predict node performance categories with an average accuracy up to 92.86%. Xue Ouyang 0003, Renyu Yang, Guogui Yang, Paul Townend, Jie Xu 0007 |
ICPADS | 5 |
| 2017 | Mitigate data skew caused stragglers through ImKP partition in MapReduceabstractSpeculative execution is the mechanism adopted by current MapReduce framework when dealing with the straggler problem, and it functions through creating redundant copies for identified stragglers. The result of the quicker task will be adopted to improve the overall job execution performance. Although proved to be effective for contention caused stragglers, speculative execution can easily meet its bottleneck when mitigating data skew caused stragglers due to its replication nature: the identical unbalanced input data will lead to a slow speculative task. The Map inputs are typically even in size according to the HDFS block configuration, therefore the skew caused stragglers happen mainly in the Reduce phase because of the unknown intermediate key distribution. In this paper, we focus on mitigating data skew caused Reduce stragglers, propose ImKP, an Intermediate Key Pre-processing framework that enables the even distributed partition for Reduce inputs. A group based ranking technique has been developed that dramatically decreases the pre-processing time, and ImKP manages to eliminate this timing overhead through parallelizing the pre-processing with the file uploading procedure (from local file system to HDFS). For jobs that take input directly from HDFS, ImKP minimizes the overhead by storing themapping result on every node within the cluster for reuse. Experiments are conducted on different datasets with various workloads. Results show that, compared to the popular hash partition, ImKP can dramatically decrease Reduce skew, achieving a 99.8% reduction in the coefficient of variation of the input sizes in average, and improve up to 29.37% job response performance. Xue Ouyang 0003, Huan Zhou 0006, Stephen J. Clement, Paul Townend, Jie Xu 0007 |
IPCCC | 4 |
| 2016 | Straggler Detection in Parallel Computing Systems through Dynamic Threshold CalculationabstractCloud computing systems face the substantial challenge of the Long Tail problem: a small subset of straggling tasks significantly impede parallel jobs completion. This behavior results in longer service response times and degraded system utilization. Speculative execution, which create task replicas at runtime, is a typical method deployed in large-scale distributed systems to tolerate stragglers. This approach defines stragglers by specifying a static threshold value, which calculates the temporal difference between an individual task and the average task progression for a job. However, specifying static threshold debilitates speculation effectiveness as it fails to consider the intrinsic diversity of job timing constraints within modern day Cloud computing systems. Capturing such heterogeneity enables the ability to impose different levels of strictness for replica creation while achieving specified levels of QoS for different application types. Furthermore, a static threshold also fails to consider system environmental constraints in terms of replication overheads and optimal system resource usage. In this paper we present an algorithm for dynamically calculating a threshold value to identify task stragglers, considering key parameters including job QoS timing constraints, task execution characteristics, and optimal system resource utilization. We study and demonstrate the effectiveness of our algorithm through simulating a number of different operational scenarios based on real production cluster data against state-of-the-art solutions. Results demonstrate that our approach is capable of creating 58.62% less replicas under high resource utilization while reducing response time up to 17.86% for idle periods compared to a static threshold. Xue Ouyang 0003, Peter Garraghan, David McKee 0001, Paul Townend, Jie Xu 0007 |
AINA | 4 |
| 2016 | Holistic data centres: Next generation data and thermal energy infrastructuresabstractDigital infrastructure is becoming more distributed and requiring more power for operation. At the same time, many countries are working to de-carbonise their energy, which will require electrical generation of heat for populated areas. What if this heat generation was combined with digital processing? Paul Townend, Jie Xu 0007, Jon Summers, Daniel Ruprecht, Harvey Thompson |
IPCCC | 1 |
| 2015 | Timely Long Tail Identification through Agent Based Monitoring and AnalyticsabstractThe increasing complexity and scale of distributed systems has resulted in the manifestation of emergent behavior which substantially affects overall system performance. A significant emergent property is that of the "Long Tail", whereby a small proportion of task stragglers significantly impact job execution completion times. To mitigate such behavior, straggling tasks occurring within the system need to be accurately identified in a timely manner. However, current approaches focus on mitigation rather than identification, which typically identify stragglers too late in the execution lifecycle. This paper presents a method and tool to identify Long Tail behavior within distributed systems in a timely manner, through a combination of online and offline analytics. This is achieved through historical analysis to profile and model task execution patterns, which then inform online analytic agents that monitor task execution at runtime. Furthermore, we provide an empirical analysis of two large-scale production Cloud data enters that demonstrate the challenge of data skew within modern distributed systems, this analysis shows that approximately 5% of task stragglers caused by data skew impact 50% of the total jobs for batch processes. Our results demonstrate that our approach is capable of identifying task stragglers less than 11% into their execution lifecycle with 98% accuracy, signifying significant improvement over current state-of-the-art practice and enables far more effective mitigation strategies in large-scale distributed systems worldwide. Peter Garraghan, Xue Ouyang 0003, Paul Townend, Jie Xu 0007 |
ISORC | 3 |
| 2014 | Fault-Tolerant Dynamic Deduplication for Utility ComputingabstractUtility computing is an increasingly important paradigm, whereby computing resources are provided on-demand as utilities. An important component of utility computing is storage, data volumes are growing rapidly, and mechanisms to mitigate this growth need to be developed. Data deduplication is a promising technique for drastically reducing the amount of data stored in such system systems, however, current approachs are static in nature, using an amount of redundancy fixed at design time. This is inappropriate for truly dynamic modern systems. We propose a real-time adaptive deduplication system for Cloud and Utility computing that monitors in real-time for changing system, user, and environmental behaviour in order to fulfill a balance between changing storage efficiency, performance, and fault tolerance requirements. We evaluate our system through simulation, with experimental results showing that our system is both efficient and sclable. We also perform experimentation to evaluate the fault tolerance of the system by measuring Mean Time to Repair (MTTR), and using these values to calculate availability of the system. The results show that higher replication levels result in higher system availability, however, the number of files in the system also effects recovery time. We show that the tradeoff between replication levels and recovery time when the system overloads needs further investigation. Waraporn Leesakul, Paul Townend, Peter Garraghan, Jie Xu 0007 |
ISORC | 2 |
| 2014 | M-VCR: Multi-View Consensus Recognition for Real-Time ExperimentationabstractA major application area in the computer vision domain is gesture recognition, requiring real-time image classification to respond to human interactions. However, current state-of-the-art high-quality algorithms for image classification do not meet many dynamic real-time requirements. This paper presents the development of M-VCR - a novel approach for improving the reliability of real-time image classification. M-VCR increases the quality of classifications under real-time constraints through the adoption of fast classification algorithms, although these algorithms individually produce lower quality results, utilisation under a 'consensus' approach can achieve results equivalent to those of much higher-quality algorithms. The proposed approach also allows for different algorithms to be utilised in parallel, building on the fault tolerance technique of N-versioning. A significant improvement in image classification is experimentally demonstrated for both the SURF and MSER feature detectors through our integration consensus approach. This improvement is delivered entirely through the integration method without requiring modification of the source algorithms being used. David McKee 0001, Paul Townend, David Webster, Jie Xu 0007 |
ISORC | 2 |
| 2014 | Towards a Virtual Integration Design and Analysis Enviroment for Automotive EngineeringabstractAs the automotive industry moves towards reduced physical prototyping it is becoming more dependent on distributed simulations. However, the current technologies do not fully enable real-time distributed simulations which involve both virtual and physical components. This paper considers the current approaches to real-time distributed simulation and proposes the use of service-orientation. The highlights and current shortfalls of current research in real-time service orientation are then identified. Finally a key area of research is focused upon requiring a fundamental change of the understanding of quality of service and capability that is necessary to enable dependable real-time service orientated architectures. David McKee 0001, David Webster, Paul Townend, Jie Xu 0007, David Battersby |
ISORC | 3 |
| 2014 | Analysis, Modeling and Simulation of Workload Patterns in a Large-Scale Utility CloudabstractUnderstanding the characteristics and patterns of workloads within a Cloud computing environment is critical in order to improve resource management and operational conditions while Quality of Service (QoS) guarantees are maintained. Simulation models based on realistic parameters are also urgently needed for investigating the impact of these workload characteristics on new system designs and operation policies. Unfortunately there is a lack of analyses to support the development of workload models that capture the inherent diversity of users and tasks, largely due to the limited availability of Cloud tracelogs as well as the complexity in analyzing such systems. In this paper we present a comprehensive analysis of the workload characteristics derived from a production Cloud data center that features over 900 users submitting approximately 25 million tasks over a time period of a month. Our analysis focuses on exposing and quantifying the diversity of behavioral patterns for users and tasks, as well as identifying model parameters and their values for the simulation of the workload created by such components. Our derived model is implemented by extending the capabilities of the CloudSim framework and is further validated through empirical comparison and statistical hypothesis tests. We illustrate several examples of this work's practical applicability in the domain of resource management and energy-efficiency. Ismael Solís Moreno, Peter Garraghan, Paul Townend, Jie Xu 0007 |
IEEE Trans. Cloud Comput. | 3 |
| 2013 | An Analysis of the Server Characteristics and Resource Utilization in Google CloudabstractUnderstanding the resource utilization and server characteristics of large-scale systems is crucial if service providers are to optimize their operations whilst maintaining Quality of Service. For large-scale data enters, identifying the characteristics of resource demand and the current availability of such resources, allows system managers to design and deploy mechanisms to improve data enter utilization and meet Service Level Agreements with their customers, as well as facilitating business expansion. In this paper, we present a large-scale analysis of server resource utilization and a characterization of a production Cloud data enter using the most recent data enter trace logs made available by Google. We present their statistical properties, and a comprehensive coarse-grain analysis of the data, including submission rates, server classification, and server resource utilization. Additionally, we perform a fine-grained analysis to quantify the resource utilization of servers wasted due to the early termination of tasks. Our results show that data enter resource utilization remains relatively stable at between 40 - 60%, that the degree of correlation between server utilization and Cloud workload environment varies by server architecture, and that the amount of resource utilization wasted varies between 4.53 - 14.22% for different server architectures. This provides invaluable real-world empirical data for Cloud researchers in many subject areas. Peter Garraghan, Paul Townend, Jie Xu 0007 |
IC2E | 2 |
| 2013 | An evaluation framework for assessing the dependability of Dynamic Binding in Service-Oriented ComputingabstractService-Oriented Computing (SOC) provides a flexible framework in which applications may be built up from services, often distributed across a network. One of the promises of SOC is that of Dynamic Binding where abstract consumer requests are bound to concrete service instances at runtime, thereby offering a high level of flexibility and adaptability. Existing research has so far focused mostly on the design and implementation of dynamic binding operations and there is little research into a comprehensive evaluation of dynamic binding systems, especially in terms of system failure and dependability. In this paper, we present a novel, extensible evaluation framework that allows for the testing and assessment of a Dynamic Binding System (DBS). Based on a fault model specially built for DBS's, we are able to insert selectively the types of fault that would affect a DBS and observe its behavior. By treating the DBS as a black box and distributing the components of the evaluation framework we are not restricted to the implementing technologies of the DBS, nor do we need to be co-located in the same environment as the DBS under test. We present the results of a series of experiments, with a focus on the interactions between a real-life DBS and the services it employs. The results on the NECTISE Software Demonstrator (NSD) system show that our proposed method and testing framework is able to trigger abnormal behavior of the NSD due to interaction faults and generate important information for improving both dependability and performance of the system under test. Anthony Sargeant, Paul Townend, Jie Xu 0007, Karim Djemame |
ISORC | 2 |
| 2013 | A novel intrusion severity analysis approach for Clouds
Junaid Arshad, Paul Townend, Jie Xu 0007 |
Future Gener. Comput. Syst. | 2 |
| 2013 | Intrusion damage assessment for multi-stage attacks for cloudsabstractClouds represent a major paradigm shift from contemporary systems, inspiring the contemporary approach to computing. They present fascinating opportunities to address dynamic user requirements with the provision of flexible computing infrastructures that are available on demand. Clouds, however, introducing novel challenges particularly with respect to security that require dedicated efforts to address them. This study is focused at one such challenge, that is, determining the extent of damage caused by an intrusion for a victim virtual machine. It has significant implications especially with respect to effective response to the intrusion. This study presents the efforts to address this challenge for Clouds in the form of a novel scheme for intrusion damage assessment for Clouds. In addition to its context‐aware operation, the scheme facilitates protection against multi‐stage attacks. The study also includes the formal specification and evaluation of the scheme, which successfully demonstrate its effectiveness to achieve rigorous damage assessment for Clouds. Junaid Arshad, Muhammad Ajmal Azad, Imran Ali Jokhio, Paul Townend |
IET Commun. | 4 |
| 2012 | Interface Refactoring in Performance-Constrained Web ServicesabstractThis paper presents the development of REF-WS an approach to enable a Web Service provider to reliably evolve their service through the application of refactoring transformations. REF-WS is intended to aid service providers, particularly in a reliability and performance constrained domain as it permits upgraded 'non-backwards compatible' services to be deployed into a performance constrained network where existing consumers depend on an older version of the service interface. In order for this to be successful, the refactoring and message mediation needs to occur without affecting functional compatibility with the services' consumers, and must operate within the performance overhead expected of the original service, introducing as little latency as possible. Furthermore, compared to a manually programmed solution, the presented approach enables the service developer to apply and parameterize refactorings with a level of confidence that they will not produce an invalid or 'corrupt' transformation of messages. This is achieved through the use of preconditions for the defined refactorings. David Webster, Paul Townend, Jie Xu 0007 |
ISORC | 2 |
| 2009 | Quantification of Security for Compute Intensive Workloads in CloudsabstractCloud computing is a promising technology to facilitate development of large-scale, on-demand, flexible computing infrastructures. However, improving dependability of cloud computing is critical for realization of its potential. In this paper, we describe our efforts to quantify security for Clouds to facilitate provision of assurance for quality of service, one of the factors contributing to dependability. This has profound implications for delivering customized security solutions such as effective intrusion prevention and detection which is the overall objective of our research. In order to demonstrate the applicability of our research, we have incorporated these requirements in the resource acquisition phase for Clouds. We also present experiments to demonstrate the effectiveness of our approach to address the random migration problem for virtualized computing environments. Junaid Arshad, Paul Townend, Jie Xu 0007 |
ICPADS | 2 |
| 2009 | Efficient and scalable search on scale-free P2P networks
Lu Liu 0001, Jie Xu 0007, Duncan Russell, Paul Townend, David Webster |
Peer-to-Peer Netw. Appl. | 4 |
| 2008 | FT-Grid: a system for achieving fault tolerance in gridsabstractAbstract The FT‐Grid system introduces a fault‐tolerance framework that allows faults occurring in service‐oriented systems to be tolerated, thus increasing the dependability of such systems. This paper presents the design, development and evaluation of FT‐Grid. We show empirical evidence of the dependability benefits offered by FT‐Grid by performing an experimental dependability analysis using fault‐injection testing performed with the WS‐FIT tool. We then illustrate a potential problem with voting‐based fault‐tolerance schemes in the service‐oriented paradigm—namely that individual channels within a fault‐tolerant system, supposed to be independent of each other, may in fact invoke common services as part of their workflow, thus increasing the potential for common‐mode failure of those channels. We propose a solution to this issue by using the technique of provenance to provide FT‐Grid with topological awareness. We implement a large experimental system, and—with the use of the Provenance Recording for Services system developed as part of the PASOA project at the University of Southampton—perform a large number of experiments that show that a topologically aware FT‐Grid system serves as a much more dependable system than any other configuration tested, while imposing a negligible timing overhead. Copyright © 2007 John Wiley & Sons, Ltd. Jie Xu 0007, Paul Townend, Nik Looker, Paul Groth |
Concurr. Comput. Pract. Exp. | 2 |
| 2005 | A Provenance-Aware Weighted Fault Tolerance Scheme for Service-Based ApplicationsabstractService-orientation has been proposed as a way of facilitating the development and integration of increasingly complex and heterogeneous system components. However, there are many new challenges to the dependability community in this new paradigm, such as how individual channels within fault-tolerant systems may invoke common services as part of their workflow, thus increasing the potential for common-mode failure. We propose a scheme that - for the first time - links the technique of provenance with that of multi-version fault tolerance. We implement a large test system and perform experiments with a single-version system, a traditional MVD system, and a provenance-aware MVD system, and compare their results. We show that for this experiment, our provenance-aware scheme results in a much more dependable system than either of the other systems tested, whilst imposing a negligible timing overhead. Paul Townend, Paul Groth, Jie Xu 0007 |
ISORC | 1 |