Paul Watson 0001

dblp:22/3634-1 · DBLP profile ↗
← Back
64ranked-venue papers
8as first author
7since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 22 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 14 · 2 first-author · 6 since 2021Software engineering, systems software and programming languages · 13 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 8 · 1 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Computer networks · 2Security and privacy · 1Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2024 Data Efficient Transformers for Wearable Sensor Analysis in Centralized and Federated Environments
abstract
Transformers have rapidly become the dominant architecture for analyzing sequential data, utilizing their self-attention mechanism to effectively capture long-term temporal patterns, outperforming recurrent-based methods across various applications. In this paper, we explore the application of transformers to wearable sensor data, focusing on the analysis of human gait, which is often complex and sensitive. We propose two novel frameworks: Data Efficient Sensor Transformer (DesT) for centralized learning and Federated Data Efficient Sensor Transformer (FeDesT) for federated learning (FL) in edge-computing environments. Both frameworks employ knowledge distillation to improve the generalization of transformers, which can be prone to over-fitting due to the limited labeled data available in wearable sensor applications. Experimental results using human gait data collected from uneven and irregular surfaces show that DesT improves the accuracy by 14.8% when compared to existing transformers. FeDesT reduces computational demands on edge devices while outperforming traditional FL methods for transformers. This work demonstrates the potential of transformers for wearable sensor data analysis in both centralized and federated contexts, particularly where privacy and computational efficiency is paramount.
Jamie McQuire, Paul Watson 0001, Nicholas G. Wright, Hugo Hiden, Michael Catt
IEEE Big Data2
2023 Mobilise-D: Experiences of Processing Large Medical Data Sets Using Cloud Computing Resources
abstract
This poster describes the data analysis infrastructure used in a large EU medical project, describes our experiences and lessons learned and presents our techniques for improving the reliability of the analysis procedure.
Hugo Hiden, Paul Watson 0001
e-Science2
2023 From eScience to Impact on the Economy and Society
abstract
eScience has had a major impact on science, engineering, healthcare and (increasingly) the humanities. However, there is enormous, largely untapped, potential to use the skills and expertise of eScience researchers to have a major impact on the economy and society. All organisations are drowning in data, but most lack the skills to extract value from it to gain insights into their work, to improve productivity and to create new profitable data-driven products and services. In contrast, eScience researchers have highly relevant experience in solving real-world problems by unlocking the value held in data. In this paper we suggest why this potential has not been fully realised and describe our efforts to address this through the creation of the UK's National Innovation Centre for Data. We draw on our experiences based on discussions with over 160 organisations, and of running 80 collaborative projects with businesses and governmental organisations over the past 4 years. Our vision for the future of eScience is of a field that is not only driving innovative research, but is also having a major, positive impact on the economy and society.
Paul Watson 0001, Barry Hodgson
e-Science1
2022 Generating Synthetic Images & Data to Improve Object Detection in CCTV Footage from Public Transport
abstract
Passenger behaviour on public transport has become a source of great interest in the wake of the COVID-19 pandemic. Operators are interested in employing new methods to monitor vehicle utilisation and passenger behaviour. One way to do this is through the use of Machine Learning, using the CCTV footage that is already being captured from the vehicles. However, one of the limitations of Machine Learning is that it requires large amounts of annotated training data, which is not always available. In this poster, we present a technique that uses 3D models to generate synthetic training images/data and discuss the effect that training with the synthetic data had on the Machine Learning models when applied to real-world CCTV footage.
Mike Simpson, Nik Khadijah Nik Aznan, John Brennan, Paul Watson 0001, Philip James 0002, Jennine Jonczyk
e-Science4
2022 The e-Science Central Study Data Platform
abstract
Presentation for IEEE e-Science 20202
Paul Watson 0001, Hugo Hiden
e-Science1
2022 On the Complexity of Object Detection on Real-world Public Transportation Images for Social Distancing Measurement
abstract
Social distancing in public spaces has become an essential aspect in helping to reduce the impact of the COVID-19 pandemic. Exploiting recent advances in machine learning, there have been many studies in the literature implementing social distancing via object detection through the use of surveillance cameras in public spaces. However, there has been no study of social distance measurement on public transport to date. The public transport setting has some unique challenges, including low-resolution images and physical camera locations that can lead to the partial occlusion of passengers, making it challenging to perform accurate detection. Thus, this paper investigates the challenges of performing accurate social distance measurements on public transportation. We benchmark several state-of-the-art object detection algorithms using real-world footage taken from the London Underground and bus network. The work highlights the complexity of performing social distancing measurements on images from current public transportation onboard cameras. Further, exploiting domain knowledge of expected passenger behaviour, we attempt to improve the quality of the detections using various strategies and show improvement over using vanilla object detection alone.
Nik Khadijah Nik Aznan, John Brennan, Daniel Bell, Jennine Jonczyk, Paul Watson 0001
IJCNN5
2021 Uneven and Irregular Surface Condition Prediction from Human Walking Data using both Centralized and Decentralized Machine Learning Approaches
abstract
Gait data collected using wearable sensors offers non-intrusive, affordable, real-time monitoring of human motion. Recognizing surface conditions from wearable sensor data has the potential to help systems discriminate between ‘poor quality’ walking data. This research investigates the predictive capabilities of machine learning models, trained on both centralized and decentralized datasets, at categorizing uneven and irregular surface conditions. The results showed that machine learning classification algorithms, trained with data originating from a single sensor positioned on the left-shank, were able to accurately discriminate between different types of surface conditions. We found the Support Vector Machine, when trained with the data centralized, had a test-set accuracy of 94%. Federated Learning offers a way to increase privacy and security for healthcare applications by avoiding the centralization of data. Our simulated federated Deep Neural Network converged to a test-accuracy of 85%, which was 8% less than the centralized counterpart.
Jamie McQuire, Paul Watson 0001, Nicholas G. Wright, Hugo Hiden, Michael Catt
BIBM2
2020 Dynamically Partitioning Workflow over Federated Clouds for Optimising the Monetary Cost and Handling Run-Time Failures
abstract
Several real-world problems in domain of healthcare, large scale scientific simulations, and manufacturing are organised as workflow applications. Efficiently managing workflow applications on the Cloud computing data-centres is challenging due to the following problems: (i) they need to perform computation over sensitive data (e.g., Healthcare workflows) hence leading to additional security and legal risks especially considering public cloud environments and (ii) the dynamism of the cloud environment can lead to several run-time problems such as data loss and abnormal termination of workflow task due to failures of computing, storage, and network services. To tackle above challenges, this paper proposes a novel workflow management framework call Deploy on Federated Cloud Framework (DoFCF) that can dynamically partition scientific workflows across federated cloud (public/private) data-centres for minimising the financial cost, adhering to security requirements, while gracefully handling run-time failures. The framework is validated in cloud simulation tool (CloudSim) as well as in a realistic workflow-based cloud platform (e-Science Central). The results showed that our approach is practical and is successful in meeting users security requirements and reduces overall cost, and dynamically adapts to the run-time failures.
Zhenyu Wen, Rawaa Qasha, Zequn Li 0002, Rajiv Ranjan 0001, Paul Watson 0001, Alexander B. Romanovsky
IEEE Trans. Cloud Comput.5
2020 Multiobjective Deployment of Data Analysis Operations in Heterogeneous IoT Infrastructure
abstract
The growth of Internet of Things (IoT) technology brings many new opportunities for applications in areas including smart healthcare, smart buildings, and smart agriculture. These applications must normally distribute the computations, required for extracting value from sensor data, over the IoT infrastructure platforms (e.g., sensors, phones, field-gateways, and clouds). This can be very challenging for IoT application developers due to the heterogeneity of the aforementioned platforms, potentially conflicting nonfunctional requirements (e.g., battery power, latency, and cost), and related deployment criteria, which is impossible to resolve manually. To address the above challenges, we have developed the PATH2iot framework that decomposes a complex IoT application into self-contained micro-operations. Based on the deployment criteria, PATH2iot automatically distributes the set of micro-operations across IoT infrastructure platforms, while respecting their run-time data and control flow dependencies. In our previous work, we have shown how to use the PATH2iot to optimize the battery life of a healthcare wearable. In this article, we describe a new research that significantly extends PATH2iot, which introduces a heuristic model capable of making optimal deployment decisions based on multiple conflicting nonfunctional requirements and selection criteria (user preferences). It does so by leveraging a well-known multicriteria decision-making method called the analytic hierarchical processes (AHP). The applicability of the deployment model is validated based on a real-world digital healthcare analytics use case. The results show that our model is able to find the optimal deployment solution for different user preferences.
Devki Nandan Jha, Peter Michalák, Zhenyu Wen, Rajiv Ranjan 0001, Paul Watson 0001
IEEE Trans. Ind. Informatics5
2019 Sharing and performance optimization of reproducible workflows in the cloud
Rawaa Qasha, Zhenyu Wen, Jacek Cala, Paul Watson 0001
Future Gener. Comput. Syst.4
2019 A note on tools and techniques for end-to-end QoS monitoring in Internet of Things
Rajiv Ranjan 0001, Ellis Solaiman, Massimo Villari, Paul Watson 0001
J. Parallel Distributed Comput.4
2018 Automating the Placement of Time Series Models for IoT Healthcare Applications
abstract
There has been a dramatic growth in the number and range of Internet of Things (IoT) sensors that generate healthcare data. These sensors stream high-dimensional time series data that must be analysed in order to provide the insights into medical conditions that can improve patient healthcare. This raises both statistical and computational challenges, including where to deploy the streaming data analytics, given that a typical healthcare IoT system will combine a highly diverse set of components with very varied computational characteristics, e.g. sensors, mobile phones and clouds. Different partitionings of the analytics across these components can dramatically affect key factors such as the battery life of the sensors, and the overall performance. In this work we describe a method for automatically partitioning stream processing across a set of components in order to optimise for a range of factors including sensor battery life and communications bandwidth. We illustrate this using our implementation of a statistical model predicting the glucose levels of type II diabetes patients in order to reduce the risk of hyperglycaemia.
Lauren Roberts, Peter Michalák, Sarah E. Heaps, Michael Trenell, Darren J. Wilkinson, Paul Watson 0001
eScience6
2018 Enabling Edge Intelligence for Activity Recognition in Smart Homes
abstract
In recent years, Edge computing has emerged as a new paradigm that can reduce communication delays over the Internet by moving computation power from far-end cloud servers to be closer to data sources. It is natural to shift the design of cloud-based IoT applications to Edge-based ones. Activity recognition in smart homes is one of the IoT applications that can benefit significantly from such a shift. In this work, we propose an Edge-based solution for addressing the activity recognition problem in smart homes from multiple perspectives, including architecture, algorithm design and system implementation. First, the Edge computing architecture is introduced and several critical management tasks are also investigated. Second, a realization of the Edge computing system is presented by using open source software and low-cost hardware. The consistency and scalability of running jobs on Edge devices are also addressed in our approach. Last, we propose a convolutional neural network model to perform activity recognition tasks on Edge devices. Preliminary experiments are conducted to compare our model with existing machine learning methods, and the results demonstrate that the performance of our model is promising.
Shaojun Zhang, Wei Li 0058, Yongwei Wu 0001, Paul Watson 0001, Albert Y. Zomaya
MASS4
2017 PATH2iot: A Holistic, Distributed Stream Processing System
abstract
The PATH2iot open-source platform presents a new approach to stream processing for Internet of Things applications by automatically partitioning and deploying the computation over the available infrastructure (e.g. cloud, field gateways and sensors) in order to meet non-functional requirements including energy, performance and security. The user gives a high-level declarative description of computation in the form of Event Processing Language queries. These are compiled, optimised, and partitioned to meet the non-functional requirements using database system techniques and cost models extended to meet the needs of IoT analytics. The paper describes the PATH2iot system, illustrated by a real-world digital healthcare analytics example, with sensor battery life as the main non-functional requirement to be optimised. It shows that the tool can automatically partition and distribute the computation across a healthcare wearable, a mobile phone and the cloud - increasing the battery life of the smart watch by 416% when compared to other possible allocations. The PATH2iot system can therefore automatically bring the benefits of fog/edge computing to IoT applications.
Peter Michalák, Paul Watson 0001
CloudCom2
2017 A Platform for the Analysis of Qualitative and Quantitative Data about the Built Environment and Its Users
abstract
There are many scenarios in which it is necessary to collect data from multiple sources in order to evaluate a system, including the collection of both quantitative data - from sensors and smart devices - and qualitative data - such as observations and interview results. However, there are currently very few systems that enable both of these data types to be combined in such a way that they can be analysed side-by-side. This paper describes an end-to-end system for the collection, analysis, storage and visualisation of qualitative and quantitative data, developed using the e-Science Central cloud analytics platform. We describe the experience of developing the system, based on a case study that involved collecting data about the built environment and its users. In this case study, data is collected from older adults living in residential care. Sensors were placed throughout the care home and smart devices were issued to the residents. This sensor data is uploaded to the analytics platform and the processed results are stored in a data warehouse, where it is integrated with qualitative data collected by healthcare and architecture researchers. Visualisations are also presented which were intended to allow the data to be explored and for potential correlations between the quantitative and qualitative data to be investigated.
Mike Simpson, Simon Woodman, Hugo Hiden, Sebastian Stein 0003, Stephen Dowsland, Mark Turner 0007, Vicki L. Hanson, Paul Watson 0001
eScience8
2017 Applications of provenance in performance prediction and data storage optimisation
Simon Woodman, Hugo Hiden, Paul Watson 0001
Future Gener. Comput. Syst.3
2017 Privacy-Aware Scheduling SaaS in High Performance Computing Environments
abstract
Hybrid clouds have gained popularity in recent times in a variety of organizations due to their ability to provide additional capacity in a public cloud, to augment private cloud capacity, when it is needed. However, scheduling distributed applications' jobs (e.g, workflow tasks) on hybrid cloud resources introduces new challenges. One key problem is the danger of exposing private data and jobs in a third-party public cloud infrastructure, for example in healthcare applications. In this article, we tackle the problem of designing workflow scheduling algorithms to meet customers' deadlines, while not compromising data and task privacy requirements. Our work is different from most studies on workflow scheduling where the main goal is to achieve a balance between desirable, yet incompatible constraints, such as meeting the deadline and/or minimizing the execution time. Although many others have addressed the trade-off between cost and time, or privacy and cost, their work still suffers from an insufficient consideration of the trade-off between privacy and time. To address such shortcomings in the literature, we present a new SaaS scheduling broker composed of MPHC-P1, MPHCP2, and MPHC-P3 policies to preserve privacy while scheduling the workflows' tasks under customers' deadlines. We evaluated our approach using real workflows running on a VMware based hybrid cloud. Results demonstrate that under our scheduling policies, MPHC-P2 and MPHC-P3 are promising in time-critical scenarios by reducing the total cost by 10-20 percent compared to alternatives. Overall, results show that our approach is efficient in reducing the cost of executing workflows while satisfying both their privacy and deadline constraints.
Shaghayegh Sharif, Paul Watson 0001, Javid Taheri, Surya Nepal, Albert Y. Zomaya
IEEE Trans. Parallel Distributed Syst.2
2017 Cost Effective, Reliable and Secure Workflow Deployment over Federated Clouds
abstract
The significant growth in cloud computing has led to increasing number of cloud providers, each offering their service under different conditions - one might be more secure whilst another might be less expensive or more reliable. At the same time user applications have become more and more complex. Often, they consist of a diverse collection of software components, and need to handle variable workloads, which poses different requirements on the infrastructure. Therefore, many organisations are considering using a combination of different clouds to satisfy these needs. It raises, however, a non-trivial issue of how to select the best combination of clouds to meet the application requirements. This paper presents a novel algorithm to deploy workflow applications on federated clouds. First, we introduce an entropy-based method to quantify the most reliable workflow deployments. Second, we apply an extension of the Bell-LaPadula Multi-Level security model to address application security requirements. Finally, we optimise deployment in terms of its entropy and also its monetary cost, taking into account the cost of computing power, data storage and inter-cloud communication. We implemented our new approach and compared it against two existing scheduling algorithms: Extended Dynamic Constraint Algorithm (EDCA) and Extended Biobjective dynamic level scheduling (EBDLS). We show that our algorithm can find deployments that are of equivalent reliability but are less expensive and meet security requirements. We have validated our solution through a set of realistic scientific workflows, using well-known cloud simulation tools (WorkflowSim and DynamicCloudSim) and a realistic cloud based data analysis system (e-Science Central).
Zhenyu Wen, Jacek Cala, Paul Watson 0001, Alexander B. Romanovsky
IEEE Trans. Serv. Comput.3
2016 Dynamic Deployment of Scientific Workflows in the Cloud Using Container Virtualization
abstract
Scientific workflows are increasingly being migrated to the Cloud. However, workflow developers face the problem of which Cloud to choose and, more importantly, how to avoid vendor lock-in. This is because there are a range of Cloud platforms, each with different functionality and interfaces. In this paper we propose a solution - a system that allows workflows to be portable across a range of Clouds. This portability is achieved through a new framework for building, dynamically deploying and enacting workflows. It combines the TOSCA specification language and container-based virtualization. TOSCA is used to build a reusable and portable description of a workflow which can be automatically deployed and enacted using Docker containers. We describe a working implementation of our framework and evaluate it using a set of existing scientific workflows that illustrate the flexibility of the proposed approach.
Rawaa Qasha, Jacek Cala, Paul Watson 0001
CloudCom3
2016 Prediction of workflow execution time using provenance traces: Practical applications in medical data processing
abstract
The use of cloud resources for processing and analysing medical data has the potential to revolutionise the treatment of a number of chronic conditions. For example, it has been shown that it is possible to manage conditions such as diabetes, obesity and cardiovascular disease by increasing the right forms of physical activity for the patient. Typically, movement data is collected for a patient over a period of several weeks using a wrist worn accelerometer. This data, however, is large and its analysis can require significant computational resources. Cloud computing offers a convenient solution as it can be paid for as needed and is capable of scaling to store and process large numbers of data sets simultaneously. However, because the charging model for the cloud represents, to some extent, an unknown cost and therefore risk to project managers, it is important to have an estimate of the likely data processing and storage costs that will be required to analyse a set of data. This could take the form of data collected from a patient in clinic or of entire cohorts of data collected from large studies. If, however, an accurate model was available that could predict the compute and storage requirements associated with a piece of analysis code, decisions could be made as to the scale of resources required in order to obtain results within a known timescale. This paper makes use of provenance and performance data collected as part of routine e-Science Central workflow executions to examine the feasibility of automatically generating predictive models for workflow execution times based solely on observed characteristics such as data volumes processed, algorithm settings and execution durations. The utility of this approach will be demonstrated via a set of benchmarking examples before being used to model workflow executions performed as part of two large medical movement analysis studies.
Hugo Hiden, Simon Woodman, Paul Watson 0001
eScience3
2016 A framework for scientific workflow reproducibility in the cloud
abstract
Workflow is a well-established means by which to capture scientific methods in an abstract graph of interrelated processing tasks. The reproducibility of scientific workflows is therefore fundamental to reproducible e-Science. However, the ability to record all the required details so as to make a workflow fully reproducible is a long-standing problem that is very difficult to solve. In this paper, we introduce an approach that integrates system description, source control, container management and automatic deployment techniques to facilitate workflow reproducibility. We have developed a framework that leverages this integration to support workflow execution, re-execution and reproducibility in the cloud and in a personal computing environment. We demonstrate the effectiveness of our approach by examining various aspects of repeatability and reproducibility on real scientific workflows. The framework allows workflow and task images to be captured automatically, which improves not only repeatability but also runtime performance. It also gives workflows portability across different cloud environments. Finally, the framework can also track changes in the development of tasks and workflows to protect them from unintentional failures.
Rawaa Qasha, Jacek Cala, Paul Watson 0001
eScience3
2016 Provenance and data differencing for workflow reproducibility analysis
abstract
Summary One of the foundations of science is that researchers must publish the methodology used to achieve their results so that others can attempt to reproduce them. This has the added benefit of allowing methods to be adopted and adapted for other purposes. In the field of e‐Science, services – often choreographed through workflow, process data to generate results. The reproduction of results is often not straightforward as the computational objects may not be made available or may have been updated since the results were generated. For example, services are often updated to fix bugs or improve algorithms. This paper addresses these problems in three ways. Firstly, it introduces a new framework to clarify the range of meanings of ‘reproducibility’. Secondly, it describes a new algorithm, PDIFF, that uses a comparison of workflow provenance traces to determine whether an experiment has been reproduced; the main innovation is that if this is not the case then the specific point(s) of divergence are identified through graph analysis, assisting any researcher wishing to understand those differences. One key feature is support for user‐defined, semantic data comparison operators. Finally, the paper describes an implementation of PDIFF that leverages the power of the e‐Science Central platform that enacts workflows in the cloud. As well as automatically generating a provenance trace for consumption by PDIFF, the platform supports the storage and reuse of old versions of workflows, data and services; the paper shows how this can be powerfully exploited to achieve reproduction and reuse. Copyright © 2013 John Wiley & Sons, Ltd.
Paolo Missier, Simon Woodman, Hugo Hiden, Paul Watson 0001
Concurr. Comput. Pract. Exp.4
2016 Formal verification of secure information flow in cloud computing
Maciej Koutny, Paul Watson 0001, Vasileios Germanos
J. Inf. Secur. Appl.3
2015 Towards Automated Workflow Deployment in the Cloud Using TOSCA
abstract
Scientific workflows play an increasingly important role in building scientific applications, while cloud computing provides on-demand access to large compute resources. Combining the two offers the potential to increase dramatically the ability to quickly extract new results from the vast amounts of scientific data now being collected. However, with the proliferation of cloud computing platforms and workflow management systems, it becomes more and more challenging to define workflows so they can reliably run in the cloud and be reused easily. This paper shows how TOSCA, a new standard for cloud service management, can be used to systematically specify the components and life cycle management of scientific workflows by mapping the basic elements of a real workflow onto entities specified by TOSCA. Ultimately, this will enable workflow definitions that are portable across clouds, resulting in the greater reusability and reproducibility of workflows.
Rawaa Qasha, Jacek Cala, Paul Watson 0001
CLOUD3
2015 Cost Effective, Reliable, and Secure Workflow Deployment over Federated Clouds
abstract
The federation of clouds can provide benefits for cloud-based applications. Different clouds have different advantages - one might be more reliable whilst another might be more secure or less expensive. However, being able to select the best combination of clouds to meet the application requirements is not trivial. This paper presents a novel algorithm to deploy workflow applications on federated clouds. Firstly, we introduce an entropy-based method to quantify the most reliable workflow deployments. Secondly, we apply an extension of the Bell-LaPadula Multi-Level security model to meet application security requirements. Finally, we optimise deployment in terms of its entropy and also its monetary cost, taking into account the price of computing power, data storage and inter-cloud communication. To evaluate the new algorithm we compared it against two existing scheduling algorithms: Dynamic Constraint Algorithm (DCA) and Biobjective dynamic level scheduling (BDLS). We show that our algorithm can find deployments that are of equivalent reliability, but are less expensive and also meet security requirements. We have validated our solution using workflows implemented in the e-Science Central cloud-based data analysis system.
Zhenyu Wen, Jacek Cala, Paul Watson 0001, Alexander B. Romanovsky
CLOUD3
2015 Monitoring of Upper Limb Rehabilitation and Recovery after Stroke: An Architecture for a Cloud-Based Therapy Platform
abstract
Amongst the therapies available to stroke sufferers, one that is gaining attention is the application of video games to encourage therapeutic movement. The Limbs Alive project at Newcastle University has developed a system that gathers therapeutic game data from patients, uses statistical tools to estimate a number of performance metrics and presents the results to patients and clinicians via web applications. This paper describes the architecture of this system and outlines the various technical challenges that were overcome, including in security and deployment.
Simon Woodman, Hugo Hiden, Mark Turner 0007, Stephen Dowsland, Paul Watson 0001
e-Science5
2015 Antares: A Scalable, Real-Time, Fault Tolerant Data Store for Spatial Analysis
abstract
The growth of mobile devices has significantly increased the velocity and volume of location-based data. Whilst there is enormous potential for applications that exploit this data in real-time, storing and querying it in real-time creates significant challenges. Traditional RDBMS systems are not sufficiently scalable, while typical cloud-based solutions such as map-reduce do not possess the capabilities required for real-time, spatial-data processing. Therefore, new approaches are needed. In this paper we explore the use of NoSQL technologies. These offer scalability, availability and fault tolerance, but -- as we show -- do not perform well with spatial data. Therefore, in this paper we address this challenge by enhancing existing spatial indexing structures with novel algorithms for inserting and searching spatial data. We have implemented this in a NoSQL solution (Antares), and evaluated it against two other NoSQL solutions, and a range of indexing structures: Kd-Tree, Quad Tree and Geohashing. The results show that Antares significantly outperforms the other approaches.
Rebecca Simmonds, Paul Watson 0001, Jonathan Halliday
SERVICES2
2014 Multi-level Security for Deploying Distributed Applications on Clouds, Devices and Things
abstract
The deployment of the components of distributed systems is now often very dynamic - server-side components are virtualised so they can be dynamically deployed on a range of platforms including public and private clouds, while users expect to be able to install clients on devices from phones to tablets. This can introduce security problems that place data at risk. This paper describes a new method for modeling the security of a distributed application and generating the set of possible deployment options that meet the overall security requirements. The model encompasses the entities that influence the security of a distributed system: data, services networks and platforms (e.g. Clouds, devices and "things"). The paper describes the method and how it can be used to answer a range of security questions, using a set of case studies including federated clouds and mobile clients.
Paul Watson 0001, Mark C. Little
CloudCom1
2014 A Scalable Method for Partitioning Workflows with Security Requirements over Federated Clouds
abstract
The significant increase in the use of cloud computing, has led to an interest in partitioning applications over a set of public and private clouds in order to meet a range of non-functional requirements including performance (for example where private cloud resources alone are insufficient), dependability (e.g. To allow the application to continue to operate even if one cloud fails) and security (for example to ensure that sensitive data is restricted to sufficiently secure clouds and networks). This paper describes a novel deployment planning algorithm to partition complex workflow-based applications over federated clouds, while meeting security requirements. The security issues are based on our previous work which extends the Bell-La Padula model to encompass cloud computing. Selecting the cheapest option for partitioning a workflow over a set of resources has been shown to be an NP-hard problem, which can take impractically long for partitioning large workflows over multiple clouds. We therefore introduce a novel adaptive partitioning algorithm to handle these large workflow applications, which significantly reduces the time required to choose a sufficiently good partitioning option. This is based on generating an initial partitioning, and then adapting it to see if a better solution can be found by bringing together on the same node services with significant communication costs. The algorithm has been implemented and evaluated by using both randomly generated and real world scientific workflows. The experiment results show that our algorithm is thousand times quicker than the exhaustive algorithm presented in our previous work. Yet, on average it generates only 25% more costly solutions. We also compared this algorithm with two other methods commonly used to partition workflows over a set of clouds.
Zhenyu Wen, Jacek Cala, Paul Watson 0001
CloudCom3
2014 Verifying Secure Information Flow in Federated Clouds
abstract
Federated cloud systems increase the reliability and reduce the cost of computational support to an organization. However, the resulting combination of secure private clouds and less secure public clouds impacts on the security requirements of the system. Therefore, applications need to be located within different clouds, which strongly affects the information flow security of the entire system. In this paper, the entities of a federated cloud system as well as the clouds are assigned security levels of a given security lattice. Then a dynamic flow sensitive security model for a federated cloud system is proposed within which the Bell-La Padula rules and cloud security rule can be captured. As a result, one can track and verify the security information flow in federated clouds. Moreover, an example is used to explain how Petri nets could be used to represent such a system, making it possible to verify secure information flow in federated clouds using the existing Petri net techniques.
Maciej Koutny, Paul Watson 0001
CloudCom3
2014 A Platform for Analysing Stream and Historic Data with Efficient and Scalable Design Patterns
abstract
Social media is an increasingly popular method for people to share information and interact with each other. Analysis of social media data has the potential to provide useful insights in a wide range of domains including social science, advertising and policing. Social media information is produced in real-time, and so analysis that can give insights into events as they occur can be particularly valuable. Similarly, analytics platforms providing low latency query responses can improve the user experience for ad-hoc data exploration on historic data sets. However, the rate at which new data is generated makes it a real challenge to design a system that can meet both of these challenges. This paper describes the deisgn and evaluation of such a system. Firstly, it describes how a meta-analysis of the types of questions that were being asked of Twitter data led to the identification of a small set of queries that could be used to answer the majority of them. Secondly, it describes the design of a scalable platform for answering these and other queries. The architecture is described: it is cloud-based, and combines both continuous query, and noSQL database technology. Evaluation results are presented which show that the system can scale to process queries on streaming data arriving at the rate of the full Twitter firehose. Experiments show that queries on large repositories of stored historic data can also be answered with low latency. Finally, we present the results of queries that combine both streaming and historic data.
Rebecca Simmonds, Paul Watson 0001, Jonathan Halliday, Paolo Missier
SERVICES2
2013 Dynamic Exception Handling for Partitioned Workflow on Federated Clouds
abstract
The aim of federated cloud computing is to allow applications to utilise a set of clouds in order to provide a better combination of properties, such as cost, security, performance and dependability, than can be achieved on a single cloud. In this paper we focus on security and dependability: introducing a new automatic method for dynamically partitioning applications across the set of clouds in an environment in which clouds can fail during workflow execution. The method deals with exceptions that occur when clouds fail, and selects the best way to repartition the workflow, whilst still meeting security requirements. This avoids the need for developers to have to code ad-hoc solutions to address cloud failure, or the alternative of simply accepting that an application will fail when a cloud fails. This paper's method builds on earlier work [1] on partitioning workflows over federated clouds to minimise cost while meeting security requirements. It extends it by pre-generating the graph of all possible ways to partition the workflow, and adding weights to the paths through the graph so that when a cloud fails, it is possible to quickly determine the cheapest possible way to make progress from that point to the completion of the workflow execution (if any path exists). The method has been implemented and evaluated through a tool which exploits e-Science Central: a portable, high-level cloud platform. The workflow application is created and distributed across a set of e-Science Central instances. By monitoring the state of each executing e-Science Central instance, the system handles exceptions as they occur at run-time. The paper describes the method and an evaluation that utilises a set of examples.
Zhenyu Wen, Paul Watson 0001
CloudCom (1)2
2013 Cloud computing for fast prediction of chemical activity
Jacek Cala, Hugo Hiden, Simon Woodman, Paul Watson 0001
Future Gener. Comput. Syst.4
2012 Formalising Workflows Partitioning over Federated Clouds: Multi-level Security and Costs
abstract
We present an abstract formalisation of federated cloud workflows using the Z notation. Various properties of interest are observed in the possible deployments symbolically calculated by the Z/EVES theorem prover. These properties are defined using rules restricting valid options in various categories like security, cost, dependability, etc.
Leo Freitas, Paul Watson 0001
SERVICES2
2012 Case for dynamic deployment in a grid-based distributed query processor
Arijit Mukherjee, Paul Watson 0001
Future Gener. Comput. Syst.2
2011 A Multi-Level Security Model for PartitioningWorkflows over Federated Clouds
abstract
Cloud computing has the potential to provide low cost, scalable computing, but cloud security is a major area of concern. Many organizations are therefore considering using a combination of a secure internal cloud, along with (what they perceive to be) less secure public clouds. However, this raises the issue of how to partition applications across a set of clouds, while meeting security requirements. Currently, this is usually done on an ad-hoc basis, which is potentially error prone, or for simplicity the whole application is deployed on a single cloud, so removing the possible performance and availability benefits of exploiting multiple clouds within a single application. This paper describes an alternative to ad-hoc approaches a method that determines all ways in which applications structured as workflows can be partitioned over the set of available clouds such that security requirements are met. The approach is based on a Multi-Level Security model that extends Bell-LaPadula to encompass cloud computing. This includes introducing workflow transformations that are needed where data is communicated between clouds. In specific cases these transformations can result in security breaches, but the paper describes how these can be detected. Once a set of valid options has been generated, a cost model is used to rank them. The method has been implemented in a tool, which is briefly described in the paper.
Paul Watson 0001
CloudCom1
2010 Automatic Software Deployment in the Azure Cloud
Jacek Cala, Paul Watson 0001
DAIS2
2010 A Tripartite Security Model for Dynamic Service-Oriented Systems Using DynaSOAr
abstract
The DynaSOAr framework presents a wholly service-oriented approach to grid and Internet-based computing that makes a clear and explicit separation of concerns between service-provision and resource-provision for each service invocation. The separation allows the dynamic deployment of code at runtime, in the form of a service implementation, between a service provider and an explicit resource provider. This paper presents work in progress towards an integrated tripartite security model and framework that enables the security constraints of each engaging party to be expressed, propagated, unified and enforced as part of a DynaSOAr service invocation.
Chris Fowler, Paul Watson 0001
ICWS2
2010 e-Science Central for CARMEN: science as a service
abstract
Abstract Scientists face many severe challenges in extracting value from the increasingly large volumes of data they generate. In this paper we describe the requirements we have derived from working across a wide range of e‐science projects. In particular, the CARMEN neuroinformatics project has exposed a range of challenges due to a need to analyse and share large volumes of data. We have identified the four key activities required by scientists with whom we work, and designed an integrated system—e‐Science Central—to provide them. This exploits three emerging technologies: software as a service to avoid the need for users to deploy and maintain any of their own software; social networking to allow users to collaborate by sharing data, services and workflows in a controlled manner and Cloud computing to provide scalable compute resources. The system can not only be used through any web browser, but also provides an API so that applications can build on the core functionality. We describe the requirements, and the design that flows from them. This includes data storage with in‐built versioning and signing, an in‐browser workflow editor and a job scheduling system that allows workflows to be run both on local ‘private’ clouds and the Microsoft Azure Cloud. Copyright © 2010 John Wiley & Sons, Ltd.
Paul Watson 0001, Hugo Hiden, Simon Woodman
Concurr. Comput. Pract. Exp.1
2009 Adaptive workload allocation in query processing in autonomous heterogeneous environments
Anastasios Gounaris, Jim Smith 0001, Norman W. Paton, Rizos Sakellariou, Alvaro A. A. Fernandes, Paul Watson 0001
Distributed Parallel Databases6
2009 The design and implementation of OGSA-DQP: A service-based distributed query processor
Steven J. Lynden, Arijit Mukherjee, Alastair C. Hume, Alvaro A. A. Fernandes, Norman W. Paton, Rizos Sakellariou, Paul Watson 0001
Future Gener. Comput. Syst.7
2008 GOLD infrastructure for virtual organizations
abstract
Abstract The paper discusses the GOLD project (Grid‐based Information Models to Support the Rapid Innovation of New High Value‐Added Chemicals) whose principal aim is to carry out research and development into enabling technologies to support the formation, operation and termination of virtual organizations. The paper discusses the outcome of this research, which is the GOLD Middleware infrastructure. The infrastructure has been implemented in the form of a set of Middleware components, which address issues such as trust, security, contract monitoring and enforcement, information management and coordination. We discuss all these issues in turn and more importantly we demonstrate how current WS standards can be used to implement these issues. In addition, the paper follows a top down approach starting with a brief outline on the architectural elements derived during the requirements engineering phase and demonstrates how these elements were mapped onto actual services that were implemented according to service‐oriented architecture principles and related technologies. Copyright © 2008 John Wiley & Sons, Ltd.
Panos Periorellis, N. Cook, Hugo Hiden, A. Conlin, M. D. Hamilton, Jiyi Wu, Jeremy W. Bryans, Xiangguo Gong, Paul Watson 0001, Allen R. Wright
Concurr. Comput. Pract. Exp.11
2008 Special Issue
abstract
The 2005 U.K. e-Science All Hands Meeting was the fourth in the series, but had a markedly different flavour to those that had gone before. The UK e-Science core programme began in 2001, under the leadership of Professor Tony Hey; hence by the 2005 meeting, many of the initial projects had ended and delivered interesting results. These spanned not just the design of the underlying computer infrastructure to support e-science, but often also the application science itself. The breadth of scientific domains had also broadened considerably by 2005. In the early days of grid computing, there was an emphasis on the use of high-performance computing facilities to support large computations—job-based grid computing was the dominant paradigm. However, the U.K. e-science programme invested heavily on applications and infrastructure to support information-driven science. Areas such as biology were facing a deluge of heterogeneous, complex data, and advances in e-science were not just of academic interest, but were absolutely necessary if the inherent value in this data was to be unlocked. By the 2005 meeting, the results of the first successful projects to achieve this were being published. One key aspect of the U.K. e-Science programme was its focus on collaboration. This was a necessity if advances were to be made: the application scientists had computational needs that were beyond the capabilities of the existing compute infrastructures, whilst computer scientists often had ideas and prototypes, but needed applications to set the requirements and provide a way to evaluate their work. Even within the computing community, collaboration was encouraged, with projects bringing together a set of sub-disciplines that had previously often only had a nodding acquaintance. A typical project might extract data from a set of databases based on some search criteria, combine and analyse the results, compare them against the existing knowledge, and then visualize any promising outputs. This required the collaboration of researchers with a diverse set of skills ranging across databases, data analysis, semantics, text mining, user interfaces and visualization. In this special issue, we have selected eight papers from the 259 submissions to the conference. They were chosen from those judged by the programme committee to be the best submissions, and we have tried to show something of the diversity of the work presented. They therefore span a set of application sciences, while the underlying infrastructure exploits a range of disciplines within computing science. They range from quantitative performance prediction in complex grid systems, through projects engaging school children in biomedical research challenges, to worldwide grid-enabled systems. We now explain why we picked each paper: Jarvis et al. 1 present some new predictive modelling to enable interactive scheduling of a complex biomedical application where the runtime is highly variable and depends on data known only at the time of job submission. Fang et al. 2 analyse the scalability of the semantics-aware service discovery engine GRIMOIRES. An important and pressing challenge in science and engineering is to engage and nurture the next generation of scientists, who are our future! Jeremy Frey leads a project described in Frey et al. 3, which aims at bringing the attention of 16–18-year-old school children to how computational drug design works, motivated by the search for better drugs to combat Malaria. Cohen et al. 4 demonstrate how advances in Grid computing technologies coupled to innovative charging models could provide a radical change in the provisioning and use of information technology across academia, business and society. Preece et al. 5 demonstrate how information quality annotations for experimental data sets can be computed and delivered using a web service by leveraging preferences provided by users against a formal domain-specific ontology. Wood et al. 6 demonstrate how computational steering can be improved by enabling users to manipulate and steer calculations by direct interaction with images or visualizations of their simulation rather than by the conventional means of varying input parameter sets. They also report on preliminary user testing of their system. Lupu et al. 7 couple wireless health monitoring sensors to a grid with the aim of enabling a patient to be monitored easily in a variety of settings. A key aspect of their system is to make it self-configuring and self-managing to minimize the need for a patient to become a technology expert in order to get the benefits of using it. Perhaps, more technology should be designed with this in mind! Grid systems are now becoming increasingly relied upon by scientists and business; if they fail then the implications can be grave. As a result, researchers are investigating methods to increase the tolerance of grid infrastructures to hardware and software failures. This makes the work of Xu et al. 8 very timely. We hope that you enjoy reading the selection of papers in this special edition and we thank the authors for allowing us to present their work.
Paul Watson 0001
Concurr. Comput. Pract. Exp.1
2007 e-Science in the Cloud with CARMEN
abstract
Summary form only given. Understanding how the brain works is a major scientific challenge which will benefit medicine, biology and computer science. It requires knowledge of how information is encoded, accessed, analysed, archived and decoded by networks of neurons. Globally, over 100,000 neuroscientists are working on this problem. However, the data that forms the basis for their work is rarely shared even though it is difficult and expensive to produce. One of the main reasons for this is that vast amounts of data are produced in a variety of formats; this is then locally described and curated. One consequence is a shortage of analysis techniques that can be applied across neuronal systems. Further, there is only a limited amount of interaction between research centres with complementary expertise. The CARMEN project (www.carmen.org.uk) is addressing these challenges. It enables data sharing, integration, and analysis supported by metadata. An expandable range of services are provided to extract value from raw and transformed data. The project's approach is to design and build a generic e-science platform in the cloud. This provides functionality to neuroscientists, who access it over the web. Scientists upload the data they generate into the system (called a CAIRN) and describe it with metadata. Internally, the CAIRN is built as a set of Web Services. These include a workflow enactment service which allows scientists to analyze data by running workflows that utilise the analysis services also held in the CAIRN. Users can browse the catalogue of existing workflows to select one that is appropriate for the task they are trying to accomplish. More expert users can build their own workflows from the available services. Even more sophisticated users can create their own services to use in workflows. A novel feature of the architecture is that the services are stored in a repository in the CAIRN and scheduled on a grid as required by the execution of workflows. This promotes the sharing of analysis services as well as data, and allows services to execute close to the data on which they operate. This is essential to avoid having to ship vast quantities (TBs) of data out of the CAIRN to the user's machine for analysis. Storing both the data and the services in the CAIRN also enables the reproducibility of analyses. This talk describes the design of the CAIRN and shows how it is used to support neuro informatics.
Paul Watson 0001
PDCAT1
2006 Topic 5: Parallel and Distributed Databases, Data Mining and Knowledge Discovery
Patrick Valduriez, Wolfgang Lehner, Domenico Talia, Paul Watson 0001
Euro-Par4
2006 Practical Adaptation to Changing Resources in Grid Query Processing
abstract
Grid computational resources, as well as being heterogeneous, may also exhibit unpredictable, volatile behaviour. Therefore, query processing on the Grid needs to be adaptive in order to cope with evolving resource characteristics, such as machine load and availability. To address this challenge in a Grid environment, the non-adaptive OGSA-DQP1 system described in [1] has been enhanced with adaptive capabilities.
Anastasios Gounaris, Norman W. Paton, Rizos Sakellariou, Alvaro A. A. Fernandes, Jim Smith 0001, Paul Watson 0001
ICDE6
2006 Dynamically Deploying Web Services on a Grid using Dynasoar
abstract
Dynasoar is an infrastructure for dynamically deploying Web services over a grid or the Internet. It enables an approach to grid computing in which distributed applications are built around services instead of jobs. Dynasoar automatically deploys a service on an available host if no existing deployments exist, or if performance requirements cannot be met by existing deployments. This is analogous to remote job scheduling, but offers the opportunity for improved performance as the cost of moving and deploying the service can be shared across the processing of many messages. A key feature of the architecture is that it makes a clear separation between Web service providers, who offer services to consumers, and host providers, who offer computational resources on which services can be deployed, and messages sent to them processed. Separating these two components and defining their interactions, opens up the opportunity for interesting new organisational/business models
Paul Watson 0001, Chris Fowler, Charles Kubicek, Arijit Mukherjee, John Colquhoun, Mark Hewitt, Savas Parastatidis
ISORC1
2006 Measuring and modelling the performance of a parallel ODMG compliant object database server
abstract
Abstract Object database management systems (ODBMSs) are now established as the database management technology of choice for a range of challenging data intensive applications. Furthermore, the applications associated with object databases typically have stringent performance requirements, and some are associated with very large data sets. An important feature for the performance of object databases is the speed at which relationships can be explored. In queries, this depends on the effectiveness of different join algorithms into which queries that follow relationships can be compiled. This paper presents a performance evaluation of the Polar parallel object database system, focusing in particular on the performance of parallel join algorithms. Polar is a parallel, shared‐nothing implementation of the Object Database Management Group (ODMG) standard for object databases. The paper presents an empirical evaluation of queries expressed in the ODMG Query Language (OQL), as well as a cost model for the parallel algebra that is used to evaluate OQL queries. The cost model is validated against the empirical results for a collection of queries using four different join algorithms, one that is value based and three that are pointer based. Copyright © 2005 John Wiley & Sons, Ltd.
Sandra de F. Mendes Sampaio, Norman W. Paton, Jim Smith 0001, Paul Watson 0001
Concurr. Comput. Pract. Exp.4
2005 A grid-based system for microbial genome comparison and analysis
abstract
Genome comparison and analysis can reveal the structures and junctions of genome sequences of different species. As more genomes are sequenced, genomic data sources are rapidly increasing such that their analysis is beyond the processing capabilities of most research institutes. The grid is a powerful solution to support large-scale genomic data processing and genome analysis. This paper presents the Microbase project that is developing a grid-based system for genome comparison and analysis, and discusses the first implementation of the system (called MicrobaseLite). MicrobaseLite uses a scalable computing environment to support computationally intensive microbial genome comparison and analysis, employing state-of-the-art technologies of Web services, notification, comparative genomics and parallel computing. Microbase will support not only system-defined genome comparison and analysis but also user-defined, remotely conceived genome analysis.
Anil Wipat, Matthew R. Pocock, Pete A. Lee, Paul Watson 0001, Keith Flanagan, James T. Worthington
CCGRID5
2005 Fault-Tolerance in Distributed Query Processing
abstract
Fault-tolerance has long been a feature of database systems, with transactions supporting the structuring of applications so as to ensure continuation of updating applications in spite of machine failures. For read-only queries the perceived wisdom has been that support for fault-tolerance is too expensive to be worthwhile. Distributed query processing is coming to be seen as a promising way of implementing applications that combine structured data and analysis operations in dynamic distributed settings such as computational grids. Such a query may be long-running and having to redo the whole query after a failure may cause problems (e.g. if the result may trigger business or safety critical activities). This work describes and evaluates a new scheme for adding fault-tolerance to distributed query processing through a rollback-recovery mechanism. The high level expression of user requests in a physical algebra offers opportunities for tuning the fault-tolerance provision so as to reduce the cost, and give better performance than employment of generic fault-tolerance mechanisms at the lowest level of query processing. This paper outlines how the publicly-available OGSA-DQP computational grid-based distributed query processing system can be modified to include support for fault-tolerance and presents a performance evaluation which includes measurements of the cost of both protocol overheads and rollback-recovery, for a set of example distributed queries.
Jim Smith 0001, Paul Watson 0001
IDEAS2
2005 The design and implementation of Grid database services in OGSA-DAI
abstract
Abstract Initially, Grid technologies were principally associated with supercomputer centres and large‐scale scientific applications in physics and astronomy. They are now increasingly seen as being relevant to many areas of e‐Science and e‐Business. The emergence of the Open Grid Services Architecture (OGSA), to complement the ongoing activity on Web Services standards, promises to provide a service‐based platform that can meet the needs of both business and scientific applications. Early Grid applications focused principally on the storage, replication and movement of file‐based data. Now the need for the full integration of database technologies with Grid middleware is widely recognized. Not only do many Grid applications already use databases for managing metadata, but increasingly many are associated with large databases of domain‐specific information (e.g. biological or astronomical data). This paper describes the design and implementation of OGSA‐DAI, a service‐based architecture for database access over the Grid. The approach involves the design of Grid Data Services that allow consumers to discover the properties of structured data stores and to access their contents. The initial focus has been on support for access to Relational and XML data, but the overall architecture has been designed to be extensible to accommodate different storage paradigms. The paper describes and motivates the design decisions that have been taken, and illustrates how the approach supports a range of application scenarios. The OGSA‐DAI software is freely available from http://www.ogsadai.org.uk . Copyright © 2005 John Wiley & Sons, Ltd.
Mario Antonioletti, Malcolm P. Atkinson 0001, Robert M. Baxter, Andrew Borley, Neil P. Chue Hong, Brian Collins, Neil Hardman, Alastair C. Hume, Alan Knox, Mike Jackson 0003, Amrey Krause, Simon Laws, James Magowan, Norman W. Paton, Dave Pearson, Tom Sugden, Paul Watson 0001, Martin D. Westhead
Concurr. Pract. Exp.17
2005 Web Service Grids: an evolutionary approach
abstract
Abstract The U.K. e‐Science Programme is a £250 million, five‐year initiative which has funded over 100 projects. These application‐led projects are underpinned by an emerging set of core middleware services that allow the coordinated, collaborative use of distributed resources. This set of middleware services runs on top of the research network and beneath the applications we call the ‘Grid’. Grid middleware is currently in transition from pre‐Web Service versions to a new version based on Web Services. Unfortunately, only a very basic set of Web Services embodied in the Web Services Interoperability proposal, WS‐I, are agreed by most IT companies. IBM and others have submitted proposals for Web Services for Grids—the Web Services ResourceFramework and Web Services Notification specifications—to the OASIS organization for standardization. This process could take up to 12 months from March 2004 and the specifications are subject to debate and potentially significant changes. Since several significant U.K. e‐Science projects come to an end before the end of this process, the U.K. needs to develop a strategy that will protect the U.K.'s investment in Grid middleware by informing the Open Middleware Infrastructure Institute's (OMII) roadmap and U.K. middleware repository in Southampton. This paper sets out an evolutionary roadmap that will allow us to capture generic middleware components from projects in a form that will facilitate migration or interoperability with the emerging Grid Web Services standards and with ongoing OGSA developments. In this paper we therefore define a set of Web Services specifications, which we call ‘WS‐I+’ to reflect the fact that this is a larger set than currently accepted by WS‐I, that we believe will enable us to achieve the twin goals of capturing these components and facilitating migration to future standards. We believe that the extra Web Services specifications we have included in WS‐I+ are both helpful in building e‐Science Grids and likely to be widely accepted. Copyright © 2005 John Wiley & Sons, Ltd.
Malcolm P. Atkinson 0001, David De Roure, Alistair N. Dunlop, Geoffrey C. Fox, Peter Henderson 0001, Anthony J. G. Hey, Norman W. Paton, Steven J. Newhouse, Savas Parastatidis, Anne E. Trefethen, Paul Watson 0001, Jim Webber
Concurr. Pract. Exp.11
2005 WS-GAF: a framework for building Grid applications using Web Services
abstract
Abstract This paper presents the motivation and design decisions for the Web Services Grid Application Framework (WS‐GAF), which is a mapping of Grid architecture requirements onto the Web Services Architecture. The goal for WS‐GAF is to describe a framework for building Grid applications that adheres to the principles of service‐oriented architectures and utilizes existing Web Services technologies. The proposed solution addresses issues including stateful interactions, logical resource naming, metadata, and lifetime management. Copyright © 2005 John Wiley & Sons, Ltd.
Savas Parastatidis, Jim Webber, Paul Watson 0001, Thomas Rischbeck
Concurr. Pract. Exp.3
2004 OGSA-DQP: A Service for Distributed Querying on the Grid
Mahmut Nedim Alpdemir, Arijit Mukherjee, Anastasios Gounaris, Norman W. Paton, Paul Watson 0001, Alvaro A. A. Fernandes, Desmond J. Fitzgerald
EDBT5
2004 Topic 5: Parallel and Distributed Databases, Data Mining and Knowledge Discovery
David B. Skillicorn, Abdelkader Hameurlain, Paul Watson 0001, Salvatore Orlando 0001
Euro-Par3
2004 The Design, Implementation and Evaluation of an ODMG Compliant, Parallel Object Database Server
Jim Smith 0001, Sandra de F. Mendes Sampaio, Paul Watson 0001, Norman W. Paton
Distributed Parallel Databases3
2003 Service-Based Distributed Querying on the Grid
Mahmut Nedim Alpdemir, Arijit Mukherjee, Norman W. Paton, Paul Watson 0001, Alvaro A. A. Fernandes, Anastasios Gounaris, Jim Smith 0001
ICSOC4
2003 Grid Data Management Systems & Services
Arun Jagatheesan, Reagan W. Moore, Norman W. Paton, Paul Watson 0001
VLDB4
2002 Speeding Up Navigational Requests in a Parallel Object Database System
Jim Smith 0001, Paul Watson 0001, Sandra de F. Mendes Sampaio, Norman W. Paton
Euro-Par2
2002 A Scalable, Multi-User VRML Server
abstract
VRML97 allows the description of dynamic worlds that can change with both the passage of time, and user interaction. Unfortunately, the current VRML usage model prevents its full potential from being realized. Initially, the whole world must be loaded into the user's desktop browser, and so large worlds can take a very long time to download and render, while a world cannot be shared among multiple users. This paper describes the design and implementation of a client-server architecture that was built to overcome these problems. The major novelty is the decoupling of VRML world execution from world rendering. Parallelism and information filtering are exploited to produce a highly scalable system that can support huge, highly active worlds, accessed simultaneously by large numbers of users. A cluster-based parallel server is responsible for maintaining the dynamic world state, and most of the world dynamics are evaluated on the server side. The server streams VRML to the client, using view frustum culling and dynamic LOD selection to reduce clients' network bandwidth, storage and rendering requirements. Clients with limited resources (e.g. wireless-connected PDAs) can therefore participate in highly complex virtual worlds. While the implementation of the design focuses on VRML worlds, the design ideas could be exploited in other types of VR system, e.g. X3D.
Thomas Rischbeck, Paul Watson 0001
VR2
2001 An Experimental Performance Evaluation of Join Algorithms for Parallel Object Databases
Sandra de F. Mendes Sampaio, Jim Smith 0001, Norman W. Paton, Paul Watson 0001
Euro-Par4
2000 Polar: An Architecture for a Parallel ODMG Compliant Object Database
abstract
Article Polar: an architecture for a parallel ODMG compliant object database Share on Authors: Jim Smith Department of Computing Science, University of Newcastle upon Tyne, Newcastle upon Tyne, NE1 7RU UK Department of Computing Science, University of Newcastle upon Tyne, Newcastle upon Tyne, NE1 7RU UKView Profile , Paul Watson Department of Computing Science, University of Newcastle upon Tyne, Newcastle upon Tyne, NE1 7RU UK Department of Computing Science, University of Newcastle upon Tyne, Newcastle upon Tyne, NE1 7RU UKView Profile , Sandra de F. Mendes Sampaio Department of Computing Science, University of Manchester, Oxford Road, Manchester, M13 9PL UK Department of Computing Science, University of Manchester, Oxford Road, Manchester, M13 9PL UKView Profile , Norman Paton Department of Computing Science, University of Manchester, Oxford Road, Manchester, M13 9PL UK Department of Computing Science, University of Manchester, Oxford Road, Manchester, M13 9PL UKView Profile Authors Info & Claims CIKM '00: Proceedings of the ninth international conference on Information and knowledge managementNovember 2000 Pages 352–359https://doi.org/10.1145/354756.354840Online:06 November 2000Publication History 12citation401DownloadsMetricsTotal Citations12Total Downloads401Last 12 Months3Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Jim Smith 0001, Paul Watson 0001, Sandra de F. Mendes Sampaio, Norman W. Paton
CIKM2
1998 Metabroker: A Generic Broker for Electronic Commerce
Steve J. Caughey, David B. Ingham, Paul Watson 0001
Comput. Networks3
1988 Flagship: A Parallel Architecture for Declarative Programming
abstract
The Flagship project aims to produce a computing technology based on the declarative style of programming. A major component of that technology is the design for a parallel machine that can efficiently utilize the implicit parallelism in declarative programs. The computational models that expose this implicit parallelism are described, and an architecture designed to use it is outlined. The operational issues, such as dynamic load balancing, that arise in such a system are discussed, and the mechanisms being used to evaluate the architecture are described.>
Ian Watson, Viv Woods, Paul Watson 0001, Richard Banach, Mark Irvine Greenberg, John Sargeant
ISCA3