Peter Kilpatrick

dblp:k/PeterKilpatrick · also P. L. Kilpatrick · DBLP profile ↗
← Back
49ranked-venue papers
0as first author
14since 2021 · last 2024
0000-0003-0818-8979ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 20 · 7 since 2021Software engineering, systems software and programming languages · 11 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 since 2021Artificial intelligence and machine learning · 4 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1
YearPublicationVenuePosition
2024 DNNShifter: An efficient DNN pruning system for edge computing
abstract
Deep neural networks (DNNs) underpin many machine learning applications. Production quality DNN models achieve high inference accuracy by training millions of DNN parameters which has a significant resource footprint. This presents a challenge for resources operating at the extreme edge of the network, such as mobile and embedded devices that have limited computational and memory resources. To address this, models are pruned to create lightweight, more suitable variants for these devices. Existing pruning methods are unable to provide similar quality models compared to their unpruned counterparts without significant time costs and overheads or are limited to offline use cases. Our work rapidly derives suitable model variants while maintaining the accuracy of the original model. The model variants can be swapped quickly when system and network conditions change to match workload demand. This paper presents DNNShifter , an end-to-end DNN training, spatial pruning, and model switching system that addresses the challenges mentioned above. At the heart of DNNShifter is a novel methodology that prunes sparse models using structured pruning - combining the accuracy-preserving benefits of unstructured pruning with runtime performance improvements of structured pruning. The pruned model variants generated by DNNShifter are smaller in size and thus faster than dense and sparse model predecessors, making them suitable for inference at the edge while retaining near similar accuracy as of the original dense model. DNNShifter generates a portfolio of model variants that can be swiftly interchanged depending on operational conditions. DNNShifter produces pruned model variants up to 93x faster than conventional training methods. Compared to sparse models, the pruned model variants are up to 5.14x smaller and have a 1.67x inference latency speedup, with no compromise to sparse model accuracy. In addition, DNNShifter has up to 11.9x lower overhead for switching models and up to 3.8x lower memory utilisation than existing approaches. DNNShifter is available for public use from https://github.com/blessonvar/DNNShifter.
Bailey J. Eccles, Philip Rodgers, Peter Kilpatrick, Ivor T. A. Spence, Blesson Varghese
Future Gener. Comput. Syst.3
2024 PiPar: Pipeline parallelism for collaborative machine learning
abstract
Collaborative machine learning (CML) techniques, such as federated learning, have been proposed to train deep learning models across multiple mobile devices and a server. CML techniques are privacy-preserving as a local model that is trained on each device instead of the raw data from the device is shared with the server. However, CML training is inefficient due to low resource utilization. We identify idling resources on the server and devices due to sequential computation and communication as the principal cause of low resource utilization. A novel framework PiPar that leverages pipeline parallelism for CML techniques is developed to substantially improve resource utilization. A new training pipeline is designed to parallelize the computations on different hardware resources and communication on different bandwidth resources, thereby accelerating the training process in CML. A low overhead automated parameter selection method is proposed to optimize the pipeline, maximizing the utilization of available resources. The experimental results confirm the validity of the underlying approach of PiPar and highlight that when compared to federated learning: (i) the idle time of the server can be reduced by up to 64.1×, and (ii) the overall training time can be accelerated by up to 34.6× under varying network conditions for a collection of six small and large popular deep neural networks and four datasets without sacrificing accuracy. It is also experimentally demonstrated that PiPar achieves performance benefits when incorporating differential privacy methods and operating in environments with heterogeneous devices and changing bandwidths.
Philip Rodgers, Peter Kilpatrick, Ivor T. A. Spence, Blesson Varghese
J. Parallel Distributed Comput.3
2024 EcoFed: Efficient Communication for DNN Partitioning-Based Federated Learning
abstract
Efficiently running federated learning (FL) on resource-constrained devices is challenging since they are required to train computationally intensive deep neural networks (DNN) independently. DNN partitioning-based FL (DPFL) has been proposed as one mechanism to accelerate training where the layers of a DNN (or computation) are offloaded from the device to the server. However, this creates significant communication overheads since the intermediate activation and gradient need to be transferred between the device and the server during training. While current research reduces the communication introduced by DNN partitioning using local loss-based methods, we demonstrate that these methods are ineffective in improving the overall efficiency (communication overhead and training speed) of a DPFL system. This is because they suffer from accuracy degradation and ignore the communication costs incurred when transferring the activation from the device to the server. This article proposesEcoFed– a communication efficient framework for DPFL systems.EcoFedeliminates the transmission of the gradient by developing pre-trained initialization of the DNN model on the device for the first time. This reduces the accuracy degradation seen in local loss-based methods. In addition,EcoFedproposes a novel replay buffer mechanism and implements a quantization-based compression technique to reduce the transmission of the activation. It is experimentally demonstrated thatEcoFedcan reduce the communication cost by up to 133× and accelerate training by up to 21× when compared to classic FL. Compared to vanilla DPFL,EcoFedachieves a 16× communication reduction and 2.86× training time speed-up.
Di Wu 0065, Rehmat Ullah 0001, Philip Rodgers, Peter Kilpatrick, Ivor T. A. Spence, Blesson Varghese
IEEE Trans. Parallel Distributed Syst.4
2023 On Overcoming HPC Challenges of Trillion-Scale Real-World Graph Datasets
abstract
Progress in High-Performance Computing in general, and High-Performance Graph Processing in particular, is highly dependent on the availability of publicly-accessible, relevant, and realistic data sets. To ensure continuation of this progress, we (i) investigate and optimize the process of generating large sequence similarity graphs as an HPC challenge and (ii) demonstrate this process in creating MS-BioGraphs, a new family of publicly available real-world edge-weighted graph datasets with up to 2.5 trillion edges, that is, 6.6 times greater than the largest graph published recently. The largest graph is created by matching (i.e., all-toall similarity aligning) 1.7 billion protein sequences. The MSBioGraphs family includes also seven subgraphs with different sizes and direction types. We describe two main challenges we faced in generating large graph datasets and our solutions, that are, (i) optimizing data structures and algorithms for this multi-step process and (ii) WebGraph parallel compression technique. The datasets are available online on https://blogs.qub.ac.uk/ DIPSA/MS-BioGraphs.
Mohsen Koohi Esfahani, Paolo Boldi, Hans Vandierendonck, Peter Kilpatrick, Sebastiano Vigna
IEEE Big Data4
2022 Dengue Fever: From Extreme Climates to Outbreak Prediction
abstract
Dengue Fever (DF) is an emerging mosquito-borne infectious disease that affects hundred of millions of people each year with considerable morbidity and mortality rates, especially for children. Together with global climate changes, it is continuously increasing in terms of number of cases and new locations. Thus, having effective early warning systems becomes an urgent need to improve disease controls and prevention. In this paper, we introduce a novel framework, called Proximity Time Ensemble, to predict DF outbreaks for multiple areas (provinces) and multiple time steps ahead, and to study the effects of climate data on DF outbreaks. PT-Ensem consists of 6 key components: (1) an event-to-event probabilistic framework to study links among extreme climate events and DF outbreaks; (2) a proximity graph that connects similar provinces; (3) an ensemble prediction technique that combines many different advanced machine learning (ML) methods to predict outbreaks within t time steps in the future using extreme climate events as model inputs; (4) a data aggregate scheme to enrich training data for each province via its neighbors in the proximity graph; (5) a proximity propagation step that propagates predicted results among similar provinces via the proximity graph until maximal agreements are reached among provinces; and (6) a time propagation step to propagate results via different predicted time steps in each province. We use PT-Ensem to predict DF outbreaks for all provinces in Vietnam using data collected from 1997-2016. Experiments show that PT-Ensem acquires significant performance boost compared to many highly-rated ML models like XGBoost, LightGBM and Catboost in the outbreak prediction task. Compared to most recent deep learning approaches like LSTM-ATT, LSTM, CNN and Transformer for predicting DF incidence, PT-Ensem also dominates in both prediction accuracy and computation times.
Son T. Mai, Ha T. Phi, Peter Kilpatrick, Hung Q. V. Nguyen, Hans Vandierendonck
ICDM4
2022 MASTIFF: structure-aware minimum spanning tree/forest
abstract
The Minimum Spanning Forest (MSF) problem finds usage in many different applications. While theoretical analysis shows that linear-time solutions exist, in practice, parallel MSF algorithms remain computationally demanding due to the continuously increasing size of data sets.
Mohsen Koohi Esfahani, Peter Kilpatrick, Hans Vandierendonck
ICS2
2022 SAPCo Sort: optimizing Degree-Ordering for Power-Law Graphs
abstract
We introduce the Structure-Aware Parallel Counting (SAPCo) Sort algorithm that optimizes performance of degree-ordering, a key operation in graph analytics. SAPCo leverages the skewed degree distribution to accelerate sorting. The evaluation for graphs of up to 3.6 billion vertices shows that SAPCo sort is, on average, 1.7-33.5 times faster than state-of-the-art sorting algorithms such as counting sort, radix sort, and sample sort.
Mohsen Koohi Esfahani, Peter Kilpatrick, Hans Vandierendonck
ISPASS2
2022 LOTUS: locality optimizing triangle counting
abstract
Triangle Counting (TC) is a basic graph mining problem with numerous applications. However, the large size of real-world graphs has a severe effect on TC performance.
Mohsen Koohi Esfahani, Peter Kilpatrick, Hans Vandierendonck
PPoPP2
2022 FedAdapt: Adaptive Offloading for IoT Devices in Federated Learning
abstract
Applying federated learning (FL) on Internet of Things (IoT) devices is necessitated by the large volumes of data they produce and growing concerns of data privacy. However, there are three challenges that need to be addressed to make FL efficient: 1) execution on devices with limited computational capabilities; 2) accounting for stragglers due to computational heterogeneity of devices; and 3) adaptation to the changing network bandwidths. This article presentsFedAdapt, an adaptive offloading FL framework to mitigate the aforementioned challenges.FedAdaptaccelerates local training in computationally constrained devices by leveraging layer offloading of deep neural networks (DNNs) to servers. Furthermore,FedAdaptadopts reinforcement learning (RL)-based optimization and clustering to adaptively identify which layers of the DNN should be offloaded for each individual device on to a server to tackle the challenges of computational heterogeneity and changing network bandwidth. The experimental studies are carried out on a lab-based testbed and it is demonstrated that by offloading a DNN from the device to the serverFedAdaptreduces the training time of a typical IoT device by over half compared to classic FL. The training time of extreme stragglers and the overall training time can be reduced by up to 57%. Furthermore, with changing network bandwidth,FedAdaptis demonstrated to reduce the training time by up to 40% when compared to classic FL, without sacrificing accuracy.
Di Wu 0065, Rehmat Ullah 0001, Paul Harvey 0002, Peter Kilpatrick, Ivor T. A. Spence, Blesson Varghese
IEEE Internet Things J.4
2021 The Case for Adaptive Deep Neural Networks in Edge Computing
abstract
Deep Neural Networks (DNNs) are an application class that benefit from being distributed across the edge and cloud. A DNN is partitioned such that specific layers of the DNN are deployed onto the edge and the cloud to meet performance and privacy objectives. However, there is limited understanding of: whether and how evolving operational conditions (increased CPU and memory utilization at the edge or reduced data transfer rates between the edge and cloud) affect the performance of already deployed DNNs, and whether a new partition configuration is required to maximize performance. A DNN that adapts to changing operational conditions is referred to as an ‘adaptive DNN’. This paper investigates whether there is a case for adaptive DNNs by considering four questions: (i) Are DNNs sensitive to operational conditions? (ii) How sensitive are DNNs to operational conditions? (iii) Do individual or a combination of operational conditions equally affect DNNs? (iv) Is DNN partitioning sensitive to hardware architectures? The exploration is carried out in the context of 8 pre-trained DNN models and the results presented are from analyzing nearly 8 million data points. The results highlight that network conditions affect DNN performance more than CPU or memory related operational conditions. Repartitioning is noted to provide a performance gain in a number of cases, but a specific trend is not noted in relation to the underlying hardware architecture. Nonetheless, the need for adaptive DNNs is confirmed.
Francis McNamee, Schahram Dustdar, Peter Kilpatrick, Weisong Shi, Ivor T. A. Spence, Blesson Varghese
CLOUD3
2021 Thrifty Label Propagation: Fast Connected Components for Skewed-Degree Graphs
abstract
Various concurrent algorithms have been proposed in the literature in recent years that mostly focus on the disjoint set approach to the Connected Components (CC) algorithm. However, these CC algorithms do not take the skewed structure of real-world graphs into account and as a result they do not benefit from common features of graph datasets to accelerate processing.We investigate the implications of the skewed degree distribution of real-world graphs on their connectivity and we use these features to introduce Thrifty Label Propagation as a structure-aware CC algorithm obtained by incorporating 4 fundamental optimization techniques in the Label Propagation CC algorithm.Our evaluation on 15 real-world graphs and 2 different processor architectures shows that Thrifty accelerates the flow of labels and processes only 1.4% of the edges of the graph.In this way, Thrifty is up to 16 × faster than state-of-the-art CC algorithms such as Afforest, Jayanti-Tarjan, and Breadth-First Search CC. In particular, Thrifty delivers 1.5 × −19.9× speedup for graph datasets larger than one billion edges.
Mohsen Koohi Esfahani, Peter Kilpatrick, Hans Vandierendonck
CLUSTER2
2021 NEUKONFIG: Reducing Edge Service Downtime When Repartitioning DNNs
abstract
Deep Neural Networks (DNNs) may be partitioned across the edge and the cloud to improve the performance efficiency of inference. DNN partitions are determined based on operational conditions such as network speed. When operational conditions change DNNs will need to be repartitioned to maintain the overall performance. However, repartitioning using existing approaches, such as Pause and Resume, will incur a service downtime on the edge. This paper presents the NEUKONFIG framework that identifies the service downtime incurred when repartitioning DNNs and proposes approaches for reducing edge service downtime. The proposed approaches are based on ‘Dynamic Switching’ in which, when the network speed changes and given an existing edge-cloud pipeline, a new edge-cloud pipeline is initialised with new DNN partitions. Incoming inference requests are switched to the new pipeline for processing data. Experimental studies are carried out on a lab-based testbed to demonstrate that Dynamic Switching reduces the downtime by at least an order of magnitude when compared to a baseline using Pause and Resume that has a downtime of 6 seconds. A trade-off in the edge service downtime and memory required is noted. The Dynamic Switching approach that requires the same amount of memory as the baseline reduces the edge service downtime to 0.6 seconds and to less than 1 millisecond in the best case when twice the amount of memory as the baseline is available.
Ayesha Abdul Majeed, Peter Kilpatrick, Ivor T. A. Spence, Blesson Varghese
IC2E2
2021 Exploiting in-Hub Temporal Locality in SpMV-based Graph Processing
abstract
The skewed degree distribution of real-world graphs is the main source of poor locality in traversing all edges of the graph, known as Sparse Matrix-Vector (SpMV) Multiplication. Conventional graph traversal methods, such as push and pull, traverse all vertices in the same manner, and we show applying a uniform traversal direction for all edges leads to sub-optimal memory locality, hence poor efficiency. This paper argues that different vertices in power-law graphs have different locality characteristics and the traversal method should be adapted to these characteristics.
Mohsen Koohi Esfahani, Peter Kilpatrick, Hans Vandierendonck
ICPP2
2021 How Do Graph Relabeling Algorithms Improve Memory Locality?
abstract
Relabeling algorithms aim to improve the poor memory locality of graph processing by changing the order of vertices. This paper analyses the functionality of three state-of-the-art relabeling algorithms: SlashBurn, GOrder, and Rabbit-Order for real-world graphs.
Mohsen Koohi Esfahani, Peter Kilpatrick, Hans Vandierendonck
ISPASS2
2020 Cross Architectural Power Modelling
abstract
Existing power modelling research focuses on the model rather than the process for developing models. An automated power modelling process that can be deployed on different processors for developing power models with high accuracy is developed. For this, (i) an automated hardware performance counter selection method that selects counters best correlated to power on both ARM and Intel processors, (ii) a noise filter based on clustering that can reduce the mean error in power models, and (iii) a two stage power model that surmounts challenges in using existing power models across multiple architectures are proposed and developed. The key results are: (i) the automated hardware performance counter selection method achieves comparable selection to the manual method reported in the literature, (ii) the noise filter reduces the mean error in power models by up to 55%, and (iii) the two stage power model can predict dynamic power with less than 8% error on both ARM and Intel processors, which is an improvement over classic models.
Peter Kilpatrick, Dimitrios S. Nikolopoulos, Blesson Varghese
CCGRID2
2020 Fast Analysis and Prediction in Large Scale Virtual Machines Resource Utilisation
abstract
Most Cloud providers running Virtual Machines (VMs) have a constant goal of preventing downtime, increas- ing performance and power management among others. The most effective way to achieve these goals is to be proactive by predicting the behaviours of the VMs. Analysing VMs is important, as it can help cloud providers gain insights to understand the needs of their customers, predict their demands, and optimise the use of resources. To manage the resources in the cloud efficiently, and to ensure the performance of cloud ser- vices, it is crucial to predict the behaviour of VMs accurately. This will also help the cloud provider improve VM placement, scheduling, consolidation, power management, etc. In this paper, we propose a framework for fast analysis and prediction in large scale VM CPU utilisation. We use a novel approach both in terms of the algorithms employed for prediction and in terms of the tools used to run these algorithms with a large dataset to deliver a solid VM CPU utilisation predictor. We processed over two million VMs from Microsoft Azure VM traces and filter out the VMs with complete one month of data which amount to 28,858VMs. The filtered VMs were subsequently used for prediction. Our Statistical analysis reveals that 94% of these VMs are predictable. Furthermore, we investigate the patterns and behaviours of those VMs and realised that most VMs have one or several spikes of which the majority are not seasonal. For all the 28,858VMs analysed and forecasted, we accurately predicted 17,523 (61%) VMs based on their CPU. We use Apache Spark for parallel and distributed processing to achieve fast processing. In terms of fast processing (execution time), on average, each VM is analysed and predicted within three seconds.
Sakil Barbhuiya, Peter Kilpatrick, Ngo Anh Vien, Dimitrios S. Nikolopoulos
CLOSER3
2020 Modelling Fog Offloading Performance
abstract
Fog computing has emerged as a computing paradigm aimed at addressing the issues of latency, bandwidth and privacy when mobile devices are communicating with remote cloud services. The concept is to offload compute services closer to the data. However many challenges exist in the realisation of this approach. During offloading, (part of) the application underpinned by the services may be unavailable, which the user will experience as down time. This paper describes work aimed at building models to allow prediction of such down time based on metrics (operational data) of the underlying and surrounding infrastructure. Such prediction would be invaluable in the context of automated Fog offloading and adaptive decision making in Fog orchestration. Models that cater for four container-based stateless and stateful offload techniques, namely Save and Load, Export and Import, Push and Pull and Live Migration, are built using four (linear and non-linear) regression techniques. Experimental results comprising over 42 million data points from multiple lab-based Fog infrastructure are presented. The results highlight that reasonably accurate predictions (measured by the coefficient of determination for regression models, mean absolute percentage error, and mean absolute error) may be obtained when considering 25 metrics relevant to the infrastructure.
Ayesha Abdul Majeed, Peter Kilpatrick, Ivor T. A. Spence, Blesson Varghese
ICFEC2
2020 Programming languages for data-Intensive HPC applications: A systematic mapping study
Vasco Amaral 0001, Beatriz Norberto, Miguel Goulão, Marco Aldinucci, Siegfried Benkner, Andrea Bracciali, Paulo Carreira 0001, Edgars Celms, Luís Correia 0001, Clemens Grelck, Helen D. Karatza, Christoph W. Kessler, Peter Kilpatrick, Hugo F. M. C. Martiniano, Ilias Mavridis, Sabri Pllana, Ana Respício, José Simão, Luís Veiga, Ari Visa
Parallel Comput.13
2018 A parallel pattern for iterative stencil + reduce
Marco Aldinucci, Marco Danelutto, Maurizio Drocco, Peter Kilpatrick, Claudia Misale, Guilherme Peretti Pezzi, Massimo Torquati
J. Supercomput.4
2017 MyMinder: A User-centric Decision Making Framework for Intercloud Migration
Esha Barlaskar, Peter Kilpatrick, Ivor T. A. Spence, Dimitrios S. Nikolopoulos
CLOSER2
2015 A Lightweight Tool for Anomaly Detection in Cloud Data Centres
abstract
Cloud data centres are critical business infrastructures and the fastest growing service providers. Detecting anomalies in Cloud data centre operation is vital. Given the vast complexity of the data centre system software stack, applications and workloads, anomaly detection is a challenging endeavour. Current tools for detecting anomalies often use machine learning techniques, application instance behaviours or system metrics distribution, which are complex to implement in Cloud computing environments as they require training, access to application-level data and complex processing. This paper presents LADT, a lightweight anomaly detection tool for Cloud data centres that uses rigorous correlation of system metrics, implemented by an efficient correlation algorithm without need for training or complex infrastructure set up. LADT is based on the hypothesis that, in an anomaly-free system, metrics from data centre host nodes and virtual machines (VMs) are strongly correlated. An anomaly is detected whenever correlation drops below a threshold value. We demonstrate and evaluate LADT using a Cloud environment, where it shows that the hosting node I/O operations per second (IOPS) are strongly correlated with the aggregated virtual machine IOPS, but this correlation vanishes when an application stresses the disk, indicating a node-level anomaly.
Sakil Barbhuiya, Zafeirios C. Papazachos, Peter Kilpatrick, Dimitrios S. Nikolopoulos
CLOSER3
2015 A Green Perspective on Structured Parallel Programming
abstract
Structured parallel programming, and in particular programming models using the algorithmic skeleton or parallel design pattern concepts, are increasingly considered to be the only viable means of supporting effective development of scalable and efficient parallel programs. Structured parallel programming models have been assessed in a number of works in the context of performance. In this paper we consider how the use of structured parallel programming models allows knowledge of the parallel patterns present to be harnessed to address both performance and energy consumption. We consider different features of structured parallel programming that may be leveraged to impact the performance/energy trade-off and we discuss a preliminary set of experiments validating our claims.
Marco Danelutto, Massimo Torquati, Peter Kilpatrick
PDP3
2014 SHEPARD: Scheduling on Heterogeneous Platforms Using Application Resource Demands
abstract
Heterogeneous computing technologies, such as multi-core CPUs, GPUs and FPGAs can provide significant performance improvements. However, developing applications for these technologies often results in coupling applications to specific devices, typically through the use of proprietary tools. This paper presents SHEPARD, a compile time and run-time framework that decouples application development from the target platform and enables run-time allocation of tasks to heterogeneous computing devices. Through the use of special annotated functions, called managed tasks, SHEPARD approximates a task's performance on available devices, and coupled with the approximation of current device demand, decides which device can satisfy the task with the lowest overall execution time. Experiments using a task parallel application, based on an in-memory database, demonstrate the opportunity for automatic run-time task allocation to achieve speed-up over a static allocation to a single specific device.
Eoghan O'Neill, John McGlone, Peter Milligan, Peter Kilpatrick
PDP4
2013 Performance models of storage contention in cloud environments
abstract
We propose simple models to predict the performance degradation of disk requests due to storage device contention in consolidated virtualized environments. Model parameters can be deduced from measurements obtained inside Virtual Machines (VMs) from a system where a single VM accesses a remote storage server. The parameterized model can then be used to predict the effect of storage contention when multiple VMs are consolidated on the same server. We first propose a trace-driven approach that evaluates a queueing network with fair share scheduling using simulation. The model parameters consider Virtual Machine Monitor level disk access optimizations and rely on a calibration technique. We further present a measurement-based approach that allows a distinct characterization of read/write performance attributes. In particular, we define simple linear prediction models for I/O request mean response times, throughputs and read/write mixes, as well as a simulation model for predicting response time distributions. We found our models to be effective in predicting such quantities across a range of synthetic and emulated application workloads.
Stephan Kraft, Giuliano Casale, Diwakar Krishnamurthy, Des Greer, Peter Kilpatrick
Softw. Syst. Model.5
2012 WIQ: Work-Intensive Query Scheduling for In-Memory Database Systems
abstract
We propose a novel admission control policy for database queries. Our methodology uses system measurements of CPU utilization and query backlogs to determine interference between queries in execution on the same database server. Query interference may arise due to the concurrent access of hardware and software resources and can affect performance in positive and negative ways. Specifically our admission control considers the mix of jobs in service and prioritizes the query classes consuming CPU resources more efficiently. The policy ignores I/O subsystems and is therefore highly appropriate for in-memory databases. We validate our approach in trace-driven simulation and show performance increases of query slowdowns and throughputs compared to first-come first-served and shortest expected processing time first scheduling. Simulation experiments are parameterized from system traces of a SAP HANA in-memory database installation with TPC-H type workloads.
Stephan Kraft, Giuliano Casale, Alin Jula, Peter Kilpatrick, Des Greer
IEEE CLOUD4
2012 An Efficient Unbounded Lock-Free Queue for Multi-core Systems
Marco Aldinucci, Marco Danelutto, Peter Kilpatrick, Massimiliano Meneghin, Massimo Torquati
Euro-Par3
2012 Parallel Patterns + Macro Data Flow for Multi-core Programming
abstract
Data flow techniques have been around since the early '70s when they were used in compilers for sequential languages. Shortly after their introduction they were also considered as a possible model for parallel computing, although the impact here was limited. Recently, however, data flow has been identified as a candidate for efficient implementation of various programming models on multi-core architectures. In most cases, however, the burden of determining data flow ``macro'' instructions is left to the programmer, while the compiler/run time system manages only the efficient scheduling of these instructions. We discuss a structured parallel programming approach supporting automatic compilation of programs to macro data flow and we show experimental results demonstrating the feasibility of the approach and the efficiency of the resulting ``object'' code on different classes of state-of-the-art multi-core architectures. The experimental results use different base mechanisms to implement the macro data flow run time support, from plain pthreads with condition variables to more modern and effective lock- and fence-free parallel frameworks. Experimental results comparing efficiency of the proposed approach with those achieved using other, more classical, parallel frameworks are also presented.
Marco Aldinucci, L. Anardu, Marco Danelutto, Massimo Torquati, Peter Kilpatrick
PDP5
2011 Accelerating Code on Multi-cores with FastFlow
Marco Aldinucci, Marco Danelutto, Peter Kilpatrick, Massimiliano Meneghin, Massimo Torquati
Euro-Par (2)3
2011 IO performance prediction in consolidated virtualized environments
abstract
We propose a trace-driven approach to predict the performance degradation of disk request response times due to storage device contention in consolidated virtualized environments. Our performance model evaluates a queueing network with fair share scheduling using trace-driven simulation. The model parameters can be deduced from measurements obtained inside Virtual Machines (VMs) from a system where a single VM accesses a remote storage server. The parameterized model can then be used to predict the effect of storage contention when multiple VMs are consolidated on the same virtualized server. The model parameter estimation relies on a search technique that tries to estimate the splitting and merging of blocks at the the Virtual Machine Monitor (VMM) level in the case of multiple competing VMs. Simulation experiments based on traces of the Postmark and FFSB disk benchmarks show that our model is able to accurately predict the impact of workload consolidation on VM disk IO response times.
Stephan Kraft, Giuliano Casale, Diwakar Krishnamurthy, Des Greer, Peter Kilpatrick
ICPE5
2009 Extending BPM Environments of Your Choice with Performance Related Decision Support
Mathias Fritzsche, Michael Picht, Wasif Gilani, Ivor T. A. Spence, T. John Brown, Peter Kilpatrick
BPM6
2009 Autonomic management of non-functional concerns in distributed & parallel application programming
abstract
An approach to the management of non-functional concerns in massively parallel and/or distributed architectures that marries parallel programming patterns with autonomic computing is presented. The necessity and suitability of the adoption of autonomic techniques are evidenced. Issues arising in the implementation of autonomic managers taking care of multiple concerns and of coordination among hierarchies of such autonomic managers are discussed. Experimental results are presented that demonstrate the feasibility of the approach.
Marco Aldinucci, Marco Danelutto, Peter Kilpatrick
IPDPS3
2009 Towards Hierarchical Management of Autonomic Components: A Case Study
abstract
We address the issue of autonomic management in hierarchical component-based distributed systems. The long term aim is to provide a modeling framework for autonomic management in which QoS goals can be defined, plans for system adaptation described and proofs of achievement of goals by (sequences of) adaptations furnished. Here we present an early step on this path. We restrict our focus to skeleton-based systems in order to exploit their well-defined structure. The autonomic cycle is described using the Orc system orchestration language while the plans are presented as structural modifications together with associated costs and benefits. A case study is presented to illustrate the interaction of managers to maintain QoS goals for throughput under varying conditions of resource availability.
Marco Aldinucci, Marco Danelutto, Peter Kilpatrick
PDP3
2008 Assessing the Reliability and Cost of Web and Grid Orchestrations
abstract
Unreliability is a characteristic feature of Web and grid based computation: a call to a Web or grid site may or may not succeed. An orchestration manager aims to control Web or grid behaviour by constructing dynamic threads which acquire and utilise appropriate resources. For example, in an orchestration, a site may be called and may fail to respond after a period of time. Subsequently, the site may be recalled, or an alternative site may be utilised. In this paper approaches to estimating the reliability and cost of complex orchestrations are proposed. It is assumed that information about site reliability can be accessed from empirically acquired information.
Alan Stewart, Maurice Clint, Terence J. Harmer, Peter Kilpatrick, Ronald H. Perrott, Joaquim Gabarró
ARES4
2008 Behavioural Skeletons in GCM: Autonomic Management of Grid Components
abstract
Autonomic management can be used to improve the QoS provided by parallel/distributed applications. We discuss behavioural skeletons introduced in earlier work: rather than relying on programmer ability to design "from scratch" efficient autonomic policies, we encapsulate general autonomic controller features into algorithmic skeletons. Then we leave to the programmer the duty of specifying the parameters needed to specialise the skeletons to the needs of the particular application at hand. This results in the programmer having the ability to fast prototype and tune distributed/parallel applications with non-trivial autonomic management capabilities. We discuss how behavioural skeletons have been implemented in the framework of GCM (the grid component model developed within the CoreGRID NoE and currently being implemented within the GridCOMP STREP project). We present results evaluating the overhead introduced by autonomic management activities as well as the overall behaviour of the skeletons. We also present results achieved with a long running application subject to autonomic management and dynamically adapting to changing features of the target architecture. Overall the results demonstrate both the feasibility of implementing autonomic control via behavioural skeletons and the effectiveness of our sample behavioural skeletons in managing the "functional replication" pattern(s).
Marco Aldinucci, Sonia Campa, Marco Danelutto, Marco Vanneschi, Peter Kilpatrick, Patrizio Dazzi, Domenico Laforenza, Nicola Tonellotto
PDP5
2008 Systematic Usage of Embedded Modelling Languages in Automated Model Transformation Chains
Mathias Fritzsche, Jendrik Johannes, Uwe Aßmann, Simon Mitschke, Wasif Gilani, Ivor T. A. Spence, T. John Brown, Peter Kilpatrick
SLE8
2007 Management in Distributed Systems: A Semi-formal Approach
Marco Aldinucci, Marco Danelutto, Peter Kilpatrick
Euro-Par3
2006 Managing Grid Computations: An ORC-Based Approach
Alan Stewart, Joaquim Gabarró, Maurice Clint, Terence J. Harmer, Peter Kilpatrick, Ronald H. Perrott
ISPA5
2006 Weaving Behavior into Feature Models for Embedded System Families
abstract
Product line software engineering depends on capturing the commonality and variability within a family of products, typically using feature modeling, and using this information to evolve a generic reference architecture for the family. For embedded systems, possible variability in hardware and operating system platforms is an added complication. The design process can be facilitated by first exploring the behavior associated with features. In this paper we outline a bidirectional feature modeling scheme that supports the capture of commonality and variability in the platform environment as well as within the required software. Additionally, 'behavior' associated with features can be included in the overall model. This is achieved by integrating the UCM path notation in a way that exploits UCM's static and dynamic stubs to capture behavioral variability and link it to the feature model structure. The resulting model is a richer source of information to support the architecture development process
T. John Brown, Rachel Gawley, Rabih Bashroush, Ivor T. A. Spence, Peter Kilpatrick, Charles Gillan
SPLC5
2005 A generic reference software architecture for load balancing over mirrored Web servers: NaSr case study
abstract
Summary form only given. With the rapid expansion of the Internet and the increasing demand on Web servers, many techniques were developed to overcome the servers' hardware performance limitation. Mirrored Web servers is one of the techniques used where a number of servers carrying the same "mirrored" set of services are deployed. Client access requests are then distributed over the set of mirrored servers to even up the load. In this paper, we present a generic reference software architecture for load balancing over mirrored Web servers. The architecture was designed adopting the latest NaSr architectural style and described using the ADLARS architecture description language. With minimal effort, different tailored product architectures can be generated from the reference architecture to serve different network protocols and server operating systems. An example product system is described and a sample Java implementation is presented.
Rabih Bashroush, Ivor T. A. Spence, Peter Kilpatrick, T. John Brown
AICCSA3
2005 ADLARS: An Architecture Description Language for Software Product Lines
abstract
Software product line (SPL) engineering has emerged to become a mature domain for maximizing reuse within the context of a family of related software products. Within the process of SPL, the variability and commonality among the different products within the scope of a family is captured and modeled into a system's `feature model'. Currently, there are no architecture description languages (ADLs) that support the relationship between the feature model domain and the system architecture domain, leaving a gap which significantly increases the complexity of analyzing the system's architecture and insuring that it complies with its set feature model and variability requirements. In this paper we present ADLARS, an architecture description language that supports the relationship between the system's feature model and the architectural structures in an attempt to alleviate the aforementioned problem. The link between the two spaces also allows the automatic generation of product architectures from the family reference architecture
Rabih Bashroush, T. John Brown, Ivor T. A. Spence, Peter Kilpatrick
SEW4
2005 Feature-Guided Architecture Development for Embedded System Families
abstract
Software product-line engineering aims to maximize reuse by exploiting the commonality within families of related systems. Its success depend on capturing the commonality and variability, and using this to evolve a reference architecture for the product family. With embedded system families, the possibility of variability in hardware and operating system platforms is an added complication. In this paper we outline a strategy for evolving reference architectures from bi-directional feature models. The proposed strategy complements information provided by the feature model with scenarios that help to elaborate feature behavior.
T. John Brown, Rabih Bashroush, Charles Gillan, Ivor T. A. Spence, Peter Kilpatrick
WICSA5
2004 A Network Architectural Style for Real-time Systems: NaSr
abstract
Inter-component communication has always been of great importance in the design of software architectures and connectors have been considered as first-class entities in many approaches by R. Allen and D. Garlan (1994), M. Shaw et al., (1995), and D. Batory and S. O'Malley (1992). We present a novel architectural style that is derived from the well-established domain of computer networks. The style adopts the inter-component communication protocol in a novel way that allows large scale software reuse. It mainly targets real-time, distributed, concurrent, and heterogeneous systems.
Rabih Bashroush, Ivor T. A. Spence, Peter Kilpatrick, T. John Brown
WICSA3
2002 Adaptable Components for Software Product Line Engineering
T. John Brown, Ivor T. A. Spence, Peter Kilpatrick, Danny Crookes
SPLC3
1996 The Tailoring of Abstract Functional Specifications of Numerical Algorithms for Sparse Data Structures through Automated Program Derivation and Transformation
abstract
The automated application of program transformations is used to derive, from abstract functional specifications of numerical mathematical algorithms, highly efficient imperative implementations tailored for execution on sequential, vector and array processors. Emphasis is placed on transformations which tailor implementations to use special programming techniques optimized for sparse matrices. We demonstrate that derived implementations attain superior execution performance than manual implementations for two significant algorithms.
Stephen Fitzpatrick, Maurice Clint, Terence J. Harmer, Peter Kilpatrick
Comput. J.4
1989 A case study in improving programming productivity on transputer networks
N. Stanley Scott, Danny Crookes, Peter Milligan, Peter Kilpatrick, Philip J. Morrow
Microprocess. Microprogramming4
1988 Network topology A critical factor in the implementation of algorithms intended for efficient execution on a transputer network
Peter Milligan, N. Stanley Scott, Danny Crookes, Peter Kilpatrick, Philip J. Morrow
Microprocess. Microprogramming4
1988 A comparison of programming paradigms for the parallel computation of racah coefficients: An application of transputers to computational atomic physics
N. Stanley Scott, Lionel C. Waring, Peter Milligan, Danny Crookes, Peter Kilpatrick, Philip J. Morrow
Microprocess. Microprogramming5
1988 An array processing language for transputer networks
Danny Crookes, Philip J. Morrow, Peter Milligan, Peter Kilpatrick, N. Stanley Scott
Parallel Comput.4
1987 Notes on implementing a language for transputer networks
Danny Crookes, Philip J. Morrow, Peter Milligan, N. Stanley Scott, Peter Kilpatrick
Microprocess. Microprogramming5