EDBT 2026 Demo / reviewers in the wild / expert
Stephen A. Jarvis
dblp:69/1268 · also Stephen Andrew Jarvis
· DBLP profile ↗
86ranked-venue papers
7as first author
5since 2021 · last 2026
0000-0002-1249-2167ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 60 · 5 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-authorComputer networks · 5Databases, data management, data science and information retrieval · 4 · 1 since 2021Artificial intelligence and machine learning · 3Graphics, computer vision, multimedia, augmented reality and games · 2Human-computer interaction and ubiquitous computing · 2Security and privacy · 1Software engineering, systems software and programming languages · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Constraint Driven Global Data Layout Optimization for Tensor Expressions
Yi Miao, Stephen A. Jarvis |
Euro-Par (1) | 2 |
| 2022 | Towards Virtual Certification of Gas Turbine Engines With Performance-Portable SimulationsabstractWe present the large-scale, computational fluid dy-namics (CFD) simulation of a full gas-turbine engine compressor, demonstrating capability towards overcoming current limitations for virtual certification of aero-engine design. The simulation is carried out through a performance portable code-base on multi-core/many-core HPC clusters with a CFD-to-CFD coupled execution, combining an industrial CFD solver linked using custom coupler software. The application innovates in its design for performance portability through the OP2 domain specific library for the CFD components, allowing the automatic generation of highly optimized platform-specific parallelizations for both multi-core (CPU) and many-core (GPU) clusters from a single high-level source. The code is used for the simulation of a 4.58B node, full-annulus 10-row production-grade test compressor (DLR's Rig250), using a coupled sliding-plane setup on the ARCHER2 and Cirrus supercomputers at EPCC. The OP2 generated multiple parallelizations, together with optimized coupler configurations on heterogeneous/hybrid settings achieve, for the first time, execution of 1 revolution in less than 6 hours on 512 nodes of ARCHER2 (65k cores), with a parallel scaling efficiency of over 80 % compared to a 107 node run. Results indicate a speed up of the CFD suite by an order of a magnitude (≈30 x) relative to current production capability. Benchmarking and performance modelling project a time-to-solution of less than 5 hours on a cluster of 488xNVIDIA V100 GPUs, about 3x-4 x speedup over CPU clusters. The work demonstrates a step-change towards achieving virtual certification of aircraft engines with the requisite fidelity and tractable time-to-solution that was previously out of reach under production settings. Gihan R. Mudalige, István Z. Reguly, Arun Prabhakar, Dario Amirante, Leigh Lapworth, Stephen A. Jarvis |
CLUSTER | 6 |
| 2022 | Developing an Unsupervised Real-Time Anomaly Detection Scheme for Time Series With Multi-SeasonalityabstractOn-line detection of anomalies in time series is a key technique used in various event-sensitive scenarios such as robotic system monitoring, smart sensor networks and data center security. However, the increasing diversity of data sources and the variety of demands make this task more challenging than ever. First, the rapid increase in unlabeled data means supervised learning is becoming less suitable in many cases. Second, a large portion of time series data have complex seasonality features. Third, on-line anomaly detection needs to be fast and reliable. In light of this, we have developed a prediction-driven, unsupervised anomaly detection scheme, which adopts a backbone model combining the decomposition and the inference of time series data. Further, we propose a novel metric, Local Trend Inconsistency (LTI), and an efficient detection algorithm that computes LTI in a real-time manner and scores each data point robustly in terms of its probability of being anomalous. We have conducted extensive experimentation to evaluate our algorithm with several datasets from both public repositories and production environments. The experimental results show that our scheme outperforms existing representative anomaly detection algorithms in terms of the commonly used metric, Area Under Curve (AUC), while achieving the desired efficiency. Wentai Wu, Ligang He, Weiwei Lin 0001, Yuhua Cui, Carsten Maple, Stephen A. Jarvis |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2021 | Predictive Analysis of Large-Scale Coupled CFD Simulations with the CPX Mini-AppabstractAs the complexity of multi-physics simulations increases, there is a need for efficient flow of information between components. Discrete ‘coupler’ codes can abstract away this process, improving solver interoperability. One such multi-physics problem is modelling the high pressure compressor of turbofan engines, where instances of rotor/stator CFD simulations are coupled. Configuring couplers and allocating resources correctly can be challenging for such problems due to the sliding interfaces between codes. In this research, we present CPX, a mini-coupler designed to model the performance behaviour of a production coupler framework at Rolls-Royce plc., used for coupling rotor/stator simulations. CPX, the first mini-coupler framework of its kind, is combined with a CFD mini-app to predict the run-time and scaling behaviour of large scale coupled CFD simulations. We demonstrate high qualitative and quantitative predictive accuracy with a less than 17 % mean error. A performance model is developed to predict the ‘optimum’ configuration of resources, and is tested to show the high accuracy of these predictions. The model is also used to project the ‘optimum’ configuration for a 6 Billion cell test case, a problem size representative of current leading-edge production workloads, on a 100,000 core cluster and a 400 GPU cluster. Further testing reveals that the ‘optimum’ configuration is unstable if not set up correctly, and therefore a trade-off needs to be made with a marginally less-than-optimal setup to ensure stability. The work illustrates the significant utility of CPX to carry out such rapid design space and run-time setup exploration studies to obtain the best performance from production CFD coupled simulations. Archie Powell, K. Choudry, Arun Prabhakar, István Z. Reguly, Dario Amirante, Stephen A. Jarvis, Gihan R. Mudalige |
HiPC | 6 |
| 2021 | SAFA: A Semi-Asynchronous Protocol for Fast Federated Learning With Low OverheadabstractFederated learning (FL) has attracted increasing attention as a promising approach to driving a vast number of end devices with artificial intelligence. However, it is very challenging to guarantee the efficiency of FL considering the unreliable nature of end devices while the cost of device-server communication cannot be neglected. In this article, we propose SAFA, a semi-asynchronous FL protocol, to address the problems in federated learning such as low round efficiency and poor convergence rate in extreme conditions (e.g., clients dropping offline frequently). We introduce novel designs in the steps of model distribution, client selection and global aggregation to mitigate the impacts of stragglers, crashes and model staleness in order to boost efficiency and improve the quality of the global model. We have conducted extensive experiments with typical machine learning tasks. The results demonstrate that the proposed protocol is effective in terms of shortening federated round duration, reducing local resource wastage, and improving the accuracy of the global model at an acceptable communication cost. Wentai Wu, Ligang He, Weiwei Lin 0001, Rui Mao 0001, Carsten Maple, Stephen A. Jarvis |
IEEE Trans. Computers | 6 |
| 2020 | An unstructured CFD mini-application for the performance prediction of a production CFD codeabstractSummary Maintaining the performance of large scientific codes is a difficult task. To aid in this task, a number of mini‐applications have been developed that are more tractable to analyze than large‐scale production codes while retaining the performance characteristics of them. These “mini‐apps” also enable faster hardware evaluation and, for sensitive commercial codes, allow evaluation of code and system changes outside of access approval processes. In this paper, we develop MG‐CFD, a mini‐application that represents a geometric multigrid, unstructured computational fluid dynamics (CFD) code, designed to exhibit similar performance characteristics without sharing commercially sensitive code. We detail our experiences of developing this application using guidelines detailed in existing research and contributing further to these. Our application is validated against the inviscid flux routine of HYDRA, a CFD code developed by Rolls‐Royce plc for turbomachinery design. This paper (1) documents the development of MG‐CFD, (2) introduces an associated performance model with which it is possible to assess the performance of HYDRA on new HPC architectures, and (3) demonstrates that it is possible to use MG‐CFD and the performance models to predict the performance of HYDRA with a mean error of 9.2% for strong‐scaling studies. Andrew Owenson, Steven A. Wright 0001, Richard A. Bunt, Yoon Ho, Matthew J. Street, Stephen A. Jarvis |
Concurr. Comput. Pract. Exp. | 6 |
| 2020 | Road and travel time cross-validation for urban modellingabstractThe physical and social processes in urban systems are inherently spatial and hence data describing them contain spatial autocorrelation (a proximity-based interdependency on a variable) that need to be accounted for. Standard k-fold cross-validation (KCV) techniques that attempt to measure the generalisation performance of machine learning and statistical algorithms are inappropriate in this setting due to their inherent i.i.d assumption, which is violated by spatial dependency. As such, more appropriate validation methods have been considered, notably blocking and spatial k-fold cross-validation (SKCV). However, the physical barriers and complex network structures which make up a city’s landscape mean that these methods are also inappropriate, largely because the travel patterns (and hence Spatial Autocorrelation (SAC)) in most urban spaces are rarely Euclidean in nature. To overcome this problem, we propose a new road distance and travel time k-fold cross-validation method, RT-KCV. We show how this outperforms the prior art in providing better estimates of the true generalisation performance to unseen data. Henry Crosby, Theodoros Damoulas, Stephen A. Jarvis |
Int. J. Geogr. Inf. Sci. | 3 |
| 2019 | Embedding road networks and travel time into distance metrics for urban modellingabstractUrban environments are restricted by various physical, regulatory and customary barriers such as buildings, one-way systems and pedestrian crossings. These features create challenges for predictive modelling in urban space, as most proximity-based models rely on Euclidean (straight line) distance metrics which, given restrictions within the urban landscape, do not fully capture spatial urban processes. Here, we argue that road distance and travel time provide effective alternatives, and we develop a new low-dimensional Euclidean distance metric based on these distances using an isomap approach. The purpose of this is to produce a valid covariance matrix for Kriging. Our primary methodological contribution is the derivation of two symmetric dissimilarity matrices (B+ and B2+), with which it is possible to compute low-dimensional Euclidean metrics for the production of a positive definite covariance matrix with commonly utilised kernels. This new method is implemented into a Kriging predictor to estimate house prices on 3,669 properties in Coventry, UK. We find that a metric estimating a combination of road distance and travel time, in both R2 and R3, produces a superior house price predictor compared with alternative state-of-the-art methods, that is, a standard Euclidean metric in RN and a non-restricted road distance metric in R2 and R3. F Henry Crosby, Theodoros Damoulas, Stephen A. Jarvis |
Int. J. Geogr. Inf. Sci. | 3 |
| 2019 | An algorithm for computing short-range forces in molecular dynamics simulations with non-uniform particle densitiesabstractWe present projection sorting, an algorithmic approach to determining pairwise short-range forces between particles in molecular dynamics simulations. We show it can be more effective than the standard approaches when particle density is non-uniform. We implement tuned versions of the algorithm in the context of a biophysical simulation of chromosome condensation, for the modern Intel Broadwell and Knights Landing architectures, across multiple nodes. We demonstrate up to 5 × overall speedup and good scaling to large problem sizes and processor counts. Timothy R. Law, Jonny Hancox, Steven A. Wright 0001, Stephen A. Jarvis |
J. Parallel Distributed Comput. | 4 |
| 2019 | The Power-optimised Software EnvelopeabstractAdvances in processor design have delivered performance improvements for decades. As physical limits are reached, refinements to the same basic technologies are beginning to yield diminishing returns. Unsustainable increases in energy consumption are forcing hardware manufacturers to prioritise energy efficiency in their designs. Research suggests that software modifications may be needed to exploit the resulting improvements in current and future hardware. New tools are required to capitalise on this new class of optimisation. In this article, we present the Power Optimised Software Envelope (POSE) model, which allows developers to assess the potential benefits of power optimisation for their applications. The POSE model is metric agnostic and in this article, we provide derivations using the established Energy-Delay Product metric and the novel Energy-Delay Sum and Energy-Delay Distance metrics that we believe are more appropriate for energy-aware optimisation efforts. We demonstrate POSE on three platforms by studying the optimisation characteristics of applications from the Mantevo benchmark suite. Our results show that the Pathfinder application has very little scope for power optimisation while TeaLeaf has the most, with all other applications in the benchmark suite falling between the two. Finally, we extend our POSE model with a formulation known as System Summary POSE—a meta-heuristic that allows developers to assess the scope a system has for energy-aware software optimisation independent of the code being run. Stephen I. Roberts, Steven A. Wright 0001, Suhaib A. Fahmy, Stephen A. Jarvis |
ACM Trans. Archit. Code Optim. | 4 |
| 2018 | BookLeaf: An Unstructured Hydrodynamics Mini-ApplicationabstractWith the age of Exascale computing causing a diversification away from traditional CPU-based homogeneous clusters, it is becoming increasingly difficult to ensure that computationally complex codes are able to run on these emerging architectures. This is especially important for large physics simulations that are themselves becoming increasingly complex and computationally expensive. One proposed solution to the problem of ensuring these applications can run on the desired architectures is to develop representative mini-applications that are simpler and so can be ported to new frameworks more easily, but which are also representative of the algorithmic and performance characteristics of the original applications. In this paper we present BookLeaf, an unstructured Arbitrary Lagrangian-Eulerian mini-application to add to the suite of representative applications developed and maintained by the UK Mini-App Consortium (UK-MAC). First, we outline the reference implementation of our application in Fortran. We then discuss a number of alternative implementations using a variety of parallel programming models and discuss the issues that arise when porting such an application to new architectures. To demonstrate our implementation, we present a study of the performance of BookLeaf on number of platforms using alternative designs, and we document a scaling study showing the behaviour of the application at scale. David Truby, Steven A. Wright 0001, Robert Kevis, Satheesh Maheswaran, Andrew Herdman, Stephen A. Jarvis |
CLUSTER | 6 |
| 2018 | Developing and Using a Geometric Multigrid, Unstructured Grid Mini-Application to Assess Many-Core ArchitecturesabstractAchieving high-performance of large scientific codes is a difficult task. This has led to the development of numerous mini-applications that are more tractable to analyse, while retaining performance characteristics of their full-sized counterparts. These "mini-apps" also enable faster hardware evaluation, and for sensitive codes allow evaluation of systems outside of access approval processes. In this paper we develop a mini-application of a geometric multigrid, unstructured grid Computational Fluid Dynamics (CFD) code, designed to exhibit similar performance characteristics without sharing code. We detail our experiences developing this application, using guidelines detailed in existing research, and contribute further additions to these to aid future mini-application developers. Our application is validated against the inviscid flux routine of HYDRA, a CFD code developed by Rolls-Royce, which confirms that the parent kernel and mini-application share fundamental causes of parallel inefficiency. We then use the mini-application to assess the impact of Intel's Knights Landing (KNL) on performance. We find that the mini-app and parent kernel continue to share scaling characteristics, however a comparison with Broadwell performance exposed significant differences between the kernels that were undetected by the validation. Andrew Owenson, Steven A. Wright 0001, Richard A. Bunt, Stephen A. Jarvis, Yoon Ho, Matthew J. Street |
PDP | 4 |
| 2017 | Achieving Performance Portability for a Heat Conduction Solver Mini-Application on Modern Multi-core SystemsabstractModernizing production-grade, often legacy applications to take advantage of modern multi-core and many-core architectures can be a difficult and costly undertaking. This is especially true currently, as it is unclear which architectures will dominate future systems. The complexity of these codes can mean that parallelisation for a given architecture requires significant re-engineering. One way to assess the benefit of such an exercise would be to use mini-applications that are representative of the legacy programs.In this paper, we investigate different implementations of TeaLeaf, a mini-application from the Mantevo suite that solves the linear heat conduction equation. TeaLeaf has been ported to use many parallel programming models, including OpenMP, CUDA and MPI among others. It has also been re-engineered to use the OPS embedded DSL and template libraries Kokkos and RAJA. We use these different implementations to assess the performance portability of each technique on modern multi-core systems.While manually parallelising the application targeting and optimizing for each platform gives the best performance, this has the obvious disadvantage that it requires the creation of different versions for each and every platform of interest. Frameworks such as OPS, Kokkos and RAJA can produce executables of the program automatically that achieve comparable portability. Based on a recently developed performance portability metric, our results show that OPS and RAJA achieve an application performance portability score of 71% and 77% respectively for this application. Richard O. Kirk, Gihan R. Mudalige, István Z. Reguly, Steven A. Wright 0001, Matt Martineau, Stephen A. Jarvis |
CLUSTER | 6 |
| 2017 | Goal-based composition of scalable hybrid analytics for heterogeneous architecturesabstractCrafting scalable analytics in order to extract actionable business intelligence is a challenging endeavour, requiring multiple layers of expertise and experience. Often, this expertise is irreconcilably split between an organisation’s engineers and subject matter domain experts. Previous approaches to this problem have relied on technically adept users with tool-specific training. Such an approach has a number of challenges: Expertise — There are few data-analytic subject domain experts with in-depth technical knowledge of compute architectures; Performance — Analysts do not generally make full use of the performance and scalability capabilities of the underlying architectures; Heterogeneity — calculating the most performant and scalable mix of real-time (on-line) and batch (off-line) analytics in a problem domain is difficult; Tools — Supporting frameworks will often direct several tasks, including, composition, planning, code generation, validation, performance tuning and analysis, but do not typically provide end-to-end solutions embedding all of these activities. In this paper, we present a novel semi-automated approach to the composition, planning, code generation and performance tuning of scalable hybrid analytics, using a semantically rich type system which requires little programming expertise from the user. This approach is the first of its kind to permit domain experts with little or no technical expertise to assemble complex and scalable analytics, for hybrid on- and off-line analytic environments, with no additional requirement for low-level engineering support. This paper describes (i) an abstract model of analytic assembly and execution, (ii) goal-based planning and (iii) code generation for hybrid on- and off-line analytics. An implementation, through a system which we call Mendeleev, is used to (iv) demonstrate the applicability of this technique through a series of case studies, where a single interface is used to create analytics that can be run simultaneously over on- and off-line environments. Finally, we (v) analyse the performance of the planner, and (vi) show that the performance of Mendeleev’s generated code is comparable with that of hand-written analytics. Peter Coetzee, Stephen A. Jarvis |
J. Parallel Distributed Comput. | 2 |
| 2017 | An Indoor Test Methodology for Solar-Powered Wireless Sensor NetworksabstractRepeatable and accurate tests are important when designing hardware and algorithms for solar-powered wireless sensor networks (WSNs). Since no two days are exactly alike with regard to energy harvesting, tests must be carried out indoors. Solar simulators are traditionally used in replicating the effects of sunlight indoors; however, solar simulators are expensive, have lighting elements that have short lifetimes, and are usually not designed to carry out the types of tests that hardware and algorithm designers require. As a result, hardware and algorithm designers use tests that are inaccurate and not repeatable (both for others and also for the designers themselves). In this article, we propose an indoor test methodology that does not rely on solar simulators. The test methodology has its basis in astronomy and photovoltaic cell design. We present a generic design for a test apparatus that can be used in carrying out the test methodology. We also present a specific design that we use in implementing an actual test apparatus. We test the efficacy of our test apparatus and, to demonstrate the usefulness of the test methodology, perform experiments akin to those required in projects involving solar-powered WSNs. Results of the said tests and experiments demonstrate that the test methodology is an invaluable tool for hardware and algorithm designers working with solar-powered WSNs. Wilson M. Tan, Paul Sullivan, Hamish Watson, Joanna Slota-Newson, Stephen A. Jarvis |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2016 | Spatially-Intensive Decision Tree Prediction of Traffic Flow across the Entire UK Road NetworkabstractThis paper introduces a novel approach to predicting UK-wide daily traffic counts on all roads in England and Wales, irrespective of sensor data availability. A key finding of this research is that many roads in a network may have no local connection, but may still share some common law, and this fact can be exploited to improve simulation. In this paper we show that: (1) Traffic counts are a function of dependant spatial, temporal and neighbourhood variables, (2) Large open-source data, such as school location and public transport hubs can, with appropriate GIS and machine learning, assist the prediction of traffic counts, (3) Real-time simulation can be scaled-up to large networks with the aid of machine learning and, (4) Such techniques can be employed in real-world tools. Validation of the proposed approach demonstrates an 88.2% prediction accuracy on traffic counts across the UK. Henry Crosby, Stephen A. Jarvis |
DS-RT | 3 |
| 2016 | A spatio-temporal, Gaussian process regression, real-estate price predictorabstractThis paper introduces a novel four-stage methodology for real-estate valuation. This research shows that space, property, economic, neighbourhood and time features are all contributing factors in producing a house price predictor in which validation shows a 96.6% accuracy on Gaussian Process Regression beating regression-kriging, random forests and an M5P-decision-tree. The output is integrated into a commercial real estate decision engine. Henry Crosby, Theodoros Damoulas, Stephen A. Jarvis |
SIGSPATIAL/GIS | 4 |
| 2016 | Predictive Evaluation of Partitioning Algorithms through Runtime ModellingabstractPerformance modelling unstructured mesh codes is a challenging process, due to the difficulty of capturing their memory access patterns, and their communication patterns at varying scale. In this paper we first develop extensions to an existing runtime performance model, aimed at overcoming the former, which we validate on up to 1,024 cores of a Haswellbased cluster, using both a geometric partitioning algorithm and ParMETIS to partition the input deck, with a maximum absolute runtime error of 12.63% and 11.55% respectively. To overcome the latter, we develop an application representative of the mesh partitioning process internal to an unstructured mesh code. This application is able to generate partitioning data that is usable with the performance model to produce predicted application runtimes within 7.31% of those produced using empirically collected data. We then demonstrate the use of the performance model by undertaking a predictive comparison among several partitioning algorithms on up to 30,000 cores. Additionally, we correctly predict the ineffectiveness of the geometric partitioning algorithm at 512 and 1024 cores. Richard A. Bunt, Steven A. Wright 0001, Stephen A. Jarvis, Yoon Ho, Matthew J. Street |
HiPC | 3 |
| 2016 | Optimisation of a Molecular Dynamics Simulation of Chromosome CondensationabstractWe present optimisations applied to a bespoke bio-physical molecular dynamics simulation designed to investigate chromosome condensation. Our primary focus is on domain-specific algorithmic improvements to determining short-range interaction forces between particles, as certain qualities of the simulation render traditional methods less effective. We implement tuned versions of the code for both traditional CPU architectures and the modern many-core architecture found in the Intel Xeon Phi coprocessor and compare their effectiveness. We achieve speed-ups starting at a factor of 10 over the original code, facilitating more detailed and larger-scale experiments. Timothy R. Law, Jonny Hancox, Tammy M. K. Cheng, Raphael A. G. Chaleil, Steven A. Wright 0001, Paul A. Bates, Stephen A. Jarvis |
SBAC-PAD | 7 |
| 2016 | Heuristic solutions to the target identifiability problem in directional sensor networksabstractExisting algorithms for orienting sensors in directional sensor networks have primarily concerned themselves with the problem of maximizing the number of covered targets, assuming that target identification is a non-issue. Such an assumption however, does not hold true in all situations. In this paper, heuristic algorithms for choosing active sensors and orienting them with the goal of balancing coverage and identifiability are presented. The performance of the algorithms are verified via extensive simulations, and shown to confer increased target identifiability compared to algorithms originally designed to simply maximize the number of targets covered. Wilson M. Tan, Stephen A. Jarvis |
J. Netw. Comput. Appl. | 2 |
| 2016 | Developing Graph-Based Co-Scheduling Algorithms on Multicore ComputersabstractIt is common that multiple cores reside on the same chip and share the on-chip cache. As a result, resource sharing can cause performance degradation of co-running jobs. Job co-scheduling is a technique that can effectively alleviate this contention and many co-schedulers have been reported in related literature. Most solutions however do not aim to find the optimal co-scheduling solution. Being able to determine the optimal solution is critical for evaluating co-scheduling systems. Moreover, most co-schedulers only consider serial jobs, and there often exist both parallel and serial jobs in real-world systems. In this paper a graph-based method is developed to find the optimal co-scheduling solution for serial jobs; the method is then extended to incorporate parallel jobs, including multi-process, and multi-threaded parallel jobs. A number of optimization measures are also developed to accelerate the solving process. Moreover, a flexible approximation technique is proposed to strike a balance between the solving speed and the solution quality. Extensive experiments are conducted to evaluate the effectiveness of the proposed co-scheduling algorithms. The results show that the proposed algorithms can find the optimal co-scheduling solution for both serial and parallel jobs. The proposed approximation technique is also shown to be flexible in the sense that we can control the solving speed by setting the requirement for the solution quality. Ligang He, Huanzhou Zhu, Stephen A. Jarvis |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2015 | Mini-App Driven Optimisation of Inertial Confinement Fusion CodesabstractIn September 2013, the large laser-based inertial confinement fusion device housed in the National Ignition Facility at Lawrence Livermore National Laboratory, was widely acclaimed to have achieved a milestone in controlled fusion -- successfully initiating a reaction that resulted in the release of more energy than the fuel absorbed. Despite this success, we remain some distance from being able to create controlled, self-sustaining fusion reactions. Inertial Confinement Fusion (ICF) represents one leading design for the generation of energy by nuclear fusion. Since the 1950s, ICF has been supported by computing simulations, providing the mathematical foundations for pulse shaping, lasers, and material shells needed to ensure effective and efficient implosion. The research presented here focuses on one such simulation code, EPOCH, a fully relativistic particle-in-cell plasma physics code, developed by a leading network of over 30 UK researchers. A significant challenge in developing large codes like EPOCH is maintaining effective scientific delivery on successive generations of high-performance computing architecture. To support this process, we adopt the use of mini-applications -- small code proxies that encapsulate important computational properties of their larger parent counterparts. Through the development of a mini-app for EPOCH (called miniEPOCH), we investigate known time-step scaling issues within EPOCH and explore possible optimisations: (i) Employing loop fission to increase levels of vectorisation, (ii) Enforcing particle ordering to allow the exploitation of domain specific knowledge and, (iii) Changing underlying data storage to improve memory locality. When applied to EPOCH, these improvements represent a 2.02× speed-up in the core algorithm and a 1.55× speed-up to the overall application runtime, when executed on EPCC's Cray XC30 ARCHER platform. Robert F. Bird, Patrick Gillies, Michael R. Bareford, J. A. Herdman, Stephen A. Jarvis |
CLUSTER | 5 |
| 2015 | Resident Block-Structured Adaptive Mesh Refinement on Thousands of Graphics Processing UnitsabstractBlock-structured adaptive mesh refinement (AMR) is a technique that can be used when solving partial differential equations to reduce the number of cells necessary to achieve the required accuracy in areas of interest. These areas (shock fronts, material interfaces, etc.) are recursively covered with finer mesh patches that are grouped into a hierarchy of refinement levels. Despite the potential for large savings in computational requirements and memory usage without a corresponding reduction in accuracy, AMR adds overhead in managing the mesh hierarchy, adding complex communication and data movement requirements to a simulation. In this paper, we describe the design and implementation of a resident GPU-based AMR library, including: the classes used to manage data on a mesh patch, the routines used for transferring data between GPUs on different nodes, and the data-parallel operators developed to coarsen and refine mesh data. We validate the performance and accuracy of our implementation using three test problems and two architectures: an 8 node cluster, and 4,196 nodes of Oak Ridge National Laboratory's Titan supercomputer. Our GPU-based AMR hydrodynamics code performs up to 4.87× faster than the CPU-based implementation, and is scalable on 4,196 K20x GPUs using a combination of MPI and CUDA. D. A. Beckingsale, Wayne P. Gaudin, Andrew Herdman, Stephen A. Jarvis |
ICPP | 4 |
| 2015 | Guest Editorial: Ubiquitous Multimedia Systems and Applications
Xiaolong Jin 0001, Ahmed Yassin Al-Dubai, Shoukat Ali, Stephen A. Jarvis |
Multim. Tools Appl. | 4 |
| 2014 | Optimizing Job Scheduling on Multicore ComputersabstractIt is common nowadays that multiple cores reside on the same chip and share the on-chip cache. Resource sharing may cause performance degradation of the co-running jobs. Job co-scheduling is a technique that can effectively alleviate the contention. Many co-schedulers have been developed in the literature, but most of them do not aim to find the optimal co-scheduling solution. Being able to determine the optimal solution is critical for evaluating co-scheduling systems. Moreover, most co-schedulers only consider serial jobs. However, there often exist both parallel and serial jobs in some situations. This paper aims to tackle these issues. In this paper, a graph-based method is developed to find the optimal co-scheduling solution for serial jobs, and then the method is extended to incorporate parallel jobs. The extensive experiments have been conducted to evaluate the effectiveness and efficiency of the proposed co-scheduling algorithms. The results show that the proposed algorithms can find the optimal co-scheduling solution for both serial and parallel jobs. Huanzhou Zhu, Ligang He, Stephen A. Jarvis |
MASCOTS | 3 |
| 2014 | Developing security-aware resource management strategies for workflows
Ligang He, Nadeem Chaudhary, Stephen A. Jarvis |
Future Gener. Comput. Syst. | 3 |
| 2014 | Developing resource consolidation frameworks for moldable virtual machines in clouds
Ligang He, Deqing Zou, Chao Chen 0011, Hai Jin 0001, Stephen A. Jarvis |
Future Gener. Comput. Syst. | 6 |
| 2014 | Towards unified secure on- and off-line analytics at scaleabstractData scientists have applied various analytic models and techniques to address the oft-cited problems of large volume, high velocity data rates and diversity in semantics. Such approaches have traditionally employed analytic techniques in a streaming or batch processing paradigm. This paper presents CRUCIBLE, a first-in-class framework for the analysis of large-scale datasets that exploits both streaming and batch paradigms in a unified manner. The CRUCIBLE framework includes a domain specific language for describing analyses as a set of communicating sequential processes, a common runtime model for analytic execution in multiple streamed and batch environments, and an approach to automating the management of cell-level security labelling that is applied uniformly across runtimes. This paper shows the applicability of CRUCIBLE to a variety of state-of-the-art analytic environments, and compares a range of runtime models for their scalability and performance against a series of native implementations. The work demonstrates the significant impact of runtime model selection, including improvements of between 2.3× and 480× between runtime models, with an average performance gap of just 14× between CRUCIBLE and a suite of equivalent native implementations. Peter Coetzee, Matthew Leeke, Stephen A. Jarvis |
Parallel Comput. | 3 |
| 2013 | Developing communication-aware service placement frameworks in the Cloud economyabstractIn a Cloud system, a number of services are often deployed with each service being hosted by a collection of Virtual Machines (VM). The services may interact with each other and the interaction patterns may be dynamic, varying according to the system information at runtime. These impose a challenge in determining the amount of resources required to deliver a desired level of QoS for each service. In this paper, we present a method to determine the sufficient number of VMs for the interacting Cloud services. The proposed method borrows the ideas from the Leontief Open Production Model in economy. Further, this paper develops a communication-aware strategy to place the VMs to Physical Machines (PM), aiming to minimize the communication costs incurred by the service interactions. The developed communication-aware placement strategy is formalized in a way that it does not need to the specific communication pattern between individual VMs. A genetic algorithm is developed to find a VM-to-PM placement with low communication costs. Simulation experiments have been conducted to evaluate the performance of the developed communication-aware placement framework. The results show that compared with the placement framework aiming to use the minimal number of PMs to host VMs, the proposed communication-aware framework is able to reduce the communication cost significantly with only a very little increase in the PM usage. Chao Chen 0011, Ligang He, Hao Chen 0002, Jianhua Sun 0002, Bo Gao 0001, Stephen A. Jarvis |
CLUSTER | 6 |
| 2013 | Exploring SIMD for Molecular Dynamics, Using Intel® Xeon® Processors and Intel® Xeon Phi CoprocessorsabstractWe analyse gather-scatter performance bottlenecks in molecular dynamics codes and the challenges that they pose for obtaining benefits from SIMD execution. This analysis informs a number of novel code-level and algorithmic improvements to Sandia's miniMD benchmark, which we demonstrate using three SIMD widths (128-, 256and 512bit). The applicability of these optimisations to wider SIMD is discussed, and we show that the conventional approach of exposing more parallelism through redundant computation is not necessarily best. In single precision, our optimised implementation is up to 5x faster than the original scalar code running on Intel®Xeon®processors with 256-bit SIMD, and adding a single Intel®Xeon Phi™coprocessor provides up to an additional 2x performance increase. These results demonstrate: (i) the importance of effective SIMD utilisation for molecular dynamics codes on current and future hardware; and (ii) the considerable performance increase afforded by the use of Intel®Xeon Phi™coprocessors for highly parallel workloads. Simon J. Pennycook, Christopher J. Hughes, Mikhail Smelyanskiy, Stephen A. Jarvis |
IPDPS | 4 |
| 2013 | Towards Automated Memory Model Generation Via Event TracingabstractThe importance of memory performance and capacity is a growing concern for high performance computing laboratories around the world. It has long been recognized that improvements in processor speed exceed the rate of improvement in dynamic random access memory speed and, as a result, memory access times can be the limiting factor in high performance scientific codes. The use of multi-core processors exacerbates this problem with the rapid growth in the number of cores not being matched by similar improvements in memory capacity, increasing the likelihood of memory contention. In this paper, we present WMTools, a lightweight memory tracing tool and analysis framework for parallel codes, which is able to identify peak memory usage and also analyse per-function memory use over time. An evaluation of WMTools, in terms of its effectiveness and also its overheads, is performed using nine established scientific applications/benchmark codes representing a variety of programming languages and scientific domains. We also show how WMTools can be used to automatically generate a parameterized memory model for one of these applications, a two-dimensional non-linear magnetohydrodynamics application, Lare2D. Through the memory model we are able to identify an unexpected growth term which becomes dominant at scale. With a refined model we are able to predict memory consumption with under 7% error. Oliver Perks, D. A. Beckingsale, Simon D. Hammond, I. Miller, J. A. Herdman, A. Vadgama, Abhir Bhalerao, Ligang He, Stephen A. Jarvis |
Comput. J. | 9 |
| 2013 | Parallel File System Analysis Through Application I/O TracingabstractInput/Output (I/O) operations can represent a significant proportion of the run-time of parallel scientific computing applications. Although there have been several advances in file format libraries, file system design and I/O hardware, a growing divergence exists between the performance of parallel file systems and the compute clusters that they support. In this paper, we document the design and application of the RIOT I/O toolkit (RIOT) being developed at the University of Warwick with our industrial partners at the Atomic Weapons Establishment and Sandia National Laboratories. We use the toolkit to assess the performance of three industry-standard I/O benchmarks on three contrasting supercomputers, ranging from a mid-sized commodity cluster to a large-scale proprietary IBM BlueGene/P system. RIOT provides a powerful framework in which to analyse I/O and parallel file system behaviour—we demonstrate, for example, the large file locking overhead of IBM's General Parallel File System, which can consume nearly 30% of the total write time in the FLASH-IO benchmark. Through I/O trace analysis, we also assess the performance of HDF-5 in its default configuration, identifying a bottleneck created by the use of suboptimal Message Passing Interface hints. Furthermore, we investigate the performance gains attributed to the Parallel Log-structured File System (PLFS) being developed by EMC Corporation and the Los Alamos National Laboratory. Our evaluation of PLFS involves two high-performance computing systems with contrasting I/O backplanes and illustrates the varied improvements to I/O that result from the deployment of PLFS (ranging from up to 25× speed-up in I/O performance on a large I/O installation to 2× speed-up on the much smaller installation at the University of Warwick). Steven A. Wright 0001, Simon D. Hammond, Simon J. Pennycook, Robert F. Bird, J. A. Herdman, I. Miller, A. Vadgama, Abhir Bhalerao, Stephen A. Jarvis |
Comput. J. | 9 |
| 2013 | An investigation of the performance portability of OpenCL
Simon J. Pennycook, Simon D. Hammond, Steven A. Wright 0001, J. A. Herdman, I. Miller, Stephen A. Jarvis |
J. Parallel Distributed Comput. | 6 |
| 2012 | Performance Analysis for Workflow Management Systems under Role-Based Authorization Control
Ligang He, Stephen A. Jarvis |
GPC | 3 |
| 2012 | Editorial Performance Modelling, Benchmarking and Simulation of High-Performance Computing SystemsabstractS.A. Jarvis; Editorial Performance Modelling, Benchmarking and Simulation of High-Performance Computing Systems, The Computer Journal, Volume 55, Issue 2, 1 Feb Stephen A. Jarvis |
Comput. J. | 1 |
| 2012 | On the Acceleration of Wavefront Applications using Distributed Many-Core ArchitecturesabstractIn this paper we investigate the use of distributed graphics processing unit (GPU)-based architectures to accelerate pipelined wavefront applications—a ubiquitous class of parallel algorithms used for the solution of a number of scientific and engineering applications. Specifically, we employ a recently developed port of the LU solver (from the NAS Parallel Benchmark suite) to investigate the performance of these algorithms on high-performance computing solutions from NVIDIA (Tesla C1060 and C2050) as well as on traditional clusters (AMD/InfiniBand and IBM BlueGene/P). Benchmark results are presented for problem classes A to C and a recently developed performance model is used to provide projections for problem classes D and E, the latter of which represents a billion-cell problem. Our results demonstrate that while the theoretical performance of GPU solutions will far exceed those of many traditional technologies, the sustained application performance is currently comparable for scientific wavefront applications. Finally, a breakdown of the GPU solution is conducted, exposing PCIe overheads and decomposition constraints. A new k-blocking strategy is proposed to improve the future performance of this class of algorithm on GPU-based architectures. Simon J. Pennycook, Simon D. Hammond, Gihan R. Mudalige, Steven A. Wright 0001, Stephen A. Jarvis |
Comput. J. | 5 |
| 2012 | Modeling and analyzing the impact of authorization on workflow executions
Ligang He, Chenlin Huang, Kewei Duan, Kenli Li 0001, Hao Chen 0002, Jianhua Sun 0002, Stephen A. Jarvis |
Future Gener. Comput. Syst. | 7 |
| 2011 | Dynamic Resource Allocation and Active Predictive Models for Enterprise Applications
Mohammad A. Alghamdi, Adam P. Chester, Ligang He, Stephen A. Jarvis |
CLOSER | 4 |
| 2011 | Modelling and analyzing the authorization and execution of video workflowsabstractIt is becoming common practice to migrate signal-based video workflows to IT-based Video workflows. Video workflows have some inherent features, including: 1) necessary human involvements in video workflows introduce security and authorization concerns; 2) the frequent change of video workflow contexts requires a flexible approach to acquiring performance data; 3) the content-centric nature of video workflows, which is in contrast to the business-centric of business workflows, requires the support of scheduled activities. This paper takes the above issues into account, proposing a novel mechanism for modeling video workflow executions in cluster-based resource pools under Role-Based Authorization Control (RBAC) schemes. The Color Timed Petri-Net (CTPN) formalism is applied to construct the models. Various types of authorization constraint are modeled in this paper, and scheduled activities are also supported in the model. There is a clear interface between workflow execution and workflow authorization modules. The constructed models are then simulated and analyzed to obtain performance data, including authorization overhead, system- and application-oriented performance. Based on the model analysis, this paper further proposes the methods to improve performance in the presence of authorization policies. This work can be used to plan system capacity subject to the authorization control, and can also be used to tune performance by changing the scheduling strategy and resource capacity when it is not possible to adjust the authorization policies. Ligang He, Chenlin Huang, Kenli Li 0001, Hao Chen 0002, Jianhua Sun 0002, Bo Gao 0001, Kewei Duan, Stephen A. Jarvis |
HiPC | 8 |
| 2011 | UK Performance Engineering Workshop 2010abstractThis special issue contains four papers selected from the 26th UK Performance Engineering Workshop (UKPEW), which was held during 8–9 July 2010 in the Department of Computer Science at the University of Warwick. UKPEW is the leading UK forum for the presentation of work relating to all aspects of performance modelling and the analysis of computer and telecommunication systems. This year's selected papers demonstrate the variety of topics on offer at UKPEW: performance issues in cloud computing, energy-aware routing in packet networks, source location privacy and multi-hop ad hoc networks for wireless networks. Stephen A. Jarvis |
Comput. J. | 1 |
| 2010 | Performance Prediction and Evaluation
Stephen A. Jarvis, Massimo Coppola, Darren J. Kerbyson |
Euro-Par (1) | 1 |
| 2009 | Predictive Simulation of HPC ApplicationsabstractThe architectures which support modern supercomputing machinery are as diverse today, as at any point during the last twenty years. The variety of processor core arrangements, threading strategies and the arrival of heterogeneous computation nodes are driving modern-day solutions to petaflop speeds. The increasing complexity of such systems, as well as codes written to take advantage of the new computational abilities, pose significant frustrations for existing techniques which aim to model and analyze the performance of such hardware and software. In this paper we demonstrate the use of post-execution analysis on trace-based profiles to support the construction of simulation-based models. This involves combining the runtime capture of call-graph information with computational timings, which in turn allows representative models of code behavior to be extracted. The main advantage of this technique is that it largely automates performance model development, a burden associated with existing techniques. We demonstrate the capabilities of our approach using both the NAS Parallel Benchmark suite and a real-world supercomputing benchmark developed by the United Kingdom Atomic Weapons Establishment. The resulting models, developed in less than two hours per code, have a good degree of predictive accuracy. We also show how one of these models can be used to explore the performance of the code on over 16,000 cores, demonstrating the scalability of our solution. Simon D. Hammond, J. A. Smith, Gihan R. Mudalige, Stephen A. Jarvis |
AINA | 4 |
| 2009 | Performance prediction for running workflows under role-based authorization mechanismsabstractWhen investigating the performance of running scientific/commercial workflows in parallel and distributed systems, we often take into account only the resources allocated to the tasks constituting the workflow, assuming that computational resources will accept the tasks and execute them to completion once the processors are available. In reality, and in particular in Grid or e-business environments, security policies may be implemented in the individual organisations in which the computational resources reside. It is therefore expedient to have methods to calculate the performance of executing workflows under security policies. Authorisation control, which specifies who is allowed to perform which tasks when, is one of the most fundamental security considerations in distributed systems such as Grids. Role-Based Access Control (RBAC), under which the users are assigned to certain roles while the roles are associated with prescribed permissions, remains one of the most popular authorisation control mechanisms. This paper presents a mechanism to theoretically compute the performance of running scientific workflows under RBAC authorisation control. Various performance metrics are calculated, including both system-oriented metrics, (such as system utilisation, throughput and mean response time) and user-oriented metrics (such as mean response time of the workflows submitted by a particular client). With this work, if a client informs an organisation of the workflows they are going to submit, the organisation is able to predict the performance of these workflows running in its local computational resources (e.g. a high-performance cluster) enforced with RBAC authorisation control, and can also report client-oriented performance to each individual user. Ligang He, Mark Calleja, Mark Hayes, Stephen A. Jarvis |
IPDPS | 4 |
| 2009 | Predictive analysis and optimisation of pipelined wavefront computationsabstractPipelined wavefront computations are a ubiquitous class of parallel algorithm used for the solution of a number of scientific and engineering applications. This paper investigates three optimisations to the generic pipelined wavefront algorithm, which are investigated through the use of predictive analytic models. The modelling of potential optimisations is supported by a recently developed reusable LogGP-based analytic performance model, which allows the speculative evaluation of each optimisation within the context of an industry-strength pipelined wavefront benchmark developed and maintained by the United Kingdom Atomic Weapons Establishment (AWE). The paper details the quantitative and qualitative benefits of: (1) parallelising computation blocks of the wavefront algorithm using OpenMP; (2) a novel restructuring/shifting of computation within the wavefront code and, (3) performing simultaneous multiple sweeps through the data grid. Gihan R. Mudalige, Simon D. Hammond, J. A. Smith, Stephen A. Jarvis |
IPDPS | 4 |
| 2009 | Peer sampling with improved accuracy
Elth Ogston, Stephen A. Jarvis |
Peer-to-Peer Netw. Appl. | 2 |
| 2009 | Connectivity-Guaranteed and Obstacle-Adaptive Deployment Schemes for Mobile Sensor NetworksabstractMobile sensors can relocate and self-deploy into a network. While focusing on the problems of coverage, existing deployment schemes largely oversimplify the conditions for network connectivity: they either assume that the communication range is large enough for sensors in geometric neighborhoods to obtain location information through local communication, or they assume a dense network that remains connected. In addition, an obstacle-free field or full knowledge of the field layout is often assumed. We present new schemes that are not governed by these assumptions, and thus adapt to a wider range of application scenarios. The schemes are designed to maximize sensing coverage and also guarantee connectivity for a network with arbitrary sensor communication/sensing ranges or node densities, at the cost of a small moving distance. The schemes do not need any knowledge of the field layout, which can be irregular and have obstacles/holes of arbitrary shape. Our first scheme is an enhanced form of the traditional virtual-force-based method, which we term the connectivity-preserved virtual force (CPVF) scheme. We show that the localized communication, which is the very reason for its simplicity, results in poor coverage in certain cases. We then describe a floor-based scheme which overcomes the difficulties of CPVF and, as a result, significantly outperforms it and other state-of-the-art approaches. Throughout the paper our conclusions are corroborated by the results from extensive simulations. Guang Tan, Stephen A. Jarvis, Anne-Marie Kermarrec |
IEEE Trans. Mob. Comput. | 2 |
| 2008 | Connectivity-Guaranteed and Obstacle-Adaptive Deployment Schemes for Mobile Sensor NetworksabstractMobile sensors can move and self-deploy into a network. While focusing on the problems of coverage, existing deployment schemes mostly over-simplify the conditions for network connectivity: they either assume that the communication range is large enough for sensors in geometric neighborhoods to obtain each other's locationby local communications, or assume a dense network that remains connected. At the same time, an obstacle-free field or full knowledge of the field layout is often assumed. We present new schemes that are not restricted by these assumptions, and thus adapt to a much wider range of application scenarios. While maximizing sensing coverage, our schemes can achieve connectivity for a network with arbitrary sensor communication/sensing ranges or node densities, at the cost of a small moving distance; the schemes do not need any knowledge of the field layout, which can be irregular and have obstacles/holes of arbitrary shape. Simulations results show that the proposed schemes achieve the targeted properties. Guang Tan, Stephen A. Jarvis, Anne-Marie Kermarrec |
ICDCS | 2 |
| 2008 | Dynamic Resource Allocation in Enterprise SystemsabstractIt is common that Internet service hosting centres use several logical pools to assign server resources to different applications, and that they try to achieve the highest total revenue by making efficient use of these resources. In this paper, multi-tiered enterprise systems are modelled as multi-class closed queueing networks, with each network station corresponding to each application tier. In such queueing networks, bottlenecks can limit overall system performance, and thus should be avoided. We propose a bottleneck-aware server switching policy, which responds to system bottlenecks and switches servers to alleviate these problems as necessary. The switching engine compares the benefits and penalties of a potential switch, and makes a decision as to whether it is likely to be worthwhile switching. We also propose a simple admission control scheme, in addition to the switching policy, to deal with system overloading and optimise the total revenue of multiple applications in the hosting centre. Performance evaluation has been done via simulation and results are compared with those from a proportional switching policy and also a system that implements no switching policy. The experimental results show that the combination of the bottleneck-aware switching policy and the admission control scheme consistently outperforms the other two policies in terms of revenue contribution. James Wen Jun Xue, Adam P. Chester, Ligang He, Stephen A. Jarvis |
ICPADS | 4 |
| 2008 | Reducing the run-time of MCMC programs by multithreading on SMP architecturesabstractThe increasing availability of multi-core and multiprocessor architectures provides new opportunities for improving the performance of many computer simulations. Markov chain Monte Carlo (MCMC) simulations are widely used for approximate counting problems, Bayesian inference and as a means for estimating very high-dimensional integrals. As such MCMC has found a wide variety of applications infields including computational biology and physics, financial econometrics, machine learning and image processing. This paper presents a new method for reducing the run-time of Markov chain Monte Carlo simulations by using SMP machines to speculatively perform iterations in parallel, reducing the runtime of MCMC programs whilst producing statistically identical results to conventional sequential implementations. We calculate the theoretical reduction in runtime that may be achieved using our technique under perfect conditions, and test and compare the method on a selection of multi-core and multi-processor architectures. Experiments are presented that show reductions in runtime of 35% using two cores and 55% using four cores. Jonathan M. R. Byrd, Stephen A. Jarvis, Abhir Bhalerao |
IPDPS | 2 |
| 2008 | A plug-and-play model for evaluating wavefront computations on parallel architecturesabstractThis paper develops a plug-and-play reusable LogGP model that can be used to predict the runtime and scaling behavior of different MPI-based pipelined wavefront applications running on modern parallel platforms with multi-core nodes. A key new feature of the model is that it requires only a few simple input parameters to project performance for wavefront codes with different structure to the sweeps in each iteration as well as different behavior during each wavefront computation and/or between iterations. We apply the model to three key benchmark applications that are used in high performance computing procurement, illustrating that the model parameters yield insight into the key differences among the codes. We also develop new, simple and highly accurate models of MPI send, receive, and group communication primitives on the dual-core Cray XT system. We validate the reusable model applied to each benchmark on up to 8192 processors on the XT3/XT4. Results show excellent accuracy for all high performance application and platform configurations that we were able to measure. Finally we use the model to assess application and hardware configurations, develop new metrics for procurement and configuration, identify bottlenecks, and assess new application design modifications that, to our knowledge, have not previously been explored. Gihan R. Mudalige, Mary K. Vernon, Stephen A. Jarvis |
IPDPS | 3 |
| 2008 | A System for Dynamic Server Allocation in Application Server ClustersabstractApplication server clusters are often used to service high-throughput web applications. In order to host more than a single application, an organisation will usually procure a separate cluster for each application. Over time the utilisation of the clusters will vary, leading to variation in the response times experienced by users of the applications. Techniques that statically assign servers to each application prevent the system from adapting to changes in the workload, and are thus susceptible to providing unacceptable levels of service. This paper investigates a system for allocating server resources to applications dynamically, thus allowing applications to automatically adapt to variable workloads. Such a scheme requires meticulous system monitoring, a method for switching application servers between \text it {server pools} and a means of calculating when a server switch should be made (balancing switching cost against perceived benefits). Experimentation is performed using such a switching system on a Web application testbed hosting two applications across eight application servers. The testbed is used to compare several theoretically derived switching policies under a variety of workloads. Recommendations are made as to the suitability of different policies under different workload conditions. Adam P. Chester, James Wen Jun Xue, Ligang He, Stephen A. Jarvis |
ISPA | 4 |
| 2008 | Performance prediction for a code with data-dependent runtimesabstractAbstract In this paper we present a predictive performance model for a key biomedical imaging application found as part of the U.K. e‐Science Information eXtraction from Images (IXI) project. This code represents a significant challenge for our existing performance prediction tools as it has internal structures that exhibit highly variable runtimes depending on qualities in the input data provided. Since the runtime can vary by more than an order of magnitude, it has been difficult to apply meaningful quality of service criteria to workflows that use this code. The model developed here is used in the context of an interactive scheduling system which provides rapid feedback to the users, allowing them to tailor their workloads to available resources or to allocate extra resources to scheduled workloads. Copyright © 2007 John Wiley & Sons, Ltd. Stephen A. Jarvis, B. P. Foley, P. J. Isitt, Daniel P. Spooner, Daniel Rueckert, Graham R. Nudd |
Concurr. Comput. Pract. Exp. | 1 |
| 2008 | A Payment-Based Incentive and Service Differentiation Scheme for Peer-to-Peer Streaming BroadcastabstractWe propose a novel payment-based incentive scheme for peer-to-peer (P2P) live media streaming. Using this approach, peers earn points by forwarding data to others. The data streaming is divided into fixed-length periods; during each of these periods, peers compete with each other for good parents (data suppliers) for the next period in a first-price-auction-like procedure using their points. We design a distributed algorithm to regulate peer competitions and consider various individual strategies for parent selection from a game-theoretic perspective. We then discuss possible strategies that can be used to maximize a peer's expected media quality by planning different bids for its substreams. Finally, in order to encourage off-session users to remain online and continue contributing to the network, we develop an optimal data forwarding strategy that allows peers to accumulate points that can be used in future services. Simulation results show that the proposed methods effectively differentiate the media qualities received by peers making different contributions (which originate from, for example, different forwarding bandwidths or servicing times) and at the same time maintain high overall system performance. Guang Tan, Stephen A. Jarvis |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2007 | A scheduling algorithm for revenue maximisation for cluster-based Internet servicesabstractThis paper proposes a new priority scheduling algorithm to maximise site revenue of session-based multi-tier Internet services in a multicluster environment. This research is part of a larger study in support of large-scale online trading systems and, as a result, this case study is chosen as a demonstrator for the techniques presented in this paper. The trading system is partitioned into a number of operations (trade, query etc.), which by their very nature are divided into orders of importance in terms of transactional response. The algorithm in this paper is based on Mean Value Analysis (MVA), which is used for the calculation of performance metrics concerning the queuing networks and workload allocation decision support in the multicluster. In addition to this, the priority assignment is based on combination of three attributes of any given request: (i) the sender class; (ii) the operation and, (iii) the status of the user’s portfolio (i.e.number of items in the user’s portfolio). A discrete event simulator has been developed to evaluate the performance of the priority scheduling scheme with different combinations of request attributes in various experimental scenarios. Our study aims to develop a dynamic scheduling policy, which takes into account real-time system parameters and optimises the site revenue. Although our priority scheduling algorithm is designed for an online trading system, it can be applied to most e-Commerce systems in which differentiated services are required. James Wen Jun Xue, Ligang He, Stephen A. Jarvis |
ICPADS | 3 |
| 2007 | Stochastic Analysis and Improvement of the Reliability of DHT-Based MulticastabstractThis paper investigates the reliability of application-level multicast based on a distributed hash table (DHT) in a highly dynamic network. Using a node residual lifetime model, we derive the stationary end-to-end delivery ratio of data streaming between a pair of nodes in the worst case, and show through numerical examples that in a practical DHT network, this ratio can be very low (e.g., less than 50%). Leveraging the property of heavy-tailed lifetime distribution, we then consider three optimizing techniques, namely senior member overlay (SMO), longer-lived neighbor selection (LNS), and reliable route selection (RRS), and present quantitative analysis of data delivery reliability under these schemes. In particular, we discuss the tradeoff between delivery ratio and the load imbalance among nodes. Simulation experiments are also used to evaluate the multicast performance under practical settings. Our model and analytic results provide useful tools for reliability analysis for other overlay-based applications (e.g., those involving persistent data transfers). Guang Tan, Stephen A. Jarvis |
INFOCOM | 2 |
| 2007 | Predicting the Effect on Performance of Container-Managed Persistence in a Distributed Enterprise ApplicationabstractContainer-managed persistence is an essential technology as it dramatically simplifies the implementation of enterprise data access. However it can also impose a significant overhead on the performance of the application at runtime. This paper presents a layered queuing performance model for predicting the effect of adding or removing container-managed persistence to a distributed enterprise application, in terms of response time and throughput performance metrics. Predictions can then be made for new server architectures - that is, server architectures for which only a small number of measurements have been made (e.g. to determine request processing speed). An experimental analysis of the model is conducted on a popular enterprise computing architecture based on IBM Websphere, using Enterprise Java Bean-based container-managed persistence as the middleware functionality. The results provide strong experimental evidence for the effectiveness of the model in terms of the accuracy of predictions, the speed with which predictions can be made and the low overhead at which the model can be rapidly parameterised. David A. Bacigalupo, James Wen Jun Xue, Simon D. Hammond, Stephen A. Jarvis, Donna Dillenberger, Graham R. Nudd |
IPDPS | 4 |
| 2007 | Distributed Broadcast Scheduling in Mobile Ad Hoc Networks with Unknown TopologiesabstractBroadcasting is a fundamental communication task in mobile ad hoc networks, and minimizing broadcasting time (or latency) is crucial to the performance ofmany applications. Extensive studies have been conducted on the minimization of broadcasting time in the context of radio networks, which are usually modeled as general graphs. In this paper, we consider how to achieve this goal with distributed algorithms based on a more realistic (and restricted) network model. We propose a randomized algorithm that completes broadcasting in O(D log(n/D)+log2 n) time, where n is the number of nodes in the network and D the eccentricity (maximum distancefrom the source node to any other node). Compared with a previous optimal algorithm that achieves the same result for general networks, our algorithm obviates the need to know the network eccentricity D beforehand We also propose a deterministic broadcasting algorithm that works in O(n) time, which is in contrast with the best known result of O(n log2 D) for general networks. Guang Tan, Stephen A. Jarvis, James Wen Jun Xue, Simon D. Hammond |
IPDPS | 2 |
| 2007 | Distributed Broadcast Scheduling in Mobile Ad Hoc Networks with Unknown TopologiesabstractBroadcasting is a fundamental communication task in mobile ad hoc networks, and minimizing broadcasting time (or latency) is crucial to the performance ofmany applications. Extensive studies have been conducted on the minimization of broadcasting time in the context of radio networks, which are usually modeled as general graphs. In this paper, we consider how to achieve this goal with distributed algorithms based on a more realistic (and restricted) network model. We propose a randomized algorithm that completes broadcasting in O(D log(n/D)+log2 n) time, where n is the number of nodes in the network and D the eccentricity (maximum distancefrom the source node to any other node). Compared with a previous optimal algorithm that achieves the same result for general networks, our algorithm obviates the need to know the network eccentricity D beforehand We also propose a deterministic broadcasting algorithm that works in O(n) time, which is in contrast with the best known result of O(n log2 D) for general networks. Guang Tan, Stephen A. Jarvis, James Wen Jun Xue, Simon D. Hammond |
IPDPS | 2 |
| 2007 | Distributed Arbitrary Segment Trees: Providing Efficient Range Query Support over Public DHT ServicesabstractIn this paper we define a Distributed Arbitrary Segment Tree (DAST), a distributed tree-like structure that layers the range query processing mechanism over public Distributed Hash Table (DHT) services. Compared with traditional segment trees, the arbitrary segment tree used by a DAST reduces the number of key-space segments that need to be maintained, which in turn results in fewer query operations and lower overheads. Moreover, considering that range queries often contain redundant entries that the clients do not need, we introduce the concept of accuracy of results (AoR) for range queries. We demonstrate that by adjusting AoR, the DHT operational overhead can be improved. DAST is implemented on a well-known public DHT service (OpenDHT) and validation through experimentation and supporting simulation is performed. The results demonstrate the effectiveness of DAST over exiting methods. Xinuo Chen, Stephen A. Jarvis |
PIMRC | 2 |
| 2007 | Improving the Fault Resilience of Overlay Multicast for Media StreamingabstractA key technical challenge for overlay multicast is that the highly dynamic multicast members can make data delivery unreliable. In this paper, we address this issue in the context of live media streaming by exploring 1) how to construct a stable multicast tree that minimizes the negative impact of frequent member departures on an existing overlay and 2) how to efficiently recover from packet errors caused by end-system or network failures. For the first problem, we identify two layout schemes for the tree nodes, namely, the bandwidth-ordered tree and the time-ordered tree, which represent two typical approaches to improving tree reliability, and conduct a stochastic analysis on their properties regarding reliability and tree depth. Based on the findings, we propose a distributed reliability-oriented switching tree (ROST) algorithm that minimizes the failure correlation among tree nodes. Compared with some commonly used distributed algorithms, the ROST algorithm significantly improves tree reliability and reduces average service delay, while incurring only a small protocol overhead; furthermore, it features a mechanism that prevents cheating or malicious behaviors in the exchange of bandwidth/time information. For the second problem, we develop a simple cooperative error recovery (CER) protocol that helps recover from packet errors efficiently. Recognizing that a single recovery source is usually incapable of providing the timely delivery of the lost data, the protocol recovers from data outages using the residual bandwidths from multiple sources, which are identified using a minimum-loss-correlation algorithm. Extensive simulations demonstrate the effectiveness of the proposed schemes Guang Tan, Stephen A. Jarvis |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2006 | Predictive Performance Analysis of a Parallel Pipelined Synchronous Wavefront Application for Commodity Processor Cluster SystemsabstractThis paper details the development and application of a model for predictive performance analysis of a pipelined synchronous wavefront application running on commodity processor cluster systems. The performance model builds on existing work (Cao et al.) by including extensions for modern commodity processor architectures. These extensions, including coarser hardware benchmarking, prove to be essential in countering the effects of modern superscalar processors (e.g. multiple operation pipelines and on-the-fly optimisations), complex memory hierarchies, and the impact of applying modern optimising compilers. The process of application modelling is also extended, combining static source code analysis with run-time profiling results for increased accuracy. The model is validated on several high performance SMP systems and the results show a high predictive accuracy (les 10% error). Additionally, the use of the performance model to speculate on the performance and scalability of this application on a hypothetical cluster with two different problem sizes is demonstrated. It is shown that such speculative techniques can be used to support system procurement, run-time verification and system maintenance and upgrading Gihan R. Mudalige, Stephen A. Jarvis, Daniel P. Spooner, Graham R. Nudd |
CLUSTER | 2 |
| 2006 | Improving the Fault Resilience of Overlay Multicast for Media StreamingabstractThis paper addresses the problem of fault resilience of overlay-based live media streaming from two aspects: (1) how to construct a stable multicast tree that minimizes the negative impact of frequent member departures on existing overlay, and (2) how to efficiently recover from packet errors caused by end-system or network failures. In particular, this paper makes two contributions: (1) A distributed Reliability-Oriented Switching Tree (ROST) algorithm that minimizes the failure correlation among tree nodes. By exploiting both bandwidth and time properties, the algorithm constructs a more reliable multicast tree than existing algorithms that solely minimize tree depth, while not compromising the quality of the tree in terms of service delay and incurring only a small protocol overhead; (2) A simple Cooperative Error Recovery (CER) protocol that helps recover from packet errors efficiently. Recognizing that a single recovery source is usually incapable of providing timely delivery of the lost data, the protocol recovers from data outages using the residual bandwidths from multiple sources, which are identified using a minimum-losscorrelation algorithm. Extensive simulations are conducted to demonstrate the effectiveness of the proposed schemes. Guang Tan, Stephen A. Jarvis, Daniel P. Spooner |
DSN | 2 |
| 2006 | Inter-Overlay Cooperation in High-Bandwidth Overlay MulticastabstractThe cooperation of end users can be exploited to boost the performance of high-bandwidth multicast. While intra-overlay cooperation, the mechanism for cooperation within a single overlay (multicast group), has been extensively studied, little attention has been paid to inter-overlay cooperation. In this paper we explore the possibility and effects of cooperation among co-existing heterogeneous overlays in the context of live media streaming, where bandwidth is the bottleneck resource. To motivate such a kind of cooperation, we design a reputation-based incentive mechanism that differentiates user' streaming qualities based on the amount of data actually forwarded by individual users. This not only stimulates users to contribute as much forwarding bandwidth as possible, but also motivates those with spare bandwidths in resource-rich overlays to find downstream users in external, often resource-poor, overlays so as to accumulate more reputation scores. Under this mechanism, an adaptive bandwidth exporting/reclaiming algorithm is developed which allows users to dynamically allocate bandwidth according to the resource availability of multiple overlays. Simulation results are reported with enhanced system performance in terms of users' average media quality Guang Tan, Stephen A. Jarvis |
ICPP | 2 |
| 2006 | Performance evaluation of scheduling applications with DAG topologies on multiclusters with independent local schedulersabstractBefore an application modelled as a directed acyclic graph (DAG) is executed on a heterogeneous system, a DAG mapping policy is often enacted. After mapping, the tasks (in the DAG-based application) to be executed at each computational resource are determined. The tasks are then sent to the corresponding resources, where they are orchestrated in the pre-designed pattern to complete the work. Most DAG mapping policies in the literature assume that each computational resource is a processing node of a single processor, i.e. the tasks mapped to a resource are to be run in sequence. Our studies demonstrate that if the resource is actually a cluster with multiple processing nodes, this assumption will cause a misperception in the tasks' execution time and execution order. This will disturb the pre-designed cooperation among tasks so that the expected performance cannot be achieved. In this paper, a DAG mapping algorithm is presented for multicluster architectures. Each constituent cluster in the multicluster is shared by background workload (from other users) and has its own independent local scheduler. The multicluster DAG mapping policy is based on theoretical analysis and its performance is evaluated through extensive experimental studies. The results show that compared with conventional DAG mapping policies, the new scheme that we present can significantly improve the scheduling performance of a DAG-based application in terms of the schedule length. Ligang He, Stephen A. Jarvis, Daniel P. Spooner, Graham R. Nudd |
IPDPS | 2 |
| 2006 | A Payment-based Incentive and Service Differentiation Mechanism for Peer-to-Peer Streaming BroadcastabstractWe proposes a novel payment-based incentive mechanism for peer-to-peer (P2P) live media streaming. Using this approach, peers earn points by forwarding data to others; the data streaming is divided into fixed length periods, during each of which peers compete with each other for good parents (data suppliers) for the next period in a first-price auction like procedure using their points. We design a distributed algorithm to regulate peer competitions, and consider various individual strategies for parent selection from a game theoretic perspective. We then discuss possible strategies that can be used to maximize a peer's expected media quality by planning different bids for its substreams. Finally, in order to encourage off-session users to keep staying online and continue contributing to the network, we develop an optimal data forwarding strategy that allows peers to accumulate points that can be used in future services. Simulations results show that proposed methods effectively differentiate the media qualities received by peers making different contributions (which originate from, for example, different forwarding band-widths or servicing times), and at the same time maintaining a high system-wide performance Guang Tan, Stephen A. Jarvis, Daniel P. Spooner |
IWQoS | 2 |
| 2006 | Performance prediction and its use in parallel and distributed computing systems
Stephen A. Jarvis, Daniel P. Spooner, Hélène N. Lim Choi Keung, Subhash Saini, Graham R. Nudd |
Future Gener. Comput. Syst. | 1 |
| 2006 | Prediction of short-lived TCP transfer latency on bandwidth asymmetric links
Guang Tan, Stephen A. Jarvis |
J. Comput. Syst. Sci. | 2 |
| 2006 | Allocating Non-Real-Time and Soft Real-Time Jobs in MulticlustersabstractThis paper addresses workload allocation techniques for two types of sequential jobs that might be found in multicluster systems, namely, non-real-time jobs and soft real-time jobs. Two workload allocation strategies, the optimized mean response time (ORT) and the optimized mean miss rate (OMR), are developed by establishing and numerically solving two optimization equation sets. The ORT strategy achieves an optimized mean response time for non-real-time jobs, while the OMR strategy obtains an optimized mean miss rate for soft real-time jobs over multiple clusters. Both strategies take into account average system behaviors (such as the mean arrival rate of jobs) in calculating the workload proportions for individual clusters and the workload allocation is updated dynamically when the change in the mean arrival rate reaches a certain threshold. The effectiveness of both strategies is demonstrated through theoretical analysis. These strategies are also evaluated through extensive experimental studies and the results show that when compared with traditional strategies, the proposed workload allocation schemes significantly improve the performance of job scheduling in multiclusters, both in terms of the mean response time (for non-real-time jobs) and the mean miss rate (for soft real-time jobs). Ligang He, Stephen A. Jarvis, Daniel P. Spooner, Donna Dillenberger, Graham R. Nudd |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2005 | Mapping DAG-based applications to multiclusters with background workloadabstractBefore an application modelled as a directed acyclic graph (DAG) is executed on a heterogeneous system, a DAG mapping policy is often enacted. After mapping, the tasks (in the DAG-based application) to be executed at each computational resource are determined. The tasks are then sent to the corresponding resources, where they are orchestrated in the pre-designed pattern to complete the work. Most DAG mapping policies in the literature assume that each computational resource is a processing node of a single processor, i.e. the tasks mapped to a resource are to be run in sequence. Our studies demonstrate that if the resource is actually a cluster with multiple processing nodes, this assumption will cause a mis-perception in the tasks' execution time and execution order. This will disturb the pre-designed cooperation among tasks so that the expected performance cannot be achieved. In this paper, a DAG mapping algorithm is presented for multicluster architectures. Each constituent cluster in the multicluster is shared by background workload (from other users) and has its own independent local scheduler. The multicluster DAG mapping policy is based on theoretical analysis and its performance is evaluated through extensive experimental studies. The results show that compared with conventional DAG mapping policies, the new scheme that we present can significantly improve the scheduling performance of a DAG-based application in terms of the schedule length. Ligang He, Stephen A. Jarvis, Daniel P. Spooner, David A. Bacigalupo, Guang Tan, Graham R. Nudd |
CCGRID | 2 |
| 2005 | Performance Analysis and Improvement of Overlay Construction for Peer-to-Peer Live Media StreamingabstractFor single-source, single-tree based peer-to-peer live media streaming, it is generally believed that a short (and wide) tree has a good comprehensive performance in terms of tree reliability and service delay. While the short tree directly benefits delay optimization, it is unclear whether such a structure maximizes tree reliability, which is sometimes more critical for a streaming Internet service. This paper studies several prevalent overlay construction algorithms in terms of (I) service reliability; (2) service delay and (3) protocol overhead. Two types of peer layout, bandwidth-ordered layout and time-ordered layout, are identified and their performance is evaluated. The analytical results show that, by appropriately placing peers according to their time properties, the tree can be much more reliable than a depth-optimized tree. We therefore propose a heap algorithm, which aims for combining the strengths of both bandwidth ordering and time ordering. It dynamically moves peers between difference layers of the tree according to a simple metric, and gradually adjusts the tree toward a layout partially ordered in time and partially ordered in bandwidth. In so doing the tree has advantages in both service reliability and delay, and maintains small protocol overheads. Extensive simulations demonstrate the effectiveness of this new algorithm. Guang Tan, Stephen A. Jarvis, Xinuo Chen, Daniel P. Spooner, Graham R. Nudd |
MASCOTS | 2 |
| 2005 | Performance-Aware Workflow Management for Grid ComputingabstractGrid middleware development has advanced rapidly over the past few years to support component-based programming models and service-oriented architectures. This is most evident with the forthcoming release of the Globus toolkit (GT4), which represents a convergence of concepts (and standards) from both the grid and web-services communities. Grid applications are increasingly modular, composed of workflow descriptions that feature both resource and application dynamism. Understanding the performance implications of scheduling grid workflows is critical in providing effective resource management and reliable service quality to users. This paper describes a series of extensions to an existing performance-aware grid management system (TITAN). These extensions provide additional support for workflow prediction and scheduling using a multi-domain performance management infrastructure. Daniel P. Spooner, Stephen A. Jarvis, Ligang He, Graham R. Nudd |
Comput. J. | 3 |
| 2005 | Performance-based middleware for Grid computingabstractAbstract This paper describes a stateful service‐oriented middleware infrastructure for the management of scientific tasks running on multi‐domain heterogeneous distributed architectures. Allocating scientific workload across multiple administrative boundaries is a key issue in Grid computing and as a result a number of supporting services including match‐making, scheduling and staging have been developed. Each of these services allows the scientist to utilize the available resources, although a sustainable level of service in such shared environments cannot always be guaranteed. A performance‐based middleware infrastructure is described in which prediction data for each scientific task are calculated, stored and published through a Globus‐based performance information service. Distributing these data allows additional performance‐based middleware services to be built, two of which are described in this paper: an intra‐domain predictive co‐scheduler and a multi‐domain workload steering system. These additional facilities significantly improve the ability of the system to meet task deadlines, as well as enhancing inter‐domain load‐balancing and system‐wide resource utilization. Copyright © 2005 John Wiley & Sons, Ltd. Graham R. Nudd, Stephen A. Jarvis |
Concurr. Pract. Exp. | 2 |
| 2005 | Grid load balancing using intelligent agents
Daniel P. Spooner, Stephen A. Jarvis, Graham R. Nudd |
Future Gener. Comput. Syst. | 3 |
| 2005 | The impact of predictive inaccuracies on execution scheduling
Stephen A. Jarvis, Ligang He, Daniel P. Spooner, Graham R. Nudd |
Perform. Evaluation | 1 |
| 2005 | An Investigation into the Application of Different Performance Prediction Methods to Distributed Enterprise Applications
David A. Bacigalupo, Stephen A. Jarvis, Ligang He, Daniel P. Spooner, Donna Dillenberger, Graham R. Nudd |
J. Supercomput. | 2 |
| 2004 | An Investigation into the Application of Different Performance Prediction Techniques to e-Commerce ApplicationsabstractSummary form only given. Predictive performance models of e-Commerce applications allows grid workload managers to provide e-Commerce clients with qualities of service (QoS) whilst making efficient use of resources. We demonstrate the use of two 'coarse-grained' modelling approaches (based on layered queuing modelling and historical performance data analysis) for predicting the performance of dynamic e-Commerce systems on heterogeneous servers. Results for a popular e-Commerce benchmark show how request response times and server throughputs can be predicted on servers with heterogeneous CPUs at different background loads. The two approaches are compared and their usefulness to grid workload management is considered. David A. Bacigalupo, Stephen A. Jarvis, Ligang He, Graham R. Nudd |
IPDPS | 2 |
| 2004 | Optimising Static Workload Allocation in MulticlustersabstractSummary form only given. Workload allocation and job dispatching are two fundamental components in static job scheduling for distributed systems. We address the static workload allocation techniques for two types of job stream in multicluster systems, namely, nonreal-time job streams and soft-real-time job streams, which request different qualities of service. Two workload allocation strategies (called ORT and OMR) are developed by establishing and numerically solving two optimisation equation sets. The ORT strategy achieves the optimised mean response time for the nonreal-time job stream; while the OMR strategy can gain the optimised mean miss rate for the soft-real-time job stream over multiple clusters (these strategies can also be applied in a single cluster system). The effectiveness of both strategies is demonstrated through theoretical analysis. The proposed workload allocation schemes are combined with two job dispatching strategies (weighted random and weighted round-robin) to generate new static job scheduling algorithms for multicluster environments. These algorithms are evaluated through extensive experimental studies and the results show that compared with static approaches without the optimisation techniques, the proposed workload allocation schemes can significantly improve the performance of static job scheduling in multiclusters, in terms of both the mean response time (for the nonreal-time jobs) and the mean miss rate (for soft-real-time jobs). Ligang He, Stephen A. Jarvis, Daniel P. Spooner, Graham R. Nudd |
IPDPS | 2 |
| 2004 | Performance-Aware Load Balancing for Multiclusters
Ligang He, Stephen A. Jarvis, David A. Bacigalupo, Daniel P. Spooner, Graham R. Nudd |
ISPA | 2 |
| 2003 | GridFlow: Workflow Management for Grid ComputingabstractGrid computing is becoming a mainstream technology for large-scale distributed resource sharing and system integration. Workflow management is emerging as one of the most important grid services. In this work, a workflow management system for grid computing, called GridFlow, is presented, including a user portal and services of both global grid workflow management and local grid sub-workflow scheduling. Simulation, execution and monitoring functionalities are provided at the global grid level, which work on top of an existing agent-based grid resource management system. At each local grid, sub-workflow scheduling and conflict management are processed on top of an existing performance prediction based task scheduling system. A fuzzy timing technique is applied to address new challenges of workflow management in a cross-domain and highly dynamic grid environment. A case study is given and corresponding results indicate that local and global grid workflow management can coordinate with each other to optimise workflow execution time and solve conflicts of interest. Stephen A. Jarvis, Subhash Saini, Graham R. Nudd |
CCGRID | 2 |
| 2003 | Dynamic Scheduling of Parallel Real-Time Jobs by Modelling Spare Capabilities in Heterogeneous ClustersabstractIn this research, a scenario is assumed where periodic real-time jobs are being run on a heterogeneous cluster of computers, and new aperiodic parallel real-time jobs, modelled by directed acyclic graphs (DAG), arrive at the system dynamically. In the scheduling scheme presented in this paper, a global scheduler situated within the cluster schedules new jobs onto the computers by modelling their spare capabilities left by existing periodic jobs. Admission control is introduced so that new jobs are rejected if their deadlines cannot be met under the precondition of still guaranteeing the real-time requirements of existing jobs. Each computer within the cluster houses a local scheduler, which uniformly schedules both periodic job instances and the subtasks in the parallel realtime jobs using an early deadline first policy. The modelling of the spare capabilities is optimal in the sense that once a new task starts running on a computer, it will utilize all the spare capability left by the periodic real-time jobs and its finish time is the earliest possible. The performance of the proposed modelling approach and scheduling scheme is evaluated by extensive simulation; results show that the system utilization is significantly enhanced, while the real-time requirements of the existing jobs remain guaranteed. Ligang He, Stephen A. Jarvis, Daniel P. Spooner, Graham R. Nudd |
CLUSTER | 2 |
| 2003 | Performance-Based Dynamic Scheduling of Hybrid Real-Time Applications on a Cluster of Hetrogeneous Workstations
Ligang He, Stephen A. Jarvis, Daniel P. Spooner, Graham R. Nudd |
Euro-Par | 2 |
| 2002 | Agent-Based Resource Management for Grid ComputingabstractIt is envisaged that the grid infrastructure will be a large-scale distributed software system that will provide high-end computational and storage capabilities to differentiated users. A number of distributed computing technologies are being applied to grid development work, including CORBA and Jini. In this work, we introduce an A4 (Agile Architecture and Autonomous Agents) methodology, which can be used for resource management for grid computing. An initial system implementation utilises the performance prediction techniques of the PACE toolkit to provide quantitative data regarding the performance of complex applications running on local grid resources. At the meta-level, a hierarchy of identical agents is used to provide an abstraction of the system architecture. Each agent is able to cooperate with other agents to provide service advertisement and discovery to schedule applications that need to utilise grid resources. A performance monitor and advisor (PMA) is in development to optimize the performance of agent behaviours. Daniel P. Spooner, James D. Turner, Stephen A. Jarvis, Darren J. Kerbyson, Subhash Saini, Graham R. Nudd |
CCGRID | 4 |
| 2002 | Portable and architecture independent parallel performance tuning using BSP
Stephen A. Jarvis, Jonathan M. D. Hill, Constantinos J. Siniolakis, Vasil P. Vasilev |
Parallel Comput. | 1 |
| 1998 | Analysing an SQL Application with a BSPlib Call-Graph Profiling Tool
Jonathan M. D. Hill, Stephen A. Jarvis, Constantinos J. Siniolakis, Vasil P. Vasilev |
Euro-Par | 2 |
| 1998 | Profiling Large-Scale Lazy Functional ProgramsabstractThe LOLITA natural language processor is an example of one of the ever-increasing number of large-scale systems written entirely in a functional programming language. The system consists of over 47,000 lines of Haskell code (excluding comments) and is able to perform a wide range of tasks such as semantic and pragmatic analysis of text, information extraction and query analysis. The efficiency of such a system is critical; interactive tasks (such as query analysis) must ensure that the user is not inconvenienced by long pauses, and batch mode tasks (such as information extraction) must ensure that an adequate throughput can be achieved. For the past three years the profiling tools supplied with GHC and HBC have been used to analyse and reason about the complexity of the LOLITA system. There have been good results, however experience has shown that in a large system the profiling life-cycle is often too long to make detailed analysis possible, and the results are often misleading. In response to these problems a profiler has been developed which allows the complete set of program costs to be recorded in so-called cost-centre stacks. These program costs are then analysed using a post-processing tool to allow the developer to explore the costs of the program in ways that are either not possible with existing tools or would require repeated compilations and executions of the program. The modifications to the Glasgow Haskell compiler based on detailed cost semantics and an efficient implementation scheme are discussed. The results of using this new profiling tool in the analysis of a number of Haskell programs are also presented. The overheads of the scheme are discussed and the benefits of this new system are considered. An outline is also given of how this approach can be modified to assist with the tracing and debugging of programs. Richard G. Morgan, Stephen A. Jarvis |
J. Funct. Program. | 2 |
| 1995 | Handling Communications in Concurrent KBS
Albert Bokma, M. Huiban, Andrew Slade, Stephen A. Jarvis |
IEA/AIE | 4 |