Steven A. Wright 0001

dblp:11/10822 · DBLP profile ↗
← Back
17ranked-venue papers
4as first author
4since 2021 · last 2026
0000-0001-7133-8533ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author
YearPublicationVenuePosition
2026 The Performance-Power Frontier: A Model-Driven Approach to Energy-Aware Application Optimisation
abstract
Energy consumption in high-performance computing (HPC) has emerged as a primary design constraint, with Tier-0 systems now demanding tens of megawatts. While users are increasingly being encouraged to minimise the environmental and operational costs of their workloads, they frequently lack the analytical frameworks required to navigate the complex trade-offs between execution time and energy efficiency. In this paper, we extend the Power-Optimised Software Envelope (POSE) model and introduce a visualisation tool designed to help performance engineers explore the energy-time design space. By aligning multi-objective optimisation metrics with institutional operational priorities, our framework enables policy-aware decision-making at the application level. We demonstrate this approach by analysing the NAS Parallel Benchmarks. Our results reveal a significant variance in energy-aware optimisation potential: while the Embarrassingly Parallel (EP) kernel shows negligible opportunity for gains (up to 0.37% improvement in Energy-Delay Summation), the Conjugate Gradient (CG) kernel offers greater scope (up to a 12.96% improvement). Furthermore, we study the benchmarks under frequency and power constraints and identify critical efficiency thresholds; for the Block Tri-Diagonal (BT) solver, configurations operating below 2.00 GHz or under a 140 W powercap fail to outperform the baseline efficiency, highlighting the risks of aggressive over-throttling.
Suryachandra Prasad Pasupuleti, Steven A. Wright 0001
ICS2
2025 The P3 Explorer: An Open Database of Performance, Portability, and Productivity
abstract
This paper introduces a web-based tool designed to organise and present visual representations of performance, portability, and productivity (P3) data from previously published scientific studies. The P3 Explorer operates as both an open repository of application performance data and a data dashboard, providing visual heuristic analyses of performance portability and developer productivity, created using Intel’s P3 Analysis library. The goal of the project is to create a community-led database of P3 studies to better inform application developers of alternative approaches to developing new applications that target high performance on diverse hardware, considering the productivity of the developer. In this paper, we evaluate our tool using a recently published study outlining a performance portable domain-specific language for particle-in-cell applications.
Steven A. Wright 0001, Zaman Lantra, Gihan R. Mudalige
PDP2
2024 OP-PIC - an Unstructured-Mesh Particle-in-Cell DSL for Developing Nuclear Fusion Simulations
abstract
Particle-in-Cell (PIC) applications form a core simulation component for designing fusion reactors and their efficient use. In this work, we introduce OP-PIC, a new embedded Domain Specific Language (DSL) for developing unstructured-mesh PIC applications. The DSL is aimed at gaining performance portability for PIC codes on current and emerging, massively parallel architectures. We investigate and bring together the state-of-the-art in PIC solver parallelization techniques, refactoring them within a multi-layered DSL. OP-PIC use source-to-source translation to generate platform-specific optimizations. These parallelizations can be reused for any application declared using the DSL’s high-level API. We showcase the performance and portability of two non-trivial PIC applications developed with OP-PIC on multiple CPU and GPU clusters, employing a number of parallelization techniques, including OpenMP, CUDA, HIP and their combinations with distributed memory (MPI) parallelization. We benchmark the OP-PIC generated code on a range of single node systems and a number of distributed-memory systems, including an AMD EPYC CPU-based HPE-Cray EX cluster, an NVIDIA V100 GPU cluster, and an AMD MI250X GPU cluster, exploring both single node and scaling performance. Results demonstrate the flexibility of the DSL to implement radically different optimizations for each platform, showing between 1.4 × to 3.5 × speed-ups with GPUs compared to CPUs on power equivalent systems and good weak-scaling to over 10 billion particles.
Zaman Lantra, Steven A. Wright 0001, Gihan R. Mudalige
ICPP2
2022 Special issue on Performance modeling, benchmarking, and simulation of high performance computing systems
Steven A. Wright 0001
Concurr. Comput. Pract. Exp.1
2020 "Clik-Thru" Terms of Service: Blockchain smart contracts to improve consumer engagement?
abstract
Technology entrepreneurship has enabled the widespread commercial adoption of internet technologies. These internet technologies have reformed consumer commercial experiences towards an online environment. The pervasiveness of the online experience raises the importance of protecting the consumer in the online context. Online services are typically delivered under “Clik-Thru” terms of service developed by the service provider alone; and accepted by the consumer with a single click and little if any consideration. The successful adoption of new internet-based technologies and commercial practices has encouraged more technology entrepreneurship in a positive feedback cycle. Efforts at improved readability are insufficient to engage consumers with these “Clik-Thru” contracts. This paper argues that some efforts at increasing consumer engagement with the “Clik-Thru” terms of service may be a useful and tractable step towards improved consumer experiences. Blockchain smart contracts appear to provide promising capabilities to enable greater consumer engagement with “Clik-Thru” contracts.
Steven A. Wright 0001
ISTAS1
2020 An unstructured CFD mini-application for the performance prediction of a production CFD code
abstract
Summary Maintaining the performance of large scientific codes is a difficult task. To aid in this task, a number of mini‐applications have been developed that are more tractable to analyze than large‐scale production codes while retaining the performance characteristics of them. These “mini‐apps” also enable faster hardware evaluation and, for sensitive commercial codes, allow evaluation of code and system changes outside of access approval processes. In this paper, we develop MG‐CFD, a mini‐application that represents a geometric multigrid, unstructured computational fluid dynamics (CFD) code, designed to exhibit similar performance characteristics without sharing commercially sensitive code. We detail our experiences of developing this application using guidelines detailed in existing research and contributing further to these. Our application is validated against the inviscid flux routine of HYDRA, a CFD code developed by Rolls‐Royce plc for turbomachinery design. This paper (1) documents the development of MG‐CFD, (2) introduces an associated performance model with which it is possible to assess the performance of HYDRA on new HPC architectures, and (3) demonstrates that it is possible to use MG‐CFD and the performance models to predict the performance of HYDRA with a mean error of 9.2% for strong‐scaling studies.
Andrew Owenson, Steven A. Wright 0001, Richard A. Bunt, Yoon Ho, Matthew J. Street, Stephen A. Jarvis
Concurr. Comput. Pract. Exp.2
2019 Performance Modeling, Benchmarking and Simulation of High Performance Computing Systems
Steven A. Wright 0001
Future Gener. Comput. Syst.1
2019 An algorithm for computing short-range forces in molecular dynamics simulations with non-uniform particle densities
abstract
We present projection sorting, an algorithmic approach to determining pairwise short-range forces between particles in molecular dynamics simulations. We show it can be more effective than the standard approaches when particle density is non-uniform. We implement tuned versions of the algorithm in the context of a biophysical simulation of chromosome condensation, for the modern Intel Broadwell and Knights Landing architectures, across multiple nodes. We demonstrate up to 5 × overall speedup and good scaling to large problem sizes and processor counts.
Timothy R. Law, Jonny Hancox, Steven A. Wright 0001, Stephen A. Jarvis
J. Parallel Distributed Comput.3
2019 The Power-optimised Software Envelope
abstract
Advances in processor design have delivered performance improvements for decades. As physical limits are reached, refinements to the same basic technologies are beginning to yield diminishing returns. Unsustainable increases in energy consumption are forcing hardware manufacturers to prioritise energy efficiency in their designs. Research suggests that software modifications may be needed to exploit the resulting improvements in current and future hardware. New tools are required to capitalise on this new class of optimisation. In this article, we present the Power Optimised Software Envelope (POSE) model, which allows developers to assess the potential benefits of power optimisation for their applications. The POSE model is metric agnostic and in this article, we provide derivations using the established Energy-Delay Product metric and the novel Energy-Delay Sum and Energy-Delay Distance metrics that we believe are more appropriate for energy-aware optimisation efforts. We demonstrate POSE on three platforms by studying the optimisation characteristics of applications from the Mantevo benchmark suite. Our results show that the Pathfinder application has very little scope for power optimisation while TeaLeaf has the most, with all other applications in the benchmark suite falling between the two. Finally, we extend our POSE model with a formulation known as System Summary POSE—a meta-heuristic that allows developers to assess the scope a system has for energy-aware software optimisation independent of the code being run.
Stephen I. Roberts, Steven A. Wright 0001, Suhaib A. Fahmy, Stephen A. Jarvis
ACM Trans. Archit. Code Optim.2
2018 BookLeaf: An Unstructured Hydrodynamics Mini-Application
abstract
With the age of Exascale computing causing a diversification away from traditional CPU-based homogeneous clusters, it is becoming increasingly difficult to ensure that computationally complex codes are able to run on these emerging architectures. This is especially important for large physics simulations that are themselves becoming increasingly complex and computationally expensive. One proposed solution to the problem of ensuring these applications can run on the desired architectures is to develop representative mini-applications that are simpler and so can be ported to new frameworks more easily, but which are also representative of the algorithmic and performance characteristics of the original applications. In this paper we present BookLeaf, an unstructured Arbitrary Lagrangian-Eulerian mini-application to add to the suite of representative applications developed and maintained by the UK Mini-App Consortium (UK-MAC). First, we outline the reference implementation of our application in Fortran. We then discuss a number of alternative implementations using a variety of parallel programming models and discuss the issues that arise when porting such an application to new architectures. To demonstrate our implementation, we present a study of the performance of BookLeaf on number of platforms using alternative designs, and we document a scaling study showing the behaviour of the application at scale.
David Truby, Steven A. Wright 0001, Robert Kevis, Satheesh Maheswaran, Andrew Herdman, Stephen A. Jarvis
CLUSTER2
2018 Developing and Using a Geometric Multigrid, Unstructured Grid Mini-Application to Assess Many-Core Architectures
abstract
Achieving high-performance of large scientific codes is a difficult task. This has led to the development of numerous mini-applications that are more tractable to analyse, while retaining performance characteristics of their full-sized counterparts. These "mini-apps" also enable faster hardware evaluation, and for sensitive codes allow evaluation of systems outside of access approval processes. In this paper we develop a mini-application of a geometric multigrid, unstructured grid Computational Fluid Dynamics (CFD) code, designed to exhibit similar performance characteristics without sharing code. We detail our experiences developing this application, using guidelines detailed in existing research, and contribute further additions to these to aid future mini-application developers. Our application is validated against the inviscid flux routine of HYDRA, a CFD code developed by Rolls-Royce, which confirms that the parent kernel and mini-application share fundamental causes of parallel inefficiency. We then use the mini-application to assess the impact of Intel's Knights Landing (KNL) on performance. We find that the mini-app and parent kernel continue to share scaling characteristics, however a comparison with Broadwell performance exposed significant differences between the kernels that were undetected by the validation.
Andrew Owenson, Steven A. Wright 0001, Richard A. Bunt, Stephen A. Jarvis, Yoon Ho, Matthew J. Street
PDP2
2017 Achieving Performance Portability for a Heat Conduction Solver Mini-Application on Modern Multi-core Systems
abstract
Modernizing production-grade, often legacy applications to take advantage of modern multi-core and many-core architectures can be a difficult and costly undertaking. This is especially true currently, as it is unclear which architectures will dominate future systems. The complexity of these codes can mean that parallelisation for a given architecture requires significant re-engineering. One way to assess the benefit of such an exercise would be to use mini-applications that are representative of the legacy programs.In this paper, we investigate different implementations of TeaLeaf, a mini-application from the Mantevo suite that solves the linear heat conduction equation. TeaLeaf has been ported to use many parallel programming models, including OpenMP, CUDA and MPI among others. It has also been re-engineered to use the OPS embedded DSL and template libraries Kokkos and RAJA. We use these different implementations to assess the performance portability of each technique on modern multi-core systems.While manually parallelising the application targeting and optimizing for each platform gives the best performance, this has the obvious disadvantage that it requires the creation of different versions for each and every platform of interest. Frameworks such as OPS, Kokkos and RAJA can produce executables of the program automatically that achieve comparable portability. Based on a recently developed performance portability metric, our results show that OPS and RAJA achieve an application performance portability score of 71% and 77% respectively for this application.
Richard O. Kirk, Gihan R. Mudalige, István Z. Reguly, Steven A. Wright 0001, Matt Martineau, Stephen A. Jarvis
CLUSTER4
2016 Predictive Evaluation of Partitioning Algorithms through Runtime Modelling
abstract
Performance modelling unstructured mesh codes is a challenging process, due to the difficulty of capturing their memory access patterns, and their communication patterns at varying scale. In this paper we first develop extensions to an existing runtime performance model, aimed at overcoming the former, which we validate on up to 1,024 cores of a Haswellbased cluster, using both a geometric partitioning algorithm and ParMETIS to partition the input deck, with a maximum absolute runtime error of 12.63% and 11.55% respectively. To overcome the latter, we develop an application representative of the mesh partitioning process internal to an unstructured mesh code. This application is able to generate partitioning data that is usable with the performance model to produce predicted application runtimes within 7.31% of those produced using empirically collected data. We then demonstrate the use of the performance model by undertaking a predictive comparison among several partitioning algorithms on up to 30,000 cores. Additionally, we correctly predict the ineffectiveness of the geometric partitioning algorithm at 512 and 1024 cores.
Richard A. Bunt, Steven A. Wright 0001, Stephen A. Jarvis, Yoon Ho, Matthew J. Street
HiPC2
2016 Optimisation of a Molecular Dynamics Simulation of Chromosome Condensation
abstract
We present optimisations applied to a bespoke bio-physical molecular dynamics simulation designed to investigate chromosome condensation. Our primary focus is on domain-specific algorithmic improvements to determining short-range interaction forces between particles, as certain qualities of the simulation render traditional methods less effective. We implement tuned versions of the code for both traditional CPU architectures and the modern many-core architecture found in the Intel Xeon Phi coprocessor and compare their effectiveness. We achieve speed-ups starting at a factor of 10 over the original code, facilitating more detailed and larger-scale experiments.
Timothy R. Law, Jonny Hancox, Tammy M. K. Cheng, Raphael A. G. Chaleil, Steven A. Wright 0001, Paul A. Bates, Stephen A. Jarvis
SBAC-PAD5
2013 Parallel File System Analysis Through Application I/O Tracing
abstract
Input/Output (I/O) operations can represent a significant proportion of the run-time of parallel scientific computing applications. Although there have been several advances in file format libraries, file system design and I/O hardware, a growing divergence exists between the performance of parallel file systems and the compute clusters that they support. In this paper, we document the design and application of the RIOT I/O toolkit (RIOT) being developed at the University of Warwick with our industrial partners at the Atomic Weapons Establishment and Sandia National Laboratories. We use the toolkit to assess the performance of three industry-standard I/O benchmarks on three contrasting supercomputers, ranging from a mid-sized commodity cluster to a large-scale proprietary IBM BlueGene/P system. RIOT provides a powerful framework in which to analyse I/O and parallel file system behaviour—we demonstrate, for example, the large file locking overhead of IBM's General Parallel File System, which can consume nearly 30% of the total write time in the FLASH-IO benchmark. Through I/O trace analysis, we also assess the performance of HDF-5 in its default configuration, identifying a bottleneck created by the use of suboptimal Message Passing Interface hints. Furthermore, we investigate the performance gains attributed to the Parallel Log-structured File System (PLFS) being developed by EMC Corporation and the Los Alamos National Laboratory. Our evaluation of PLFS involves two high-performance computing systems with contrasting I/O backplanes and illustrates the varied improvements to I/O that result from the deployment of PLFS (ranging from up to 25× speed-up in I/O performance on a large I/O installation to 2× speed-up on the much smaller installation at the University of Warwick).
Steven A. Wright 0001, Simon D. Hammond, Simon J. Pennycook, Robert F. Bird, J. A. Herdman, I. Miller, A. Vadgama, Abhir Bhalerao, Stephen A. Jarvis
Comput. J.1
2013 An investigation of the performance portability of OpenCL
Simon J. Pennycook, Simon D. Hammond, Steven A. Wright 0001, J. A. Herdman, I. Miller, Stephen A. Jarvis
J. Parallel Distributed Comput.3
2012 On the Acceleration of Wavefront Applications using Distributed Many-Core Architectures
abstract
In this paper we investigate the use of distributed graphics processing unit (GPU)-based architectures to accelerate pipelined wavefront applications—a ubiquitous class of parallel algorithms used for the solution of a number of scientific and engineering applications. Specifically, we employ a recently developed port of the LU solver (from the NAS Parallel Benchmark suite) to investigate the performance of these algorithms on high-performance computing solutions from NVIDIA (Tesla C1060 and C2050) as well as on traditional clusters (AMD/InfiniBand and IBM BlueGene/P). Benchmark results are presented for problem classes A to C and a recently developed performance model is used to provide projections for problem classes D and E, the latter of which represents a billion-cell problem. Our results demonstrate that while the theoretical performance of GPU solutions will far exceed those of many traditional technologies, the sustained application performance is currently comparable for scientific wavefront applications. Finally, a breakdown of the GPU solution is conducted, exposing PCIe overheads and decomposition constraints. A new k-blocking strategy is proposed to improve the future performance of this class of algorithm on GPU-based architectures.
Simon J. Pennycook, Simon D. Hammond, Gihan R. Mudalige, Steven A. Wright 0001, Stephen A. Jarvis
Comput. J.4