Stephen Booth

dblp:89/708 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
0since 2021 · last 2017
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Parallel and multicore computing · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing › parallel libraries
communication library
0.012000
Single sided MPI implementations for SUN MPI · SC 2000
Parallel and multicore computing › parallel programming models
message passing
0.012000
Single sided MPI implementations for SUN MPI · SC 2000
Parallel and multicore computing › parallel programming models › message passing
MPI one-sided communication
0.012000
Single sided MPI implementations for SUN MPI · SC 2000
Parallel and multicore computing › parallel computing › parallel communication
shared-memory communication
0.012000
Single sided MPI implementations for SUN MPI · SC 2000

Methods — techniques the papers use, named apart from their topics

type packing · 0.0caching · 0.0
YearPublicationVenuePosition
2017 Generalisation of recursive doubling for AllReduce: Now with simulation
abstract
The performance of AllReduce is crucial at scale. The recursive doubling with pairwise exchange algorithm theoretically achieves O (log 2 N ) scaling for short messages with N peers, but is limited by improvements in network latency. A multi-way exchange can be implemented using message pipelining, which is easier to improve than latency. Using our method, recursive multiplying, we show reductions in execution time of between 8% and 40% of AllReduce on a Cray XC30 over recursive doubling. Using a custom simulator we further explore the dynamics of recursive multiplying.
Martin Ruefenacht, Mark Bull, Stephen Booth
Parallel Comput.3
2016 Generalisation of Recursive Doubling for AllReduce
abstract
The performance of AllReduce is crucial at scale. The recursive doubling with pairwise exchange algorithm theoretically achieves O(log2 N) scaling for short messages with N peers, but is limited by improvements in network latency. A multi-way exchange can be implemented using message pipelining, which is easier to improve than latency. Using our method, recursive multiplying, we show reductions in execution time of between 8% and 40% of AllReduce on a Cray XC30 over recursive doubling.
Martin Ruefenacht, Mark Bull, Stephen Booth
EuroMPI3
2013 McMPI: a managed-code MPI library in pure C#
abstract
This paper presents McMPI, an entirely new MPI library written in C# using only safe managed-code, and performance results from low-level benchmarks demonstrating ping-pong latency and bandwidth comparable with MS-MPI and MPICH2.
Daniel J. Holmes, Stephen Booth
EuroMPI2
2011 Parallel Optimisation Strategies for Fusion Codes
abstract
We have previously documented the on-going work in the EUFORIA project to parallelise and optimise European fusion simulation codes, see. This involves working with a wide range of codes to try and address any performance and scaling issues that these codes have. However, as no two simulation codes are exactly the same, it is very hard to apply exactly the same approach to optimising a disparate range of codes. Indeed, it can be seen from that the codes investigated range in terms of performance and ability from well-optimised, highly parallelised codes, to serial or poorly performing codes. After analysing, optimising, and parallelising a range of codes it is, actually, possible to discern a number of distinct optimisation techniques or approaches/strategies that can be used to improve the performance or scaling of a parallel simulation code. This paper outlines the distinct approaches that we have identified, highlighting their benefits and drawbacks, giving an overview of the type of work that is often attempted for fusion simulation code optimisation.
Adrian Jackson, Fiona Reid, Stephen Booth, Joachim Hein, Jan Westerholm, Mats Aspnäs, Miquel Catala, Alejandro Soba
PDP3
2005 HPCx: towards capability computing
abstract
Abstract We introduce HPCx—the U.K.'s new National HPC Service—which aims to deliver a world‐class service for capability computing to the U.K. scientific community. HPCx is targeting an environment that will both result in world‐leading science and address the challenges involved in scaling existing codes to the capability levels required. Close working relationships with scientific consortia and user groups throughout the research process will be a central feature of the service. A significant number of key user applications have already been ported to the system. We present initial benchmark results from this process and discuss the optimization of the codes and the performance levels achieved on HPCx in comparison with other systems. We find a range of performance with some algorithms scaling far better than others. Copyright © 2005 John Wiley & Sons, Ltd.
Mike Ashworth, Ian J. Bush, Martyn F. Guest, Andrew G. Sunderland, Stephen Booth, Joachim Hein, Lorna Smith, Kevin Stratford, Alessandro Curioni
Concurr. Pract. Exp.5
2001 Optimising the MPI Library for the T3E
Stephen Booth
Euro-Par1
2000 Single sided MPI implementations for SUN MPI
abstract
This paper describes an implementation of generic MPI-2 single sided communications for SUN-MPI. Our implementation is layered on top of point-to-point MPI communications and therefore can be adapted to other MPI implementations. The code is designed to co-exist with other MPI-2 single sided implementations (for example direct use of shared memory) providing a generic fall-back implementation for those communication paths where an optimised single-sided implementation is not available. MPI-2 single sided communications require the transfer of data-type information as well as user data. We describe a type packing and caching mechanism used to optimise the transfer of data-type information. The performance of this implementation is measured in comparison to equivalent point to point operations and the shared memory implementation provided by SUN.
Stephen Booth, Fernando Elson Mourão
SC1