EDBT 2026 Demo / reviewers in the wild / expert
Stephen Booth
dblp:89/708
· DBLP profile ↗
7ranked-venue papers
2as first author
0since 2021 · last 2017
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Parallel and multicore computing · 100% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Parallel and multicore computing › parallel libraries
communication library |
0.0 | 1 | 2000 | Single sided MPI implementations for SUN MPI · SC 2000 |
Parallel and multicore computing › parallel programming models
message passing |
0.0 | 1 | 2000 | Single sided MPI implementations for SUN MPI · SC 2000 |
Parallel and multicore computing › parallel programming models › message passing
MPI one-sided communication |
0.0 | 1 | 2000 | Single sided MPI implementations for SUN MPI · SC 2000 |
Parallel and multicore computing › parallel computing › parallel communication
shared-memory communication |
0.0 | 1 | 2000 | Single sided MPI implementations for SUN MPI · SC 2000 |
Methods — techniques the papers use, named apart from their topics
type packing · 0.0caching · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2017 | Generalisation of recursive doubling for AllReduce: Now with simulationabstractThe performance of AllReduce is crucial at scale. The recursive doubling with pairwise exchange algorithm theoretically achieves O (log 2 N ) scaling for short messages with N peers, but is limited by improvements in network latency. A multi-way exchange can be implemented using message pipelining, which is easier to improve than latency. Using our method, recursive multiplying, we show reductions in execution time of between 8% and 40% of AllReduce on a Cray XC30 over recursive doubling. Using a custom simulator we further explore the dynamics of recursive multiplying. Martin Ruefenacht, Mark Bull, Stephen Booth |
Parallel Comput. | 3 |
| 2016 | Generalisation of Recursive Doubling for AllReduceabstractThe performance of AllReduce is crucial at scale. The recursive doubling with pairwise exchange algorithm theoretically achieves O(log2 N) scaling for short messages with N peers, but is limited by improvements in network latency. A multi-way exchange can be implemented using message pipelining, which is easier to improve than latency. Using our method, recursive multiplying, we show reductions in execution time of between 8% and 40% of AllReduce on a Cray XC30 over recursive doubling. Martin Ruefenacht, Mark Bull, Stephen Booth |
EuroMPI | 3 |
| 2013 | McMPI: a managed-code MPI library in pure C#abstractThis paper presents McMPI, an entirely new MPI library written in C# using only safe managed-code, and performance results from low-level benchmarks demonstrating ping-pong latency and bandwidth comparable with MS-MPI and MPICH2. Daniel J. Holmes, Stephen Booth |
EuroMPI | 2 |
| 2011 | Parallel Optimisation Strategies for Fusion CodesabstractWe have previously documented the on-going work in the EUFORIA project to parallelise and optimise European fusion simulation codes, see. This involves working with a wide range of codes to try and address any performance and scaling issues that these codes have. However, as no two simulation codes are exactly the same, it is very hard to apply exactly the same approach to optimising a disparate range of codes. Indeed, it can be seen from that the codes investigated range in terms of performance and ability from well-optimised, highly parallelised codes, to serial or poorly performing codes. After analysing, optimising, and parallelising a range of codes it is, actually, possible to discern a number of distinct optimisation techniques or approaches/strategies that can be used to improve the performance or scaling of a parallel simulation code. This paper outlines the distinct approaches that we have identified, highlighting their benefits and drawbacks, giving an overview of the type of work that is often attempted for fusion simulation code optimisation. Adrian Jackson, Fiona Reid, Stephen Booth, Joachim Hein, Jan Westerholm, Mats Aspnäs, Miquel Catala, Alejandro Soba |
PDP | 3 |
| 2005 | HPCx: towards capability computingabstractAbstract We introduce HPCx—the U.K.'s new National HPC Service—which aims to deliver a world‐class service for capability computing to the U.K. scientific community. HPCx is targeting an environment that will both result in world‐leading science and address the challenges involved in scaling existing codes to the capability levels required. Close working relationships with scientific consortia and user groups throughout the research process will be a central feature of the service. A significant number of key user applications have already been ported to the system. We present initial benchmark results from this process and discuss the optimization of the codes and the performance levels achieved on HPCx in comparison with other systems. We find a range of performance with some algorithms scaling far better than others. Copyright © 2005 John Wiley & Sons, Ltd. Mike Ashworth, Ian J. Bush, Martyn F. Guest, Andrew G. Sunderland, Stephen Booth, Joachim Hein, Lorna Smith, Kevin Stratford, Alessandro Curioni |
Concurr. Pract. Exp. | 5 |
| 2001 | Optimising the MPI Library for the T3E
Stephen Booth |
Euro-Par | 1 |
| 2000 | Single sided MPI implementations for SUN MPIabstractThis paper describes an implementation of generic MPI-2 single sided communications for SUN-MPI. Our implementation is layered on top of point-to-point MPI communications and therefore can be adapted to other MPI implementations. The code is designed to co-exist with other MPI-2 single sided implementations (for example direct use of shared memory) providing a generic fall-back implementation for those communication paths where an optimised single-sided implementation is not available. MPI-2 single sided communications require the transfer of data-type information as well as user data. We describe a type packing and caching mechanism used to optimise the transfer of data-type information. The performance of this implementation is measured in comparison to equivalent point to point operations and the shared memory implementation provided by SUN. Stephen Booth, Fernando Elson Mourão |
SC | 1 |