J. Mark Bull

dblp:30/4850 · also Jonathan Mark Bull · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
1since 2021 · last 2026
0000-0002-4186-584XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 3 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Performance modeling and evaluation · 56% Parallel and multicore computing · 36% Memory systems · 8%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Performance modeling and evaluation
benchmarking
0.012001
A parallel java grande benchmark suite · SC 2001
Performance modeling and evaluation › benchmarking
parallel benchmark suites
0.012001
A parallel java grande benchmark suite · SC 2001
Parallel and multicore computing
parallel programming models
0.012001
A parallel java grande benchmark suite · SC 2001
Parallel and multicore computing › parallel programming models
message passing
0.012001
A parallel java grande benchmark suite · SC 2001
Memory systems
shared memory
0.012001
A parallel java grande benchmark suite · SC 2001

Methods — techniques the papers use, named apart from their topics

java threads · 0.0MPJ · 0.0JOMP · 0.0
YearPublicationVenuePosition
2026 Optimizing MPI collective operations for intra-node shared memory exploitation and oversubscription
abstract
Node sizes in multicore clusters are becoming larger, so applications should exploit the shared memory inside a node, to potentially reduce communication latencies compared to network communications. The Message Passing Interface library (MPI) serves as the de facto standard for parallel applications on distributed memory environments. The appearance of the MPI-3 shared memory extension has made the idea of optimizing the MPI collective operations at the intra-node level both attractive and portable. Taking advantage of this facility, we present a hierarchical design of the algorithms for MPI_Allreduce and MPI_Bcast collective operations, which we name Fullpar. The proposal is based on partitioning the messages and exploiting concurrency between network communication and shared-memory operations at intra-node level in the message dissemination process. Previously proposed hierarchical collective algorithms do not exploit all the available parallelism and/or the direct implementation of shared-memory at intra-node level. Furthermore, we explore the application of oversubscription, to enhance resource utilization. Oversubscribing CPUs offers opportunities to optimize resource allocation, as well as making CPUs available to other applications. We implement our proposal on top of the Intel MPI and OpenMPI libraries using the MPI profiling (PMPI) mechanism, and carry out evaluations on platforms with different architectures characteristics such as the node size and interconnection network. Experimental results show that the proposed approach often achieves lower execution times than the evaluated alternatives for medium and large message sizes. Furthermore, the introduction of oversubscription is shown to have almost no overhead, especially for the broadcast of large messages.
Gladys Utrera, J. Mark Bull
Parallel Comput.2
2020 Evaluating Worksharing Tasks on Distributed Environments
abstract
Hybrid programming is a promising approach to exploit clusters of multicore systems. Our focus is on the combination of MPI and tasking. This hybrid approach combines the low-latency and high throughput of MPI with the flexibility of tasking models and their inherent ability to handle load imbalance. However, combining tasking with standard MPI implementations can be a challenge. The Task-Aware MPI library (TAMPI) eases the development of applications combining tasking with MPI. TAMPI enables developers to overlap computation and communication phases by relying on the tasking data-flow execution model. Using this approach, the original computation that was distributed in many different MPI ranks is grouped together in fewer MPI ranks, and split into several tasks per rank. Nevertheless, programmers must be careful with task granularity. Too fine-grained tasks introduce too much overhead, while too coarse-grained tasks lead to lack of parallelism. An adequate granularity may not always exist, especially in distributed environments where the same amount of work is distributed among many more cores. Worksharing tasks are a special kind of tasks, recently proposed, that internally leverage worksharing techniques. By doing so, a single worksharing task may run in several cores concurrently. Nonetheless, the task management costs remain the same than a regular task. In this work, we study the combination of worksharing tasks and TAMPI on distributed environments using two well known mini-apps: HPCCG and LULESH. Our results show significant improvements using worksharing tasks compared to regular tasks, and to other state-of-the-art alternatives such as OpenMP worksharing.
Marcos Maronas, Xavier Teruel, J. Mark Bull, Eduard Ayguadé, Vicenç Beltran 0001
CLUSTER3
2019 iPregel: Vertex-centric programmability vs memory efficiency and performance, why choose?
abstract
The vertex-centric programming model, designed to improve the programmability in graph processing application writing, has attracted great attention over the years. Multiple shared memory frameworks that have implemented the vertex-centric interface all expose a common tradeoff: programmability against memory efficiency and performance. Our approach consists in preserving vertex-centric programmability, while implementing optimisations missing from FemtoGraph, developing new ones and designing these so they are transparent to a user’s application code, hence not impacting programmability. We therefore implemented our own shared memory vertex-centric framework iPregel, relying on in-memory storage and synchronous execution. In this paper, we evaluate it against FemtoGraph, whose characteristics are identical, but also an asynchronous counterpart GraphChi and the vertex-subset-centric framework Ligra. Our experiments include three of the most popular vertex-centric benchmark applications over 4 real-world publicly accessible graphs, which cover all orders of magnitude between a million to a billion edges. We then measure the execution time and the peak memory usage. Finally, we evaluate the programmability of each framework by comparing it against the original Pregel, Google’s closed-source implementation that started the whole area of vertex-centric programming. Experiments demonstrate that iPregel, like FemtoGraph, does not sacrifice vertex-centric programmability for additional performance and memory efficiency optimisations, which contrasts with GraphChi and Ligra. Sacrificing vertex-centric programmability allowed the latter to benefit from substantial performance and memory efficiency gains. However, experiments demonstrate that iPregel is up to 2300 times faster than FemtoGraph, as well as generating a memory footprint up to 100 times smaller. These results greatly change the situation; Ligra and GraphChi are up to 17,000 and 700 times faster than FemtoGraph but, when comparing against iPregel, this maximum speed-up drops to 10. Furthermore, on PageRank, it is iPregel that proves to be the fastest overall. When it comes to memory efficiency, the same observation applies; Ligra and GraphChi are 100 and 50 times lighter than FemtoGraph, but iPregel nullifies these benefits: it provides the same memory efficiency as Ligra and even proves to be 3 to 6 times lighter than GraphChi on average. In other words, iPregel demonstrates that preserving vertex-centric programmability is not incompatible with a competitive performance and memory efficiency.
Ludovic Anthony Richard Capelli, Zhenjiang Hu 0002, Timothy A. K. Zakian, Nick Brown 0002, J. Mark Bull
Parallel Comput.5
2003 Benchmarking Java against C and Fortran for scientific applications
abstract
Abstract Increasing interest is being shown in the use of Java for scientific applications. The Java Grande benchmark suite was designed with such applications primarily in mind. The perceived lack of performance of Java still deters many potential users, despite recent advances in just‐in‐time and adaptive compilers. There are, however, few benchmark results available comparing Java to more traditional languages such as C and Fortran. To address this issue, a subset of the Java Grande benchmarks has been re‐written in C and Fortran allowing direct performance comparisons between the three languages. The performance of a range of Java execution environments, C and Fortran compilers have been tested across a number of platforms using the suite. These demonstrate that on some platforms (notable Intel Pentium) the performance gap is now quite small. Copyright © 2003 John Wiley & Sons, Ltd.
J. Mark Bull, Lorna A. Smith, Carwyn Ball, Linday Pottage, Robin Freeman
Concurr. Comput. Pract. Exp.1
2001 Topic 02: Performance Evaluation and Prediction
Allen D. Malony, Graham D. Riley, Bernd Mohr, J. Mark Bull, Tomàs Margalef
Euro-Par4
2001 Performance Analysis Tools for Parallel Java Applications on Shared-memory Systems
abstract
In this paper we describe an instrumentation environment for the performance analysis and visualization of parallel applications written in JOMP, an OpenMP-like interface for Java. The environment includes two complementary approaches. The first one has been designed to provide a detailed analysis of the parallel behavior at the JOMP programming model level. At this level, the user is faced with parallel, work-sharing and synchronization constructs, which are the core of JOMP. The second mechanism has been designed to support an in-depth analysis of the threaded execution inside the Java virtual machine (JVM). At this level of analysis, the user is faced with the supporting threads layer monitors and conditional variables. The paper discusses the implementation of both mechanisms and evaluates the overhead incurred by them.
Jordi Guitart, Jordi Torres, Eduard Ayguadé, J. Mark Bull
ICPP4
2001 A parallel java grande benchmark suite
abstract
Increasing interest is being shown in the use of Java for large scale or Grande applications. This new use of Java places specific demands on the Java execution environments that can be tested using the Java Grande benchmark suite [5], [6], [7]. The large processing requirements of Grande applications makes parallelisation of interest. A suite of parallel benchmarks has been developed from the serial Java Grande benchmark suite, using three parallel programming models: Java native threads, MPJ (a message passing interface) and JOMP (a set of OpenMP-like directives). The contents of the suite are described, and results presented for a number of platforms.
L. A. Smith, J. Mark Bull, Jan Obdrzálek
SC2
2001 An OpenMP-like interface for parallel programming in Java
abstract
Abstract This paper describes the definition and implementation of an OpenMP‐like set of directives and library routines for shared memory parallel programming in Java. A specification of the directives and routines is proposed and discussed. A prototype implementation, consisting of a compiler and a runtime library, both written entirely in Java, is presented, which implements most of the proposed specification. Some preliminary performance results are reported. Copyright © 2001 John Wiley & Sons, Ltd.
Mark Kambites, Jan Obdrzálek, J. Mark Bull
Concurr. Comput. Pract. Exp.3
2000 A benchmark suite for high performance Java
abstract
Increasing interest is being shown in the use of Java for large scale or Grande applications. This new use of Java places specific demands on the Java execution environments that could be tested and compared using a standard benchmark suite. We describe the design and implementation of such a suite, paying particular attention to Java-specific issues. Sample results are presented for a number of implementations of the Java Virtual Machine (JVM). Copyright © 2000 John Wiley & Sons, Ltd.
J. Mark Bull, L. A. Smith, Martin D. Westhead, D. S. Henty, R. A. Davey
Concurr. Pract. Exp.1
1998 Feedback Guided Dynamic Loop Scheduling: Algorithms and Experiments
J. Mark Bull
Euro-Par1
1997 Performance Improvement through Overhead Analysis: A Case Study in Molecular Dynamics
abstract
Article Free Access Share on Performance improvement through overhead analysis: a case study in molecular dynamics Authors: Graham D. Riley Centre for Novel Computing, University of Manchester, Oxford Road, Manchester M13 9PL, United Kingdom Centre for Novel Computing, University of Manchester, Oxford Road, Manchester M13 9PL, United KingdomView Profile , J. Mark Bull Centre for Novel Computing, University of Manchester, Oxford Road, Manchester M13 9PL, United Kingdom Centre for Novel Computing, University of Manchester, Oxford Road, Manchester M13 9PL, United KingdomView Profile , John R. Gurd Centre for Novel Computing, University of Manchester, Oxford Road, Manchester M13 9PL, United Kingdom Centre for Novel Computing, University of Manchester, Oxford Road, Manchester M13 9PL, United KingdomView Profile Authors Info & Claims ICS '97: Proceedings of the 11th international conference on SupercomputingJuly 1997 Pages 36–43https://doi.org/10.1145/263580.263589Published:11 July 1997Publication History 12citation261DownloadsMetricsTotal Citations12Total Downloads261Last 12 Months6Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Graham D. Riley, J. Mark Bull, John R. Gurd
International Conference on Supercomputing2
1994 Parallelisation of the SDEM distinct element stress analysis code on the KSR-1
abstract
The SDEM code models systems of interacting blocks of rock using the distinct element (DE) method, which represents these systems as discontinuums with each block acting under Newton's laws of motion. The data structures associated with the DE method make the task of obtaining performance gains through vectorisation difficult. Typical systems, however, contain thousands of blocks and there is the potential to perform calculations associated with groups of blocks in parallel.
Gregory K. Egan, Graham D. Riley, J. Mark Bull
International Conference on Supercomputing3