Brian R. Toonen

dblp:29/3542 · DBLP profile ↗
← Back
13ranked-venue papers
1as first author
1since 2021 · last 2025
0009-0004-8562-3203ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Parallel and multicore computing · 43% High-performance computing · 26% Distributed systems · 11%
Computer networks
1 paper
Internet architecture and protocols · 100%

Topics — the 17 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing › parallel programming models
message passing
0.122008
Distributed mpi cross-site run performance using mpig · HPDC 2008
Supporting efficient execution in heterogeneous distributed computing environments with cactus and globus · SC 2001
Parallel and multicore computing › parallel programming models › message passing
MPI implementation
0.112008
Distributed mpi cross-site run performance using mpig · HPDC 2008
Parallel and multicore computing › parallel computing
distributed execution
0.012001
Supporting efficient execution in heterogeneous distributed computing environments with cactus and globus · SC 2001
Distributed systems
grid computing
0.012001
Supporting efficient execution in heterogeneous distributed computing environments with cactus and globus · SC 2001
Parallel and multicore computing
MPI
0.012001
Interfacing Parallel Jobs to Process Managers · HPDC 2001
Internet architecture and protocols › quality of service
differentiated services
0.012000
MPICH-GQ: Quality-of-Service for Message Passing Programs · SC 2000
Internet architecture and protocols
quality of service
0.012000
MPICH-GQ: Quality-of-Service for Message Passing Programs · SC 2000
High-performance computing
performance optimization at scale
0.012000
MPICH-GQ: Quality-of-Service for Message Passing Programs · SC 2000
High-performance computing › scientific computing systems
molecular dynamics simulation
0.012008
Distributed mpi cross-site run performance using mpig · HPDC 2008
Cloud and datacenter computing › resource allocation › multi-resource allocation
resource co-allocation
0.012008
Distributed mpi cross-site run performance using mpig · HPDC 2008
Cloud and datacenter computing
resource management
0.012008
Distributed mpi cross-site run performance using mpig · HPDC 2008
High-performance computing
scientific computing
0.012008
Distributed mpi cross-site run performance using mpig · HPDC 2008
Memory systems › cache coherence
cache coherence protocol
0.011997
Relaxed Consistency and Coherence Granularity in DSM Systems: A Performance Evaluation · PPoPP 1997
Memory systems › cache coherence
coherence granularity
0.011997
Relaxed Consistency and Coherence Granularity in DSM Systems: A Performance Evaluation · PPoPP 1997
Distributed systems
consistency models
0.011997
Relaxed Consistency and Coherence Granularity in DSM Systems: A Performance Evaluation · PPoPP 1997
Memory systems › shared memory
distributed shared memory
0.011997
Relaxed Consistency and Coherence Granularity in DSM Systems: A Performance Evaluation · PPoPP 1997
Computational science and engineering › computational physics
numerical relativity
0.012001
Supporting efficient execution in heterogeneous distributed computing environments with cactus and globus · SC 2001

Methods — techniques the papers use, named apart from their topics

message passing · 0.1non-blocking communication · 0.1adaptive parameter tuning · 0.1traffic shaping · 0.1qos reservation · 0.1API design · 0.0workload characterization · 0.0simulation · 0.0
YearPublicationVenuePosition
2025 A Big Data Approach for Efficient Processing of Machine Operational Data
Eric Pershey, Ben Lenard, Brian R. Toonen, Peter Upton, Alexander Rasin
SSDBM3
2008 Distributed mpi cross-site run performance using mpig
abstract
Large scale supercomputing applications typically run on clusters using vendor message passing libraries, limiting the application to the availability of memory and CPU resources on that single machine. The ability to run inter-cluster parallel code is attractive since it allows the consolidation of multiple large scale resources for computational simulations not possible on a single machine, and it also allows the conglomeration of small subsets of CPU cores for rapid turnaround, for example, in the case of high-availability computing. MPIg is a grid-enabled implementation of the Message Passing Interface (MPI), extending the MPICH implementation of MPI to use Globus Toolkit services such as resource allocation and authentication. To achieve co-availability of resources, HARC, the Highly-Available Resource Co-allocator, is used. Here we examine two applications using MPIg: LAMMPS (Large-scale Atomic/ Molecular Massively Parallel Simulator), is used with a replica exchange molecular dynamics approach to enhance binding affinity calculations in HIV drug research, and HemeLB, which is a lattice-Boltzmann solver designed to address fluid flow in geometries such as the human cerebral vascular system. The cross-site scalability of both these applications is tested and compared to single-machine performance. In HemeLB, communication costs are hidden by effectively overlapping non-blocking communication with computation, essentially scaling linearly across multiple sites, and LAMMPS scales almost as well when run between two significantly geographically separated sites as it does at a single site.
Steven Manos, Marco D. Mazzeo, Owain Kenway, Peter V. Coveney, Nicholas T. Karonis, Brian R. Toonen
HPDC6
2007 Implementation of Distributed Loop Scheduling Schemes on the TeraGrid
abstract
Grid computing can be used for high performance computations. However, a serious difficulty in concurrent programming of such heterogeneous systems is how to deal with scheduling and load balancing of such systems which may consist of heterogeneous computers on different sites. Distributed scheduling schemes suitable for parallel loops with independent iterations on heterogeneous computer clusters have been proposed and analyzed in the past. In this article, we implement the previous schemes in MPICH-G2 and MPIg on the TeraGrid. We present performance results for three loop scheduling schemes on single and multi-site TeraGrid clusters.
Satish Penmatsa, Anthony T. Chronopoulos, Nicholas T. Karonis, Brian R. Toonen
IPDPS4
2005 Implementing MPI-IO atomic mode without file system support
abstract
The ROMIO implementation of the MPI-IO standard provides a portable infrastructure for use on top of any number of different underlying storage targets. These different targets vary widely in their capabilities, and in some cases, additional effort is needed within ROMIO to support the complete MPI-IO semantics. One aspect of the interface that can be problematic to implement is the MPI-IO atomic mode. This mode requires enforcing strict consistency semantics. For some file systems, native locks may be used to enforce these semantics, but not all file systems have lock support. In this work, we describe two algorithms for implementing efficient mutex locks using MPI-1 and MPI-2 capabilities. We then show how these algorithms may be used to implement a portable MPI-IO atomic mode for ROMIO. We evaluate the performance of these algorithms and show that they impose little additional overhead on the system. Because of the low-overhead nature of these algorithms, they are likely useful in a variety of situations where distributed locks are needed in the MPI-2 environment.
Robert B. Ross, Robert Latham, William Gropp, Rajeev Thakur, Brian R. Toonen
CCGRID5
2004 Implementing MPI on the BlueGene/L Supercomputer
Gheorghe Almási 0001, Charles Archer, José G. Castaños, C. Christopher Erway, Philip Heidelberger, Xavier Martorell, José E. Moreira, Kurt W. Pinnow, Joe Ratterman, Nils Smeds, Burkhard D. Steinmacher-Burow, William Gropp, Brian R. Toonen
Euro-Par13
2004 Design and Implementation of MPICH2 over InfiniBand with RDMA Support
abstract
Summary form only given. For several years, MPI has been the de facto standard for writing parallel applications. One of the most popular MPI implementations is MPICH. Its successor, MPICH2, features a completely new design that provides more performance and flexibility. To ensure portability, it has a hierarchical structure based on which porting can be done at different levels. In this paper, we present our experiences in designing and implementing MPICH2 over InfiniBand. Because of its high performance and open standard, InfiniBand is gaining popularity in the area of high-performance computing. Our study focuses on optimizing the performance of MPl-1 functions in MPICH2. One of our objectives is to exploit remote direct memory access (RDMA) in InfiniBand to achieve high performance. We have based our design on the RDMA channel interface provided by MP1CH2, which encapsulates architecture-dependent communication functionalities into a very small set of functions. Starting with a basic design, we apply different optimizations and also propose a zero-copy-based design. We characterize the impact of our optimizations and designs using microbenchmarks. We have also performed an application-level evaluation using the NAS parallel benchmarks. Our optimized MPICH2 implementation achieves 7.6/spl mu/s latency and 857 MB/s bandwidth, which are close to the raw performance of the underlying InfiniBand layer. Our study shows that the RDMA channel interface in MPICH2 provides a simple, yet powerful, abstraction that enables implementations with high performance by exploiting RDMA operations in InfiniBand. To the best of our knowledge, this is the first high-performance design and implementation ofMPICH2 on InfiniBand using RDMA support.
Jiuxing Liu, Weihang Jiang, Pete Wyckoff, Dhabaleswar K. Panda 0001, David Ashton, Darius Buntinas, William Gropp, Brian R. Toonen
IPDPS8
2003 MPICH-G2: A Grid-enabled implementation of the Message Passing Interface
Nicholas T. Karonis, Brian R. Toonen, Ian T. Foster
J. Parallel Distributed Comput.2
2001 Interfacing Parallel Jobs to Process Managers
abstract
A variety of projects worldwide are developing what we call "heterogeneous MPI". These MPI implementations are designed to operate on multiple computers, perhaps of different types, ranging in complexity from a set of desktop workstations to several supercomputers connected via a wide area network. These considerations led us to investigate the feasibility of defining a common API that could be used within MPI implementations to access process startup, initialization, monitoring, and control functions provided by an underlying process management system. If various MPI implementations coded to that API, one could then develop multiple "process management" modules that could be reused within different MPI implementations, thus allowing partitioning of effort between different development groups. In pursuit of this goal, we have designed such an API, which we call BNR. The major goals of the BNR interface are outlined.
Brian R. Toonen, David Ashton, Ewing L. Lusk, Ian T. Foster, William Gropp, Edgar Gabriel, Ralph M. Butler, Nicholas T. Karonis
HPDC1
2001 Supporting efficient execution in heterogeneous distributed computing environments with cactus and globus
abstract
Improvements in the performance of processors and networks make it both feasible and interesting to treat collections of workstations, servers, clusters, and supercomputers as integrated computational resources, or Grids. However, the highly heterogeneous and dynamic nature of such Grids can make application development difficult. Here we describe an architecture and prototype implementation for a Grid-enabled computational framework based on Cactus, the MPICH-G2 Grid-enabled message-passing library, and a variety of specialized features to support efficient execution in Grid environments. We have used this framework to perform record-setting computations in numerical relativity, running across four supercomputers and achieving scaling of 88% (1140 CPU's) and 63% (1500 CPUs). The problem size we were able to compute was about five times larger than any other previous run. Further, we introduce and demonstrate adaptive methods that automatically adjust computational parameters during run time, to increase dramatically the efficiency of a distributed Grid simulation, without modification of the application and without any knowledge of the underlying network connecting the distributed computers.
Gabrielle Allen, Thomas Dramlitsch, Ian T. Foster, Nicholas T. Karonis, Matei Ripeanu, Edward Seidel, Brian R. Toonen
SC7
2000 MPICH-GQ: Quality-of-Service for Message Passing Programs
abstract
Parallel programmers typically assume that all resources required for a program’s execution are dedicated to that purpose. However, in local and wide area networks, contention for shared networks, CPUs, and I/O systems can result in significant variations in availability, with consequent adverse effects on overall performance. We describe a new message-passing architecture, MPICH-GQ, that uses quality of service (QoS) mechanisms to manage contention and hence improve performance of message passing interface (MPI) applications. MPICH-GQ combines new QoS specification, traffic shaping, QoS reservation, and QoS implementation techniques to deliver QoS capabilities to the high-bandwidth bursty flows, complex structures, and reliable protocols used in high-performance applications-characteristics very different from the low-bandwidth, constant bit-rate media flows and unreliable protocols for which QoS mechanisms were designed. Results obtained on a differentiated services testbed demonstrate our ability to maintain application performance in the face of heavy network contention.
Alain J. Roy, Ian T. Foster, William Gropp, Nicholas T. Karonis, Volker Sander, Brian R. Toonen
SC6
1998 A computational framework for telemedicine
Ian T. Foster, Gregor von Laszewski, George K. Thiruvathukal, Brian R. Toonen
Future Gener. Comput. Syst.4
1997 Relaxed Consistency and Coherence Granularity in DSM Systems: A Performance Evaluation
abstract
During the past few years, two main approaches have been taken to improve the performance of software shared memory implementations: relaxing consistency models and providing fine-grained access control. Their performance tradeoffs, however, we not well understood. This paper studies these tradeoffs on a platform that provides access control in hardware but runs coherence protocols in software, We compare the performance of three protocols across four coherence granularities, using 12 applications on a 16-node cluster of workstations. Our results show that no single combination of protocol and granularity performs best for all the applications. The combination of a sequentially consistent (SC) protocol and fine granularity works well with 7 of the 12 applications. The combination of a multiple-writer, home-based lazy release consistency (HLRC) protocol and page granularity works well with 8 out of the 12 applications. For applications that suffer performance losses in moving to coarser granularity under sequential consistency, the performance can usually be regained quite effectively using relaxed protocols, particularly HLRC. We also find that the HLRC protocol performs substantially better than a single-writer lazy release consistent (SW-LRC) protocol at coase granularity for many irregular applications. For our applications and platform, when we use the original versions of the applications ported directly from hardware-coherent shared memory, we find that the SC protocol with 256-byte granularity performs best on average. However, when the best versions of the applications are compared, the balance shifts in favor of HLRC at page granularity.
Yuanyuan Zhou 0001, Liviu Iftode, Jaswinder Pal Singh, Kai Li 0001, Brian R. Toonen, Ioannis Schoinas, Mark D. Hill, David A. Wood 0001
PPoPP5
1995 Design and Performance of a Scalable Parallel Community Climate Model
John B. Drake, Ian T. Foster, John Michalakes, Brian R. Toonen, Patrick H. Worley
Parallel Comput.4