EDBT 2026 Demo / reviewers in the wild / expert
Brian R. Toonen
dblp:29/3542
· DBLP profile ↗
13ranked-venue papers
1as first author
1since 2021 · last 2025
0009-0004-8562-3203ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Parallel and multicore computing · 43% High-performance computing · 26% Distributed systems · 11% | |
| Computer networks
1 paper |
Internet architecture and protocols · 100% |
Topics — the 17 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Parallel and multicore computing › parallel programming models
message passing |
0.1 | 2 | 2008 | Distributed mpi cross-site run performance using mpig · HPDC 2008 Supporting efficient execution in heterogeneous distributed computing environments with cactus and globus · SC 2001 |
Parallel and multicore computing › parallel programming models › message passing
MPI implementation |
0.1 | 1 | 2008 | Distributed mpi cross-site run performance using mpig · HPDC 2008 |
Parallel and multicore computing › parallel computing
distributed execution |
0.0 | 1 | 2001 | Supporting efficient execution in heterogeneous distributed computing environments with cactus and globus · SC 2001 |
Distributed systems
grid computing |
0.0 | 1 | 2001 | Supporting efficient execution in heterogeneous distributed computing environments with cactus and globus · SC 2001 |
Parallel and multicore computing
MPI |
0.0 | 1 | 2001 | Interfacing Parallel Jobs to Process Managers · HPDC 2001 |
Internet architecture and protocols › quality of service
differentiated services |
0.0 | 1 | 2000 | MPICH-GQ: Quality-of-Service for Message Passing Programs · SC 2000 |
Internet architecture and protocols
quality of service |
0.0 | 1 | 2000 | MPICH-GQ: Quality-of-Service for Message Passing Programs · SC 2000 |
High-performance computing
performance optimization at scale |
0.0 | 1 | 2000 | MPICH-GQ: Quality-of-Service for Message Passing Programs · SC 2000 |
High-performance computing › scientific computing systems
molecular dynamics simulation |
0.0 | 1 | 2008 | Distributed mpi cross-site run performance using mpig · HPDC 2008 |
Cloud and datacenter computing › resource allocation › multi-resource allocation
resource co-allocation |
0.0 | 1 | 2008 | Distributed mpi cross-site run performance using mpig · HPDC 2008 |
Cloud and datacenter computing
resource management |
0.0 | 1 | 2008 | Distributed mpi cross-site run performance using mpig · HPDC 2008 |
High-performance computing
scientific computing |
0.0 | 1 | 2008 | Distributed mpi cross-site run performance using mpig · HPDC 2008 |
Memory systems › cache coherence
cache coherence protocol |
0.0 | 1 | 1997 | Relaxed Consistency and Coherence Granularity in DSM Systems: A Performance Evaluation · PPoPP 1997 |
Memory systems › cache coherence
coherence granularity |
0.0 | 1 | 1997 | Relaxed Consistency and Coherence Granularity in DSM Systems: A Performance Evaluation · PPoPP 1997 |
Distributed systems
consistency models |
0.0 | 1 | 1997 | Relaxed Consistency and Coherence Granularity in DSM Systems: A Performance Evaluation · PPoPP 1997 |
Memory systems › shared memory
distributed shared memory |
0.0 | 1 | 1997 | Relaxed Consistency and Coherence Granularity in DSM Systems: A Performance Evaluation · PPoPP 1997 |
Computational science and engineering › computational physics
numerical relativity |
0.0 | 1 | 2001 | Supporting efficient execution in heterogeneous distributed computing environments with cactus and globus · SC 2001 |
Methods — techniques the papers use, named apart from their topics
message passing · 0.1non-blocking communication · 0.1adaptive parameter tuning · 0.1traffic shaping · 0.1qos reservation · 0.1API design · 0.0workload characterization · 0.0simulation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Big Data Approach for Efficient Processing of Machine Operational Data
Eric Pershey, Ben Lenard, Brian R. Toonen, Peter Upton, Alexander Rasin |
SSDBM | 3 |
| 2008 | Distributed mpi cross-site run performance using mpigabstractLarge scale supercomputing applications typically run on clusters using vendor message passing libraries, limiting the application to the availability of memory and CPU resources on that single machine. The ability to run inter-cluster parallel code is attractive since it allows the consolidation of multiple large scale resources for computational simulations not possible on a single machine, and it also allows the conglomeration of small subsets of CPU cores for rapid turnaround, for example, in the case of high-availability computing. MPIg is a grid-enabled implementation of the Message Passing Interface (MPI), extending the MPICH implementation of MPI to use Globus Toolkit services such as resource allocation and authentication. To achieve co-availability of resources, HARC, the Highly-Available Resource Co-allocator, is used. Here we examine two applications using MPIg: LAMMPS (Large-scale Atomic/ Molecular Massively Parallel Simulator), is used with a replica exchange molecular dynamics approach to enhance binding affinity calculations in HIV drug research, and HemeLB, which is a lattice-Boltzmann solver designed to address fluid flow in geometries such as the human cerebral vascular system. The cross-site scalability of both these applications is tested and compared to single-machine performance. In HemeLB, communication costs are hidden by effectively overlapping non-blocking communication with computation, essentially scaling linearly across multiple sites, and LAMMPS scales almost as well when run between two significantly geographically separated sites as it does at a single site. Steven Manos, Marco D. Mazzeo, Owain Kenway, Peter V. Coveney, Nicholas T. Karonis, Brian R. Toonen |
HPDC | 6 |
| 2007 | Implementation of Distributed Loop Scheduling Schemes on the TeraGridabstractGrid computing can be used for high performance computations. However, a serious difficulty in concurrent programming of such heterogeneous systems is how to deal with scheduling and load balancing of such systems which may consist of heterogeneous computers on different sites. Distributed scheduling schemes suitable for parallel loops with independent iterations on heterogeneous computer clusters have been proposed and analyzed in the past. In this article, we implement the previous schemes in MPICH-G2 and MPIg on the TeraGrid. We present performance results for three loop scheduling schemes on single and multi-site TeraGrid clusters. Satish Penmatsa, Anthony T. Chronopoulos, Nicholas T. Karonis, Brian R. Toonen |
IPDPS | 4 |
| 2005 | Implementing MPI-IO atomic mode without file system supportabstractThe ROMIO implementation of the MPI-IO standard provides a portable infrastructure for use on top of any number of different underlying storage targets. These different targets vary widely in their capabilities, and in some cases, additional effort is needed within ROMIO to support the complete MPI-IO semantics. One aspect of the interface that can be problematic to implement is the MPI-IO atomic mode. This mode requires enforcing strict consistency semantics. For some file systems, native locks may be used to enforce these semantics, but not all file systems have lock support. In this work, we describe two algorithms for implementing efficient mutex locks using MPI-1 and MPI-2 capabilities. We then show how these algorithms may be used to implement a portable MPI-IO atomic mode for ROMIO. We evaluate the performance of these algorithms and show that they impose little additional overhead on the system. Because of the low-overhead nature of these algorithms, they are likely useful in a variety of situations where distributed locks are needed in the MPI-2 environment. Robert B. Ross, Robert Latham, William Gropp, Rajeev Thakur, Brian R. Toonen |
CCGRID | 5 |
| 2004 | Implementing MPI on the BlueGene/L Supercomputer
Gheorghe Almási 0001, Charles Archer, José G. Castaños, C. Christopher Erway, Philip Heidelberger, Xavier Martorell, José E. Moreira, Kurt W. Pinnow, Joe Ratterman, Nils Smeds, Burkhard D. Steinmacher-Burow, William Gropp, Brian R. Toonen |
Euro-Par | 13 |
| 2004 | Design and Implementation of MPICH2 over InfiniBand with RDMA SupportabstractSummary form only given. For several years, MPI has been the de facto standard for writing parallel applications. One of the most popular MPI implementations is MPICH. Its successor, MPICH2, features a completely new design that provides more performance and flexibility. To ensure portability, it has a hierarchical structure based on which porting can be done at different levels. In this paper, we present our experiences in designing and implementing MPICH2 over InfiniBand. Because of its high performance and open standard, InfiniBand is gaining popularity in the area of high-performance computing. Our study focuses on optimizing the performance of MPl-1 functions in MPICH2. One of our objectives is to exploit remote direct memory access (RDMA) in InfiniBand to achieve high performance. We have based our design on the RDMA channel interface provided by MP1CH2, which encapsulates architecture-dependent communication functionalities into a very small set of functions. Starting with a basic design, we apply different optimizations and also propose a zero-copy-based design. We characterize the impact of our optimizations and designs using microbenchmarks. We have also performed an application-level evaluation using the NAS parallel benchmarks. Our optimized MPICH2 implementation achieves 7.6/spl mu/s latency and 857 MB/s bandwidth, which are close to the raw performance of the underlying InfiniBand layer. Our study shows that the RDMA channel interface in MPICH2 provides a simple, yet powerful, abstraction that enables implementations with high performance by exploiting RDMA operations in InfiniBand. To the best of our knowledge, this is the first high-performance design and implementation ofMPICH2 on InfiniBand using RDMA support. Jiuxing Liu, Weihang Jiang, Pete Wyckoff, Dhabaleswar K. Panda 0001, David Ashton, Darius Buntinas, William Gropp, Brian R. Toonen |
IPDPS | 8 |
| 2003 | MPICH-G2: A Grid-enabled implementation of the Message Passing Interface
Nicholas T. Karonis, Brian R. Toonen, Ian T. Foster |
J. Parallel Distributed Comput. | 2 |
| 2001 | Interfacing Parallel Jobs to Process ManagersabstractA variety of projects worldwide are developing what we call "heterogeneous MPI". These MPI implementations are designed to operate on multiple computers, perhaps of different types, ranging in complexity from a set of desktop workstations to several supercomputers connected via a wide area network. These considerations led us to investigate the feasibility of defining a common API that could be used within MPI implementations to access process startup, initialization, monitoring, and control functions provided by an underlying process management system. If various MPI implementations coded to that API, one could then develop multiple "process management" modules that could be reused within different MPI implementations, thus allowing partitioning of effort between different development groups. In pursuit of this goal, we have designed such an API, which we call BNR. The major goals of the BNR interface are outlined. Brian R. Toonen, David Ashton, Ewing L. Lusk, Ian T. Foster, William Gropp, Edgar Gabriel, Ralph M. Butler, Nicholas T. Karonis |
HPDC | 1 |
| 2001 | Supporting efficient execution in heterogeneous distributed computing environments with cactus and globusabstractImprovements in the performance of processors and networks make it both feasible and interesting to treat collections of workstations, servers, clusters, and supercomputers as integrated computational resources, or Grids. However, the highly heterogeneous and dynamic nature of such Grids can make application development difficult. Here we describe an architecture and prototype implementation for a Grid-enabled computational framework based on Cactus, the MPICH-G2 Grid-enabled message-passing library, and a variety of specialized features to support efficient execution in Grid environments. We have used this framework to perform record-setting computations in numerical relativity, running across four supercomputers and achieving scaling of 88% (1140 CPU's) and 63% (1500 CPUs). The problem size we were able to compute was about five times larger than any other previous run. Further, we introduce and demonstrate adaptive methods that automatically adjust computational parameters during run time, to increase dramatically the efficiency of a distributed Grid simulation, without modification of the application and without any knowledge of the underlying network connecting the distributed computers. Gabrielle Allen, Thomas Dramlitsch, Ian T. Foster, Nicholas T. Karonis, Matei Ripeanu, Edward Seidel, Brian R. Toonen |
SC | 7 |
| 2000 | MPICH-GQ: Quality-of-Service for Message Passing ProgramsabstractParallel programmers typically assume that all resources required for a program’s execution are dedicated to that purpose. However, in local and wide area networks, contention for shared networks, CPUs, and I/O systems can result in significant variations in availability, with consequent adverse effects on overall performance. We describe a new message-passing architecture, MPICH-GQ, that uses quality of service (QoS) mechanisms to manage contention and hence improve performance of message passing interface (MPI) applications. MPICH-GQ combines new QoS specification, traffic shaping, QoS reservation, and QoS implementation techniques to deliver QoS capabilities to the high-bandwidth bursty flows, complex structures, and reliable protocols used in high-performance applications-characteristics very different from the low-bandwidth, constant bit-rate media flows and unreliable protocols for which QoS mechanisms were designed. Results obtained on a differentiated services testbed demonstrate our ability to maintain application performance in the face of heavy network contention. Alain J. Roy, Ian T. Foster, William Gropp, Nicholas T. Karonis, Volker Sander, Brian R. Toonen |
SC | 6 |
| 1998 | A computational framework for telemedicine
Ian T. Foster, Gregor von Laszewski, George K. Thiruvathukal, Brian R. Toonen |
Future Gener. Comput. Syst. | 4 |
| 1997 | Relaxed Consistency and Coherence Granularity in DSM Systems: A Performance EvaluationabstractDuring the past few years, two main approaches have been taken to improve the performance of software shared memory implementations: relaxing consistency models and providing fine-grained access control. Their performance tradeoffs, however, we not well understood. This paper studies these tradeoffs on a platform that provides access control in hardware but runs coherence protocols in software, We compare the performance of three protocols across four coherence granularities, using 12 applications on a 16-node cluster of workstations. Our results show that no single combination of protocol and granularity performs best for all the applications. The combination of a sequentially consistent (SC) protocol and fine granularity works well with 7 of the 12 applications. The combination of a multiple-writer, home-based lazy release consistency (HLRC) protocol and page granularity works well with 8 out of the 12 applications. For applications that suffer performance losses in moving to coarser granularity under sequential consistency, the performance can usually be regained quite effectively using relaxed protocols, particularly HLRC. We also find that the HLRC protocol performs substantially better than a single-writer lazy release consistent (SW-LRC) protocol at coase granularity for many irregular applications. For our applications and platform, when we use the original versions of the applications ported directly from hardware-coherent shared memory, we find that the SC protocol with 256-byte granularity performs best on average. However, when the best versions of the applications are compared, the balance shifts in favor of HLRC at page granularity. Yuanyuan Zhou 0001, Liviu Iftode, Jaswinder Pal Singh, Kai Li 0001, Brian R. Toonen, Ioannis Schoinas, Mark D. Hill, David A. Wood 0001 |
PPoPP | 5 |
| 1995 | Design and Performance of a Scalable Parallel Community Climate Model
John B. Drake, Ian T. Foster, John Michalakes, Brian R. Toonen, Patrick H. Worley |
Parallel Comput. | 4 |