Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Scott B. Baden

dblp:74/1242 · DBLP profile ↗
← Back
31ranked-venue papers
5as first author
0since 2021 · last 2019
0000-0002-5479-8199ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 26 · 3 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
10 papers
Parallel and multicore computing · 33% High-performance computing · 27% Cloud and datacenter computing · 21%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 26 heaviest of 28, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing › datacenter operations
node failure prediction
0.312018
Doomsday: predicting which node will fail when on supercomputers · SC 2018
High-performance computing › system resilience
supercomputer reliability
0.312018
Doomsday: predicting which node will fail when on supercomputers · SC 2018
Distributed systems › communication optimization
communication-computation overlap
0.222012
Bamboo: translating MPI applications to a latency-tolerant, data-driven form · SC 2012
Communication overlap in multi-tier parallel algorithms · SC 1998
Parallel and multicore computing
parallel programming models
0.232006
Poster reception - Asynchronous programming with Tarragon · SC 2006
Short Paper: Asynchronous programming with Tarragon · HPDC 2006
A Programming Methodology for Dual-Tier Multicomputers · IEEE Trans. Software Eng. 2000
Compilers and program optimization › program transformation
source-to-source transformation
0.112012
Bamboo: translating MPI applications to a latency-tolerant, data-driven form · SC 2012
Parallel and multicore computing › parallel computation models
data-driven execution
0.112012
Bamboo: translating MPI applications to a latency-tolerant, data-driven form · SC 2012
Parallel and multicore computing › concurrent programming › concurrency model
actor-based programming
0.122006
Poster reception - Asynchronous programming with Tarragon · SC 2006
Short Paper: Asynchronous programming with Tarragon · HPDC 2006
Distributed systems › fault tolerance › proactive fault tolerance
failure prediction
0.112018
Doomsday: predicting which node will fail when on supercomputers · SC 2018
Parallel and multicore computing › parallel programming runtimes
runtime systems and scheduling
0.112006
Poster reception - Asynchronous programming with Tarragon · SC 2006
Performance modeling and evaluation
simulation
0.112006
Short Paper: Asynchronous programming with Tarragon · HPDC 2006
Cloud and datacenter computing
workload shifting
0.112006
Poster reception - Asynchronous programming with Tarragon · SC 2006
High-performance computing
scientific computing systems
0.122006
SCALLOP: A Highly Scalable Parallel Poisson Solver in Three Dimensions · SC 2003
Poster reception - Asynchronous programming with Tarragon · SC 2006
Parallel and multicore computing › parallel programming models › message passing
MPI application optimization
0.012012
Bamboo: translating MPI applications to a latency-tolerant, data-driven form · SC 2012
High-performance computing
parallel numerical algorithms
0.012003
SCALLOP: A Highly Scalable Parallel Poisson Solver in Three Dimensions · SC 2003
High-performance computing › numerical linear algebra › linear solver
poisson solver
0.012003
SCALLOP: A Highly Scalable Parallel Poisson Solver in Three Dimensions · SC 2003
Parallel and multicore computing
parallel programming models and runtimes
0.011998
Communication overlap in multi-tier parallel algorithms · SC 1998
High-performance computing
performance optimization at scale
0.011998
Communication overlap in multi-tier parallel algorithms · SC 1998
Distributed systems › operating system support › interprocess communication
asynchronous communication
0.012006
Poster reception - Asynchronous programming with Tarragon · SC 2006
Processor architecture and microarchitecture › clustered architecture
cluster assignment
0.011997
Parallel Cluster Identification for Multidimensional Lattices · IEEE Trans. Parallel Distributed Syst. 1997
Parallel and multicore computing › parallel graph algorithms
connected components
0.011997
Parallel Cluster Identification for Multidimensional Lattices · IEEE Trans. Parallel Distributed Syst. 1997
Parallel and multicore computing › load balancing
dynamic load balancing
0.011996
Dynamic Partitioning of Non-Uniform Structured Workloads with Spacefilling Curves · IEEE Trans. Parallel Distributed Syst. 1996
Parallel and multicore computing › parallel computing
parallel scientific computing
0.011996
Dynamic Partitioning of Non-Uniform Structured Workloads with Spacefilling Curves · IEEE Trans. Parallel Distributed Syst. 1996
Parallel and multicore computing
task partitioning
0.011996
Dynamic Partitioning of Non-Uniform Structured Workloads with Spacefilling Curves · IEEE Trans. Parallel Distributed Syst. 1996
High-performance computing › scientific computing systems
adaptive mesh refinement
0.011995
A Parallel Software Infrastructure for Structured Adaptive Mesh Methods · SC 1995
Parallel and multicore computing › parallel computation models
bulk synchronous parallel
0.011998
Communication overlap in multi-tier parallel algorithms · SC 1998
Computational science and engineering › statistical physics
statistical physics simulation
0.011997
Parallel Cluster Identification for Multidimensional Lattices · IEEE Trans. Parallel Distributed Syst. 1997

Methods — techniques the papers use, named apart from their topics

split-phase communication · 0.3source-to-source translation · 0.3over-decomposition · 0.1metadata-guided scheduling · 0.1actor model · 0.1asynchronous collective communication · 0.0SPMD control flow · 0.0hierarchical communication model · 0.0aggregate section moves · 0.0locality optimization · 0.0distributed-memory parallelization · 0.0distributed memory parallelization · 0.0
YearPublicationVenuePosition
2019 UPC++: A High-Performance Communication Framework for Asynchronous Computation
abstract
UPC++ is a C++ library that supports high-performance computation via an asynchronous communication framework. This paper describes a new incarnation that differs substantially from its predecessor, and we discuss the reasons for our design decisions. We present new design features, including future-based asynchrony management, distributed objects, and generalized Remote Procedure Call (RPC). We show microbenchmark performance results demonstrating that one-sided Remote Memory Access (RMA) in UPC++ is competitive with MPI-3 RMA; on a Cray XC40 UPC++ delivers up to a 25% improvement in the latency of blocking RMA put, and up to a 33% bandwidth improvement in an RMA throughput test. We showcase the benefits of UPC++ with irregular applications through a pair of application motifs, a distributed hash table and a sparse solver component. Our distributed hash table in UPC++ delivers near-linear weak scaling up to 34816 cores of a Cray XC40. Our UPC++ implementation of the sparse solver component shows robust strong scaling up to 2048 cores, where it outperforms variants communicating using MPI by up to 3.1x. UPC++ encourages the use of aggressive asynchrony in low overhead RMA and RPC, improving programmer productivity and delivering high performance in irregular applications.
John Bachan, Scott B. Baden, Steven Hofmeyr, Mathias Jacquelin, Amir Kamil, Dan Bonachea, Paul Hargrove, Hadia Ahmed
IPDPS2
2018 Doomsday: predicting which node will fail when on supercomputers
Anwesha Das 0001, Frank Mueller 0001, Paul Hargrove, Eric Roman, Scott B. Baden
SC5
2017 Toucan - A Translator for Communication Tolerant MPI Applications
abstract
We discuss early results with Toucan, a source-to-source translator that automatically restructures C/C++ MPI applications tooverlap communication with computation. We co-designed the translator and runtime system to enable dynamic, dependence-driven execution of MPI applications, and require only a modest amount of programmer annotation. Co-design was essential to realizing overlap through dynamic code block reordering and avoiding the limitations of static code relocation and inlining. We demonstrate that Toucan hides significant communication in four representative applications running on up to 24Kcores of NERSC's Edison platform. Using Toucan, we have hidden from 33% to 85% of the communication overhead, with performance meeting or exceeding that of painstakingly hand-written overlap variants.
Sergio M. Martin, Marsha J. Berger, Scott B. Baden
IPDPS3
2017 Automatic translation of MPI source into a latency-tolerant, data-driven form
Tan Nguyen 0001, Pietro Cicotti, Eric J. Bylaska, Daniel J. Quinlan, Scott B. Baden
J. Parallel Distributed Comput.5
2015 LU Factorization: Towards Hiding Communication Overheads with a Lookahead-Free Algorithm
abstract
Lookahead is a well-known technique for masking communication in matrix factorization, but at the cost of complicating application software. We present a new approach, based on automated code-restructuring, that realizes the benefits of lookahead while avoiding the complications. We apply our technique to HPL, the Linpack benchmark used to assess the performance of supercomputers. Starting with the simpler non-lookahead version of the application, we are able to meet the performance of lookahead on the Stampede mainframe.
Tan Nguyen 0001, Scott B. Baden
CLUSTER2
2014 Effective multi-GPU communication using multiple CUDA streams and threads
abstract
In the context of multiple GPUs that share the same PCIe bus, we propose a new communication scheme that leads to a more effective overlap of communication and computation. Multiple CUDA streams and OpenMP threads are adopted so that data can simultaneously be sent and received. A representative 3D stencil example is used to demonstrate the effectiveness of our scheme. We compare the performance of our new scheme with an MPI-based state-of-the-art scheme. Results show that our approach outperforms the state-of-the-art scheme, being up to 1.85× faster. However, our performance results also indicate that the current underlying PCIe bus architecture needs improvements to handle the future scenario of many GPUs per node.
Mohammed Sourouri, Tor Gillberg, Scott B. Baden, Xing Cai
ICPADS3
2013 A software-based dynamic-warp scheduling approach for load-balancing the Viola-Jones face detection algorithm on GPUs
Tan Nguyen 0001, Daniel Hefenbrock, Jason Oberg, Ryan Kastner, Scott B. Baden
J. Parallel Distributed Comput.5
2012 Bamboo: translating MPI applications to a latency-tolerant, data-driven form
abstract
We present Bamboo, a custom source-to-source translator that transforms MPI C source into a data-driven form that automatically overlaps communication with available computation. Running on up to 98304 processors of NERSC's Hopper system, we observe that Bamboo's overlap capability speeds up MPI implementations of a 3D Jacobi iterative solver and Cannon's matrix multiplication. Bamboo's generated code meets or exceeds the performance of hand optimized MPI, which includes split-phase coding, the method classically employed to hide communication. We achieved our results with only modest amounts of programmer annotation and no intrusive reprogramming of the original application source.
Tan Nguyen 0001, Pietro Cicotti, Eric J. Bylaska, Dan Quinlan, Scott B. Baden
SC5
2011 The Saaz Framework for Turbulent Flow Queries
abstract
In many respects, numerical simulations involving solutions to partial differential equations have replaced physical experimentation. However, few tools are available to sift through the deluge of data. We present Saaz, a query framework to analyze the simulation results of multi-scale physical phenomena which admit mathematical rules for characterizing features of interest. Saaz provides high-level primitives that free the domain-scientist to concentrate more on scientific discovery and less on code implementation and maintenance. It supports user-defined domain-specific query operations which may be subsequently composed into more complex queries. While Saaz supports offline processing of queries, we explore here the online capabilities by attaching Saaz to a running simulation, improving the simulation's effective temporal resolution. We discuss analysis for a computational fluid dynamics simulation of turbulent flow running on a cluster.
Alden King, Eric Arobone, Scott B. Baden, Sutanu Sarkar
eScience3
2011 Mint: realizing CUDA performance in 3D stencil methods with annotated C
abstract
We present Mint, a programming model that enables the non-expert to enjoy the performance benefits of hand coded CUDA without becoming entangled in the details. Mint targets stencil methods, which are an important class of scientific applications. We have implemented the Mint programming model with a source-to-source translator that generates optimized CUDA C from traditional C source. The translator relies on annotations to guide translation at a high level. The set of pragmas is small, and the model is compact and simple. Yet, Mint is able to deliver performance competitive with painstakingly hand-optimized CUDA. We show that, for a set of widely used stencil kernels, Mint realized 80% of the performance obtained from aggressively optimized CUDA on the 200 series NVIDIA GPUs. Our optimizations target three dimensional kernels, which present a daunting array of optimizations.
Didem Unat, Xing Cai, Scott B. Baden
ICS3
2010 Source-to-Source Optimization of CUDA C for GPU Accelerated Cardiac Cell Modeling
Fred V. Lionetti, Andrew D. McCulloch, Scott B. Baden
Euro-Par (1)3
2010 Accelerating Viola-Jones Face Detection to FPGA-Level Using GPUs
abstract
Face detection is an important aspect for biometrics, video surveillance and human computer interaction. We present a multi-GPU implementation of the Viola-Jones face detection algorithm that meets the performance of the fastest known FPGA implementation. The GPU design offers far lower development costs, but the FPGA implementation consumes less power. We discuss the performance programming required to realize our design, and describe future research directions.
Daniel Hefenbrock, Jason Oberg, Nhat Thanh, Ryan Kastner, Scott B. Baden
FCCM5
2009 An Adaptive Sub-sampling Method for In-memory Compression of Scientific Data
abstract
A current challenge in scientific computing is how to curb the growth of simulation datasets without losing valuable information. While wavelet based methods are popular, they require that data be decompressed before it can analyzed, for example, when identifying time-dependent structures in turbulent flows. We present adaptive coarsening, an adaptive subsampling compression strategy that enables the compressed data product to be directly manipulated in memory without requiring costly decompression.We demonstrate compression factors of up to 8 in turbulent flow simulations in three dimensions.Our compression strategy produces a non-progressive multiresolution representation, subdividing the dataset into fixed sized regions and compressing each region independently.
Didem Unat, Theodore Hromadka III, Scott B. Baden
DCC3
2008 Advancing supercomputer performance through interconnection topology synthesis
abstract
In today’s many-core era, the interconnection networks have been the key factor that dominates the performance of a computer system. In this paper, we propose a design flow to discover the best topology in terms of the communication latency and physical constraints. First a set of representative candidate topologies are generated for the interconnection networks among computing chips; then an efficient multi-commodity flow algorithm is devised to evaluate the performance. The experiments show that the best topologies identified by our algorithm can achieve better average latency compared to the existing networks.
Yi Zhu 0002, Michael B. Taylor, Scott B. Baden, Chung-Kuan Cheng
ICCAD3
2007 Toward Petascale Simulation of Cellular Microphysiology
abstract
MCell is a Monte Carlo simulator of cell microphysiology, and the scalable variant can be used to study challenging problems of interest to the biological community. MCell can currently model a single synapse out of thousands on a single cell. Petascale technology will enable significant advances in the ability to treat larger structures involving many synapses, with correspondingly more complex behavior. However, there are significant challenges to scaling MCell across two orders of magnitude in performance: increased communication delays and uneven workload concentrations. We discuss software solutions currently under investigation that will accompany us on the path to petascale cell microphysiology.
Scott B. Baden, Terrence J. Sejnowski, Thomas M. Bartol, Joel R. Stiles
BIBE1
2006 Short Paper: Asynchronous programming with Tarragon
abstract
Tarragon is an actor-based programming model and library for implementing latency tolerant asynchronous event driven simulations. It is novel in its support for meta data describing run time virtualized process structures, which may be optimized as a free-standing object. We demonstrate early results with a synthetic benchmark, and observe that Tarragon can mask communication costs with ongoing computation
Pietro Cicotti, Scott B. Baden
HPDC2
2006 Poster reception - Asynchronous programming with Tarragon
abstract
Tarragon is an actor-based programming model and library for implementing parallel scientific applications requiring fine grain asynchronous communication. Tarragon raises the level of abstraction by encapsulating run-time services that mange the actor semantics. The workload is over-decomposed into many virtual processes called WorkUnits. WorkUnits can become ready for execution after receiving input; scheduling and communication services coordinate WorkUnit execution and management. In order to maintain balanced workloads, Tarragon automatically monitors workload distribution and redistributes as needed. Tarragon is novel in its support for meta data describing run-time virtual process structures used to manage actor semantics. This meta data may be used to guide run time services policies in order to optimize performance. We are currently applying Tarragon to the MCell cell microphysiology simulator and are considering other applications as well, such as sparse matrix linear algebra.
Pietro Cicotti, Scott B. Baden
SC2
2005 Building an XQuery interpreter in a compiler construction course
abstract
For two years, we have been teaching a quarter-long compiler construction course where students implement an interpreter for a variant of the XML query language XQuery. Our goal is to motivate students' interest in the course by exposing them to an interesting and powerful new language which they see as relevant to potential future experiences. In this paper, we first explain the workings of the course itself, and then describe some pedagogically interesting variants of the XQuery language. We close with a discussion of challenges faced and conclusions.
Sara Miner More, Tim Pevzner, Alin Deutsch, Scott B. Baden, Paul Kube
SIGCSE4
2004 A Large Scale Monte Carlo Simulator for Cellular Microphysiology
abstract
Summary form only given. Biological structures are extremely complex at the cellular level. The MCell project has been highly successful in simulating the microphysiology of systems of modest size, but many larger problems require too much storage and computation time to be simulated on a single workstation. MCell-K, a new parallel variant of MCell, has been implemented using the KeLP framework and is running on NPACl's Blue Horizon. MCell-K not only produces validated results consistent with the serial version of MCell but does so with unprecedented scalability. We have thus found a level of description and a way to simulate cellular systems that can approach the complexity of nature on its own terms. At the heart of MCell is a 3D random walk that models diffusion using a Monte Carlo method. We discuss two challenging issues that arose in parallelizing the diffusion process - detecting time-step termination efficiently and performing parallel diffusion of particles in a biophysically accurate way. We explore the scalability limits of the present parallel algorithm and discuss ways to improve upon these limits.
Gregory T. Balls, Scott B. Baden, Tilman Kispersky, Thomas M. Bartol, Terrence J. Sejnowski
IPDPS2
2003 SCALLOP: A Highly Scalable Parallel Poisson Solver in Three Dimensions
abstract
SCALLOP is a highly scalable solver and library for elliptic partial differential equations on regular block-structured domains. SCALLOP avoids high communication overheads algorithmically by taking advantage of the locality properties inherent to solutions to elliptic PDEs. Communication costs are small, on the order of a few percent of the total running time on up to 1024 processors of NPACI's and NERSC's IBM Power-3 SP sytems. SCALLOP trades off numerical overheads against communication. These numerical overheads are independent of the number of processors for a wide range of problem sizes. SCALLOP is implicitly designed for infinite domain (free space) boundary conditions, but the algorithm can be reformulated to accommodate other boundary conditions. The SCALLOP library is built on top of the KeLP programming system and runs on a variety of platforms.
Gregory T. Balls, Scott B. Baden, Phillip Colella
SC2
2001 Topic 10: Parallel Programming: Models, Methods and Programming Languages
Scott B. Baden, Paul H. J. Kelly, Sergei Gorlatch, Calvin Lin
Euro-Par1
2001 Parallel Software Abstractions for Structured Adaptive Mesh Methods
Scott R. Kohn, Scott B. Baden
J. Parallel Distributed Comput.2
2000 Programming Languages, Models, and Methods
Paul H. J. Kelly, Sergei Gorlatch, Scott B. Baden, Vladimir Getov
Euro-Par3
2000 A Programming Methodology for Dual-Tier Multicomputers
abstract
Hierarchically organized ensembles of shared memory multiprocessors possess a richer and more complex model of locality than previous generation multicomputers with single processor nodes. These dual-tier computers introduce many new factors into the programmer's performance model. We present a methodology for implementing block-structured numerical applications on dual-tier computers and a run-time infrastructure, called KeLP2, that implements the methodology. KeLP2 supports two levels of locality and parallelism via hierarchical SPMD control flow, run-time geometric meta-data, and asynchronous collective communication. KeLP applications can effectively overlap communication with computation under conditions where nonblocking point-to-point message passing fails to do so. KeLP's abstractions hide considerable detail without sacrificing performance and dual-tier applications written in KeLP consistently outperform equivalent single-tier implementations written in MPI. We describe the KeLP2 model and show how it facilitates the implementation of five block-structured applications specially formulated to hide communication latency on dual-tiered architectures. We support our arguments with empirical data from applications running on various single- and dual-tier multicomputers. KeLP2 supports a migration path from single-tier to dual-tier platforms and we illustrate this capability with a detailed programming example.
Scott B. Baden, Stephen J. Fink
IEEE Trans. Software Eng.1
1999 Multiple data parallelism with HPF and KeLP
John H. Merlin, Scott B. Baden, Stephen J. Fink, Barbara M. Chapman
Future Gener. Comput. Syst.2
1998 Communication overlap in multi-tier parallel algorithms
abstract
Hierarchically organized multicomputers such as SMP clusters offer new opportunities and new challenges for high-performance computation, but realizing their full potential remains a formidable task. We present a hierarchical model of communication targeted to block- structured, bulk-synchronous applications running on dedicated clusters of symmetric multiprocessors. Our model supports node-level rather processor-level communication as the fundamental operation, and is optimized for aggregate patterns of regular section moves rather than point-to-point messages. These two capabilities work synergistically. They provide flexibility in overlapping communication and overcome deficiencies in the underlying communication layer on systems where inter-node communication bandwidth is at a premium. We have implemented our communication model in the KeLP2.0 run time library. We present empirical results for five applications running on a cluster of Digital AlphaServer 2100's. Four of the applications were able to overlap communication on a system which does not support overlap via non-blocking message passing using MPI. Overall performance improvements due to our overlap strategy ranged from 12% to 28%.
Scott B. Baden, Stephen J. Fink
SC1
1998 Efficient Run-Time Support for Irregular Block-Structured Applications
Stephen J. Fink, Scott B. Baden, Scott R. Kohn
J. Parallel Distributed Comput.2
1997 Parallel Cluster Identification for Multidimensional Lattices
abstract
The cluster identification problem is a variant of connected component labeling that arises in cluster algorithms for spin models in statistical physics. We present a multidimensional version of K.P. Belkhale and P. Banerjee's quad algorithm (1992) for connected component labeling on distributed memory parallel computers. Our extension abstracts away extraneous spatial connectivity information in more than two dimensions, simplifying implementation for higher dimensionality. We identify two types of locality present in cluster configurations, and present optimizations to exploit locality for better performance. Performance results from 2D, 3D, and 4D Ising model simulations with Swendson-Wang dynamics show that the optimizations improve performance by 20-80 percent.
Stephen J. Fink, Craig Huston, Scott B. Baden, Karl Jansen
IEEE Trans. Parallel Distributed Syst.3
1996 Dynamic Partitioning of Non-Uniform Structured Workloads with Spacefilling Curves
abstract
We discuss inverse spacefilling partitioning (ISP), a partitioning strategy for non-uniform scientific computations running on distributed memory MIMD parallel computers. We consider the case of a dynamic workload distributed on a uniform mesh, and compare ISP against orthogonal recursive bisection (ORE) and a median of medians variant of ORE, ORB-MM. We present two results. First, ISP and ORB-MM are superior to ORE in rendering balanced workloads-because they are more fine-grained-and incur communication overheads that are comparable to ORE. Second, ISP is more attractive than ORB-MM from a software engineering standpoint because it avoids elaborate bookkeeping. Whereas ISP partitionings can be described succinctly as logically contiguous segments of the line, ORB-MM's partitionings are inherently unstructured. We describe the general d-dimensional ISP algorithm and report empirical results with two- and three-dimensional, non-hierarchical particle methods.
John R. Pilkington, Scott B. Baden
IEEE Trans. Parallel Distributed Syst.2
1995 A Parallel Software Infrastructure for Structured Adaptive Mesh Methods
abstract
Structured adaptive mesh algorithms dynamically allocate computational resources to accurately resolve interesting portions of a numerical calculation. Such methods are difficult to implement and parallelize because they rely on dynamic, irregular data structures. We have developed an efficient, portable, parallel software infrastructure for adaptive mesh methods; our software provides computational scientists with high-level facilities that hide low-level details of parallelism and resource management. We have applied our software infrastructure to the solution of adaptive eigenvalue problems arising in materials design. We describe our software infrastructure and analyze its performance. We also present computational results which indicate that the uniformity restrictions imposed by a data parallel Fortran implementation of a structured adaptive mesh application would significantly impact performance. 1 Introduction The accurate solution of many problems in science and engineering ...
Scott R. Kohn, Scott B. Baden
SC2
1995 Portable Parallel Programming of Numerical Problems under the LPAR System
Scott B. Baden, Scott R. Kohn
J. Parallel Distributed Comput.1