EDBT 2026 Demo / reviewers in the wild / expert
Phillip M. Dickens
dblp:46/3540
· DBLP profile ↗
14ranked-venue papers
11as first author
0since 2021 · last 2015
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 11 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Storage systems · 55% High-performance computing · 24% Performance modeling and evaluation · 12% | |
| Computer networks
1 paper |
Transport protocols and congestion control · 87% Network performance modeling · 13% |
Topics — the 11 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Storage systems
i/o optimization |
0.1 | 1 | 2009 | Y-lib: a user level library to increase the performance of MPI-IO in a lustre file system environment · HPDC 2009 |
Storage systems › file systems › distributed file system
parallel file system |
0.1 | 1 | 2009 | Y-lib: a user level library to increase the performance of MPI-IO in a lustre file system environment · HPDC 2009 |
High-performance computing
parallel i/o |
0.1 | 1 | 2009 | Y-lib: a user level library to increase the performance of MPI-IO in a lustre file system environment · HPDC 2009 |
Transport protocols and congestion control
TCP performance |
0.0 | 1 | 2002 | An Evaluation of Object-Based Data Transfers on High Performance Networks · HPDC 2002 |
Transport protocols and congestion control
transport protocols |
0.0 | 1 | 2002 | An Evaluation of Object-Based Data Transfers on High Performance Networks · HPDC 2002 |
Interconnection networks and networks-on-chip
high-speed networks |
0.0 | 1 | 2002 | An Evaluation of Object-Based Data Transfers on High Performance Networks · HPDC 2002 |
Storage systems › file systems › distributed file system › parallel file system
lustre file system |
0.0 | 1 | 2009 | Y-lib: a user level library to increase the performance of MPI-IO in a lustre file system environment · HPDC 2009 |
Performance modeling and evaluation › simulation › architectural simulation
execution-driven simulation |
0.0 | 1 | 1996 | Parallelized Direct Execution Simulation of Message-Passing Parallel Programs · IEEE Trans. Parallel Distributed Syst. 1996 |
Performance modeling and evaluation › performance prediction
parallel program performance prediction |
0.0 | 1 | 1996 | Parallelized Direct Execution Simulation of Message-Passing Parallel Programs · IEEE Trans. Parallel Distributed Syst. 1996 |
Performance modeling and evaluation
simulation |
0.0 | 1 | 1996 | Parallelized Direct Execution Simulation of Message-Passing Parallel Programs · IEEE Trans. Parallel Distributed Syst. 1996 |
Parallel and multicore computing › parallel programming models › message passing
message-passing parallel programs |
0.0 | 1 | 1996 | Parallelized Direct Execution Simulation of Message-Passing Parallel Programs · IEEE Trans. Parallel Distributed Syst. 1996 |
Methods — techniques the papers use, named apart from their topics
benchmarking · 0.2parallelized simulation · 0.0discrete-event simulation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2015 | A Performance and Scalability Analysis of the MPI Based Tools Utilized in a Large Ice Sheet Model Executing in a Multicore Environment
Phillip M. Dickens |
ICA3PP (4) | 1 |
| 2010 | A high performance implementation of MPI-IO for a Lustre file system environmentabstractAbstract It is often the case that MPI‐IO performs poorly in a Lustre file system environment, although the reasons for such performance have heretofore not been well understood. We hypothesize that such performance is a direct result of the fundamental assumptions upon which most parallel I/O optimizations are based. In particular, it is almost universally believed that parallel I/O performance is optimized when aggregator processes perform large, contiguous I/O operations in parallel. Our research, however, shows that this approach can actually provide the worst performance in a Lustre environment, and that the best performance may be obtained by performing a large number of small, non‐contiguous I/O operations. In this paper, we provide empirical results demonstrating these non‐intuitive results and explore the reasons for such unexpected performance. We present our solution to the problem, which is embodied in a user‐level library termed Y‐Lib, which redistributes the data in a way that conforms much more closely with the Lustre storage architecture than does the data redistribution pattern employed by MPI‐IO. We provide a large body of experimental results, taken across two large‐scale Lustre installations, demonstrating that Y‐Lib outperforms MPI‐IO by up to 36% on one system and 1000% on the other. We discuss the factors that impact the performance improvement obtained by Y‐Lib, which include the number of aggregator processes and Object Storage Devices, as well as the power of the system's communications infrastructure. We also show that the optimal data redistribution pattern for Y‐Lib is dependent upon these same factors. Copyright © 2009 John Wiley & Sons, Ltd. Phillip M. Dickens, Jeremy Logan |
Concurr. Comput. Pract. Exp. | 1 |
| 2009 | Y-lib: a user level library to increase the performance of MPI-IO in a lustre file system environmentabstractIt is widely known that MPI-IO performs poorly in a Lustre file system environment, although the reasons for such performance are currently not well understood. The research presented in this paper strongly supports our hypothesis that MPI-IO performs poorly in this environment because of the fundamental assumptions upon which most parallel I/O optimizations are based. In particular, it is almost universally believed that parallel I/O performance is optimized when aggregator processes perform large, contiguous I/O operations in parallel. Our research shows that this approach generally provides the worst performance in a Lustre environment, and that the best performance is often obtained when the aggregator processes perform a large number of small, non-contiguous I/O operations. Phillip M. Dickens, Jeremy Logan |
HPDC | 1 |
| 2008 | Towards an understanding of the performance of MPI-IO in Lustre file systemsabstractLustre is becoming an increasingly important file system for large-scale computing clusters. The problem, however, is that many data-intensive applications use MPI-IO for their I/O requirements, and MPI-IO performs poorly in a Lustre file system environment. While this poor performance has been well documented, the reasons for such performance are currently not well understood. Our research suggests that the primary performance issues have to do with the assumptions underpinning most of the parallel I/O optimizations implemented in MPI-IO, which do not appear to hold in a Lustre environment. Perhaps the most important assumption is that optimal performance is obtained by performing large, contiguous I/O operations. However, the research results presented in this poster show that this is often the worst approach to take in a Lustre file system. In fact, we found that the best performance is often achieved when each process performs a series of smaller, non-contiguous I/O requests. In this poster, we provide experimental results supporting these non-intuitive ideas, and provide alternative approaches that significantly enhance the performance of MPI-IO in a Lustre file system. Jeremy Logan, Phillip M. Dickens |
CLUSTER | 2 |
| 2005 | Towards a Bayesian Statistical Model for the Classification of the Causes of Data Loss
Phillip M. Dickens, Jeffery Peden |
HPCC | 1 |
| 2004 | Classifiers for the causes of data loss using packet-loss signaturesabstractA necessary step in the development of next-generation congestion control mechanisms is the ability to accurately classify the root cause(s) of observed data loss and to develop responses tailored to the particular cause. Toward this end, we are developing a classification mechanism based on the collection and analysis of what we term packet-loss signatures, which describe the patterns of packet loss in the current transmission window. We are exploring the application of complexity theory to the problem of learning the underlying structure (or lack thereof) of these signatures, and studying the relationship between such underlying structure and the system conditions responsible for its generation. In this paper, we describe the algorithm for determining the complexity of packet-loss signatures, show how complexity measures can be mapped to the underlying causes of packet loss, and provide experimental results demonstrating the effectiveness of our approach. Phillip M. Dickens, Jay W. Larson |
CCGRID | 1 |
| 2004 | Diagnostics for Causes of Packet Loss in a High Performance Data Transfer SystemabstractSummary form only given. As computational grids become an increasingly dominant force in the high-performance computing arena, the problem of efficiently transferring very large data sets, across geographically distributed computing resources, becomes increasingly difficult and important. Current approaches view the problem largely, if not exclusively, as a network-level problem. Thus all packet loss is interpreted and treated as a network congestion event, limiting the ability to detect or react to changes in the end-to-end system. We believe that a new approach to this problem is worth pursuing, and we are investigating techniques that can differentiate between data loss caused by contention in the network and loss caused by contention for shared CPU resources at the communication endpoints. The approach is to collect and analyze what we term packet-loss signatures that describe the patterns of packet-loss in the current transmission window. We analyze these signatures using Fourier analysis and symbolic dynamics, and present a simple set of experiments demonstrating the effectiveness of this approach. Phillip M. Dickens, Jay W. Larson, David M. Nicol |
IPDPS | 1 |
| 2003 | FOBS: A Lightweight Communication Protocol for Grid Computing
Phillip M. Dickens |
Euro-Par | 1 |
| 2002 | An Evaluation of Object-Based Data Transfers on High Performance NetworksabstractWe describe FOBS: a simple user-level communication protocol designed to take advantage of the available bandwidth in a high-bandwidth, high-delay network environment. We compare the performance of FOBS with that of TCP both with and without the so-called Large Window extensions designed to improve the performance of TCP in this type of network environment. It is shown that FOBS can obtain on the order of 90% of the available bandwidth across both short and long high-performance network connections. In the case of the long haul connection, this represents a bandwidth that is 1.8 times higher than that of the optimized TCP algorithm. Also, we demonstrate that the additional traffic placed on the network due to the greedy nature of the algorithm is quite reasonable, representing approximately 3% of the total data transferred. Phillip M. Dickens, William Gropp |
HPDC | 1 |
| 2001 | High-performance file I/O in Java: Existing approaches and bulk I/O extensionsabstractAbstract There is a growing interest in using Java as the language for developing high‐performance computing applications. To be successful in the high‐performance computing domain, however, Java must not only be able to provide high computational performance, but also high‐performance I/O. In this paper, we first examine several approaches that attempt to provide high‐performance I/O in Java—many of which are not obvious at first glance—and evaluate their performance on two parallel machines, the IBM SP and the SGI Origin2000. We then propose extensions to the Java I/O library that address the deficiencies in the Java I/O API and improve performance dramatically. The extensions add bulk (array) I/O operations to Java, thereby removing much of the overhead currently associated with array I/O in Java. We have implemented the extensions in two ways: in a standard JVM using the Java Native Interface (JNI) and in a high‐performance parallel dialect of Java called Titanium. We describe the two implementations and present performance results that demonstrate the benefits of the proposed extensions. Copyright © 2001 John Wiley & Sons, Ltd. Dan Bonachea, Phillip M. Dickens, Rajeev Thakur |
Concurr. Comput. Pract. Exp. | 2 |
| 2001 | Evaluation of Collective I/O Implementations on Parallel Architectures
Phillip M. Dickens, Rajeev Thakur |
J. Parallel Distributed Comput. | 1 |
| 2000 | Java on networks of workstations (JavaNOW): a parallel computing framework inspired by Linda and the Message Passing Interface (MPI)abstractNetworks of workstations are a dominant force in the distributed computing arena, due primarily to the excellent price/performance ratio of such systems when compared to traditionally massively parallel architectures. It is therefore critical to develop programming languages and environments that can help harness the raw computational power available on these systems. In this article, we present JavaNOW (Java on Networks of Workstations), a Java-based framework for parallel programming on networks of workstations. It creates a virtual parallel machine similar to the MPI (Message Passing Interface) model, and provides distributed associative shared memory similar to the Linda memory model but with a richer set of primitive operations. JavaNOW provides a simple yet powerful framework for performing computation on networks of workstations. In addition to the Linda memory model, it provides for shared objects, implicit multithreading, implicit synchronization, object dataflow, and collective communications similar to those defined in MPI. JavaNOW is also a component of the Computational Neighborhood, a Java-enabled suite of services for desktop computational sharing. The intent of JavaNOW is to present an environment for parallel computing that is both expressive and reliable and ultimately can deliver good to excellent performance. As JavaNOW is a work in progress, this article emphasizes the expressive potential of the JavaNOW environment and presents preliminary performance results only. Copyright © 2000 John Wiley & Sons, Ltd. George K. Thiruvathukal, Phillip M. Dickens, Shahzad Bhatti |
Concurr. Pract. Exp. | 2 |
| 1998 | A Performance Study of Two-Phase I/O
Phillip M. Dickens |
Euro-Par | 1 |
| 1996 | Parallelized Direct Execution Simulation of Message-Passing Parallel ProgramsabstractAs massively parallel computers proliferate, there is growing interest in finding ways by which performance of massively parallel codes can be efficiently predicted. This problem arises in diverse contexts such as parallelizing compilers, parallel performance monitoring, and parallel algorithm development. In this paper, we describe one solution where one directly executes the application code, but uses a discrete-event simulator to model details of the presumed parallel machine, such as operating system and communication network behavior. Because this approach is computationally expensive, we are interested in its own parallelization, specifically the parallelization of the discrete-event simulator. We describe methods suitable for parallelized direct execution simulation of message-passing parallel programs, and report on the performance of such a system, LAPSE (Large Application Parallel Simulation Environment), we have built on the Intel Paragon. On all codes measured to date, LAPSE predicts performance well, typically within 10% relative error. Depending on the nature of the application code, we have observed low slowdowns (relative to natively executing code) and high relative speedups using up to 64 processors. Phillip M. Dickens, Philip Heidelberger, David M. Nicol |
IEEE Trans. Parallel Distributed Syst. | 1 |