Kei Davis

dblp:49/4295 · DBLP profile ↗
← Back
27ranked-venue papers
3as first author
0since 2021 · last 2020
0000-0002-4134-1798ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 24 · 1 first-authorSoftware engineering, systems software and programming languages · 3 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
12 papers
Storage systems · 32% High-performance computing · 28% Memory systems · 26%
Software engineering, system software, and programming languages
2 papers
Operating systems · 100%

Topics — the 30 heaviest of 39, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing
scientific computing systems
0.422019
Persistent Octrees for Parallel Mesh Refinement through Non-Volatile Byte-Addressable Memory · IEEE Trans. Parallel Distributed Syst. 2019
OVERTURE: An Object-Oriented Framework for High Performance Scientific Computing · SC 1998
High-performance computing › scientific computing systems
adaptive mesh refinement
0.412019
Persistent Octrees for Parallel Mesh Refinement through Non-Volatile Byte-Addressable Memory · IEEE Trans. Parallel Distributed Syst. 2019
Memory systems › non-volatile memory › persistent memory
byte-addressable persistent memory
0.412019
Persistent Octrees for Parallel Mesh Refinement through Non-Volatile Byte-Addressable Memory · IEEE Trans. Parallel Distributed Syst. 2019
Distributed systems › fault tolerance
failure recovery
0.412019
Persistent Octrees for Parallel Mesh Refinement through Non-Volatile Byte-Addressable Memory · IEEE Trans. Parallel Distributed Syst. 2019
Memory systems
non-volatile memory
0.412019
Persistent Octrees for Parallel Mesh Refinement through Non-Volatile Byte-Addressable Memory · IEEE Trans. Parallel Distributed Syst. 2019
Storage systems
storage reliability
0.412019
Persistent Octrees for Parallel Mesh Refinement through Non-Volatile Byte-Addressable Memory · IEEE Trans. Parallel Distributed Syst. 2019
Storage systems › i/o optimization
collective i/o
0.212013
Orthrus: a framework for implementing high-performance collective I/O in the multicore clusters · HPDC 2013
Storage systems › i/o optimization › i/o prefetching
disk prefetching
0.212013
A Prefetching Scheme Exploiting both Data Layout and Access History on Disk · ACM Trans. Storage 2013
High-performance computing
parallel i/o
0.212013
Orthrus: a framework for implementing high-performance collective I/O in the multicore clusters · HPDC 2013
Storage systems › file systems › distributed file system
parallel file system
0.222013
IOrchestrator: Improving the Performance of Multi-node I/O Systems via Inter-Server Coordination · SC 2010
Orthrus: a framework for implementing high-performance collective I/O in the multicore clusters · HPDC 2013
High-performance computing › supercomputing
supercomputing systems
0.122008
Entering the petaflop era: the architecture and performance of Roadrunner · SC 2008
A Performance and Scalability Analysis of the BlueGene/L Architecture · SC 2004
Storage systems
shared storage
0.112011
QoS support for end users of I/O-intensive applications using shared storage systems · SC 2011
Storage systems › storage performance
storage quality of service
0.112011
QoS support for end users of I/O-intensive applications using shared storage systems · SC 2011
Memory systems › non-volatile memory
persistent data structures
0.112019
Persistent Octrees for Parallel Mesh Refinement through Non-Volatile Byte-Addressable Memory · IEEE Trans. Parallel Distributed Syst. 2019
Memory systems › cache management › chip multiprocessor cache management
cache coordination
0.112010
Improving Networked File System Performance Using a Locality-Aware Cooperative Cache Protocol · IEEE Trans. Computers 2010
Memory systems › cache management › storage caching
cooperative caching
0.112010
Improving Networked File System Performance Using a Locality-Aware Cooperative Cache Protocol · IEEE Trans. Computers 2010
Storage systems › file systems
distributed file system
0.112010
Improving Networked File System Performance Using a Locality-Aware Cooperative Cache Protocol · IEEE Trans. Computers 2010
Storage systems › i/o architecture
i/o subsystem
0.112010
IOrchestrator: Improving the Performance of Multi-node I/O Systems via Inter-Server Coordination · SC 2010
Hardware accelerators and domain-specific architectures › many-core accelerator
cell broadband engine
0.112008
Entering the petaflop era: the architecture and performance of Roadrunner · SC 2008
Storage systems › buffer management
buffer cache management
0.112007
Coordinated Multilevel Buffer Cache Management with Consistent Access Locality Quantification · IEEE Trans. Computers 2007
Memory systems
cache management
0.112007
Coordinated Multilevel Buffer Cache Management with Consistent Access Locality Quantification · IEEE Trans. Computers 2007
Memory systems › cache management
cache replacement
0.112007
Coordinated Multilevel Buffer Cache Management with Consistent Access Locality Quantification · IEEE Trans. Computers 2007
Storage systems › i/o optimization
i/o prefetching
0.112007
DiskSeen: Exploiting Disk Layout and Access History to Enhance I/O Prefetch · USENIX ATC 2007
Operating systems › resource management › storage management
file systems
0.122013
A Prefetching Scheme Exploiting both Data Layout and Access History on Disk · ACM Trans. Storage 2013
DiskSeen: Exploiting Disk Layout and Access History to Enhance I/O Prefetch · USENIX ATC 2007
Distributed systems › communication optimization
communication-computation overlap
0.112006
MPI tools and performance studies - Quantifying the potential benefit of overlapping communication and computation in large-scale scientific applications · SC 2006
High-performance computing
performance optimization at scale
0.112006
MPI tools and performance studies - Quantifying the potential benefit of overlapping communication and computation in large-scale scientific applications · SC 2006
High-performance computing › supercomputing
bluegene/l
0.012004
A Performance and Scalability Analysis of the BlueGene/L Architecture · SC 2004
High-performance computing
i/o intensive applications
0.012011
QoS support for end users of I/O-intensive applications using shared storage systems · SC 2011
Distributed systems
distributed coordination
0.012010
IOrchestrator: Improving the Performance of Multi-node I/O Systems via Inter-Server Coordination · SC 2010
Performance modeling and evaluation › benchmarking
microbenchmarking
0.012008
Entering the petaflop era: the architecture and performance of Roadrunner · SC 2008

Methods — techniques the papers use, named apart from their topics

feature-directed sampling · 0.4erasure coding · 0.4prefetching · 0.3access pattern tracking · 0.3performance modeling · 0.1qos-aware scheduling · 0.1trace-driven simulation · 0.1reuse distance analysis · 0.1i/o orchestration · 0.1microbenchmarks · 0.1finite volume · 0.0finite difference · 0.0composite overlapping grid · 0.0adaptive mesh refinement · 0.0
YearPublicationVenuePosition
2020 On the memory attribution problem: A solution and case study using MPI
abstract
Summary As parallel applications running on large‐scale computing systems become increasingly memory constrained, the ability to attribute memory usage to the various components of the application is becoming increasingly important. We present the design and implementation of memnesia, a novel memory usage profiler for parallel and distributed message‐passing applications. Our approach captures both application– and message‐passing library–specific memory usage statistics from unmodified binaries dynamically linked to a message‐passing communication library. Using microbenchmarks and proxy applications, we evaluated our profiler across three Message Passing Interface (MPI) implementations and two hardware platforms. The results show that our approach and the corresponding implementation can accurately quantify memory resource usage as a function of time, scale, communication workload, and software or hardware system architecture, clearly distinguishing between application and MPI library memory usage at a per‐process level. With this new capability, we show that job size, communication workload, and hardware/software architecture influence peak runtime memory usage. In practice, this tool provides a potentially valuable source of information for application developers seeking to measure and optimize memory usage.
Samuel K. Gutierrez, Dorian C. Arnold, Kei Davis, Patrick S. McCormick
Concurr. Comput. Pract. Exp.3
2019 Persistent Octrees for Parallel Mesh Refinement through Non-Volatile Byte-Addressable Memory
abstract
Adaptive mesh refinement based on octree data structures has enabled efficient simulations of complex physical phenomena. Existing meshing algorithms were proposed with the assumption that computer memory is volatile. Consequently, for failure recovery, in-core algorithms need to save memory states as snapshots with slow file I/O, while out-of-core algorithms store octants on disk for persistence. However, neither was designed to best exploit the unique characteristics of non-volatile byte-addressable memory (NVBM). We propose a novel data structure, the Distributed Persistent Merged octree (DPM-octree), for both meshing and in-memory storage of persistent octrees using NVBM. DPM-octree is a multi-version data structure that can recover from failures using an earlier persistent version stored in NVBM. In addition, we design a feature-directed sampling approach to help dynamically transform the DPM-octree layout for reducing NVBM-induced memory write latency. DPM-octree uses parity trees which are created using erasure coding and stored in NVBM to support low-latency in-memory octant recovery after data loss. DPM-octree has been successfully integrated with the Gerris software for simulation of fluid dynamics. Our experimental results with real-world scientific workloads show that DPM-octree scales up to 1.1 billion mesh elements with 1,000 processors on the Titan supercomputer.
Bao Nguyen, Hua Tan, Kei Davis, Xuechen Zhang 0001
IEEE Trans. Parallel Distributed Syst.3
2017 Accommodating Thread-Level Heterogeneity in Coupled Parallel Applications
abstract
Hybrid parallel program models that combine message passing and multithreading (MP+MT) are becoming more popular, extending the basic message passing (MP) model that uses single-threaded processes for both inter- and intra-node parallelism. A consequence is that coupled parallel applications increasingly comprise MP libraries together with MP+MT libraries with differing preferred degrees of threading, resulting in thread-level heterogeneity. Retroactively matching threading levels between independently developed and maintained libraries is difficult; the challenge is exacerbated because contemporary parallel job launchers provide only static resource binding policies over entire application executions. A standard approach for accommodating thread-level heterogeneity is to under-subscribe compute resources such that the library with the highest degree of threading per process has one processing element per thread. This results in libraries with fewer threads per process utilizing only a fraction of the available compute resources. We present and evaluate a novel approach for accommodating thread-level heterogeneity. Our approach enables full utilization of all available compute resources throughout an application's execution by providing programmable facilities to dynamically reconfigure runtime environments for compute phases with differing threading factors and memory affinities. We show that our approach can improve overall application performance by up to 5.8× in real-world production codes. Furthermore, the practicality and utility of our approach has been demonstrated by continuous production use for over one year, and by more recent incorporation into a number of production codes.
Samuel K. Gutierrez, Kei Davis, Dorian C. Arnold, Randal S. Baker, Robert W. Robey, Patrick S. McCormick, Daniel Holladay, Jon A. Dahl, Joe Zerr, Florian Weik, Christoph Junghans
IPDPS2
2013 Orthrus: a framework for implementing high-performance collective I/O in the multicore clusters
Xuechen Zhang 0001, Jianqiang Ou, Kei Davis, Song Jiang 0001
HPDC3
2013 iBridge: Improving Unaligned Parallel File Access with Solid-State Drives
abstract
When files are striped in a parallel I/O system, requests to the files are decomposed into a number of sub-requests that are distributed over multiple servers. If a request is not aligned with the striping pattern such decomposition can make the first and last sub-requests much smaller than the striping unit. Because hard-disk-based servers can be much less efficient in serving small requests than large ones, the system exhibits heterogeneity in serving sub-requests of different sizes, and the net throughput of the entire system can be severely degraded by the inefficiency of serving the smaller requests, or fragments. Because a request is not considered complete until its slowest sub-request is, the penalty is yet greater for synchronous requests. To make the situation even worse, the larger the request, or the more data servers the requested data is striped over, the larger the detrimental performance effect of serving fragments can be. This effect can become the Achilles' heel of a parallel I/O system performance seeking scalability with large sequential accesses. In this paper we propose iBridge, a scheme that uses solid-state drives to serve request fragments and thereby bridge the performance gap between serving fragments and serving large sub-requests. We have implemented iBridge in the PVFS file system. Our experimental results with representative MPI-IO benchmarks show that iBridge can significantly improve the I/O throughput of storage systems, especially for large requests with fragments.
Xuechen Zhang 0001, Kei Davis, Song Jiang 0001
IPDPS3
2013 Synergistic coupling of SSD and hard disk for QoS-aware virtual memory
abstract
With significant advantages in capacity, power consumption, and price, solid state disk (SSD) has good potential to be employed as an extension of DRAM (memory), such that applications with large working sets could run efficiently on a modestly configured system. While initial results reported in recent works show promising prospects for this use of SSD by incorporating it into the management of virtual memory, frequent writes from write-intensive programs could quickly wear out SSD, making the idea less practical. We propose a scheme, HybridSwap, that integrates a hard disk with an SSD for virtual memory management, synergistically achieving the advantages of both. In addition, HybridSwap can constrain performance loss caused by swapping according to user-specified QoS requirements.
Xuechen Zhang 0001, Kei Davis, Song Jiang 0001
ISPASS3
2013 A Prefetching Scheme Exploiting both Data Layout and Access History on Disk
abstract
Prefetching is an important technique for improving effective hard disk performance. A prefetcher seeks to accurately predict which data will be requested and load it ahead of the arrival of the corresponding requests. Current disk prefetch policies in major operating systems track access patterns at the level of file abstraction. While this is useful for exploiting application-level access patterns, for two reasons file-level prefetching cannot realize the full performance improvements achievable by prefetching. First, certain prefetch opportunities can only be detected by knowing the data layout on disk, such as the contiguous layout of file metadata or data from multiple files. Second, nonsequential access of disk data (requiring disk head movement) is much slower than sequential access, and the performance penalty for mis-prefetching a randomly located block, relative to that of a sequential block, is correspondingly greater.
Song Jiang 0001, Xiaoning Ding, Yuehai Xu, Kei Davis
ACM Trans. Storage4
2012 iHarmonizer: Improving the Disk Efficiency of I/O-intensive Multithreaded Codes
abstract
Challenged by serious power and thermal constraints and limited by available instruction-level parallelism, processor designs have evolved to multi-core architectures. These architectures, many augmented with native simultaneous multithreading, are driving software developers to use multithreaded programs to exploit thread-level parallelism. While multithreading is well known to introduce concerns of data dependency and CPU load balance, less known is that the uncertainty of relative progress of thread execution can cause patterns of I/O requests, issued by different threads, to be effectively random and so significantly degrade hard-disk efficiency. This effect can severely offset the performance gains from parallel execution, especially for I/O-intensive programs. Retaining the benefits of multithreading while not losing I/O efficiency is an urgent and challenging problem. We propose a user-level scheme, iHarmonizer, to streamline the servicing of I/O requests from multiple threads in the Open MP programs. Specifically, we use the compiler to insert code into Open MP programs so that data usage can be transmitted at run time to a supporting run-time library that prefetches data in a disk friendly way and coordinates threads' execution according to the availability of their requested data. Transparent to the programmer, iHarmonizer makes a multithreaded program I/O efficient while maintaining the benefits of parallelism. Our experiments show that iHarmonizer can significantly speed up the execution of a representative set of I/O-intensive scientific benchmarks.
Kei Davis, Yuehai Xu, Song Jiang 0001
IPDPS2
2012 Opportunistic Data-driven Execution of Parallel Programs for Efficient I/O Services
abstract
A parallel system relies on both process scheduling and I/O scheduling for efficient use of resources, and a program's performance hinges on the resource on which it is bottlenecked. Existing process schedulers and I/O schedulers are independent. However, when the bottleneck is I/O, there is an opportunity to alleviate it via cooperation between the I/O and process schedulers: the service efficiency of I/O requests can be highly dependent on their issuance order, which in turn is heavily influenced by process scheduling. We propose a data-driven program execution mode in which process scheduling and request issuance are coordinated to facilitate effective I/O scheduling for high disk efficiency. Our implementation, Dual Par, uses process suspension and resumption, as well as pre-execution and prefetching techniques, to provide a pool of pre-sorted requests to the I/O scheduler. This data-driven execution mode is enabled when I/O is detected to be the bottleneck, otherwise the program runs in the normal computation-driven mode. Dual Par is implemented in the MPICH2 MPI-IO library for MPI programs to coordinate I/O service and process execution. Our experiments on a 120-node cluster using the PVFS2 file system show that Dual Par can increase system I/O throughput by 31% on average, compared to existing MPI-IO with or without using collective I/O.
Xuechen Zhang 0001, Kei Davis, Song Jiang 0001
IPDPS2
2012 iTransformer: Using SSD to Improve Disk Scheduling for High-performance I/O
abstract
The parallel data accesses inherent to large-scale data-intensive scientific computing require that data servers handle very high I/O concurrency. Concurrent requests from different processes or programs to hard disk can cause disk head thrashing between different disk regions, resulting in unacceptably low I/O performance. Current storage systems either rely on the disk scheduler at each data server, or use SSD as storage, to minimize this negative performance effect. However, the ability of the scheduler to alleviate this problem by scheduling requests in memory is limited by concerns such as long disk access times, and potential loss of dirty data with system failure. Meanwhile, SSD is too expensive to be widely used as the major storage device in the HPC environment. We propose iTransformer, a scheme that employs a small SSD to schedule requests for the data on disk. Being less space constrained than with more expensive DRAM, iTransformer can buffer larger amounts of dirty data before writing it back to the disk, or prefetch a larger volume of data in a batch into the SSD. In both cases high disk efficiency can be maintained even for concurrent requests. Furthermore, the scheme allows the scheduling of requests in the background to hide the cost of random disk access behind serving process requests. Finally, as a non-volatile memory, concerns about the quantity of dirty data are obviated. We have implemented iTransformer in the Linux kernel and tested it on a large cluster running PVFS2. Our experiments show that iTransformer can improve the I/O throughput of the cluster by 35% on average for MPI/IO benchmarks of various data access patterns.
Xuechen Zhang 0001, Kei Davis, Song Jiang 0001
IPDPS2
2011 QoS support for end users of I/O-intensive applications using shared storage systems
abstract
While the performance of compute-bound applications can be effectively guaranteed with techniques such as space sharing or QoS-aware process scheduling, it remains a challenge to meet QoS requirements for end users of I/O-intensive applications using shared storage systems because of the difficulty of differentiating I/O services for different applications with individual quality requirements. Furthermore, it is difficult for end users to accurately specify performance goals to the storage system using I/O-related metrics such as request latency or throughput. As access patterns, request rates, and the system workload change in time, a fixed I/O performance goal, such as bounds on throughput or latency, can be expensive to achieve and may not provide performance guarantees such as bounded program execution time.
Xuechen Zhang 0001, Kei Davis, Song Jiang 0001
SC2
2010 IOrchestrator: Improving the Performance of Multi-node I/O Systems via Inter-Server Coordination
abstract
A cluster of data servers and a parallel file system are often used to provide high-throughput I/O service to parallel programs running on a compute cluster. To exploit I/O parallelism parallel file systems stripe file data across the data servers. While this practice is effective in serving asynchronous requests, it may break individual program's spatial locality, which can seriously degrade I/O performance when the data servers concurrently serve synchronous requests from multiple I/O-intensive programs. In this paper we propose a scheme, IOrchestrator, to improve I/O performance of multi-node storage systems by orchestrating I/O services among programs when such inter-data-server coordination is dynamically determined to be cost effective. We have implemented IOrchestrator in the PVFS2 parallel file system. Our experiments with representative parallel benchmarks show that IOrchestrator can significantly improve I/O performance-- by up to a factor of 2.5--delivered by a cluster of data servers servicing concurrently-running parallel programs. Notably, we have not observed any scenarios in which the use of IOrchestrator causes substantial performance degradation.
Xuechen Zhang 0001, Kei Davis, Song Jiang 0001
SC2
2010 Improving Networked File System Performance Using a Locality-Aware Cooperative Cache Protocol
abstract
In a distributed environment, the utilization of file buffer caches in different clients may greatly vary. Cooperative caching has been proposed to increase cache utilization by coordinating the shared usage of distributed caches. It allows clients that would more greatly benefit from larger caches to forward data objects to peer clients with relatively underutilized caches. To support such coordination, global cache utilization must be dynamically evaluated. This, in turn, requires an effective analysis of application data access patterns. Existing coordination protocols are demonstrably suboptimal in this respect, exhibiting inefficient memory utilization and undue interference among clients. We propose a locality-aware cooperative caching protocol, called LAC, that is based on analysis and manipulation of data block reuse distance to effectively predict cache utilization and the probability of data reuse at each client. Using a dynamically adaptive synchronization technique, we keep local information up to date and consistently comparable across clients. The system is highly scalable in the sense that global coordination is achieved without centralized control. We have conducted thorough trace-driven simulation experiments to assess the performance differences between LAC and various existing protocols representative of the general class. Using a realistic and representative cost model, we show that the LAC protocol significantly and consistently outperforms existing cooperative caching protocols, demonstrating high and balanced utilization of caches across all clients. In our experiments, LAC reduces block access time by up to 36 percent, with an average of 31 percent, over the system without peer cache coordination, and reduces block access time by up to 22 percent, with an average of 13 percent, over the best performer of the existing protocols.
Song Jiang 0001, Xuechen Zhang 0001, Kei Davis
IEEE Trans. Computers4
2009 Performance modeling in action: Performance prediction of a Cray XT4 system during upgrade
abstract
We present predictive performance models of two of the petascale applications, S3D and GTC, from the DOE Office of Science workload. We outline the development of these models and demonstrate their validation on an Opteron/Infiniband cluster and the pre-upgrade ORNL Jaguar system (Cray XT3/XT4). Given the high accuracy of the full application models, we predict the performance of the Jaguar system after the upgrade of its nodes, and subsequently compare this to the actual performance of the upgraded system. We then analyze the performance of the system based on the models to quantify bottlenecks and potential optimizations. Finally, the models are used to quantify the benefits of alternative node allocation strategies.
Kevin J. Barker, Kei Davis, Darren J. Kerbyson
IPDPS2
2009 Making resonance a common case: A high-performance implementation of collective I/O on parallel file systems
abstract
Collective I/O is a widely used technique to improve I/O performance in parallel computing. It can be implemented as a client-based or as a server-based scheme. The client-based implementation is more widely adopted in the MPIIO software such as ROMIO because of its independence from the storage system configuration and its greater portability. However, existing implementations of client-side collective I/O do not consider the actual pattern of file striping over multiple I/O nodes in the storage system. This can cause a large number of requests for non-sequential data at I/O nodes, substantially degrading I/O performance. Investigating a surprisingly high I/O throughput achieved when there is an accidental match between a particular request pattern and the data striping pattern on the I/O nodes, we reveal the resonance phenomenon as the cause. Exploiting readily available information on data striping from the metadata server in popular file systems such as PVFS2 and Lustre, we design a new collective I/O implementation technique, named as resonant I/O, that makes resonance a common case. Resonant I/O rearranges requests from multiple MPI processes according to the presumed data layout on the disks of I/O nodes so that non-sequential access of disk data can be turned into sequential access, significantly improving I/O performance without compromising the independence of a client-based implementation. We have implemented our design in ROMIO. Our experimental results on a small- and medium-scale cluster show that the scheme can increase I/O throughput for some commonly used parallel I/O benchmarks such as mpi-io-test and ior-mpi-io over the existing implementation of ROMIO by up to 157%, with no scenario demonstrating significantly decreased performance.
Xuechen Zhang 0001, Song Jiang 0001, Kei Davis
IPDPS3
2008 Experiences in scaling scientific applications on current-generation quad-core processors
abstract
In this work we present an initial performance evaluation of AMD and Intel's first quad-core processor offerings: the AMD Barcelona and the Intel Xeon X7350. We examine the suitability of these processors in quad-socket compute nodes as building blocks for large-scale scientific computing clusters. Our analysis of intra-processor and intra-node scalability of microbenchmarks and a range of large- scale scientific applications indicates that quad-core processors can deliver an improvement in performance of up to 4x per processor but is heavily dependent on the workload being processed. While the Intel processor has a higher clock rate and peak performance, the AMD processor has higher memory bandwidth and intra-node scalability. The scientific applications we analyzed exhibit a range of performance improvements from only 3x up to the full 16x speed-up over a single core. Also, we note that the maximum node performance is not necessarily achieved by using all 16 cores.
Kevin J. Barker, Kei Davis, Adolfy Hoisie, Darren J. Kerbyson, Michael Lang 0003, Scott Pakin, José Carlos Sancho
IPDPS2
2008 Entering the petaflop era: the architecture and performance of Roadrunner
abstract
Roadrunner is a 1.38 Pflop/s-peak (double precision) hybrid-architecture supercomputer developed by LANL and IBM. It contains 12,240 IBM PowerXCell 8i processors and 12,240 AMD Opteron cores in 3,060 compute nodes. Roadrunner is the first supercomputer to run Linpack at a sustained speed in excess of 1 Pflop/s. In this paper we present a detailed architectural description of Roadrunner and a detailed performance analysis of the system. A case study of optimizing the MPI-based application Sweep3D to exploit Roadrunner's hybrid architecture is also included. The performance of Sweep3D is compared to that of the code on a previous implementation of the Cell Broadband Engine architecture-the Cell BE-and on multi-core processors. Using validated performance models combined with Roadrunner-specific microbenchmarks we identify performance issues in the early pre-delivery system and infer how well the final Roadrunner configuration will perform once the system software stack has matured.
Kevin J. Barker, Kei Davis, Adolfy Hoisie, Darren J. Kerbyson, Michael Lang 0003, Scott Pakin, José Carlos Sancho
SC2
2007 DiskSeen: Exploiting Disk Layout and Access History to Enhance I/O Prefetch
Xiaoning Ding, Song Jiang 0001, Feng Chen 0005, Kei Davis, Xiaodong Zhang 0001
USENIX ATC4
2007 Coordinated Multilevel Buffer Cache Management with Consistent Access Locality Quantification
Song Jiang 0001, Kei Davis, Xiaodong Zhang 0001
IEEE Trans. Computers2
2006 MPI tools and performance studies - Quantifying the potential benefit of overlapping communication and computation in large-scale scientific applications
abstract
The design and implementation of a high performance communication network are critical factors in determining the performance and cost-effectiveness of a largescale computing system. The major issues center on the trade-off between the network cost and the impact of latency and bandwidth on application performance. One promising technique for extracting maximum application performance given limited network resources is based on overlapping computation with communication, which partially or entirely hides communication delays. While this approach is not new, there are few studies that quantify the potential benefit of such overlapping for large-scale production scientific codes. We address this with an empirical method combined with a network model to quantify the potential overlap in several codes and examine the possible performance benefit. Our results demonstrate, for the codes examined, that a high potential tolerance to network latency and bandwidth exists because of a high degree of potential overlap. Moreover, our results indicate that there is often no need to use finegrained communication mechanisms to achieve this benefit, since the major source of potential overlap is found in independent work--computation on which pending messages does not depend. This allows for a potentially significant relaxation of network requirements without a consequent degradation of application performance.
José Carlos Sancho, Kevin J. Barker, Darren J. Kerbyson, Kei Davis
SC4
2004 Designing Parallel Operating Systems via Parallel Programming
Eitan Frachtenberg, Kei Davis, Fabrizio Petrini, Juan Fernández Peinador, José Carlos Sancho
Euro-Par2
2004 System-Level Fault-Tolerance in Large-Scale Parallel Machines with Buffered Coscheduling
abstract
Summary form only given. As the number of processors for multiteraflop systems grows to tens of thousands, with proposed petaflops systems likely to contain hundreds of thousands of processors, the assumption of fully reliable hardware has been abandoned. Although the mean time between failures for the individual components can be very high, the large total component count will inevitably lead to frequent failures. It is therefore of paramount importance to develop new software solutions to deal with the unavoidable reality of hardware faults. We will first describe the nature of the failures of current large-scale machines, and extrapolate these results to future machines. Based on this preliminary analysis we will present a new technology that we are currently developing, buffered coscheduling, which seeks to implement fault tolerance at the operating system level. Major design goals include dynamic reallocation of resources to allow continuing execution in the presence of hardware failures, very high scalability, high efficiency (low overhead), and transparency - requiring no changes to user applications. Preliminary results show that this is attainable with current hardware.
Fabrizio Petrini, Kei Davis, José Carlos Sancho
IPDPS2
2004 A Performance and Scalability Analysis of the BlueGene/L Architecture
abstract
Based on a set of measurements done on the 512-node 500MHz prototype and early results on a 2048 node 700MHz BlueGene/L machine at IBM Watson, we present a performance and scalability analysis of the architecture from low-level characteristics to large-scale applications. In addition, we present predictions using our models for the performance of two representative applications from the ASC² workload on the full BlueGene/L configuration of 64K nodes. We have compared the measured values for several of the benchmarks in our suite against the predicted numbers from our performance models. In general, the error bars were relatively low. A comparison between the performance of BlueGene/L and the ASCI Q, the largest supercomputer in the US, is presented, also based on our predictive performance models.
Kei Davis, Adolfy Hoisie, Greg Johnson, Darren J. Kerbyson, Michael Lang 0003, Scott Pakin, Fabrizio Petrini
SC1
2000 The Multi-architecture Performance of the Parallel Functional Language GP H (Research Note)
Philip W. Trinder, Hans-Wolfgang Loidl, Ed. Barry Jr., Kei Davis, Kevin Hammond, Ulrike Klusik, Simon L. Peyton Jones, Álvaro J. Rebón Portillo
Euro-Par4
1998 OVERTURE: An Object-Oriented Framework for High Performance Scientific Computing
abstract
The Overture Framework is an object-oriented environment for solving PDEs on serial and parallel architectures. It is a collection of C++ libraries that enables the use of finite difference and finite volume methods at a level that hides the details of the associated data structures, as well as the details of the parallel implementation. It is based on the A++/P++ array class library and is designed for solving problems on a structured grid or a collection of structured grids. In particular, it can use curvilinear grids, adaptive mesh refinement and the composite overlapping grid methods to represent problems with complex moving geometry. This paper introduces Overture, its motivation, and specifically the aspects of the design central to portability and high performance. In particular we focus on the mechanisms within Overture that permit a hierarchy of abstractions and those mechanisms which permit their efficiency on advanced serial and parallel architectures. We expect that these same mechanisms will become increasingly important within other object-oriented frameworks in the future.
Federico Bassetti, David L. Brown, Kei Davis, William D. Henshaw, Daniel J. Quinlan
SC3
1994 PERs from Projections for Binding-Time Analysis
Kei Davis
PEPM1
1993 Higher-order Binding-time Analysis
abstract
The partial evaluation process requires a binding-time analysis. Binding-time analysis seeks to determine which parts of a program's result is determined when some part of the input is known. Domain projections provide a very general way to encode a description of which parts of a data structure are static (known), and which are dynamic (not static). For first-order functional languages Launchbury [Lau91a] has developed an abstract interpretation technique for binding-time analysis in which the basic abstract value is a projection. Unfortunately this technique does not generalise easily to higher-order languages. This paper develops such a generalisation: a projection-based abstract interpretation suitable for higher-order binding-time analysis.
Kei Davis
PEPM1