EDBT 2026 Demo / reviewers in the wild / expert
Kei Davis
dblp:49/4295
· DBLP profile ↗
27ranked-venue papers
3as first author
0since 2021 · last 2020
0000-0002-4134-1798ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 24 · 1 first-authorSoftware engineering, systems software and programming languages · 3 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
12 papers |
Storage systems · 32% High-performance computing · 28% Memory systems · 26% | |
| Software engineering, system software, and programming languages
2 papers |
Operating systems · 100% |
Topics — the 30 heaviest of 39, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
High-performance computing
scientific computing systems |
0.4 | 2 | 2019 | Persistent Octrees for Parallel Mesh Refinement through Non-Volatile Byte-Addressable Memory · IEEE Trans. Parallel Distributed Syst. 2019 OVERTURE: An Object-Oriented Framework for High Performance Scientific Computing · SC 1998 |
High-performance computing › scientific computing systems
adaptive mesh refinement |
0.4 | 1 | 2019 | Persistent Octrees for Parallel Mesh Refinement through Non-Volatile Byte-Addressable Memory · IEEE Trans. Parallel Distributed Syst. 2019 |
Memory systems › non-volatile memory › persistent memory
byte-addressable persistent memory |
0.4 | 1 | 2019 | Persistent Octrees for Parallel Mesh Refinement through Non-Volatile Byte-Addressable Memory · IEEE Trans. Parallel Distributed Syst. 2019 |
Distributed systems › fault tolerance
failure recovery |
0.4 | 1 | 2019 | Persistent Octrees for Parallel Mesh Refinement through Non-Volatile Byte-Addressable Memory · IEEE Trans. Parallel Distributed Syst. 2019 |
Memory systems
non-volatile memory |
0.4 | 1 | 2019 | Persistent Octrees for Parallel Mesh Refinement through Non-Volatile Byte-Addressable Memory · IEEE Trans. Parallel Distributed Syst. 2019 |
Storage systems
storage reliability |
0.4 | 1 | 2019 | Persistent Octrees for Parallel Mesh Refinement through Non-Volatile Byte-Addressable Memory · IEEE Trans. Parallel Distributed Syst. 2019 |
Storage systems › i/o optimization
collective i/o |
0.2 | 1 | 2013 | Orthrus: a framework for implementing high-performance collective I/O in the multicore clusters · HPDC 2013 |
Storage systems › i/o optimization › i/o prefetching
disk prefetching |
0.2 | 1 | 2013 | A Prefetching Scheme Exploiting both Data Layout and Access History on Disk · ACM Trans. Storage 2013 |
High-performance computing
parallel i/o |
0.2 | 1 | 2013 | Orthrus: a framework for implementing high-performance collective I/O in the multicore clusters · HPDC 2013 |
Storage systems › file systems › distributed file system
parallel file system |
0.2 | 2 | 2013 | IOrchestrator: Improving the Performance of Multi-node I/O Systems via Inter-Server Coordination · SC 2010 Orthrus: a framework for implementing high-performance collective I/O in the multicore clusters · HPDC 2013 |
High-performance computing › supercomputing
supercomputing systems |
0.1 | 2 | 2008 | Entering the petaflop era: the architecture and performance of Roadrunner · SC 2008 A Performance and Scalability Analysis of the BlueGene/L Architecture · SC 2004 |
Storage systems
shared storage |
0.1 | 1 | 2011 | QoS support for end users of I/O-intensive applications using shared storage systems · SC 2011 |
Storage systems › storage performance
storage quality of service |
0.1 | 1 | 2011 | QoS support for end users of I/O-intensive applications using shared storage systems · SC 2011 |
Memory systems › non-volatile memory
persistent data structures |
0.1 | 1 | 2019 | Persistent Octrees for Parallel Mesh Refinement through Non-Volatile Byte-Addressable Memory · IEEE Trans. Parallel Distributed Syst. 2019 |
Memory systems › cache management › chip multiprocessor cache management
cache coordination |
0.1 | 1 | 2010 | Improving Networked File System Performance Using a Locality-Aware Cooperative Cache Protocol · IEEE Trans. Computers 2010 |
Memory systems › cache management › storage caching
cooperative caching |
0.1 | 1 | 2010 | Improving Networked File System Performance Using a Locality-Aware Cooperative Cache Protocol · IEEE Trans. Computers 2010 |
Storage systems › file systems
distributed file system |
0.1 | 1 | 2010 | Improving Networked File System Performance Using a Locality-Aware Cooperative Cache Protocol · IEEE Trans. Computers 2010 |
Storage systems › i/o architecture
i/o subsystem |
0.1 | 1 | 2010 | IOrchestrator: Improving the Performance of Multi-node I/O Systems via Inter-Server Coordination · SC 2010 |
Hardware accelerators and domain-specific architectures › many-core accelerator
cell broadband engine |
0.1 | 1 | 2008 | Entering the petaflop era: the architecture and performance of Roadrunner · SC 2008 |
Storage systems › buffer management
buffer cache management |
0.1 | 1 | 2007 | Coordinated Multilevel Buffer Cache Management with Consistent Access Locality Quantification · IEEE Trans. Computers 2007 |
Memory systems
cache management |
0.1 | 1 | 2007 | Coordinated Multilevel Buffer Cache Management with Consistent Access Locality Quantification · IEEE Trans. Computers 2007 |
Memory systems › cache management
cache replacement |
0.1 | 1 | 2007 | Coordinated Multilevel Buffer Cache Management with Consistent Access Locality Quantification · IEEE Trans. Computers 2007 |
Storage systems › i/o optimization
i/o prefetching |
0.1 | 1 | 2007 | DiskSeen: Exploiting Disk Layout and Access History to Enhance I/O Prefetch · USENIX ATC 2007 |
Operating systems › resource management › storage management
file systems |
0.1 | 2 | 2013 | A Prefetching Scheme Exploiting both Data Layout and Access History on Disk · ACM Trans. Storage 2013 DiskSeen: Exploiting Disk Layout and Access History to Enhance I/O Prefetch · USENIX ATC 2007 |
Distributed systems › communication optimization
communication-computation overlap |
0.1 | 1 | 2006 | MPI tools and performance studies - Quantifying the potential benefit of overlapping communication and computation in large-scale scientific applications · SC 2006 |
High-performance computing
performance optimization at scale |
0.1 | 1 | 2006 | MPI tools and performance studies - Quantifying the potential benefit of overlapping communication and computation in large-scale scientific applications · SC 2006 |
High-performance computing › supercomputing
bluegene/l |
0.0 | 1 | 2004 | A Performance and Scalability Analysis of the BlueGene/L Architecture · SC 2004 |
High-performance computing
i/o intensive applications |
0.0 | 1 | 2011 | QoS support for end users of I/O-intensive applications using shared storage systems · SC 2011 |
Distributed systems
distributed coordination |
0.0 | 1 | 2010 | IOrchestrator: Improving the Performance of Multi-node I/O Systems via Inter-Server Coordination · SC 2010 |
Performance modeling and evaluation › benchmarking
microbenchmarking |
0.0 | 1 | 2008 | Entering the petaflop era: the architecture and performance of Roadrunner · SC 2008 |
Methods — techniques the papers use, named apart from their topics
feature-directed sampling · 0.4erasure coding · 0.4prefetching · 0.3access pattern tracking · 0.3performance modeling · 0.1qos-aware scheduling · 0.1trace-driven simulation · 0.1reuse distance analysis · 0.1i/o orchestration · 0.1microbenchmarks · 0.1finite volume · 0.0finite difference · 0.0composite overlapping grid · 0.0adaptive mesh refinement · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | On the memory attribution problem: A solution and case study using MPIabstractSummary As parallel applications running on large‐scale computing systems become increasingly memory constrained, the ability to attribute memory usage to the various components of the application is becoming increasingly important. We present the design and implementation of memnesia, a novel memory usage profiler for parallel and distributed message‐passing applications. Our approach captures both application– and message‐passing library–specific memory usage statistics from unmodified binaries dynamically linked to a message‐passing communication library. Using microbenchmarks and proxy applications, we evaluated our profiler across three Message Passing Interface (MPI) implementations and two hardware platforms. The results show that our approach and the corresponding implementation can accurately quantify memory resource usage as a function of time, scale, communication workload, and software or hardware system architecture, clearly distinguishing between application and MPI library memory usage at a per‐process level. With this new capability, we show that job size, communication workload, and hardware/software architecture influence peak runtime memory usage. In practice, this tool provides a potentially valuable source of information for application developers seeking to measure and optimize memory usage. Samuel K. Gutierrez, Dorian C. Arnold, Kei Davis, Patrick S. McCormick |
Concurr. Comput. Pract. Exp. | 3 |
| 2019 | Persistent Octrees for Parallel Mesh Refinement through Non-Volatile Byte-Addressable MemoryabstractAdaptive mesh refinement based on octree data structures has enabled efficient simulations of complex physical phenomena. Existing meshing algorithms were proposed with the assumption that computer memory is volatile. Consequently, for failure recovery, in-core algorithms need to save memory states as snapshots with slow file I/O, while out-of-core algorithms store octants on disk for persistence. However, neither was designed to best exploit the unique characteristics of non-volatile byte-addressable memory (NVBM). We propose a novel data structure, the Distributed Persistent Merged octree (DPM-octree), for both meshing and in-memory storage of persistent octrees using NVBM. DPM-octree is a multi-version data structure that can recover from failures using an earlier persistent version stored in NVBM. In addition, we design a feature-directed sampling approach to help dynamically transform the DPM-octree layout for reducing NVBM-induced memory write latency. DPM-octree uses parity trees which are created using erasure coding and stored in NVBM to support low-latency in-memory octant recovery after data loss. DPM-octree has been successfully integrated with the Gerris software for simulation of fluid dynamics. Our experimental results with real-world scientific workloads show that DPM-octree scales up to 1.1 billion mesh elements with 1,000 processors on the Titan supercomputer. Bao Nguyen, Hua Tan, Kei Davis, Xuechen Zhang 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2017 | Accommodating Thread-Level Heterogeneity in Coupled Parallel ApplicationsabstractHybrid parallel program models that combine message passing and multithreading (MP+MT) are becoming more popular, extending the basic message passing (MP) model that uses single-threaded processes for both inter- and intra-node parallelism. A consequence is that coupled parallel applications increasingly comprise MP libraries together with MP+MT libraries with differing preferred degrees of threading, resulting in thread-level heterogeneity. Retroactively matching threading levels between independently developed and maintained libraries is difficult; the challenge is exacerbated because contemporary parallel job launchers provide only static resource binding policies over entire application executions. A standard approach for accommodating thread-level heterogeneity is to under-subscribe compute resources such that the library with the highest degree of threading per process has one processing element per thread. This results in libraries with fewer threads per process utilizing only a fraction of the available compute resources. We present and evaluate a novel approach for accommodating thread-level heterogeneity. Our approach enables full utilization of all available compute resources throughout an application's execution by providing programmable facilities to dynamically reconfigure runtime environments for compute phases with differing threading factors and memory affinities. We show that our approach can improve overall application performance by up to 5.8× in real-world production codes. Furthermore, the practicality and utility of our approach has been demonstrated by continuous production use for over one year, and by more recent incorporation into a number of production codes. Samuel K. Gutierrez, Kei Davis, Dorian C. Arnold, Randal S. Baker, Robert W. Robey, Patrick S. McCormick, Daniel Holladay, Jon A. Dahl, Joe Zerr, Florian Weik, Christoph Junghans |
IPDPS | 2 |
| 2013 | Orthrus: a framework for implementing high-performance collective I/O in the multicore clusters
Xuechen Zhang 0001, Jianqiang Ou, Kei Davis, Song Jiang 0001 |
HPDC | 3 |
| 2013 | iBridge: Improving Unaligned Parallel File Access with Solid-State DrivesabstractWhen files are striped in a parallel I/O system, requests to the files are decomposed into a number of sub-requests that are distributed over multiple servers. If a request is not aligned with the striping pattern such decomposition can make the first and last sub-requests much smaller than the striping unit. Because hard-disk-based servers can be much less efficient in serving small requests than large ones, the system exhibits heterogeneity in serving sub-requests of different sizes, and the net throughput of the entire system can be severely degraded by the inefficiency of serving the smaller requests, or fragments. Because a request is not considered complete until its slowest sub-request is, the penalty is yet greater for synchronous requests. To make the situation even worse, the larger the request, or the more data servers the requested data is striped over, the larger the detrimental performance effect of serving fragments can be. This effect can become the Achilles' heel of a parallel I/O system performance seeking scalability with large sequential accesses. In this paper we propose iBridge, a scheme that uses solid-state drives to serve request fragments and thereby bridge the performance gap between serving fragments and serving large sub-requests. We have implemented iBridge in the PVFS file system. Our experimental results with representative MPI-IO benchmarks show that iBridge can significantly improve the I/O throughput of storage systems, especially for large requests with fragments. Xuechen Zhang 0001, Kei Davis, Song Jiang 0001 |
IPDPS | 3 |
| 2013 | Synergistic coupling of SSD and hard disk for QoS-aware virtual memoryabstractWith significant advantages in capacity, power consumption, and price, solid state disk (SSD) has good potential to be employed as an extension of DRAM (memory), such that applications with large working sets could run efficiently on a modestly configured system. While initial results reported in recent works show promising prospects for this use of SSD by incorporating it into the management of virtual memory, frequent writes from write-intensive programs could quickly wear out SSD, making the idea less practical. We propose a scheme, HybridSwap, that integrates a hard disk with an SSD for virtual memory management, synergistically achieving the advantages of both. In addition, HybridSwap can constrain performance loss caused by swapping according to user-specified QoS requirements. Xuechen Zhang 0001, Kei Davis, Song Jiang 0001 |
ISPASS | 3 |
| 2013 | A Prefetching Scheme Exploiting both Data Layout and Access History on DiskabstractPrefetching is an important technique for improving effective hard disk performance. A prefetcher seeks to accurately predict which data will be requested and load it ahead of the arrival of the corresponding requests. Current disk prefetch policies in major operating systems track access patterns at the level of file abstraction. While this is useful for exploiting application-level access patterns, for two reasons file-level prefetching cannot realize the full performance improvements achievable by prefetching. First, certain prefetch opportunities can only be detected by knowing the data layout on disk, such as the contiguous layout of file metadata or data from multiple files. Second, nonsequential access of disk data (requiring disk head movement) is much slower than sequential access, and the performance penalty for mis-prefetching a randomly located block, relative to that of a sequential block, is correspondingly greater. Song Jiang 0001, Xiaoning Ding, Yuehai Xu, Kei Davis |
ACM Trans. Storage | 4 |
| 2012 | iHarmonizer: Improving the Disk Efficiency of I/O-intensive Multithreaded CodesabstractChallenged by serious power and thermal constraints and limited by available instruction-level parallelism, processor designs have evolved to multi-core architectures. These architectures, many augmented with native simultaneous multithreading, are driving software developers to use multithreaded programs to exploit thread-level parallelism. While multithreading is well known to introduce concerns of data dependency and CPU load balance, less known is that the uncertainty of relative progress of thread execution can cause patterns of I/O requests, issued by different threads, to be effectively random and so significantly degrade hard-disk efficiency. This effect can severely offset the performance gains from parallel execution, especially for I/O-intensive programs. Retaining the benefits of multithreading while not losing I/O efficiency is an urgent and challenging problem. We propose a user-level scheme, iHarmonizer, to streamline the servicing of I/O requests from multiple threads in the Open MP programs. Specifically, we use the compiler to insert code into Open MP programs so that data usage can be transmitted at run time to a supporting run-time library that prefetches data in a disk friendly way and coordinates threads' execution according to the availability of their requested data. Transparent to the programmer, iHarmonizer makes a multithreaded program I/O efficient while maintaining the benefits of parallelism. Our experiments show that iHarmonizer can significantly speed up the execution of a representative set of I/O-intensive scientific benchmarks. Kei Davis, Yuehai Xu, Song Jiang 0001 |
IPDPS | 2 |
| 2012 | Opportunistic Data-driven Execution of Parallel Programs for Efficient I/O ServicesabstractA parallel system relies on both process scheduling and I/O scheduling for efficient use of resources, and a program's performance hinges on the resource on which it is bottlenecked. Existing process schedulers and I/O schedulers are independent. However, when the bottleneck is I/O, there is an opportunity to alleviate it via cooperation between the I/O and process schedulers: the service efficiency of I/O requests can be highly dependent on their issuance order, which in turn is heavily influenced by process scheduling. We propose a data-driven program execution mode in which process scheduling and request issuance are coordinated to facilitate effective I/O scheduling for high disk efficiency. Our implementation, Dual Par, uses process suspension and resumption, as well as pre-execution and prefetching techniques, to provide a pool of pre-sorted requests to the I/O scheduler. This data-driven execution mode is enabled when I/O is detected to be the bottleneck, otherwise the program runs in the normal computation-driven mode. Dual Par is implemented in the MPICH2 MPI-IO library for MPI programs to coordinate I/O service and process execution. Our experiments on a 120-node cluster using the PVFS2 file system show that Dual Par can increase system I/O throughput by 31% on average, compared to existing MPI-IO with or without using collective I/O. Xuechen Zhang 0001, Kei Davis, Song Jiang 0001 |
IPDPS | 2 |
| 2012 | iTransformer: Using SSD to Improve Disk Scheduling for High-performance I/OabstractThe parallel data accesses inherent to large-scale data-intensive scientific computing require that data servers handle very high I/O concurrency. Concurrent requests from different processes or programs to hard disk can cause disk head thrashing between different disk regions, resulting in unacceptably low I/O performance. Current storage systems either rely on the disk scheduler at each data server, or use SSD as storage, to minimize this negative performance effect. However, the ability of the scheduler to alleviate this problem by scheduling requests in memory is limited by concerns such as long disk access times, and potential loss of dirty data with system failure. Meanwhile, SSD is too expensive to be widely used as the major storage device in the HPC environment. We propose iTransformer, a scheme that employs a small SSD to schedule requests for the data on disk. Being less space constrained than with more expensive DRAM, iTransformer can buffer larger amounts of dirty data before writing it back to the disk, or prefetch a larger volume of data in a batch into the SSD. In both cases high disk efficiency can be maintained even for concurrent requests. Furthermore, the scheme allows the scheduling of requests in the background to hide the cost of random disk access behind serving process requests. Finally, as a non-volatile memory, concerns about the quantity of dirty data are obviated. We have implemented iTransformer in the Linux kernel and tested it on a large cluster running PVFS2. Our experiments show that iTransformer can improve the I/O throughput of the cluster by 35% on average for MPI/IO benchmarks of various data access patterns. Xuechen Zhang 0001, Kei Davis, Song Jiang 0001 |
IPDPS | 2 |
| 2011 | QoS support for end users of I/O-intensive applications using shared storage systemsabstractWhile the performance of compute-bound applications can be effectively guaranteed with techniques such as space sharing or QoS-aware process scheduling, it remains a challenge to meet QoS requirements for end users of I/O-intensive applications using shared storage systems because of the difficulty of differentiating I/O services for different applications with individual quality requirements. Furthermore, it is difficult for end users to accurately specify performance goals to the storage system using I/O-related metrics such as request latency or throughput. As access patterns, request rates, and the system workload change in time, a fixed I/O performance goal, such as bounds on throughput or latency, can be expensive to achieve and may not provide performance guarantees such as bounded program execution time. Xuechen Zhang 0001, Kei Davis, Song Jiang 0001 |
SC | 2 |
| 2010 | IOrchestrator: Improving the Performance of Multi-node I/O Systems via Inter-Server CoordinationabstractA cluster of data servers and a parallel file system are often used to provide high-throughput I/O service to parallel programs running on a compute cluster. To exploit I/O parallelism parallel file systems stripe file data across the data servers. While this practice is effective in serving asynchronous requests, it may break individual program's spatial locality, which can seriously degrade I/O performance when the data servers concurrently serve synchronous requests from multiple I/O-intensive programs. In this paper we propose a scheme, IOrchestrator, to improve I/O performance of multi-node storage systems by orchestrating I/O services among programs when such inter-data-server coordination is dynamically determined to be cost effective. We have implemented IOrchestrator in the PVFS2 parallel file system. Our experiments with representative parallel benchmarks show that IOrchestrator can significantly improve I/O performance-- by up to a factor of 2.5--delivered by a cluster of data servers servicing concurrently-running parallel programs. Notably, we have not observed any scenarios in which the use of IOrchestrator causes substantial performance degradation. Xuechen Zhang 0001, Kei Davis, Song Jiang 0001 |
SC | 2 |
| 2010 | Improving Networked File System Performance Using a Locality-Aware Cooperative Cache ProtocolabstractIn a distributed environment, the utilization of file buffer caches in different clients may greatly vary. Cooperative caching has been proposed to increase cache utilization by coordinating the shared usage of distributed caches. It allows clients that would more greatly benefit from larger caches to forward data objects to peer clients with relatively underutilized caches. To support such coordination, global cache utilization must be dynamically evaluated. This, in turn, requires an effective analysis of application data access patterns. Existing coordination protocols are demonstrably suboptimal in this respect, exhibiting inefficient memory utilization and undue interference among clients. We propose a locality-aware cooperative caching protocol, called LAC, that is based on analysis and manipulation of data block reuse distance to effectively predict cache utilization and the probability of data reuse at each client. Using a dynamically adaptive synchronization technique, we keep local information up to date and consistently comparable across clients. The system is highly scalable in the sense that global coordination is achieved without centralized control. We have conducted thorough trace-driven simulation experiments to assess the performance differences between LAC and various existing protocols representative of the general class. Using a realistic and representative cost model, we show that the LAC protocol significantly and consistently outperforms existing cooperative caching protocols, demonstrating high and balanced utilization of caches across all clients. In our experiments, LAC reduces block access time by up to 36 percent, with an average of 31 percent, over the system without peer cache coordination, and reduces block access time by up to 22 percent, with an average of 13 percent, over the best performer of the existing protocols. Song Jiang 0001, Xuechen Zhang 0001, Kei Davis |
IEEE Trans. Computers | 4 |
| 2009 | Performance modeling in action: Performance prediction of a Cray XT4 system during upgradeabstractWe present predictive performance models of two of the petascale applications, S3D and GTC, from the DOE Office of Science workload. We outline the development of these models and demonstrate their validation on an Opteron/Infiniband cluster and the pre-upgrade ORNL Jaguar system (Cray XT3/XT4). Given the high accuracy of the full application models, we predict the performance of the Jaguar system after the upgrade of its nodes, and subsequently compare this to the actual performance of the upgraded system. We then analyze the performance of the system based on the models to quantify bottlenecks and potential optimizations. Finally, the models are used to quantify the benefits of alternative node allocation strategies. Kevin J. Barker, Kei Davis, Darren J. Kerbyson |
IPDPS | 2 |
| 2009 | Making resonance a common case: A high-performance implementation of collective I/O on parallel file systemsabstractCollective I/O is a widely used technique to improve I/O performance in parallel computing. It can be implemented as a client-based or as a server-based scheme. The client-based implementation is more widely adopted in the MPIIO software such as ROMIO because of its independence from the storage system configuration and its greater portability. However, existing implementations of client-side collective I/O do not consider the actual pattern of file striping over multiple I/O nodes in the storage system. This can cause a large number of requests for non-sequential data at I/O nodes, substantially degrading I/O performance. Investigating a surprisingly high I/O throughput achieved when there is an accidental match between a particular request pattern and the data striping pattern on the I/O nodes, we reveal the resonance phenomenon as the cause. Exploiting readily available information on data striping from the metadata server in popular file systems such as PVFS2 and Lustre, we design a new collective I/O implementation technique, named as resonant I/O, that makes resonance a common case. Resonant I/O rearranges requests from multiple MPI processes according to the presumed data layout on the disks of I/O nodes so that non-sequential access of disk data can be turned into sequential access, significantly improving I/O performance without compromising the independence of a client-based implementation. We have implemented our design in ROMIO. Our experimental results on a small- and medium-scale cluster show that the scheme can increase I/O throughput for some commonly used parallel I/O benchmarks such as mpi-io-test and ior-mpi-io over the existing implementation of ROMIO by up to 157%, with no scenario demonstrating significantly decreased performance. Xuechen Zhang 0001, Song Jiang 0001, Kei Davis |
IPDPS | 3 |
| 2008 | Experiences in scaling scientific applications on current-generation quad-core processorsabstractIn this work we present an initial performance evaluation of AMD and Intel's first quad-core processor offerings: the AMD Barcelona and the Intel Xeon X7350. We examine the suitability of these processors in quad-socket compute nodes as building blocks for large-scale scientific computing clusters. Our analysis of intra-processor and intra-node scalability of microbenchmarks and a range of large- scale scientific applications indicates that quad-core processors can deliver an improvement in performance of up to 4x per processor but is heavily dependent on the workload being processed. While the Intel processor has a higher clock rate and peak performance, the AMD processor has higher memory bandwidth and intra-node scalability. The scientific applications we analyzed exhibit a range of performance improvements from only 3x up to the full 16x speed-up over a single core. Also, we note that the maximum node performance is not necessarily achieved by using all 16 cores. Kevin J. Barker, Kei Davis, Adolfy Hoisie, Darren J. Kerbyson, Michael Lang 0003, Scott Pakin, José Carlos Sancho |
IPDPS | 2 |
| 2008 | Entering the petaflop era: the architecture and performance of RoadrunnerabstractRoadrunner is a 1.38 Pflop/s-peak (double precision) hybrid-architecture supercomputer developed by LANL and IBM. It contains 12,240 IBM PowerXCell 8i processors and 12,240 AMD Opteron cores in 3,060 compute nodes. Roadrunner is the first supercomputer to run Linpack at a sustained speed in excess of 1 Pflop/s. In this paper we present a detailed architectural description of Roadrunner and a detailed performance analysis of the system. A case study of optimizing the MPI-based application Sweep3D to exploit Roadrunner's hybrid architecture is also included. The performance of Sweep3D is compared to that of the code on a previous implementation of the Cell Broadband Engine architecture-the Cell BE-and on multi-core processors. Using validated performance models combined with Roadrunner-specific microbenchmarks we identify performance issues in the early pre-delivery system and infer how well the final Roadrunner configuration will perform once the system software stack has matured. Kevin J. Barker, Kei Davis, Adolfy Hoisie, Darren J. Kerbyson, Michael Lang 0003, Scott Pakin, José Carlos Sancho |
SC | 2 |
| 2007 | DiskSeen: Exploiting Disk Layout and Access History to Enhance I/O Prefetch
Xiaoning Ding, Song Jiang 0001, Feng Chen 0005, Kei Davis, Xiaodong Zhang 0001 |
USENIX ATC | 4 |
| 2007 | Coordinated Multilevel Buffer Cache Management with Consistent Access Locality Quantification
Song Jiang 0001, Kei Davis, Xiaodong Zhang 0001 |
IEEE Trans. Computers | 2 |
| 2006 | MPI tools and performance studies - Quantifying the potential benefit of overlapping communication and computation in large-scale scientific applicationsabstractThe design and implementation of a high performance communication network are critical factors in determining the performance and cost-effectiveness of a largescale computing system. The major issues center on the trade-off between the network cost and the impact of latency and bandwidth on application performance. One promising technique for extracting maximum application performance given limited network resources is based on overlapping computation with communication, which partially or entirely hides communication delays. While this approach is not new, there are few studies that quantify the potential benefit of such overlapping for large-scale production scientific codes. We address this with an empirical method combined with a network model to quantify the potential overlap in several codes and examine the possible performance benefit. Our results demonstrate, for the codes examined, that a high potential tolerance to network latency and bandwidth exists because of a high degree of potential overlap. Moreover, our results indicate that there is often no need to use finegrained communication mechanisms to achieve this benefit, since the major source of potential overlap is found in independent work--computation on which pending messages does not depend. This allows for a potentially significant relaxation of network requirements without a consequent degradation of application performance. José Carlos Sancho, Kevin J. Barker, Darren J. Kerbyson, Kei Davis |
SC | 4 |
| 2004 | Designing Parallel Operating Systems via Parallel Programming
Eitan Frachtenberg, Kei Davis, Fabrizio Petrini, Juan Fernández Peinador, José Carlos Sancho |
Euro-Par | 2 |
| 2004 | System-Level Fault-Tolerance in Large-Scale Parallel Machines with Buffered CoschedulingabstractSummary form only given. As the number of processors for multiteraflop systems grows to tens of thousands, with proposed petaflops systems likely to contain hundreds of thousands of processors, the assumption of fully reliable hardware has been abandoned. Although the mean time between failures for the individual components can be very high, the large total component count will inevitably lead to frequent failures. It is therefore of paramount importance to develop new software solutions to deal with the unavoidable reality of hardware faults. We will first describe the nature of the failures of current large-scale machines, and extrapolate these results to future machines. Based on this preliminary analysis we will present a new technology that we are currently developing, buffered coscheduling, which seeks to implement fault tolerance at the operating system level. Major design goals include dynamic reallocation of resources to allow continuing execution in the presence of hardware failures, very high scalability, high efficiency (low overhead), and transparency - requiring no changes to user applications. Preliminary results show that this is attainable with current hardware. Fabrizio Petrini, Kei Davis, José Carlos Sancho |
IPDPS | 2 |
| 2004 | A Performance and Scalability Analysis of the BlueGene/L ArchitectureabstractBased on a set of measurements done on the 512-node 500MHz prototype and early results on a 2048 node 700MHz BlueGene/L machine at IBM Watson, we present a performance and scalability analysis of the architecture from low-level characteristics to large-scale applications. In addition, we present predictions using our models for the performance of two representative applications from the ASC² workload on the full BlueGene/L configuration of 64K nodes. We have compared the measured values for several of the benchmarks in our suite against the predicted numbers from our performance models. In general, the error bars were relatively low. A comparison between the performance of BlueGene/L and the ASCI Q, the largest supercomputer in the US, is presented, also based on our predictive performance models. Kei Davis, Adolfy Hoisie, Greg Johnson, Darren J. Kerbyson, Michael Lang 0003, Scott Pakin, Fabrizio Petrini |
SC | 1 |
| 2000 | The Multi-architecture Performance of the Parallel Functional Language GP H (Research Note)
Philip W. Trinder, Hans-Wolfgang Loidl, Ed. Barry Jr., Kei Davis, Kevin Hammond, Ulrike Klusik, Simon L. Peyton Jones, Álvaro J. Rebón Portillo |
Euro-Par | 4 |
| 1998 | OVERTURE: An Object-Oriented Framework for High Performance Scientific ComputingabstractThe Overture Framework is an object-oriented environment for solving PDEs on serial and parallel architectures. It is a collection of C++ libraries that enables the use of finite difference and finite volume methods at a level that hides the details of the associated data structures, as well as the details of the parallel implementation. It is based on the A++/P++ array class library and is designed for solving problems on a structured grid or a collection of structured grids. In particular, it can use curvilinear grids, adaptive mesh refinement and the composite overlapping grid methods to represent problems with complex moving geometry. This paper introduces Overture, its motivation, and specifically the aspects of the design central to portability and high performance. In particular we focus on the mechanisms within Overture that permit a hierarchy of abstractions and those mechanisms which permit their efficiency on advanced serial and parallel architectures. We expect that these same mechanisms will become increasingly important within other object-oriented frameworks in the future. Federico Bassetti, David L. Brown, Kei Davis, William D. Henshaw, Daniel J. Quinlan |
SC | 3 |
| 1994 | PERs from Projections for Binding-Time Analysis
Kei Davis |
PEPM | 1 |
| 1993 | Higher-order Binding-time AnalysisabstractThe partial evaluation process requires a binding-time analysis. Binding-time analysis seeks to determine which parts of a program's result is determined when some part of the input is known. Domain projections provide a very general way to encode a description of which parts of a data structure are static (known), and which are dynamic (not static). For first-order functional languages Launchbury [Lau91a] has developed an abstract interpretation technique for binding-time analysis in which the basic abstract value is a projection. Unfortunately this technique does not generalise easily to higher-order languages. This paper develops such a generalisation: a projection-based abstract interpretation suitable for higher-order binding-time analysis. Kei Davis |
PEPM | 1 |